Embedded Web Servers in Spring Boot: Tomcat, Jetty, Undertow
Configure embedded servers in Spring Boot: compare Tomcat, Jetty, and Undertow, tune thread pools, enable access logs, and switch implementations.
Configure embedded servers in Spring Boot: compare Tomcat, Jetty, and Undertow, tune thread pools, enable access logs, and switch implementations. The guide uses practical examples to explain switching the embedded server, excluding the default and shows how to apply the ideas in a Spring Boot project. It closes with common pitfalls and production checks so you can apply the pattern with fewer surprises.
Embedded Web Servers in Spring Boot
Introduction
Switching the Embedded Server
Excluding the Default
spring-boot-starter-web pulls in Tomcat by default. To swap it out, exclude Tomcat and add your preferred server’s starter.
Maven:
<dependencies>
<!-- Exclude Tomcat -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-web</artifactId>
<exclusions>
<exclusion>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-tomcat</artifactId>
</exclusion>
</exclusions>
</dependency>
<!-- Use Jetty instead -->
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-jetty</artifactId>
</dependency>
</dependencies>
The same pattern works for Undertow, just swap the artifact ID to spring-boot-starter-undertow.
Dependency Fallback
Spring Boot’s dependency management means you do not need to specify versions when using its curated BOM. The framework pins compatible versions of each server. You can verify what gets pulled in with mvn dependency:tree | grep jetty or similar.
If you need a specific server version that differs from Spring Boot’s default, you can override the managed version property. For example, to use Jetty 12.x when Spring Boot manages 11.x, set jetty.version in your pom.xml. This works but introduces risk: Spring Boot’s test suite may not have covered that specific version combination, so do your own regression testing.
Server Configuration Properties
All three servers share common configuration properties under the server prefix. Spring Boot normalizes these across implementations.
Common Properties
server:
port: 8080
servlet:
context-path: /api
shutdown: graceful # graceful shutdown, not immediate
compression:
enabled: true
mime-types: text/html,text/xml,text/plain,text/css,application/javascript
min-response-size: 1024
# Connection timeout (all servers)
server:
connection-timeout: 20000 # ms, Tomcat/Undertow; Jetty uses idle-timeout
# Graceful shutdown
server:
shutdown: graceful
spring:
lifecycle:
timeout-per-shutdown-phase: 30s
Tomcat-Specific Tuning
server:
tomcat:
threads:
max: 200
min-spare: 10
max-connections: 10000
accept-count: 100 # backlog queue size when all threads busy
connection-timeout: 20000 # ms to wait for request line + headers
idle-timeout: 60000 # ms for keep-alive connections
max-http-form-post-size: 2MB
accesslog:
enabled: true
directory: /var/log/tomcat
prefix: access_log
pattern: '%h %l %u %t "%r" %s %b %Dms'
Jetty-Specific Tuning
server:
jetty:
threads:
max: 200
min: 10
idle-timeout: 60000 # ms before an idle thread is shrunk
max-connections: 10000 # maximum queued connections
acceptors: -1 # -1 = auto-detect based on OS
selectors: -1
connection-idle-timeout: 30000 # ms before closing idle connections
Undertow-Specific Tuning
server:
undertow:
threads:
io: 4 # I/O threads (one per CPU core is typical)
worker: 200 # worker threads for request handling
max-connections: 10000
buffer-size: 16384 # bytes, must be power of 2
buffers-per-region: 20
direct-buffers: true # use off-heap memory for buffers
accesslog:
enabled: true
directory: /var/log/undertow
pattern: '%h %l %u %t "%r" %s %b "%{i,Referer}"'
Thread Pool Architecture
Each server handles threads differently, and understanding the difference shapes how you tune them.
graph LR
A[Incoming Request] --> B[Acceptor Thread<br/>Tomcat: Acceptor<br/>Jetty: Acceptor<br/>Undertow: XNIO Listener]
B --> C[Thread Pool<br/>Worker Threads]
C --> D[Application Code<br/>Servlet / Filter]
D --> E[Response]
F[I/O Threads<br/>Undertow only] --> C
style B stroke:#00fff9,color:#000
style C stroke:#ff00ff,color:#fff
Tomcat uses a pure thread-per-connection model historically but switched to NIO by default. It has a fixed thread pool for request handling and a separate acceptor pool for incoming connections.
Jetty uses a more flexible model with a thread pool for computation and a set of acceptors/selectors for I/O. Its QueuedThreadPool is the default and behaves well under load: tasks queue rather than reject when the pool is saturated.
Undertow splits threads into I/O threads and worker threads. I/O threads handle the NIO event loop and should match your CPU core count. Worker threads execute the actual request handling and should match your expected concurrency. Separating these means I/O operations never block request processing threads.
Choosing Thread Pool Sizes
A common starting point:
| Server | I/O Threads | Worker Threads |
|---|---|---|
| Tomcat | N/A (handled internally) | 200 max |
| Jetty | N/A | 200 max |
| Undertow | CPU cores | 200 max |
These are not constants. If your application does blocking work (database calls, external HTTP), you need more worker threads. If it is purely non-blocking (reactive stack), fewer workers suffice.
Monitor actual thread usage with JMX or Micrometer. If activeCount consistently approaches maxThreads, you are saturating the pool.
Access Logging
Access logs are essential for debugging, auditing, and traffic analysis. All three servers support them.
Tomcat
server:
tomcat:
accesslog:
enabled: true
directory: /var/log/myapp
prefix: access
suffix: .log
pattern: '%h %l %u %t "%r" %s %b %Dms "%{i,Referer}" "%{i,User-Agent}"'
file-date-format: .yyyy-MM-dd
rotate: true
Jetty
Jetty logs via Logback or Log4j2 by configuring the NCSARequestLog. Configuration is more involved than Tomcat:
import org.eclipse.jetty.server.NCSARequestLog;
import org.springframework.boot.web.embedded.jetty.JettyServletWebServerFactory;
import org.springframework.boot.web.server.WebServerFactoryCustomizer;
@Component
public class JettyAccessLogCustomizer implements WebServerFactoryCustomizer<JettyServletWebServerFactory> {
@Override
public void customize(JettyServletWebServerFactory factory) {
NCSARequestLog requestLog = new NCSARequestLog("/var/log/jetty/access.log");
requestLog.setAppend(true);
requestLog.setExtended(false);
requestLog.setLogCookies(false);
requestLog.setLogTimeZone("UTC");
factory.setRequestLog(requestLog);
}
}
Undertow
Undertow’s access log is configured via the UndertowServletWebServerFactory:
server:
undertow:
accesslog:
enabled: true
directory: /var/log/undertow
prefix: access
suffix: .log
pattern: '%h %l %u %t "%r" %s %b "%{i,Referer}"'
rotate: true
Undertow’s access log configuration is the most straightforward of the three.
When to Use / When NOT to Use
When to Use Tomcat
Tomcat is the right choice when you want the most straightforward, well-documented servlet container experience. It is the Spring Boot default for a reason: the community has accumulated years of operational knowledge around it, tuning guides are plentiful, and most CI/CD pipelines have first-class Tomcat support. Use Tomcat for standard microservices where there is nothing unusual about the workload.
When NOT to Use Tomcat
Tomcat is a thread-per-connection model under the hood (even with NIO, worker threads are tied to connections for servlet execution). If you are building a massively concurrent service handling tens of thousands of simultaneous idle connections (long-polling, WebSocket, SSE), Tomcat’s thread model becomes a liability. At that point, Undertow’s separated I/O and worker thread model will serve you better.
When to Use Jetty
Jetty earns its place in memory-constrained environments. A Jetty-based Spring Boot JAR starts with a smaller heap footprint than the equivalent Tomcat build, which matters in Kubernetes environments with tight memory limits. Jetty also has a loyal following in the Eclipse community and among teams that have been operating it in production for years.
When NOT to Use Jetty
Jetty’s access logging requires Java code to configure, unlike Tomcat and Undertow which support YAML-based access log configuration. If your ops team expects YAML-driven operational configuration and does not want to maintain a custom WebServerFactoryCustomizer, Jetty becomes an operational burden rather than a simplification.
When to Use Undertow
Undertow shines for high-concurrency workloads where connections are long-lived and mostly idle. Its separated I/O thread pool (one per CPU core for NIO) means it never blocks acceptors waiting on slow clients. If you are building services that hold open connections for WebSocket, Server-Sent Events, or long-poll patterns, Undertow’s I/O model is purpose-built for exactly that.
When NOT to Use Undertow
Undertow’s buffer size must be a power of 2, and misconfiguration causes silent failures at startup. If your team does not want to maintain yet another server-specific configuration constraint, Tomcat avoids this entirely. Undertow also has a smaller community than Tomcat, which means fewer tuning guides and fewer people available to debug unusual behavior.
Trade-Off Table
| Dimension | Tomcat | Jetty | Undertow |
|---|---|---|---|
| Startup memory | Medium (~60MB heap) | Low (~40MB heap) | Medium (~50MB heap) |
| Thread model | Thread pool (worker) | QueuedThreadPool | Separate I/O + worker threads |
| NIO support | Yes (since 7.x) | Yes | Yes |
| Access log via YAML | Yes | No (requires Java code) | Yes |
| WebSocket/SSE support | Good | Good | Excellent |
| Throughput (raw) | Good | Good | Best historically |
| Operational knowledge | Widely available | Moderate | Less common |
| Graceful shutdown | Yes | Yes | Yes |
| Max connections | max-connections |
max-connections |
max-connections |
Failure Scenarios
Application starts but immediately returns 503
- Cause: All threads in the pool are busy and the backlog queue (
accept-countin Tomcat) is full. - Fix: Increase
server.tomcat.threads.max, increaseserver.tomcat.max-connections, or scale horizontally.
OutOfMemoryError: Unable to create new native thread
- Cause: Thread pool size exceeds what the OS allows per process. Common on containers with low ulimits.
- Fix: Reduce max threads, reduce container memory limits, or increase the OS thread limit (
ulimit -u).
Access logs not writing
- Cause: Log directory does not exist or is not writable by the process.
- Fix: Create the directory and set permissions, or change the log path to a writable location.
Jetty access log requires custom code
- Cause: Unlike Tomcat and Undertow, Jetty does not support access log configuration via
application.ymlalone. - Fix: Implement
WebServerFactoryCustomizer<JettyServletWebServerFactory>to configure theNCSARequestLogprogrammatically.
Undertow buffer size error at startup
- Cause:
buffer-sizemust be a power of 2 (e.g., 16384, not 15000). - Fix: Use a valid power-of-2 value for
server.undertow.buffer-size.
Security Notes
-
Disable server header: By default, servers expose their identity in responses (
Server: Apache-Coyote/1.1). Hide this in production:server: tomcat: remoteip: protocol-header: X-Forwarded-Proto # plus add server.use-forward-headers=trueFor all servers, set
server.include-native-tomcat-jni=disabledif you do not need JNI features. -
Limit request size: Prevent denial-of-service via large payloads:
server: tomcat: max-http-form-post-size: 2MB # or equivalent for Jetty/Undertow -
Connection timeouts: Set
connection-timeoutto prevent slow-client attacks tying up threads.
Common Pitfalls / Anti-Patterns
- Setting max threads too high: More threads means more context switching. The sweet spot is workload-dependent: start conservative and increase while monitoring.
- Ignoring the backlog queue:
accept-count(Tomcat) controls how many connections queue when all threads are busy. If your load balancer times out before connections reach your app, increase this. - Jetty access log complexity: Teams expect YAML configuration and are surprised when Jetty requires Java code for access logging.
- Undertow buffer size not power of 2: A common startup error: always verify buffer sizes are
1024,2048,4096,8192,16384, etc. - Forgetting graceful shutdown: Without
server.shutdown=gracefulandspring.lifecycle.timeout-per-shutdown-phase, in-flight requests die when you deploy.
Observability Checklist
- Verify thread pool actual utilization with JMX/Micrometer (not just configured size)
- Enable access logs in production for all environments
- Confirm
server.shutdown=gracefuland set a reasonable shutdown timeout - Check
accept-count/ backlog queue size matches your load balancer’s retry behavior - Set
server.include-native-tomcat-jni=disabledif JNI is not needed - Validate
max-http-form-post-sizeis set to prevent large-payload DoS - Confirm log directories exist and are writable before deployment
- Test graceful shutdown behavior with a load test before going to production
Quick Recap Checklist
- Switch servers by excluding
spring-boot-starter-tomcatand addingspring-boot-starter-jettyorspring-boot-starter-undertow - Tomcat threads are configured via
server.tomcat.threads.max - Jetty uses
QueuedThreadPoolwhich queues tasks when saturated - Undertow splits I/O threads (CPU-bound) from worker threads (blocking work)
- Access logging is YAML-configurable for Tomcat and Undertow, requires Java code for Jetty
-
accept-count(Tomcat) controls the connection backlog queue size -
buffer-sizein Undertow must always be a power of 2 - Set
server.shutdown=gracefuland aspring.lifecycle.timeout-per-shutdown-phasevalue - Monitor actual thread pool utilization via JMX, not just configured values
Interview Questions
You exclude the default server starter and add the one you want. In Maven, add an <exclusions> block under spring-boot-starter-web to remove spring-boot-starter-tomcat, then add either spring-boot-starter-jetty or spring-boot-starter-undertow as a separate dependency. Spring Boot's dependency management ensures compatible versions. No code changes are required—the framework auto-configures whichever server is on the classpath.
I/O threads run the NIO event loop and handle non-blocking I/O events—they never block. Worker threads execute the actual servlet code, filters, and application logic, which may involve blocking calls like JDBC or external HTTP. By separating them, Undertow ensures that a slow database query does not block other requests from being accepted. The typical ratio is one I/O thread per CPU core, with many more worker threads to handle concurrent blocking work.
When all worker threads are busy, new incoming connections queue in the backlog. In Tomcat this is controlled by accept-count, in Jetty by max-connections, and in Undertow by similar settings. If the queue fills before a thread becomes available, new connections are refused. This interacts directly with your load balancer's behavior—if the load balancer times out waiting for a connection while the queue is full, requests fail with 503s. Sizing this correctly matters for handling traffic spikes without losing requests.
Set server.shutdown=graceful and spring.lifecycle.timeout-per-shutdown-phase=30s. When Spring Boot receives a shutdown signal, it stops accepting new requests and waits up to the timeout for in-flight requests to complete before halting the JVM. Without this, a deployment kills requests mid-flight—users see 503s, database transactions roll back unexpectedly, and your logs show broken requests. It is especially important for services handling writes or long-running operations.
Undertow is well-suited here because its separated I/O and worker thread model handles non-blocking I/O efficiently. However, if you are using the reactive stack with spring-boot-starter-webflux, Reactor Netty is already the default and preferred server—it is built for non-blocking I/O from the ground up. The distinction matters: Undertow wins for servlet-based (blocking) workloads where you want efficient NIO, while Netty wins for purely reactive stacks.
Unlike Tomcat's thread pool which can reject tasks when full, Jetty's QueuedThreadPool queues tasks when all threads are busy. This means a Jetty-based application will continue accepting connections up to the backlog limit rather than immediately rejecting them with a 503. The queue has a maximum size; once full, new connections are refused. This behavior is more graceful under sudden traffic spikes but requires careful monitoring—queued tasks consume memory and can lead to OOM if the queue grows unbounded.
Jetty has the smallest startup footprint at approximately 40MB heap, making it ideal for Kubernetes environments with tight memory limits. Tomcat sits at approximately 60MB heap by default. Undertow is in the middle at approximately 50MB heap. These numbers vary based on configured thread pool sizes, enabled features, and the application itself. In containerized environments where you pay for memory, Jetty's lower baseline can directly translate to cost savings through smaller resource requests.
Undertow uses direct memory buffers (off-heap) for I/O operations, and the underlying buffer pool implementation requires buffer sizes to be powers of 2 for efficient memory alignment and allocation. If you set a non-power-of-2 value like 15000 bytes, Undertow throws an exception at startup and refuses to start. Valid values are 512, 1024, 2048, 4096, 8192, 16384, or larger powers of 2. This is a common misconfiguration that teams encounter when porting settings from other servers that do not have this constraint.
Tomcat and Undertow both support YAML-based access log configuration—enable the access log, set the directory, prefix, pattern, and rotation in application.yml and you are done. Jetty requires Java code: you must implement WebServerFactoryCustomizer<JettyServletWebServerFactory> and configure the NCSARequestLog programmatically. For teams that manage configuration as code (YAML/GitOps), Jetty's approach is more cumbersome. For teams that value operational flexibility, Jetty's programmatic configuration offers more customization options.
max-threads defines how many worker threads handle requests simultaneously. max-connections controls how many connections the server accepts at the protocol level (NIO selector). accept-count is the OS-level backlog queue for connections that arrive when all worker threads are busy. The relationship: connections up to max-connections are accepted and distributed to worker threads. Once max-connections is reached, new connections queue in the accept-count backlog. If that fills, new connections are refused. A common mistake is setting accept-count lower than the load balancer's retry timeout, causing 503s during traffic spikes.
Spring Boot's WebServerFactoryCustomizer mechanism auto-detects whichever server is on the classpath and configures it using properties under server.* plus server-specific prefixes like server.tomcat.*, server.jetty.*, and server.undertow.*. You can take full control by defining your own WebServerFactoryCustomizer<T> beans, which run after the auto-configuration. This is how Jetty access logging is implemented in Spring Boot—by customizing the factory before it creates the web server. The auto-configuration respects any customizations you provide through these customizers.
Thread-per-request (blocking) models allocate a thread for the entire duration of a request—even when that thread is waiting on I/O like a database call or HTTP response. At scale, this consumes threads rapidly and can lead to thread exhaustion. NIO event loops use a small number of threads to handle many connections by registering interest in I/O events and only using threads when actual work needs processing. Tomcat uses NIO but still allocates worker threads per connection for servlet execution. Undertow separates I/O threads (event loop) from worker threads (blocking work) giving you the best of both: efficient I/O handling with the ability to do blocking work without starving the event loop.
Use a tool like wrk, hey, or Apache Bench to generate load. Measure requests per second, latency percentiles (p50, p95, p99), and error rates under sustained concurrent load. For a fair comparison, use identical thread pool sizes and JVM settings across servers. Run warm-up iterations before measuring since JIT compilation affects early results. Undertow historically shows best raw throughput for connection-intensive workloads, but the differences are often dwarfed by application-level bottlenecks (database queries, external HTTP calls). Profile with JMC or async-profiler to identify whether the server or the application is the bottleneck.
Two properties work together: server.shutdown=graceful stops accepting new connections and waits for in-flight requests to complete, and spring.lifecycle.timeout-per-shutdown-phase=30s sets the maximum wait time. If a request exceeds this timeout, Spring Boot terminates it forcibly. Without graceful shutdown configured, a SIGTERM during deployment kills in-flight requests immediately, returning 503s to clients and potentially rolling back database transactions. Always test shutdown behavior with a load test before production deployment.
All three servers support the same Spring Boot SSL properties: server.ssl.enabled=true, server.ssl.key-store, server.ssl.key-store-password, server.ssl.key-alias. Under the hood, Spring Boot configures the SSL context on each server's factory. For more advanced SSL needs, you can use WebServerFactoryCustomizer to access the underlying server's SSL API directly—for example, to configure client certificates or specific cipher suites. Tomcat offers additional properties like server.tomcat.remoteip.protocol-header for behind-proxy SSL termination scenarios.
All three servers support WebSocket through the servlet specification (JSR-356) and their own proprietary APIs. Undertow has the most efficient WebSocket implementation for servlet-based applications because its I/O threads never block—long-lived WebSocket connections hold minimal resources. Tomcat's WebSocket support is solid and well-documented. Jetty offers excellent WebSocket support with additional features like the Jetty WebSocket API which provides more control than the servlet standard. For Server-Sent Events (SSE), all three work well since SSE is just HTTP with a long-lived response stream.
In Kubernetes environments, Jetty's lower memory footprint directly reduces the minimum heap requirement and thus the resource request you must declare. Undertow's separate I/O thread model works well when you have varying workloads since I/O threads stay active even if worker threads saturate. In environments with CPU limits, Undertow's I/O thread tuning becomes critical since I/O threads should match core count. For local development, Tomcat's more verbose logging and familiar error messages make debugging easier. For FIPS-compliant environments, Tomcat has better support for JSSE and FIPS-certified crypto modules.
Enable JMX and expose the thread pool metrics via Micrometer to Prometheus or your monitoring system. Key metrics to watch: activeCount (threads doing work), poolSize (current threads), queueSize (queued tasks for Jetty). If activeCount consistently approaches maxThreads while queue grows, you are saturated. In Tomcat, check ThreadPool JMX beans. In Undertow, monitor both I/O and worker thread pools separately—running out of I/O threads blocks new connections entirely. Set alerts when active threads exceed 80% of max for proactive alerting before 503s occur.
For GraalVM native image builds (Spring Native), the embedded server must be compatible with compile-time reflection analysis. Tomcat requires the most extensive native image configuration due to its complex class hierarchy. Jetty and Undertow generally require less configuration for native builds. Undertow's modular architecture maps well to GraalVM's reachability metadata. When building native images, you must also configure the server's native libraries (JNI, SSL) explicitly. Spring Boot 3.x improved native image support significantly, but the server choice still affects build complexity and image size.
In Tomcat, max-connections controls the NIO selector's limit—connections beyond this queue into the OS-level accept-count backlog. In Jetty, max-connections is the maximum number of connections queued at the accept layer before refusing new connections. In Undertow, max-connections similarly controls the number of concurrent connections at the XNIO level. The semantics are similar but the internal mechanisms differ. Jetty additionally has acceptors and selectors count properties that control how many threads handle accepting and selecting on connections respectively.
Further Reading
- Spring Boot Reference Documentation: Embedded Web Servers
- Tomcat Architecture Documentation
- Jetty Documentation: Threading Model
- Undertow Documentation: Architecture
- Spring Boot 3.x Migration Guide: Servlet Containers
- Understanding Jetty’s QueuedThreadPool
- Undertow I/O and Worker Thread Configuration
Conclusion
Spring Boot’s embedded servlet containers—Tomcat, Jetty, and Undertow—each represent a different point on the trade-off spectrum between operational familiarity, memory efficiency, and I/O model sophistication. Tomcat remains the default because it is the most widely understood and has the shallowest learning curve. Jetty earns its place in memory-constrained container environments where a smaller heap footprint directly reduces infrastructure cost. Undertow is the choice for high-concurrency workloads involving long-lived connections or non-blocking I/O patterns.
All three servers expose their configuration through Spring Boot’s normalized property prefixes, so switching between them requires only dependency changes—no code rewrites. This portability means you can start with Tomcat for simplicity and migrate to Undertow for performance reasons without touching application logic. The thread pool architecture differs meaningfully across implementations: Tomcat and Jetty use pooled worker threads, while Undertow separates I/O threads from worker threads entirely. Understanding this distinction matters for capacity planning and for debugging saturation issues in production.
Graceful shutdown is the production feature most teams underconfigure. Without it, every deployment kills in-flight requests and potentially rolls back database transactions. Set server.shutdown=graceful and size spring.lifecycle.timeout-per-shutdown-phase to at least 2x your p99 request duration. Access logging is equally non-negotiable in production—Tomcat and Undertow make this straightforward via YAML, while Jetty requires a programmatic customizer.
When evaluating which server to use, start with Tomcat unless you have a specific reason not to. If your Kubernetes memory limits are tight, try Jetty. If you are building a reactive non-blocking service or one that handles WebSockets and Server-Sent Events at scale, Undertow’s separated thread model will serve you better.
Category
Related Posts
Spring Boot Build Tools: Maven & Gradle
Configure Maven and Gradle for Spring Boot projects—plugins, dependency management, packaging JARs and WARs, and build automation essentials.
JUnit 5 & Jupiter: Lifecycle, Nested & Parameterized Tests
Explore JUnit 5 Jupiter features: master test lifecycle annotations, organize tests with @Nested, and parameterize tests with @CsvSource and @MethodSource.
Mockito: Mocking, Stubbing, Verifications, Spy vs Mock
Learn Mockito fundamentals: create mocks and spies, stub behavior with when().thenReturn(), verify interactions, and choose between Spy vs Mock wisely.