Skip to content

Optimized(undertow): process requests optimizations cutting allocations and response body writes NIO - #896

Open
GoodforGod wants to merge 9 commits into
masterfrom
feature/http-server-undertow-io-optimization
Open

GoodforGod wants to merge 9 commits into
masterfrom
feature/http-server-undertow-io-optimization

Conversation

@GoodforGod

Copy link
Copy Markdown
Contributor

EN

Description

Improved the Undertow HTTP server so that each request is processed on the connection's virtual thread while the response is sent and the exchange is completed on the connection's XNIO I/O thread. Previously the virtual thread wrote the response through Undertow's blocking output stream and completed the exchange itself, which makes the I/O thread busy-spin in HttpReadListener.handleEvent whenever the next pipelined or keep-alive request arrives before completion finishes. Response bodies are now materialized adaptively: small bodies are sent as one buffer and large ones are streamed through a bounded asynchronous pipe with backpressure. Controller, routing, telemetry, and configuration contracts stay the same.

Intention & Motivation

Improved throughput and CPU efficiency of the default Undertow server under high concurrency, driven by profiling of the scenario at c=256. The blocking write path spent 60.7% of all CPU in Undertow's CAS spin loop, capping throughput at ~85k req/s, while the XNIO-owned send path reaches 109–119k req/s with 121–125 µs CPU per request on the same hardware.


RU

Описание

Улучшен HTTP-сервер на Undertow: запрос обрабатывается на виртуальном потоке соединения, а отправка ответа и завершение обмена выполняются на потоке ввода-вывода XNIO этого соединения. Раньше виртуальный поток сам писал ответ через блокирующий поток вывода Undertow и завершал обмен, из-за чего поток XNIO активно крутился в HttpReadListener.handleEvent, если следующий запрос приходил раньше, чем завершался предыдущий. Тела ответов теперь формируются адаптивно: небольшие отправляются одним буфером, крупные передаются потоком через ограниченный асинхронный канал с обратным давлением. Контракты контроллеров, маршрутизации, телеметрии и конфигурации не меняются.

Намерение и мотивация

Улучшена пропускная способность и эффективность использования CPU стандартного сервера Undertow при высокой конкурентности по результатам профилирования сценария. Путь с блокирующей записью тратил 60.7% всего CPU на цикл ожидания Undertow и упирался примерно в 85 тысяч запросов в секунду, тогда как отправка на потоке XNIO даёт 109–119 тысяч запросов в секунду при 121–125 мкс CPU на запрос на том же оборудовании.


Changelog

  • Improved Undertow request handling: routing, the controller call, and response body materialization run on the per-connection virtual thread, while response headers, Sender.send, body close, and endExchange run on the connection's XNIO I/O thread; the cross-thread exchange completion that made HttpReadListener.handleEvent spin is gone.
  • Improved response body writing: bodies with full content available are sent directly; other HttpBodyOutput bodies are buffered up to 64 KiB and sent as one buffer, larger ones are streamed in 16 KiB chunks through a bounded pipe (4 in-flight chunks) that parks only the producing virtual thread on socket backpressure; HEAD responses send no body, and a body that fails mid-write is sent truncated and the connection is closed.
  • W3C trace context is extracted from requests and injected into responses only when telemetry is enabled; with noop telemetry no extraction, injection, or exchange completion listener is performed per request.
  • Fixed DefaultFullHttpBody.write for non-array (direct) buffers throwing BufferUnderflowException when the remaining size is not a multiple of the 1 KiB copy buffer.
  • Refactored KoraVirtualThreadDispatchHttpHandler into KoraVirtualThreadPerConnectionDispatchHttpHandler with unchanged per-connection virtual-thread dispatch semantics.
  • Refactored HttpClientRequestMapperModule form and JSON mapper factories to return HttpClientRequestMapper<T> instead of concrete mapper types.

Design

The request lifecycle is split into two phases owned by different threads. KoraRequestProcessingHttpHandler.handleRequest runs on the per-connection virtual thread created by KoraVirtualThreadPerConnectionDispatchHttpHandler. It binds UndertowContext, MDC, and OpentelemetryContext scoped values, routes the request, invokes the controller, and turns the HttpServerResponse into an immutable ProcessedResponse. That record carries status, headers, content type, content, length, body, and observation. It is then handed to the exchange's I/O thread with exchange.getIoThread().execute(processedResponse).

On the I/O thread ProcessedResponse writes the status and headers, including Server when headerServerNameEnabled is set, calls exchange.getResponseSender().send(content, this), and completes the exchange in its own IoCallback. It implements Runnable, IoCallback, and ExchangeCompletionListener itself, so the send path does not allocate a dispatch lambda, an I/O callback, and a completion listener per request. Because only the I/O thread ever transitions Undertow's request state, the I/O thread never waits on a virtual thread that the OS may have descheduled mid-transition.

Body materialization depends on what the body can offer:

  • getFullContentIfAvailable() non-null: the buffer is sent as is, with Content-Length.
  • Otherwise HttpBodyOutput.write() runs on the virtual thread against AdaptiveBodyOutputStream. Up to 64 KiB stays in one growable array and is sent once; crossing 64 KiB starts AsyncBodyPipe, which begins sending on the I/O thread immediately and accepts 16 KiB chunks with at most 4 pending, parking the virtual thread when the socket is slow.
  • HEAD: headers and declared length only, no body.
  • Handler failure: 500 plaintext with the error message; failure after bytes were produced: the response is sent truncated and the connection is closed so the client cannot mistake it for a complete body.

Measured on at concurrency 256 against an external PostgreSQL host (8 cores / 16 threads), with interleaved runs:

Response path req/s CPU / request
Blocking write + exchange completion on the virtual thread 83–87k 158–171 µs
This PR: send + completion on XNIO 109–116k 121–125 µs
This PR + Hikari maxPoolSize=128, ioThreads=12, carrier parallelism 12 118–119k 103 µs

The blocking-write numbers were measured with the same blocking-write model on the sibling VT io write optimized branch. In the blocking-write path, JFR CPU-time sampling attributes 60.7% of all CPU to the self time of HttpReadListener.handleEvent. The XNIO-owned path keeps Undertow under 5% and spends most of its CPU in kernel socket I/O. No virtual-thread pinning and no GC pressure were observed in either path. All 33 http-server-undertow module tests pass.

Buffer small bodies on request VTs and stream large bodies through a
bounded XNIO pipe. Keep send, close, and exchange completion on I/O
threads to avoid buffer-pool contention and unbounded memory use.
Streaming responses larger than the socket buffer aborted mid-body.
AsyncBodyPipe started the I/O-thread consumer while the producer was
still inside handleRequest, leaving the exchange IN_CALL | DISPATCHED.
Undertow forbids resuming writes in that state, so AsyncSenderImpl threw
UT000146 on the first partial write. Small bodies never hit it because
their writes complete synchronously.

handleRequest now only claims the dispatch slot and defers the request to
the dispatch task, which Connectors.executeRootHandler runs after
setInCall(false)/unDispatch(). Async writes are legal for the whole
response, and AsyncBodyPipe no longer needs its keep-alive dispatch hack.

Also fixes a double observeError: once the pipe reported its terminal
failure, the producer's next offer() threw and prepareResponse counted
that echo as a second error.

- Cache lowercase header name -> interned HttpString. Kora lowercases
  every header name while Undertow's Headers cache is keyed by the
  canonical spelling, so tryFromString() missed on every header of every
  response and allocated an HttpString plus its backing array. Also
  covers traceparent/tracestate on the injection path.
- Reuse exchange.getResponseSender() instead of allocating a second
  AsyncSenderImpl per response (3.7% of baseline allocation).
- Let ProcessedResponse be its own Runnable, IoCallback and
  ExchangeCompletionListener, removing three allocations per request.
- Size the adaptive body buffer from the declared content length, keeping
  256 bytes when it is unknown. A 1 KiB default made
  AdaptiveBodyOutputStream the largest byte[] source in the process and
  inflated total allocation by 18% on /db.
- Hoist the W3C propagator and response HeaderMap, switch on reserved
  header names, drop the redundant leading MDC.clear().

Wire the new tracing contextPropagation flag through the handler so W3C
extract/inject can be disabled independently of telemetry.

Fix a BufferUnderflowException in DefaultFullHttpBody.write() for
non-heap-backed buffers with under 1024 bytes remaining.
- RawHttpClient in the kit's test fixtures: a small socket client for keep-alive, pipelining, reading headers separately, and abrupt disconnects, which OkHttp can't do. The kit also gets a port() accessor.
- DefaultFullHttpBodyTest in http-common (28 tests): the direct-buffer fix, at sizes around the 1 KiB copy buffer.
@GoodforGod GoodforGod added module: http Related to HTTP module optimization Optimization change that doesn't change behavior or bring something new labels Sep 26, 2026
@github-actions

Copy link
Copy Markdown

Dependency Update Report

Update level: patch

Found 1 dependency updates.

gradle/libs.versions.toml

  • s3client-aws (software.amazon.awssdk:s3, inline:254): 2.55.5 -> 2.55.6

@GoodforGod
GoodforGod requested a review from Squiry September 26, 2026 20:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module: http Related to HTTP module optimization Optimization change that doesn't change behavior or bring something new

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant