Optimized(undertow): process requests optimizations cutting allocations and response body writes NIO - #896
Open
GoodforGod wants to merge 9 commits into
Open
Optimized(undertow): process requests optimizations cutting allocations and response body writes NIO#896GoodforGod wants to merge 9 commits into
GoodforGod wants to merge 9 commits into
Conversation
Buffer small bodies on request VTs and stream large bodies through a bounded XNIO pipe. Keep send, close, and exchange completion on I/O threads to avoid buffer-pool contention and unbounded memory use.
Streaming responses larger than the socket buffer aborted mid-body. AsyncBodyPipe started the I/O-thread consumer while the producer was still inside handleRequest, leaving the exchange IN_CALL | DISPATCHED. Undertow forbids resuming writes in that state, so AsyncSenderImpl threw UT000146 on the first partial write. Small bodies never hit it because their writes complete synchronously. handleRequest now only claims the dispatch slot and defers the request to the dispatch task, which Connectors.executeRootHandler runs after setInCall(false)/unDispatch(). Async writes are legal for the whole response, and AsyncBodyPipe no longer needs its keep-alive dispatch hack. Also fixes a double observeError: once the pipe reported its terminal failure, the producer's next offer() threw and prepareResponse counted that echo as a second error. - Cache lowercase header name -> interned HttpString. Kora lowercases every header name while Undertow's Headers cache is keyed by the canonical spelling, so tryFromString() missed on every header of every response and allocated an HttpString plus its backing array. Also covers traceparent/tracestate on the injection path. - Reuse exchange.getResponseSender() instead of allocating a second AsyncSenderImpl per response (3.7% of baseline allocation). - Let ProcessedResponse be its own Runnable, IoCallback and ExchangeCompletionListener, removing three allocations per request. - Size the adaptive body buffer from the declared content length, keeping 256 bytes when it is unknown. A 1 KiB default made AdaptiveBodyOutputStream the largest byte[] source in the process and inflated total allocation by 18% on /db. - Hoist the W3C propagator and response HeaderMap, switch on reserved header names, drop the redundant leading MDC.clear(). Wire the new tracing contextPropagation flag through the handler so W3C extract/inject can be disabled independently of telemetry. Fix a BufferUnderflowException in DefaultFullHttpBody.write() for non-heap-backed buffers with under 1024 bytes remaining.
- RawHttpClient in the kit's test fixtures: a small socket client for keep-alive, pipelining, reading headers separately, and abrupt disconnects, which OkHttp can't do. The kit also gets a port() accessor. - DefaultFullHttpBodyTest in http-common (28 tests): the direct-buffer fix, at sizes around the 1 KiB copy buffer.
Dependency Update ReportUpdate level: Found
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
EN
Description
Improved the Undertow HTTP server so that each request is processed on the connection's virtual thread while the response is sent and the exchange is completed on the connection's XNIO I/O thread. Previously the virtual thread wrote the response through Undertow's blocking output stream and completed the exchange itself, which makes the I/O thread busy-spin in
HttpReadListener.handleEventwhenever the next pipelined or keep-alive request arrives before completion finishes. Response bodies are now materialized adaptively: small bodies are sent as one buffer and large ones are streamed through a bounded asynchronous pipe with backpressure. Controller, routing, telemetry, and configuration contracts stay the same.Intention & Motivation
Improved throughput and CPU efficiency of the default Undertow server under high concurrency, driven by profiling of the scenario at c=256. The blocking write path spent 60.7% of all CPU in Undertow's CAS spin loop, capping throughput at ~85k req/s, while the XNIO-owned send path reaches 109–119k req/s with 121–125 µs CPU per request on the same hardware.
RU
Описание
Улучшен HTTP-сервер на Undertow: запрос обрабатывается на виртуальном потоке соединения, а отправка ответа и завершение обмена выполняются на потоке ввода-вывода XNIO этого соединения. Раньше виртуальный поток сам писал ответ через блокирующий поток вывода Undertow и завершал обмен, из-за чего поток XNIO активно крутился в
HttpReadListener.handleEvent, если следующий запрос приходил раньше, чем завершался предыдущий. Тела ответов теперь формируются адаптивно: небольшие отправляются одним буфером, крупные передаются потоком через ограниченный асинхронный канал с обратным давлением. Контракты контроллеров, маршрутизации, телеметрии и конфигурации не меняются.Намерение и мотивация
Улучшена пропускная способность и эффективность использования CPU стандартного сервера Undertow при высокой конкурентности по результатам профилирования сценария. Путь с блокирующей записью тратил 60.7% всего CPU на цикл ожидания Undertow и упирался примерно в 85 тысяч запросов в секунду, тогда как отправка на потоке XNIO даёт 109–119 тысяч запросов в секунду при 121–125 мкс CPU на запрос на том же оборудовании.
Changelog
Sender.send, body close, andendExchangerun on the connection's XNIO I/O thread; the cross-thread exchange completion that madeHttpReadListener.handleEventspin is gone.HttpBodyOutputbodies are buffered up to 64 KiB and sent as one buffer, larger ones are streamed in 16 KiB chunks through a bounded pipe (4 in-flight chunks) that parks only the producing virtual thread on socket backpressure;HEADresponses send no body, and a body that fails mid-write is sent truncated and the connection is closed.DefaultFullHttpBody.writefor non-array (direct) buffers throwingBufferUnderflowExceptionwhen the remaining size is not a multiple of the 1 KiB copy buffer.KoraVirtualThreadDispatchHttpHandlerintoKoraVirtualThreadPerConnectionDispatchHttpHandlerwith unchanged per-connection virtual-thread dispatch semantics.HttpClientRequestMapperModuleform and JSON mapper factories to returnHttpClientRequestMapper<T>instead of concrete mapper types.Design
The request lifecycle is split into two phases owned by different threads.
KoraRequestProcessingHttpHandler.handleRequestruns on the per-connection virtual thread created byKoraVirtualThreadPerConnectionDispatchHttpHandler. It bindsUndertowContext, MDC, andOpentelemetryContextscoped values, routes the request, invokes the controller, and turns theHttpServerResponseinto an immutableProcessedResponse. That record carries status, headers, content type, content, length, body, and observation. It is then handed to the exchange's I/O thread withexchange.getIoThread().execute(processedResponse).On the I/O thread
ProcessedResponsewrites the status and headers, includingServerwhenheaderServerNameEnabledis set, callsexchange.getResponseSender().send(content, this), and completes the exchange in its ownIoCallback. It implementsRunnable,IoCallback, andExchangeCompletionListeneritself, so the send path does not allocate a dispatch lambda, an I/O callback, and a completion listener per request. Because only the I/O thread ever transitions Undertow's request state, the I/O thread never waits on a virtual thread that the OS may have descheduled mid-transition.Body materialization depends on what the body can offer:
getFullContentIfAvailable()non-null: the buffer is sent as is, withContent-Length.HttpBodyOutput.write()runs on the virtual thread againstAdaptiveBodyOutputStream. Up to 64 KiB stays in one growable array and is sent once; crossing 64 KiB startsAsyncBodyPipe, which begins sending on the I/O thread immediately and accepts 16 KiB chunks with at most 4 pending, parking the virtual thread when the socket is slow.HEAD: headers and declared length only, no body.500plaintext with the error message; failure after bytes were produced: the response is sent truncated and the connection is closed so the client cannot mistake it for a complete body.Measured on at concurrency 256 against an external PostgreSQL host (8 cores / 16 threads), with interleaved runs:
maxPoolSize=128,ioThreads=12, carrier parallelism 12The blocking-write numbers were measured with the same blocking-write model on the sibling VT io write optimized branch. In the blocking-write path, JFR CPU-time sampling attributes 60.7% of all CPU to the self time of
HttpReadListener.handleEvent. The XNIO-owned path keeps Undertow under 5% and spends most of its CPU in kernel socket I/O. No virtual-thread pinning and no GC pressure were observed in either path. All 33http-server-undertowmodule tests pass.