From c2d69c4f3d4186b4ee0844cbd87273cec6426704 Mon Sep 17 00:00:00 2001 From: stopachka Date: Mon, 28 Sep 2026 09:33:54 -0700 Subject: [PATCH 1/4] Reclaim native allocator pages and retain API failover capacity --- server/.ebextensions/resources.config | 2 +- server/docker-compose.yml | 3 ++- server/infra/README.md | 39 ++++++++++++++++++++++++--- 3 files changed, 39 insertions(+), 5 deletions(-) diff --git a/server/.ebextensions/resources.config b/server/.ebextensions/resources.config index 320193cc6f..bc0dbfec2b 100644 --- a/server/.ebextensions/resources.config +++ b/server/.ebextensions/resources.config @@ -1,6 +1,6 @@ option_settings: aws:autoscaling:asg: - MinSize: '1' + MinSize: '2' MaxSize: '2' Cooldown: '60' aws:autoscaling:launchconfiguration: diff --git a/server/docker-compose.yml b/server/docker-compose.yml index 035f0700c6..c676d03a16 100644 --- a/server/docker-compose.yml +++ b/server/docker-compose.yml @@ -16,7 +16,8 @@ services: # `web` depends_on `vector` so that we don't lose logs on docker compose down depends_on: - vector - command: sh -c 'java $${JAVA_OPTS} --add-modules java.se --add-exports java.base/jdk.internal.ref=ALL-UNNAMED --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/sun.nio.ch=ALL-UNNAMED --add-opens java.management/sun.management=ALL-UNNAMED --add-opens jdk.management/com.sun.management.internal=ALL-UNNAMED -XX:+UnlockDiagnosticVMOptions -XX:+DebugNonSafepoints -Djdk.attach.allowAttachSelf -Djava.awt.headless=true -Dcom.sun.management.jmxremote.port=10001 -Dcom.sun.management.jmxremote.rmi.port=10001 -Dcom.sun.management.jmxremote.authenticate=false -Dcom.sun.management.jmxremote.ssl=false -XX:StartFlightRecording=disk=true,maxsize=1G,name=instant,filename=instant_$(date +%s%N).jfr -Dcom.sun.management.jmxremote.local.only=false -Djava.rmi.server.hostname=localhost -server -jar target/instant-standalone.jar' + # Return freed native allocator pages to the OS; JAVA_OPTS can override these defaults. + command: sh -c 'java -XX:TrimNativeHeapInterval=60000 -Xlog:trimnative=info $${JAVA_OPTS} --add-modules java.se --add-exports java.base/jdk.internal.ref=ALL-UNNAMED --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/sun.nio.ch=ALL-UNNAMED --add-opens java.management/sun.management=ALL-UNNAMED --add-opens jdk.management/com.sun.management.internal=ALL-UNNAMED -XX:+UnlockDiagnosticVMOptions -XX:+DebugNonSafepoints -Djdk.attach.allowAttachSelf -Djava.awt.headless=true -Dcom.sun.management.jmxremote.port=10001 -Dcom.sun.management.jmxremote.rmi.port=10001 -Dcom.sun.management.jmxremote.authenticate=false -Dcom.sun.management.jmxremote.ssl=false -XX:StartFlightRecording=disk=true,maxsize=1G,name=instant,filename=instant_$(date +%s%N).jfr -Dcom.sun.management.jmxremote.local.only=false -Djava.rmi.server.hostname=localhost -server -jar target/instant-standalone.jar' vector: # Mirrored from timberio/vector:0.56.0-alpine into our private ECR so diff --git a/server/infra/README.md b/server/infra/README.md index 2b6de5628d..4d4fe6295f 100644 --- a/server/infra/README.md +++ b/server/infra/README.md @@ -25,11 +25,15 @@ stall rather than one busy sample. With complete telemetry, the 3-of-5 rule counts three breaching minutes, which need not be consecutive. With missing samples, CloudWatch can evaluate older data and alarm after a breach followed by gaps ([AWS behavior](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/alarms-and-missing-data.html)). -The alarms share the existing scale-up policy. Normal capacity is one to two -instances. Adding a spare does not free a leaking JVM's heap; application health +The alarms share the existing scale-up policy. The bundle keeps a minimum of two +instances so a host failure does not leave all traffic waiting for a replacement +to boot. The default maximum is also two; scale-out requires a higher maximum. +Adding a spare does not reclaim memory from the other JVM; application health checks still handle sustained unresponsiveness. -`jvm_autoscaling.yaml` checks scale-in once per minute. Over the last fifteen +`jvm_autoscaling.yaml` checks scale-in once per minute and respects the group's +minimum capacity. With the default minimum and maximum of two, it leaves both +instances running. When capacity exceeds the minimum, over the last fifteen complete minutes it requires average CPU below 30%, each JVM's average GC pressure below 6%, each JVM up for at least an hour, and the sum of the JVMs' median heap pressure below 80%, assuming equal 90 GiB heap limits. One JVM @@ -41,6 +45,35 @@ does not: it reports one impacted instance for much of the day without a request-level cause, and the group's ELB health check already replaces instances that fail. +## Native memory + +The API container enables `-XX:TrimNativeHeapInterval=60000` on Corretto 26 +with glibc. A dedicated JVM thread periodically returns free native allocator +pages to the OS. This does not collect the Java heap or reclaim live native +allocations. `-Xlog:trimnative=info` records the reclaimed memory and trim +duration. Heap sizing still comes from `JAVA_OPTS`. + +On September 28, two successive hosts reached about 120.7 GiB RSS on a +123.1 GiB machine before becoming unresponsive. Available memory fell to roughly +100 MiB, while the last heap-pressure readings were below 50%. Both stalls +coincided with sustained root-disk reads at 125 MiB/s. Heap and GC alarms do not +cover this host-memory failure. A later native trim on the surviving process +returned about 4.6 GiB without restarting Java, confirming that freed allocator +pages can consume substantial headroom. + +After a rollout, check trim reclamation and duration alongside host +`MemAvailable`, memory pressure, request latency, and completion of the scheduled +backup. Trimming can contend with native allocations, and it cannot bound live +native memory growth. The two-instance minimum provides spare serving capacity; +it does not protect against simultaneous failures or replace these checks. + +To disable periodic trimming, append `-XX:TrimNativeHeapInterval=0` to the +existing `JAVA_OPTS` and roll the configuration. The environment options follow +the container defaults, so they take precedence. Preserve the heap settings and +two-instance minimum when rolling back only trimming. + +## Deployment + The controller invokes the full down-policy ARN with cooldown; IAM restricts execution to the exact group ARN. The old low-CPU alarms have no direct scaling actions, so a missing controller keeps extra capacity running. From 4b69e56b9eb3d0504f6994a2ebe6dc9695b41864 Mon Sep 17 00:00:00 2001 From: stopachka Date: Mon, 28 Sep 2026 10:24:15 -0700 Subject: [PATCH 2/4] Remove two-host minimum from memory investigation --- server/.ebextensions/resources.config | 2 +- server/infra/README.md | 18 +++++++----------- 2 files changed, 8 insertions(+), 12 deletions(-) diff --git a/server/.ebextensions/resources.config b/server/.ebextensions/resources.config index bc0dbfec2b..320193cc6f 100644 --- a/server/.ebextensions/resources.config +++ b/server/.ebextensions/resources.config @@ -1,6 +1,6 @@ option_settings: aws:autoscaling:asg: - MinSize: '2' + MinSize: '1' MaxSize: '2' Cooldown: '60' aws:autoscaling:launchconfiguration: diff --git a/server/infra/README.md b/server/infra/README.md index 4d4fe6295f..403cbefd6f 100644 --- a/server/infra/README.md +++ b/server/infra/README.md @@ -25,15 +25,11 @@ stall rather than one busy sample. With complete telemetry, the 3-of-5 rule counts three breaching minutes, which need not be consecutive. With missing samples, CloudWatch can evaluate older data and alarm after a breach followed by gaps ([AWS behavior](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/alarms-and-missing-data.html)). -The alarms share the existing scale-up policy. The bundle keeps a minimum of two -instances so a host failure does not leave all traffic waiting for a replacement -to boot. The default maximum is also two; scale-out requires a higher maximum. -Adding a spare does not reclaim memory from the other JVM; application health +The alarms share the existing scale-up policy. Normal capacity is one to two +instances. Adding a spare does not free a leaking JVM's heap; application health checks still handle sustained unresponsiveness. -`jvm_autoscaling.yaml` checks scale-in once per minute and respects the group's -minimum capacity. With the default minimum and maximum of two, it leaves both -instances running. When capacity exceeds the minimum, over the last fifteen +`jvm_autoscaling.yaml` checks scale-in once per minute. Over the last fifteen complete minutes it requires average CPU below 30%, each JVM's average GC pressure below 6%, each JVM up for at least an hour, and the sum of the JVMs' median heap pressure below 80%, assuming equal 90 GiB heap limits. One JVM @@ -64,13 +60,13 @@ pages can consume substantial headroom. After a rollout, check trim reclamation and duration alongside host `MemAvailable`, memory pressure, request latency, and completion of the scheduled backup. Trimming can contend with native allocations, and it cannot bound live -native memory growth. The two-instance minimum provides spare serving capacity; -it does not protect against simultaneous failures or replace these checks. +native memory growth. The allocation owner must be measured if the process +continues growing after freed pages have been returned. To disable periodic trimming, append `-XX:TrimNativeHeapInterval=0` to the existing `JAVA_OPTS` and roll the configuration. The environment options follow -the container defaults, so they take precedence. Preserve the heap settings and -two-instance minimum when rolling back only trimming. +the container defaults, so they take precedence. Preserve the heap settings +when rolling back only trimming. ## Deployment From 8b3fbc3d198f74f1abac36bbf9723e4422633c45 Mon Sep 17 00:00:00 2001 From: stopachka Date: Mon, 28 Sep 2026 11:05:53 -0700 Subject: [PATCH 3/4] Bound temporary NIO buffer retention after large socket writes --- server/docker-compose.yml | 4 +-- server/infra/README.md | 66 +++++++++++++++++++++++++-------------- 2 files changed, 44 insertions(+), 26 deletions(-) diff --git a/server/docker-compose.yml b/server/docker-compose.yml index c676d03a16..0a15db270d 100644 --- a/server/docker-compose.yml +++ b/server/docker-compose.yml @@ -16,8 +16,8 @@ services: # `web` depends_on `vector` so that we don't lose logs on docker compose down depends_on: - vector - # Return freed native allocator pages to the OS; JAVA_OPTS can override these defaults. - command: sh -c 'java -XX:TrimNativeHeapInterval=60000 -Xlog:trimnative=info $${JAVA_OPTS} --add-modules java.se --add-exports java.base/jdk.internal.ref=ALL-UNNAMED --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/sun.nio.ch=ALL-UNNAMED --add-opens java.management/sun.management=ALL-UNNAMED --add-opens jdk.management/com.sun.management.internal=ALL-UNNAMED -XX:+UnlockDiagnosticVMOptions -XX:+DebugNonSafepoints -Djdk.attach.allowAttachSelf -Djava.awt.headless=true -Dcom.sun.management.jmxremote.port=10001 -Dcom.sun.management.jmxremote.rmi.port=10001 -Dcom.sun.management.jmxremote.authenticate=false -Dcom.sun.management.jmxremote.ssl=false -XX:StartFlightRecording=disk=true,maxsize=1G,name=instant,filename=instant_$(date +%s%N).jfr -Dcom.sun.management.jmxremote.local.only=false -Djava.rmi.server.hostname=localhost -server -jar target/instant-standalone.jar' + # Bound cached NIO copy buffers and return freed native pages; JAVA_OPTS can override these defaults. + command: sh -c 'java -Djdk.nio.maxCachedBufferSize=131072 -XX:TrimNativeHeapInterval=60000 -Xlog:trimnative=info $${JAVA_OPTS} --add-modules java.se --add-exports java.base/jdk.internal.ref=ALL-UNNAMED --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/sun.nio.ch=ALL-UNNAMED --add-opens java.management/sun.management=ALL-UNNAMED --add-opens jdk.management/com.sun.management.internal=ALL-UNNAMED -XX:+UnlockDiagnosticVMOptions -XX:+DebugNonSafepoints -Djdk.attach.allowAttachSelf -Djava.awt.headless=true -Dcom.sun.management.jmxremote.port=10001 -Dcom.sun.management.jmxremote.rmi.port=10001 -Dcom.sun.management.jmxremote.authenticate=false -Dcom.sun.management.jmxremote.ssl=false -XX:StartFlightRecording=disk=true,maxsize=1G,name=instant,filename=instant_$(date +%s%N).jfr -Dcom.sun.management.jmxremote.local.only=false -Djava.rmi.server.hostname=localhost -server -jar target/instant-standalone.jar' vector: # Mirrored from timberio/vector:0.56.0-alpine into our private ECR so diff --git a/server/infra/README.md b/server/infra/README.md index 403cbefd6f..4ba1a732d7 100644 --- a/server/infra/README.md +++ b/server/infra/README.md @@ -43,30 +43,48 @@ instances that fail. ## Native memory -The API container enables `-XX:TrimNativeHeapInterval=60000` on Corretto 26 -with glibc. A dedicated JVM thread periodically returns free native allocator -pages to the OS. This does not collect the Java heap or reclaim live native -allocations. `-Xlog:trimnative=info` records the reclaimed memory and trim -duration. Heap sizing still comes from `JAVA_OPTS`. - -On September 28, two successive hosts reached about 120.7 GiB RSS on a -123.1 GiB machine before becoming unresponsive. Available memory fell to roughly -100 MiB, while the last heap-pressure readings were below 50%. Both stalls -coincided with sustained root-disk reads at 125 MiB/s. Heap and GC alarms do not -cover this host-memory failure. A later native trim on the surviving process -returned about 4.6 GiB without restarting Java, confirming that freed allocator -pages can consume substantial headroom. - -After a rollout, check trim reclamation and duration alongside host -`MemAvailable`, memory pressure, request latency, and completion of the scheduled -backup. Trimming can contend with native allocations, and it cannot bound live -native memory growth. The allocation owner must be measured if the process -continues growing after freed pages have been returned. - -To disable periodic trimming, append `-XX:TrimNativeHeapInterval=0` to the -existing `JAVA_OPTS` and roll the configuration. The environment options follow -the container defaults, so they take precedence. Preserve the heap settings -when rolling back only trimming. +The API sets `-Djdk.nio.maxCachedBufferSize=131072`. When NIO writes a heap +buffer to a socket, it copies the data into a temporary native buffer. Corretto +26 caches these buffers on long-lived platform/carrier threads without a size +limit by default. Undertow's gathering writes can leave many large buffers in +each IO thread's cache after the WebSocket messages finish. These buffers are +outside the Java heap and are excluded from the direct-buffer MXBean and +`MaxDirectMemorySize` accounting. + +The 128 KiB cap applies to each cached buffer. It preserves reuse of ordinary +Java socket buffers while freeing larger temporary buffers after I/O. It does +not limit message sizes or the memory needed by writes in progress. The cache +can hold up to 1,024 entries per thread; this is not a process-wide native +memory limit. The property is read at NIO initialization, so changes require a +JVM restart. + +Two September 28 hosts reached about 120.7 GiB RSS on 123.1 GiB machines while +heap pressure remained below 50%. Subsequent profiling identified repeated +34.6 MiB native allocations in `Util.getTemporaryDirectBuffer` during Undertow +WebSocket writes. A read-only cache census found 4.97 GiB retained on the +surviving host, including 4.96 GiB on its 32 IO threads. The newer host already +held 1.67 GiB. The failed processes were unavailable for a cache census, so +these measurements do not retrospectively assign every byte of their RSS. + +The container also sets `-XX:TrimNativeHeapInterval=60000` with +`-Xlog:trimnative=info`. A dedicated JVM thread returns freed glibc pages to the +OS every minute and logs reclamation and duration. A previous trim reclaimed +4.6 GiB on the surviving process. This complements the cache cap: trimming +cannot reclaim buffers that the NIO cache still owns. Heap sizing remains in +`JAVA_OPTS`. + +After rollout, observe host RSS and `MemAvailable`, trim duration, request and +reactivity latency, GC pressure, and large-message traffic through a full backup +cycle. Backpressured large writes can allocate/free temporary buffers repeatedly; +the cache cap therefore trades some allocation work for bounded retention. +Local socket tests establish payload correctness and memory reclamation, not +production latency bounds. + +`JAVA_OPTS` follows these defaults, allowing an explicit override. To restore +the original NIO cache behavior, append +`-Djdk.nio.maxCachedBufferSize=9223372036854775807`. To disable trimming, append +`-XX:TrimNativeHeapInterval=0`. Roll the configuration and preserve the other JVM +options, including heap settings. ## Deployment From 3a57848451595b0d438d6f391041092e6b897784 Mon Sep 17 00:00:00 2001 From: stopachka Date: Mon, 28 Sep 2026 11:06:01 -0700 Subject: [PATCH 4/4] Clarify final heap pressure observations --- server/infra/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/server/infra/README.md b/server/infra/README.md index 4ba1a732d7..66604814b9 100644 --- a/server/infra/README.md +++ b/server/infra/README.md @@ -59,7 +59,7 @@ memory limit. The property is read at NIO initialization, so changes require a JVM restart. Two September 28 hosts reached about 120.7 GiB RSS on 123.1 GiB machines while -heap pressure remained below 50%. Subsequent profiling identified repeated +their last reported heap pressure was below 50%. Profiling identified repeated 34.6 MiB native allocations in `Util.getTemporaryDirectBuffer` during Undertow WebSocket writes. A read-only cache census found 4.97 GiB retained on the surviving host, including 4.96 GiB on its 32 IO threads. The newer host already