Describe the issue
Since 1.29.0, any process whose command line is longer than roughly 4KB dies with SIGSEGV (exit 139) inside OrtEnvironment.getEnvironment(), before any session or model is involved. 1.30.0 is still affected. 1.28.0 and earlier are fine. The JVM writes no hs_err file, so the process just disappears.
This is a second, independent trigger of the 1.29.0 POSIX telemetry regression reported in #32173. The fix that closed that issue (#32226, null popen() pipe in shell-less containers) does not touch this path, which is why 1.30.0 still crashes.
The very common real-world hit is any Java process launched from an IDE (IntelliJ, Eclipse, VS Code) or by Gradle/Maven test runners: they pass the full explicit -classpath, which is easily 10-20KB. From the IDE every model load crashes; the same code launched from a shaded jar works, which makes this look like an environment problem rather than an ORT regression. It is not Java-specific: any host (Python, .NET, C++) with a long argv will hit it.
Root cause (from strace + gdb + source, details below): the 1DS SDK (cpp_client_telemetry, pinned to v3.10.173.1 in cmake/deps.txt) registers the appId field as
// lib/pal/posix/sysinfo_sources.cpp, sysinfo_sources_impl::sysinfo_sources_impl()
add("appId", {"/proc/self/cmdline", "(.*)[ ]*.*[\n]*"});
and sysinfo_sources::fetch() reads the entire file and evaluates that selector with std::regex_search. libstdc++'s std::regex is a recursive backtracking matcher that recurses (at least) once per input character for .* (GCC PR 86164), so stack usage grows linearly with the length of /proc/self/cmdline. With the JVM's default 1MB thread stack the limit is ~4KB of command line; with -Xss8m it is ~15-30KB; with -Xss64m a 100KB command line passes. The fault happens on the calling thread inside OrtEnv creation, so nothing in the application can catch it.
Evidence
Minimal Java program (Repro.java below) that only calls OrtEnvironment.getEnvironment(), run with an otherwise unused -Dpad=<N bytes of 'x'> to lengthen the command line. All runs on the same machine, same JDK, official Maven Central artifacts.
| Artifact |
-Dpad size (total cmdline) |
Extra |
Result |
onnxruntime_gpu 1.28.0 |
none / 15,000 B |
|
OK |
onnxruntime_gpu 1.29.0 |
15,000 B |
|
SIGSEGV, exit 139 |
onnxruntime_gpu 1.30.0 |
none |
|
OK |
onnxruntime_gpu 1.30.0 |
1,000 / 2,000 / 3,000 B (~3.2KB total) |
|
OK |
onnxruntime_gpu 1.30.0 |
4,000 / 6,000 / 8,000 / 15,000 B |
|
SIGSEGV, exit 139 |
onnxruntime (CPU) 1.30.0 |
3,000 B |
|
OK |
onnxruntime (CPU) 1.30.0 |
4,000 / 6,000 / 8,000 B |
|
SIGSEGV, exit 139 |
onnxruntime_gpu 1.30.0 |
15,000 B |
ORT_DISABLE_TELEMETRY=1 |
OK |
onnxruntime_gpu 1.30.0 |
15,000 B |
-Xss8m |
OK |
onnxruntime_gpu 1.30.0 |
30,000 / 60,000 / 100,000 B |
-Xss8m |
SIGSEGV |
onnxruntime_gpu 1.30.0 |
100,000 B |
-Xss64m |
OK |
strace -f of the crashing run, after libonnxruntime.so is loaded (this is exactly the sysinfo_sources_impl constructor's file list, in order):
openat(AT_FDCWD, "/etc/os-release", O_RDONLY) = 8
openat(AT_FDCWD, "/etc/os-release", O_RDONLY) = 8
openat(AT_FDCWD, "/etc/os-release", O_RDONLY) = 8
openat(AT_FDCWD, "/etc/machine-id", O_RDONLY) = 8
openat(AT_FDCWD, "/proc/self/cmdline", O_RDONLY) = 8
read(8, "java\0-Dpad=xxxxxxxxxxxxxxxxxx"..., 8191) = 8191
read(8, "xxxxxxxxxxxxxxxxxxxxxxxxxxxxx"..., 8191) = 6974
read(8, "", 8191) = 0
--- SIGSEGV {si_signo=SIGSEGV, si_code=SEGV_ACCERR, si_addr=0x76dce4d03fdf} ---
--- SIGSEGV {si_signo=SIGSEGV, si_code=SI_KERNEL, si_addr=NULL} ---
+++ killed by SIGSEGV (core dumped) +++
gdb on the fatal signal (JVM's own benign SIGSEGVs passed through): a 10,600+ frame backtrace consisting of one three-address cycle inside libonnxruntime.so, bottoming out in the JNI entry point. The shipped .so is stripped so the frames have no names, but the depth, the 3-frame cycle and the /proc/self/cmdline read immediately before match std::__detail::_Executor::_M_dfs recursion.
#0 0x00007fff8e1fb11b in ?? () from libonnxruntime.so
#1 0x00007fff8e1fba85 in ?? () from libonnxruntime.so
#2 0x00007fff8e1fb22a in ?? () from libonnxruntime.so
#3 0x00007fff8e1fb53a in ?? () from libonnxruntime.so
#4 0x00007fff8e1fba85 in ?? () from libonnxruntime.so
#5 0x00007fff8e1fb22a in ?? () from libonnxruntime.so
... (same three addresses repeating)
#10609 0x00007fff8e1fba85 in ?? () from libonnxruntime.so
#10610 0x00007fff8e1fb22a in ?? () from libonnxruntime.so
#10611 0x00007fff8e1fb4de in ?? () from libonnxruntime.so
...
#10630 0x00007fff8d1401ed in ?? () from libonnxruntime.so
#10631 0x00007fffefef34a9 in Java_ai_onnxruntime_OrtEnvironment_createHandle__JILjava_lang_String_2 () from libonnxruntime4j_jni.so
To reproduce
Repro.java:
import ai.onnxruntime.OrtEnvironment;
public class Repro {
public static void main(String[] args) throws Exception {
System.out.println("[stage] calling OrtEnvironment.getEnvironment()...");
System.out.flush();
OrtEnvironment env = OrtEnvironment.getEnvironment();
System.out.println("[stage] OK, ORT version " + env.getVersion());
}
}
JAR=~/.m2/repository/com/microsoft/onnxruntime/onnxruntime_gpu/1.30.0/onnxruntime_gpu-1.30.0.jar # or onnxruntime (CPU) 1.30.0, same result
javac -cp $JAR Repro.java
PAD=$(head -c 15000 /dev/zero | tr '\0' x)
java -cp "$JAR:." Repro # OK
java -Dpad=$PAD -cp "$JAR:." Repro # SIGSEGV, exit 139, no hs_err
ORT_DISABLE_TELEMETRY=1 java -Dpad=$PAD -cp "$JAR:." Repro # OK
java -Xss8m -Dpad=$PAD -cp "$JAR:." Repro # OK (until the cmdline grows further)
No model is needed; the crash is in OrtEnv creation.
Suggested fix
The appId selector should not run a backtracking regex over an unbounded input. /proc/self/cmdline is NUL-separated, so the regex (.*)[ ]*.*[\n]* is not even meaningful for it. Options, cheapest first:
- In ORT's
cmake/patches/cpp_client_telemetry/cpp_client_telemetry.patch, change the appId source to selector "*" (raw read, no regex) and truncate at the first '\0' in fetch() / the consumer, or cap the bytes read from /proc/self/cmdline (e.g. first 4KB) before applying any regex.
- Same change upstream in
microsoft/cpp_client_telemetry (the regex is still present at HEAD).
- Longer term,
fetch() should not use std::regex on file contents of unbounded size at all (the /proc/version and /etc/os-release selectors are also .*-based, just on inputs that are small in practice).
Independently of the fix, it would help to document ORT_DISABLE_TELEMETRY=1 in the Java/Linux docs, since the failure mode is a silent process death with no hs_err and no stderr output.
Urgency
High for Java users on Linux: every IDE launch and most build-tool test runs crash on 1.29.0/1.30.0 with no diagnostics. We are pinned to 1.28.0 because of this, which also blocks us from picking up the CUDA arena fix from #29589 that we were waiting for (#29351).
Platform
Linux
OS Version
Ubuntu (glibc 2.43), kernel 7.0.0; JDK: Amazon Corretto 26.0.1 (also reproduced on 25). Same result under IntelliJ, Maven surefire and a plain shell as long as the command line is long enough.
ONNX Runtime Installation
Released Package (Maven Central: com.microsoft.onnxruntime:onnxruntime_gpu and com.microsoft.onnxruntime:onnxruntime)
ONNX Runtime Version or Commit ID
1.29.0, 1.30.0 (1.28.0 and 1.27.0 unaffected)
ONNX Runtime API
Java
Architecture
X64
Execution Provider
Default CPU (crash is in OrtEnv creation, before any EP is selected; both the CPU and GPU packages reproduce)
Execution Provider Library Version
n/a
Describe the issue
Since 1.29.0, any process whose command line is longer than roughly 4KB dies with SIGSEGV (exit 139) inside
OrtEnvironment.getEnvironment(), before any session or model is involved. 1.30.0 is still affected. 1.28.0 and earlier are fine. The JVM writes nohs_errfile, so the process just disappears.This is a second, independent trigger of the 1.29.0 POSIX telemetry regression reported in #32173. The fix that closed that issue (#32226, null
popen()pipe in shell-less containers) does not touch this path, which is why 1.30.0 still crashes.The very common real-world hit is any Java process launched from an IDE (IntelliJ, Eclipse, VS Code) or by Gradle/Maven test runners: they pass the full explicit
-classpath, which is easily 10-20KB. From the IDE every model load crashes; the same code launched from a shaded jar works, which makes this look like an environment problem rather than an ORT regression. It is not Java-specific: any host (Python, .NET, C++) with a longargvwill hit it.Root cause (from strace + gdb + source, details below): the 1DS SDK (
cpp_client_telemetry, pinned to v3.10.173.1 incmake/deps.txt) registers theappIdfield asand
sysinfo_sources::fetch()reads the entire file and evaluates that selector withstd::regex_search. libstdc++'sstd::regexis a recursive backtracking matcher that recurses (at least) once per input character for.*(GCC PR 86164), so stack usage grows linearly with the length of/proc/self/cmdline. With the JVM's default 1MB thread stack the limit is ~4KB of command line; with-Xss8mit is ~15-30KB; with-Xss64ma 100KB command line passes. The fault happens on the calling thread insideOrtEnvcreation, so nothing in the application can catch it.Evidence
Minimal Java program (
Repro.javabelow) that only callsOrtEnvironment.getEnvironment(), run with an otherwise unused-Dpad=<N bytes of 'x'>to lengthen the command line. All runs on the same machine, same JDK, official Maven Central artifacts.-Dpadsize (total cmdline)onnxruntime_gpu1.28.0onnxruntime_gpu1.29.0onnxruntime_gpu1.30.0onnxruntime_gpu1.30.0onnxruntime_gpu1.30.0onnxruntime(CPU) 1.30.0onnxruntime(CPU) 1.30.0onnxruntime_gpu1.30.0ORT_DISABLE_TELEMETRY=1onnxruntime_gpu1.30.0-Xss8monnxruntime_gpu1.30.0-Xss8monnxruntime_gpu1.30.0-Xss64mstrace -fof the crashing run, afterlibonnxruntime.sois loaded (this is exactly thesysinfo_sources_implconstructor's file list, in order):gdb on the fatal signal (JVM's own benign SIGSEGVs passed through): a 10,600+ frame backtrace consisting of one three-address cycle inside
libonnxruntime.so, bottoming out in the JNI entry point. The shipped.sois stripped so the frames have no names, but the depth, the 3-frame cycle and the/proc/self/cmdlineread immediately before matchstd::__detail::_Executor::_M_dfsrecursion.To reproduce
Repro.java:No model is needed; the crash is in
OrtEnvcreation.Suggested fix
The
appIdselector should not run a backtracking regex over an unbounded input./proc/self/cmdlineis NUL-separated, so the regex(.*)[ ]*.*[\n]*is not even meaningful for it. Options, cheapest first:cmake/patches/cpp_client_telemetry/cpp_client_telemetry.patch, change theappIdsource to selector"*"(raw read, no regex) and truncate at the first'\0'infetch()/ the consumer, or cap the bytes read from/proc/self/cmdline(e.g. first 4KB) before applying any regex.microsoft/cpp_client_telemetry(the regex is still present at HEAD).fetch()should not usestd::regexon file contents of unbounded size at all (the/proc/versionand/etc/os-releaseselectors are also.*-based, just on inputs that are small in practice).Independently of the fix, it would help to document
ORT_DISABLE_TELEMETRY=1in the Java/Linux docs, since the failure mode is a silent process death with no hs_err and no stderr output.Urgency
High for Java users on Linux: every IDE launch and most build-tool test runs crash on 1.29.0/1.30.0 with no diagnostics. We are pinned to 1.28.0 because of this, which also blocks us from picking up the CUDA arena fix from #29589 that we were waiting for (#29351).
Platform
Linux
OS Version
Ubuntu (glibc 2.43), kernel 7.0.0; JDK: Amazon Corretto 26.0.1 (also reproduced on 25). Same result under IntelliJ, Maven surefire and a plain shell as long as the command line is long enough.
ONNX Runtime Installation
Released Package (Maven Central:
com.microsoft.onnxruntime:onnxruntime_gpuandcom.microsoft.onnxruntime:onnxruntime)ONNX Runtime Version or Commit ID
1.29.0, 1.30.0 (1.28.0 and 1.27.0 unaffected)
ONNX Runtime API
Java
Architecture
X64
Execution Provider
Default CPU (crash is in
OrtEnvcreation, before any EP is selected; both the CPU and GPU packages reproduce)Execution Provider Library Version
n/a