Repository navigation
fix(executorch): initialize a CUDA context for default-stream execution - #4779
Merged
lanluo-nvidia merged 1 commit intoOct 6, 2026
Merged
Conversation
shoumikhin
force-pushed
the
executorch-thread-cuda-context
branch
2 times, most recently
from
October 4, 2026 15:02
10bb698 to
d267a5d
Compare
shoumikhin
force-pushed
the
executorch-thread-cuda-context
branch
from
October 4, 2026 21:06
d267a5d to
430937e
Compare
shoumikhin
marked this pull request as ready for review
October 4, 2026 21:07
shoumikhin
force-pushed
the
executorch-thread-cuda-context
branch
3 times, most recently
from
October 6, 2026 04:15
e3e32a7 to
f664ab9
Compare
Running a TensorRT program through ExecuTorch can fail on a worker thread that has not used CUDA yet, even when it works on the loading thread. TensorRT 11.3 looks up default-stream contexts through the CUDA driver, and that lookup needs a current context. This affects engines loaded with shared activation scratch disabled, and engines loaded with it enabled that need no scratch. Both skip the pool's capture check, which otherwise makes a context current. With no caller stream, cudaStreamLegacy, or cudaStreamPerThread, their first enqueue can fail. For a default stream, execute now calls cudaFree(nullptr) after the capture refusal and the input and output checks, before claiming scratch or enqueueing. This makes the primary context current only when none is current, preserving a context the caller set. Caller-created streams are unchanged. A failed call logs the CUDA error, clears it, and returns Error::Internal before enqueueing. Sticky device faults can survive the clear. Capture stays unsupported. A call on any stream must not overlap a Global capture on any thread, or a ThreadLocal capture on the calling thread; the README and header now say so for every engine. The header also notes that NotSupported can mean a refused output resize. One limit remains: a fresh thread can reject an input or output mapped from another GPU, because pointer validation runs before context setup. Tested on Linux x86_64 with TensorRT 11.3: the shared scratch backend test file passes 36 of 36. The new scratch-free engine test fails three times out of three when only that path is skipped. The fresh-thread tests fail without the fix. With TensorRT 11.1, all 36 pass, and the fresh-thread failure does not reproduce.
shoumikhin
force-pushed
the
executorch-thread-cuda-context
branch
from
October 6, 2026 18:37
f664ab9 to
1ee0626
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Running a TensorRT program through ExecuTorch can fail on a worker thread that has not used CUDA yet, even when it works on the loading thread. TensorRT 11.3 looks up default-stream contexts through the CUDA driver, and that lookup needs a current context.
This affects engines loaded with shared activation scratch disabled and engines loaded with it enabled that need no scratch. Both skip the pool's capture check, which otherwise makes a context current. With no caller stream,
cudaStreamLegacy, orcudaStreamPerThread, their first enqueue can fail.For a default stream,
execute()callscudaFree(nullptr)after capture refusal and input/output validation, before claiming scratch or enqueueing. This makes the primary context current only when no context is current, preserving a context the caller set. Explicitnullptrgets the call too; caller-created streams are unchanged. A failed call logs CUDA's error, clears it, and returnsError::Internalbefore enqueueing. Sticky device faults can survive the clear.Capture remains unsupported. Calls on any stream must not overlap a Global capture on any thread or a ThreadLocal capture on the calling thread. This includes option-on engines that need no scratch. A ThreadLocal capture on another thread is unaffected. The header also clarifies that
NotSupportedcan mean a refused output resize, not just a pool capture refusal.One existing limit remains: a fresh thread can reject an input or output mapped from another GPU because pointer validation runs before context setup. Such buffers still need the engine's context made current on that thread before execution.
Tests cover each default-stream selection from fresh threads, including option-on scratch-free engines and pooled engines, caller-context preservation, capture refusal, and rejection of bad input. On Linux x86_64 with TensorRT 11.3, the shared scratch backend test file passes 36 of 36, and the new scratch-free test fails three times out of three when only that path is skipped. The explicit
nullptrrow also passes without this fix. With TensorRT 11.1, all 36 pass, and the fresh-thread regression does not reproduce.