Skip to content

fix(executorch): initialize a CUDA context for default-stream execution - #4779

Merged
lanluo-nvidia merged 1 commit into
pytorch:mainfrom
shoumikhin:executorch-thread-cuda-context
Oct 6, 2026
Merged

lanluo-nvidia merged 1 commit into
pytorch:mainfrom
shoumikhin:executorch-thread-cuda-context

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Running a TensorRT program through ExecuTorch can fail on a worker thread that has not used CUDA yet, even when it works on the loading thread. TensorRT 11.3 looks up default-stream contexts through the CUDA driver, and that lookup needs a current context.

This affects engines loaded with shared activation scratch disabled and engines loaded with it enabled that need no scratch. Both skip the pool's capture check, which otherwise makes a context current. With no caller stream, cudaStreamLegacy, or cudaStreamPerThread, their first enqueue can fail.

For a default stream, execute() calls cudaFree(nullptr) after capture refusal and input/output validation, before claiming scratch or enqueueing. This makes the primary context current only when no context is current, preserving a context the caller set. Explicit nullptr gets the call too; caller-created streams are unchanged. A failed call logs CUDA's error, clears it, and returns Error::Internal before enqueueing. Sticky device faults can survive the clear.

Capture remains unsupported. Calls on any stream must not overlap a Global capture on any thread or a ThreadLocal capture on the calling thread. This includes option-on engines that need no scratch. A ThreadLocal capture on another thread is unaffected. The header also clarifies that NotSupported can mean a refused output resize, not just a pool capture refusal.

One existing limit remains: a fresh thread can reject an input or output mapped from another GPU because pointer validation runs before context setup. Such buffers still need the engine's context made current on that thread before execution.

Tests cover each default-stream selection from fresh threads, including option-on scratch-free engines and pooled engines, caller-context preservation, capture refusal, and rejection of bad input. On Linux x86_64 with TensorRT 11.3, the shared scratch backend test file passes 36 of 36, and the new scratch-free test fails three times out of three when only that path is skipped. The explicit nullptr row also passes without this fix. With TensorRT 11.1, all 36 pass, and the fresh-thread regression does not reproduce.

@meta-cla meta-cla Bot added the cla signed label Oct 4, 2026
@github-actions github-actions Bot added the component: api [C++] Issues re: C++ API label Oct 4, 2026
@github-actions
github-actions Bot requested a review from narendasan October 4, 2026 13:54
@shoumikhin
shoumikhin force-pushed the executorch-thread-cuda-context branch 2 times, most recently from 10bb698 to d267a5d Compare October 4, 2026 15:02
@shoumikhin shoumikhin changed the title fix(executorch): make a CUDA context current before enqueue on the per-thread stream fix(executorch): make execute work on a thread that has not used CUDA yet Oct 4, 2026
@shoumikhin
shoumikhin force-pushed the executorch-thread-cuda-context branch from d267a5d to 430937e Compare October 4, 2026 21:06
@github-actions github-actions Bot added the component: tests Issues re: Tests label Oct 4, 2026
@shoumikhin
shoumikhin marked this pull request as ready for review October 4, 2026 21:07
@shoumikhin
shoumikhin force-pushed the executorch-thread-cuda-context branch 3 times, most recently from e3e32a7 to f664ab9 Compare October 6, 2026 04:15
Running a TensorRT program through ExecuTorch can fail on a worker thread
that has not used CUDA yet, even when it works on the loading thread.
TensorRT 11.3 looks up default-stream contexts through the CUDA driver,
and that lookup needs a current context.

This affects engines loaded with shared activation scratch disabled, and
engines loaded with it enabled that need no scratch. Both skip the pool's
capture check, which otherwise makes a context current. With no caller
stream, cudaStreamLegacy, or cudaStreamPerThread, their first enqueue can
fail.

For a default stream, execute now calls cudaFree(nullptr) after the
capture refusal and the input and output checks, before claiming scratch
or enqueueing. This makes the primary context current only when none is
current, preserving a context the caller set. Caller-created streams are
unchanged. A failed call logs the CUDA error, clears it, and returns
Error::Internal before enqueueing. Sticky device faults can survive the
clear.

Capture stays unsupported. A call on any stream must not overlap a Global
capture on any thread, or a ThreadLocal capture on the calling thread;
the README and header now say so for every engine. The header also notes
that NotSupported can mean a refused output resize.

One limit remains: a fresh thread can reject an input or output mapped
from another GPU, because pointer validation runs before context setup.

Tested on Linux x86_64 with TensorRT 11.3: the shared scratch backend
test file passes 36 of 36. The new scratch-free engine test fails three
times out of three when only that path is skipped. The fresh-thread tests
fail without the fix. With TensorRT 11.1, all 36 pass, and the
fresh-thread failure does not reproduce.
@shoumikhin
shoumikhin force-pushed the executorch-thread-cuda-context branch from f664ab9 to 1ee0626 Compare October 6, 2026 18:37
@shoumikhin shoumikhin changed the title fix(executorch): make execute work on a thread that has not used CUDA yet fix(executorch): initialize a CUDA context for default-stream execution Oct 6, 2026
@lanluo-nvidia
lanluo-nvidia merged commit a512a9f into pytorch:main Oct 6, 2026
49 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants