Skip to content

feat(executorch): share TensorRT engines across loads - #4778

Merged
lanluo-nvidia merged 2 commits into
pytorch:mainfrom
shoumikhin:executorch-share-engines
Oct 6, 2026
Merged

lanluo-nvidia merged 2 commits into
pytorch:mainfrom
shoumikhin:executorch-share-engines

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Loading the same TensorRT engine more than once used to keep a separate copy of its weights for each load. This change shares one engine across loads with matching engine hashes, sizes, devices and requested weight streaming budgets. Sharing also works across different programs that carry the same engine. Each handle keeps its own execution context, buffers and lock, and the last handle releases the engine.

The use_shared_engines boolean load option defaults to true. C++ callers can pass false through Module::load to keep a module's engines private. Passing this load-only key to executorch::runtime::set_option returns Error::InvalidArgument and explains where to pass it. Python bindings do not expose load options, so Python loads use sharing.

Matching uses a non-cryptographic hash and size, without comparing the bytes. Every program in the process must be trusted. Racing loads may temporarily deserialize separate copies before sharing the published engine. The first publisher's budget stays in effect, including an automatic budget chosen during a race. Use an explicit budget for a predictable request.

The first commit extracts budget application without changing behavior. The second adds sharing and validates budgets before deserialization. Code that accesses EngineHandle must rebuild because its layout changes.

Built and ran the C++ shared-engine and scratch suites on Linux x86_64 with two GPUs, including five shuffled runs and 20 recovery repeats. Removing sharing, recovery or device separation made the matching tests fail. The two load-only-option tests fail before the fix and pass after it. The Python artifact guards passed on macOS; 11 Linux-only cases skipped. The recovery hook is linked only into tests, not release backend libraries. The concurrent execution cases check outputs, not overlap.

A pooled-scratch concurrency check intermittently returned one wrong output, also without engine sharing. Reruns passed, but its cause remains unresolved.

@meta-cla meta-cla Bot added the cla signed label Oct 4, 2026
@github-actions github-actions Bot added component: tests Issues re: Tests component: api [C++] Issues re: C++ API labels Oct 4, 2026
@github-actions
github-actions Bot requested a review from lanluo-nvidia October 4, 2026 13:54
@shoumikhin
shoumikhin force-pushed the executorch-share-engines branch 3 times, most recently from 56b0354 to 81987f8 Compare October 4, 2026 21:06
@shoumikhin
shoumikhin marked this pull request as ready for review October 4, 2026 21:07
@shoumikhin
shoumikhin force-pushed the executorch-share-engines branch 2 times, most recently from ef32a21 to 396213e Compare October 5, 2026 22:38
…unction

init applied the budget inline, between deserializing the engine and
creating its execution context. Move that block, unchanged, into
apply_weight_streaming_budget, called at the same point, so the next
change can run it once per engine rather than once per handle.

No behavior change.

Tested the 29 shared scratch backend cases on Linux x86_64 with an H100,
TensorRT 11.1 and CUDA 12.8. Every case passed, including allocation
failure and input copy cleanup.
@shoumikhin
shoumikhin force-pushed the executorch-share-engines branch from 396213e to 97b2195 Compare October 6, 2026 05:14
Loading the same engine more than once used to keep separate copies of
its weights. Share one engine across loads with matching engine hashes,
sizes, devices and requested weight streaming budgets, including loads
from different programs. Each handle keeps its own execution context,
buffers and lock. The last handle releases the engine.

The use_shared_engines boolean is a load option, true by default. C++
callers can pass false through Module::load to keep a module private.
set_option rejects this load-only key without applying other options in
the same batch. Python bindings do not expose backend load options.

Matching uses a non-cryptographic hash and size without comparing the
bytes. Every program in the process must be trusted. The map lock never
covers TensorRT calls or logging. Racing loads may deserialize separate
copies before one publishes. A failed load checks for an engine published
in the meantime; otherwise it returns the failure. The first published
engine keeps its budget, including an automatic budget chosen during a
race. An explicit budget gives a predictable request.

Budget validation precedes deserialization. If both are invalid, the
budget error is reported first. Code using EngineHandle must rebuild
because its layout changes. A separate test-only runtime accessor supports
allocator failure injection without interposing private TensorRT symbols.

Tested: Linux x86_64 with two GPUs, 23 shared-engine cases, five shuffled
runs, 20 recovery repetitions and 29 scratch cases. Both load-only-option
tests fail before the fix. Removing sharing, recovery or device separation
makes the matching tests fail. Python artifact guards: 122 pass on macOS,
11 Linux-only cases skip. Concurrent execution tests check outputs, not
overlap.

A pooled-scratch concurrency check intermittently returned one wrong
output, also without engine sharing. Reruns passed, but its cause remains
unresolved.
@shoumikhin
shoumikhin force-pushed the executorch-share-engines branch from 97b2195 to 015b0ed Compare October 6, 2026 17:08
@shoumikhin shoumikhin changed the title feat(executorch): share one TensorRT engine between loads of the same program feat(executorch): share TensorRT engines across loads Oct 6, 2026
@lanluo-nvidia lanluo-nvidia added this to the v2.15.0 milestone Oct 6, 2026

@lanluo-nvidia lanluo-nvidia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@github-actions
github-actions Bot requested a review from lanluo-nvidia October 6, 2026 20:40

@lanluo-nvidia lanluo-nvidia left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@lanluo-nvidia
lanluo-nvidia merged commit ec59f97 into pytorch:main Oct 6, 2026
192 of 233 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants