Repository navigation
feat(executorch): share TensorRT engines across loads - #4778
Merged
Merged
Conversation
shoumikhin
force-pushed
the
executorch-share-engines
branch
3 times, most recently
from
October 4, 2026 21:06
56b0354 to
81987f8
Compare
shoumikhin
marked this pull request as ready for review
October 4, 2026 21:07
shoumikhin
force-pushed
the
executorch-share-engines
branch
2 times, most recently
from
October 5, 2026 22:38
ef32a21 to
396213e
Compare
…unction init applied the budget inline, between deserializing the engine and creating its execution context. Move that block, unchanged, into apply_weight_streaming_budget, called at the same point, so the next change can run it once per engine rather than once per handle. No behavior change. Tested the 29 shared scratch backend cases on Linux x86_64 with an H100, TensorRT 11.1 and CUDA 12.8. Every case passed, including allocation failure and input copy cleanup.
shoumikhin
force-pushed
the
executorch-share-engines
branch
from
October 6, 2026 05:14
396213e to
97b2195
Compare
Loading the same engine more than once used to keep separate copies of its weights. Share one engine across loads with matching engine hashes, sizes, devices and requested weight streaming budgets, including loads from different programs. Each handle keeps its own execution context, buffers and lock. The last handle releases the engine. The use_shared_engines boolean is a load option, true by default. C++ callers can pass false through Module::load to keep a module private. set_option rejects this load-only key without applying other options in the same batch. Python bindings do not expose backend load options. Matching uses a non-cryptographic hash and size without comparing the bytes. Every program in the process must be trusted. The map lock never covers TensorRT calls or logging. Racing loads may deserialize separate copies before one publishes. A failed load checks for an engine published in the meantime; otherwise it returns the failure. The first published engine keeps its budget, including an automatic budget chosen during a race. An explicit budget gives a predictable request. Budget validation precedes deserialization. If both are invalid, the budget error is reported first. Code using EngineHandle must rebuild because its layout changes. A separate test-only runtime accessor supports allocator failure injection without interposing private TensorRT symbols. Tested: Linux x86_64 with two GPUs, 23 shared-engine cases, five shuffled runs, 20 recovery repetitions and 29 scratch cases. Both load-only-option tests fail before the fix. Removing sharing, recovery or device separation makes the matching tests fail. Python artifact guards: 122 pass on macOS, 11 Linux-only cases skip. Concurrent execution tests check outputs, not overlap. A pooled-scratch concurrency check intermittently returned one wrong output, also without engine sharing. Reruns passed, but its cause remains unresolved.
shoumikhin
force-pushed
the
executorch-share-engines
branch
from
October 6, 2026 17:08
97b2195 to
015b0ed
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Loading the same TensorRT engine more than once used to keep a separate copy of its weights for each load. This change shares one engine across loads with matching engine hashes, sizes, devices and requested weight streaming budgets. Sharing also works across different programs that carry the same engine. Each handle keeps its own execution context, buffers and lock, and the last handle releases the engine.
The
use_shared_enginesboolean load option defaults to true. C++ callers can pass false throughModule::loadto keep a module's engines private. Passing this load-only key toexecutorch::runtime::set_optionreturnsError::InvalidArgumentand explains where to pass it. Python bindings do not expose load options, so Python loads use sharing.Matching uses a non-cryptographic hash and size, without comparing the bytes. Every program in the process must be trusted. Racing loads may temporarily deserialize separate copies before sharing the published engine. The first publisher's budget stays in effect, including an automatic budget chosen during a race. Use an explicit budget for a predictable request.
The first commit extracts budget application without changing behavior. The second adds sharing and validates budgets before deserialization. Code that accesses
EngineHandlemust rebuild because its layout changes.Built and ran the C++ shared-engine and scratch suites on Linux x86_64 with two GPUs, including five shuffled runs and 20 recovery repeats. Removing sharing, recovery or device separation made the matching tests fail. The two load-only-option tests fail before the fix and pass after it. The Python artifact guards passed on macOS; 11 Linux-only cases skipped. The recovery hook is linked only into tests, not release backend libraries. The concurrent execution cases check outputs, not overlap.
A pooled-scratch concurrency check intermittently returned one wrong output, also without engine sharing. Reruns passed, but its cause remains unresolved.