Repository navigation
perf(executorch): store each engine once, as named data, when saving a .pte - #4821
Draft
cehongwang wants to merge 1 commit into
Draft
cehongwang wants to merge 1 commit into
cehongwang wants to merge 1 commit into
Conversation
…a .pte Saving a .pte kept two host copies of every engine until the file was written: the serialized plan, registered on the exported program as the engine's buffer, and the delegate blob, which preprocess built by copying the plan in after the header. With the resource partitioner the save then held two copies of all engines at once and became the export's peak. preprocess now hands the engine to ExecuTorch as named data, keyed by its SHA-256, and the blob carries only the header and metadata, with an engine_key naming the entry. ExecuTorch stores named data only as bytes, so _resolve_engine_tensor converts TensorRT's serialized plan to bytes once, while TensorRT's buffer is the only other copy, and preprocess reuses that same object. Identical engines in different methods share one entry. At load, TensorRTBackend::init reads the engine from the program's named data map when the header names one, passes it to the shared-engine cache or load_engine as before, and frees the host copy once the engine is on the GPU. Inline blobs load as before. The header parser rejects a blob that has both or neither. On pi0.5 (bf16, 9 engines at cpu_memory_budget=8 GiB, on top of the getitem fix in the resource partitioner), the save's peak host RSS drops from 22.1 to 14.0 GiB, which is the export's peak.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
.ptekept two host copies of every engine until the file was written: the serialized plan, registered on the exported program as the engine's buffer, and the delegate blob, whichpreprocessbuilt by copying the plan in after the header. With the resource partitioner the save held two copies of all engines at once and became the export's peak.TensorRTBackend.preprocessnow hands the engine to ExecuTorch as named data (NamedDataStore, keytensorrt_engine_<sha256>, 16-byte aligned). The blob carries only the header and metadata, plus anengine_keynaming the entry. Identical engines in different methods share one entry.bytes, so_resolve_engine_tensorconverts TensorRT's serialized plan tobytesonce, while TensorRT's buffer is the only other copy, andpreprocessreuses that same object (engine_bytes_tensor/_trt_engine_bytes).TensorRTBackend::initreads the engine from the program'sNamedDataMapwhen the header names one, passes it toacquire_shared_engine(feat(executorch): share TensorRT engines across loads #4778) orload_engineas before, and frees the host copy once the engine is on the GPU, before building the execution context. Inline blobs load as before.TensorRTBlobHeader::parserejects a blob that has both an inline engine and a key, or neither.An earlier version of this change, written against
mainbefore #4795, kept a registry from engine to module so the save could read the module's retainedserialized_engine. #4795 removes that retained plan, so this version serializes from the C++ engine (serialized_engine_tensor()) and needs no registry.Memory saved
On the pi0.5 ExecuTorch + TensorRT export, this PR lowers the peak host memory from 17.8 GiB to 11.6 GiB (−6.2 GiB) when the model is partitioned (13 engines). For a single engine the peak is set by the engine build and doesn't change, but the save holds 4.9 GiB less (13.5 → 8.6 GiB).
Measured with
pi05_executorch_tensorrt.py, the lerobot pi0.5 export script (bf16, A40). Peak host RSS of the exporting process tree, GiB:.ptecpu_memory_budget=4 GiB), basecpu_memory_budget=4 GiB), this PRWith partitioning, the save held two copies of every engine and was the export's peak; now it holds one and stays level with compile. With a single engine the save peak barely moves: it serializes the engine from the GPU (one copy), then converts it to
bytes(a second), so both exist for a moment; afterwards only one is held instead of two. With partitioning each engine is converted in turn, so that moment is small. Removing it entirely needs the engine streamed to the file (for example through ExecuTorch'sFileBackedData), which I'd leave to a follow-up.A second export script (lerobot's own
PI05Policy, bf16 vision kept in fp32, 9 engines atcpu_memory_budget=8 GiB) shows the same: save peak 22.1 → 14.0 GiB, which is its export peak.How these numbers were measured
Same script, checkpoint, environment and command for both rows of each pair; only
PYTHONPATHdiffers (this branch vs. its base,narendasan/compile-mem-plan-passthrough).TORCHTRT_ENABLE_BUILDER_MALLOC_TRIM=1, TensorRT 11.0, torch 2.13. The partitioned runs also carry #4783 (keepgetitemwith its multi-output producer when splitting), without which pi0.5 fails to partition, and callrelease_host_and_device_memory()before compile. My ExecuTorch Python predates the device-resident I/O config the script uses, so the wrapper drops that config (I/O stays on the host; the tensors are KB-sized) and skips the script's post-save check of it; compile and save are otherwise unchanged. A sidecar samples the process tree's RSS (psutil) every 0.2 s. "Save" runs from where RSS bottoms out after the last engine build to the moment the.pteis written. The script doesn't compare against eager; the second script's partitioned exports match eager at cosine 0.999979 with and without this PR.Compatibility
.pteexported before and after..ptewith a named engine at load, with a clear error rather than undefined behavior:TR02magic so older runtimes refuse it up front. Reviewers: would you like a new magic here too, or is theengine_keyfield enough?Tests
test_backend.pypreprocess tests updated for the named-data layout; newtest_engine_info_accessors.py::test_rewrite_stages_the_engine_once_as_byteschecks that the staged buffer and thebytespreprocessreads share memory.ReadsTheEngineKeyOfAnEngineStoredAsNamedDataintest_executorch_blob_header.cpp; 48/48 pass.examples/executorch_reference_runner. Both produce identical outputs that match eager within fp16 tolerance; the.ptesizes are equal (32.1 MiB).tests/py/dynamo/executorchandtests/py/dynamo/runtime/test_compile_memory.py: the same tests pass and fail with and without this PR. The failures are environmental in my setup (an older ExecuTorch Python than the pin):test_zero_copy_kv.py(missinget_copy._h2d_copy/propagate_device_config), two intest_cuda_partitioner_composition.py, andtest_compile_memory.py::test_setup_does_not_copy_the_plan.test_load_compatibility.pywas not run, since it imports the editable install alongside the branch under test.pre-commit:typosflags two existing lines this PR doesn't change (backend.py"mis-bind",test_backend.py"mis-ordered").Stack
6 of 6. Base: #4796.