Skip to content

perf(executorch): store each engine once, as named data, when saving a .pte - #4821

Draft
cehongwang wants to merge 1 commit into
narendasan/compile-mem-plan-passthroughfrom
cehongw/executorch-engine-named-data
Draft

cehongwang wants to merge 1 commit into
narendasan/compile-mem-plan-passthroughfrom
cehongw/executorch-engine-named-data

Conversation

@cehongwang

@cehongwang cehongwang commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Saving a .pte kept two host copies of every engine until the file was written: the serialized plan, registered on the exported program as the engine's buffer, and the delegate blob, which preprocess built by copying the plan in after the header. With the resource partitioner the save held two copies of all engines at once and became the export's peak.
  • TensorRTBackend.preprocess now hands the engine to ExecuTorch as named data (NamedDataStore, key tensorrt_engine_<sha256>, 16-byte aligned). The blob carries only the header and metadata, plus an engine_key naming the entry. Identical engines in different methods share one entry.
  • ExecuTorch stores named data only as bytes, so _resolve_engine_tensor converts TensorRT's serialized plan to bytes once, while TensorRT's buffer is the only other copy, and preprocess reuses that same object (engine_bytes_tensor / _trt_engine_bytes).
  • TensorRTBackend::init reads the engine from the program's NamedDataMap when the header names one, passes it to acquire_shared_engine (feat(executorch): share TensorRT engines across loads #4778) or load_engine as before, and frees the host copy once the engine is on the GPU, before building the execution context. Inline blobs load as before. TensorRTBlobHeader::parse rejects a blob that has both an inline engine and a key, or neither.

An earlier version of this change, written against main before #4795, kept a registry from engine to module so the save could read the module's retained serialized_engine. #4795 removes that retained plan, so this version serializes from the C++ engine (serialized_engine_tensor()) and needs no registry.

Memory saved

On the pi0.5 ExecuTorch + TensorRT export, this PR lowers the peak host memory from 17.8 GiB to 11.6 GiB (−6.2 GiB) when the model is partitioned (13 engines). For a single engine the peak is set by the engine build and doesn't change, but the save holds 4.9 GiB less (13.5 → 8.6 GiB).

Measured with pi05_executorch_tensorrt.py, the lerobot pi0.5 export script (bf16, A40). Peak host RSS of the exporting process tree, GiB:

Compile peak Save peak Held when written Export peak .pte
13 engines (cpu_memory_budget=4 GiB), base 10.7 17.8 17.8 17.8 7.44 GB
13 engines (cpu_memory_budget=4 GiB), this PR 11.2 11.6 11.2 11.6 7.44 GB
1 engine, base 28.2 13.6 13.5 28.2 5.27 GB
1 engine, this PR 28.2 13.3 8.6 28.2 5.27 GB

With partitioning, the save held two copies of every engine and was the export's peak; now it holds one and stays level with compile. With a single engine the save peak barely moves: it serializes the engine from the GPU (one copy), then converts it to bytes (a second), so both exist for a moment; afterwards only one is held instead of two. With partitioning each engine is converted in turn, so that moment is small. Removing it entirely needs the engine streamed to the file (for example through ExecuTorch's FileBackedData), which I'd leave to a follow-up.

A second export script (lerobot's own PI05Policy, bf16 vision kept in fp32, 9 engines at cpu_memory_budget=8 GiB) shows the same: save peak 22.1 → 14.0 GiB, which is its export peak.

How these numbers were measured

Same script, checkpoint, environment and command for both rows of each pair; only PYTHONPATH differs (this branch vs. its base, narendasan/compile-mem-plan-passthrough). TORCHTRT_ENABLE_BUILDER_MALLOC_TRIM=1, TensorRT 11.0, torch 2.13. The partitioned runs also carry #4783 (keep getitem with its multi-output producer when splitting), without which pi0.5 fails to partition, and call release_host_and_device_memory() before compile. My ExecuTorch Python predates the device-resident I/O config the script uses, so the wrapper drops that config (I/O stays on the host; the tensors are KB-sized) and skips the script's post-save check of it; compile and save are otherwise unchanged. A sidecar samples the process tree's RSS (psutil) every 0.2 s. "Save" runs from where RSS bottoms out after the last engine build to the moment the .pte is written. The script doesn't compare against eager; the second script's partitioned exports match eager at cosine 0.999979 with and without this PR.

Compatibility

  • A runtime with this change loads both formats: tested with the reference runner on a .pte exported before and after.
  • A runtime without it rejects a .pte with a named engine at load, with a clear error rather than undefined behavior:
    TensorRT: IRuntime::deserializeCudaEngine: Error Code 3: API Usage Error (Cannot deserialize with an empty memory buffer...)
    TensorRTBackend::init: failed to deserialize TensorRT engine
    
    Aliased I/O introduced the TR02 magic so older runtimes refuse it up front. Reviewers: would you like a new magic here too, or is the engine_key field enough?

Tests

  • Python: test_backend.py preprocess tests updated for the named-data layout; new test_engine_info_accessors.py::test_rewrite_stages_the_engine_once_as_bytes checks that the staged buffer and the bytes preprocess reads share memory.
  • C++: new ReadsTheEngineKeyOfAnEngineStoredAsNamedData in test_executorch_blob_header.cpp; 48/48 pass.
  • End to end: a small fp16 MLP exported with this branch and with its base, loaded and run through examples/executorch_reference_runner. Both produce identical outputs that match eager within fp16 tolerance; the .pte sizes are equal (32.1 MiB).
  • tests/py/dynamo/executorch and tests/py/dynamo/runtime/test_compile_memory.py: the same tests pass and fail with and without this PR. The failures are environmental in my setup (an older ExecuTorch Python than the pin): test_zero_copy_kv.py (missing et_copy._h2d_copy / propagate_device_config), two in test_cuda_partitioner_composition.py, and test_compile_memory.py::test_setup_does_not_copy_the_plan. test_load_compatibility.py was not run, since it imports the editable install alongside the branch under test.
  • pre-commit: typos flags two existing lines this PR doesn't change (backend.py "mis-bind", test_backend.py "mis-ordered").

Stack

6 of 6. Base: #4796.

…a .pte

Saving a .pte kept two host copies of every engine until the file was
written: the serialized plan, registered on the exported program as the
engine's buffer, and the delegate blob, which preprocess built by copying
the plan in after the header. With the resource partitioner the save then
held two copies of all engines at once and became the export's peak.

preprocess now hands the engine to ExecuTorch as named data, keyed by its
SHA-256, and the blob carries only the header and metadata, with an
engine_key naming the entry. ExecuTorch stores named data only as bytes, so
_resolve_engine_tensor converts TensorRT's serialized plan to bytes once,
while TensorRT's buffer is the only other copy, and preprocess reuses that
same object. Identical engines in different methods share one entry.

At load, TensorRTBackend::init reads the engine from the program's named
data map when the header names one, passes it to the shared-engine cache or
load_engine as before, and frees the host copy once the engine is on the
GPU. Inline blobs load as before. The header parser rejects a blob that has
both or neither.

On pi0.5 (bf16, 9 engines at cpu_memory_budget=8 GiB, on top of the
getitem fix in the resource partitioner), the save's peak host RSS drops
from 22.1 to 14.0 GiB, which is the export's peak.
@meta-cla meta-cla Bot added the cla signed label Oct 8, 2026
@github-actions github-actions Bot added component: tests Issues re: Tests component: api [Python] Issues re: Python API component: api [C++] Issues re: C++ API labels Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cla signed component: api [C++] Issues re: C++ API component: api [Python] Issues re: Python API component: tests Issues re: Tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant