Problem
sim_backend=mujoco (mjbatch) declares an in-process HOST_BRIDGE tensor lifecycle, but the canonical SAC g1_walk_flat/mujoco owner did not declare env.tensor_runtime. The CUDA public plane was therefore reached only when a caller also passed training.inference_transport=cuda; without that extra override the composed owner fell back to the CPU TorchEnv plane. This is inconsistent with the mjwarp owner contract and leaves the mjbatch TorchEnv route dependent on transport selection rather than backend ownership.
Evidence
Current head, 2048 envs, local RTX 4090:
- Runner-injected CUDA transport produced the expected packed path and one packed read/reset per phase.
- Owner-only composition defaulted the same env to the CPU public plane unless transport was also overridden.
- Active collector benchmark after making the owner declaration explicit stayed at ~96.7k env/s, showing this is a contract cleanup, not a performance fallback.
Delivered behavior
- The SAC
g1_walk_flat/mujoco owner selects env.tensor_runtime: true with tensor_runtime_device: null.
- The runner/rank resolver remains responsible for filling the rank-local CUDA device.
- The owner documents mjbatch semantics: CPU physics stays authoritative, with one packed boundary per control/read/reset phase.
Audit status
No direct private backend-state access exists under src/unilab (rg '_backend\\._[A-Za-z]' returns no hot-path hits). The remaining backend interactions are public SimBackend APIs. mjbatch's public get_sensor_data/body-state calls inside MuJoCoHostBridgeTransferPlan are the CPU-side pack sources, not extra H2D transfers.
Measured 100 vector steps at 2048 envs:
- SAC/G1 walk: one full packed H2D per read plus one selected H2D on reset, and one packed control D2H; ~2/3 of steps also have one packed reset D2H.
- FlashSAC/G1 motion: exactly one packed control D2H and one full packed state H2D per step, plus one packed reset D2H and selected post-reset H2D on reset steps.
These are the declared HOST_BRIDGE boundaries, not scattered per-sensor transfers.
Acceptance criteria
Scope boundary
No new backend capability or transfer mechanism is introduced. GPU-resident backends must not copy this HOST_BRIDGE owner comment; their owner contracts remain DIRECT/DEVICE_RESIDENT.
Problem
sim_backend=mujoco(mjbatch) declares an in-processHOST_BRIDGEtensor lifecycle, but the canonical SACg1_walk_flat/mujocoowner did not declareenv.tensor_runtime. The CUDA public plane was therefore reached only when a caller also passedtraining.inference_transport=cuda; without that extra override the composed owner fell back to the CPU TorchEnv plane. This is inconsistent with the mjwarp owner contract and leaves the mjbatch TorchEnv route dependent on transport selection rather than backend ownership.Evidence
Current head, 2048 envs, local RTX 4090:
Delivered behavior
g1_walk_flat/mujocoowner selectsenv.tensor_runtime: truewithtensor_runtime_device: null.Audit status
No direct private backend-state access exists under
src/unilab(rg '_backend\\._[A-Za-z]'returns no hot-path hits). The remaining backend interactions are publicSimBackendAPIs. mjbatch's publicget_sensor_data/body-state calls insideMuJoCoHostBridgeTransferPlanare the CPU-side pack sources, not extra H2D transfers.Measured 100 vector steps at 2048 envs:
These are the declared HOST_BRIDGE boundaries, not scattered per-sensor transfers.
Acceptance criteria
g1_walk_flat/mujococomposes to the packed tensor runtime without a transport override.Scope boundary
No new backend capability or transfer mechanism is introduced. GPU-resident backends must not copy this HOST_BRIDGE owner comment; their owner contracts remain DIRECT/DEVICE_RESIDENT.