You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
perf: end-to-end MJWarp training throughput gate and tensor-native closeout
Objective
Close the remaining #1811 runtime work using end-to-end training throughput as the decisive performance gate. The canonical production behavior is the real uv run train path—not a collector-only microbenchmark and not a final single iteration.
sum(env steps represented by the final 20 completed training iterations)
/
sum(wall time of those iterations)
Rules:
It is an aggregate-rate over one bounded window, not the mean of per-iteration rates.
The window is 20 iterations: large enough to suppress single-iteration jitter, small enough not to extend development loops.
Startup, backend materialization, Torch compilation, graph capture, replay fill, and first-update warm-up are excluded.
Existing final_env_steps_per_sec is diagnostic only and is not acceptance evidence.
Collector scripts under scripts/benchmark are diagnostic attribution tools only. They must never replace the end-to-end train gate.
Current local feasibility evidence
Local RTX 4090, 160 iterations, 2048 envs, no playback/ONNX, CUDA_VISIBLE_DEVICES unset:
Workload
tail steps/s
final-iter steps/s
FlashSAC / g1_motion_tracking / MJWarp
104,652
104,797
SAC / g1_walk_flat / MJWarp
96,383
82,203
This proves the metric is feasible and demonstrates why the final iteration is unsuitable: SAC's final iteration reports 82k while its bounded tail is 96k.
Artifacts: /tmp/flashsac-tail-contract/run_summary.json and /tmp/sac-tail-contract/run_summary.json.
Acceptance targets
Local workstation
Local results establish feasibility and guide optimization.
FlashSAC motion already exceeds 100k steps/s.
SAC walk must reach ≥100,000 tail steps/s.
592 / RTX 5090 server
Both canonical workloads must reach:
tail_env_steps_per_sec >= 100,000
A performance claim requires repeated interleaved end-to-end training runs under the same contract. A single run is not acceptance evidence.
Backend and architecture scope
Active runtime scope:
mjwarp
mujoco
Shelved for this effort:
Genesis
all other adapters
Non-negotiable runtime invariants
The following must remain true after every change:
Manager-Based API is the sole training entry point.
No direct environment mode.
No task-owned direct runtime registration.
No rollback to a generic NumPy Manager.
No rollback to a NumPy RNG owner.
The Manager-owned Torch RNG remains authoritative.
No re-introduction of env.tensor_runtime / env.tensor_runtime_device.
Placement is derived from actual public env carrier placement and backend capabilities.
No backend-private probing from task/env/training code.
Backend behavior must remain behind declared SimBackend public capabilities/APIs.
No compatibility layer or legacy NumPy owner path for MJWarp.
Tensor-native is the only MJWarp execution path.
NumPy is permitted only at the explicit MuJoCo HOST_BRIDGE D2H/H2D boundary.
MJWarp owners must not use NumPy state carriers, NumPy reset composers, or host RNG reset paths.
Required tensor-native migration
A. g1_walk_flat host-composer removal
This is the primary architectural blocker and the nearest local performance gap.
Remove the NumPy host owner paths from the canonical MJWarp SAC owner:
UniformVelocityCommand / G1VelocityCommand
remove NumPy command mirrors such as vel_command_b, vel_command_w, heading_target, heading_error;
remove selected-row D2H and _host_uniform reset sampling;
execute resampling, dead-zone filtering, and update using Torch rows and Manager Torch RNG.
G1GaitPhase.reset
remove selected-row D2H and NumPy RNG sampling;
sample phases on the public device using Manager Torch RNG.
Reset events
replace canonical reset_root_state_uniform and reset_scene_to_default usage with their tensor event peers;
remove the origins_tensor.abs().max().item() hot-path synchronization by making origin applicability a cold-path owner fact or otherwise device-resident branch.
G1PenaltyCurriculum / EpisodeLengthTracker
keep episode-length statistics and scaling device-resident during steady training;
convert to host scalar only at the unavoidable reward-weight mutation boundary, and avoid doing so on every reset.
G1 walk sensor/reward owners
remove capability fallbacks that silently select a NumPy host reader on MJWarp;
fail closed if a required sensor is absent from the public tensor read plan.
B. MJWarp motion owner residual cleanup
The FlashSAC motion owner is already largely tensor-native, but closeout must verify and, where needed, remove:
legacy MotionCommand base-class NumPy methods still reachable from an MJWarp owner;
capability fallbacks that select host readers on a DEVICE_RESIDENT backend;
remaining selected-reset NumPy row mirrors;
reward/log host publication outside the interval-controlled collector contract.
C. Legacy path removal
No dormant MJWarp NumPy fallback should remain:
remove unreachable NumPy execution branches for MJWarp owners rather than preserving them as compatibility;
do not add an owner-level tensor/NumPy execution switch;
do not preserve “legacy behavior” in config comments or runtime dispatch.
MuJoCo HOST_BRIDGE remains a production backend, so its public NumPy boundary is not a fallback and is not deleted.
Benchmark governance
scripts/benchmark must be updated to remain useful but subordinate to the end-to-end gate:
mirror the canonical two-workload set;
exclude warmup/compile;
report a bounded tail aggregate rate;
preserve detailed phase attribution for optimization decisions;
clearly label results as diagnostic, not acceptance evidence.
Any benchmark change must not lengthen the canonical loop merely to produce a more impressive number.
Required validations
Each child PR must run the smallest relevant checks, and the final head must run the full gate.
Direct changes to local sibling repositories are authorized:
unisim
unilab_rl
Sibling changes require their own reviewable PR, tests, and linked evidence; they must not become hidden dependencies of a UniLab PR.
Maintain ADR-0012 and the support matrix as scope and behavior change.
Do not close this issue merely because local throughput improves. Server repeated end-to-end evidence and the tensor-native architecture boundary are both mandatory.
Success criteria
MJWarp tensor-native is the sole execution path for both canonical tasks.
No canonical MJWarp owner uses a NumPy host composer, host RNG reset path, or hidden NumPy state mirror.
Local SAC / g1_walk_flat / MJWarp reaches ≥100,000 tail steps/s.
592 / RTX 5090 repeated runs show both canonical tasks at ≥100,000 tail steps/s.
Diagnostic benchmarks mirror the end-to-end contract and clearly remain non-acceptance tools.
make check, make test, and make test-all pass on the final head.
Both canonical 160-iteration training runs complete with normal shutdown and no residual collector process.
perf: end-to-end MJWarp training throughput gate and tensor-native closeout
Objective
Close the remaining #1811 runtime work using end-to-end training throughput as the decisive performance gate. The canonical production behavior is the real
uv run trainpath—not a collector-only microbenchmark and not a final single iteration.Canonical workloads
Both workloads are mandatory and run on MJWarp:
g1_motion_trackingg1_walk_flatMeasurement contract
The acceptance metric is the new runner-owned:
from unilab-rl #74.
Definition:
Rules:
final_env_steps_per_secis diagnostic only and is not acceptance evidence.scripts/benchmarkare diagnostic attribution tools only. They must never replace the end-to-end train gate.Current local feasibility evidence
Local RTX 4090, 160 iterations, 2048 envs, no playback/ONNX,
CUDA_VISIBLE_DEVICESunset:g1_motion_tracking/ MJWarpg1_walk_flat/ MJWarpThis proves the metric is feasible and demonstrates why the final iteration is unsuitable: SAC's final iteration reports 82k while its bounded tail is 96k.
Artifacts:
/tmp/flashsac-tail-contract/run_summary.jsonand/tmp/sac-tail-contract/run_summary.json.Acceptance targets
Local workstation
592 / RTX 5090 server
Both canonical workloads must reach:
A performance claim requires repeated interleaved end-to-end training runs under the same contract. A single run is not acceptance evidence.
Backend and architecture scope
Active runtime scope:
mjwarpmujocoShelved for this effort:
Non-negotiable runtime invariants
The following must remain true after every change:
env.tensor_runtime/env.tensor_runtime_device.SimBackendpublic capabilities/APIs.Required tensor-native migration
A.
g1_walk_flathost-composer removalThis is the primary architectural blocker and the nearest local performance gap.
Remove the NumPy host owner paths from the canonical MJWarp SAC owner:
UniformVelocityCommand/G1VelocityCommandvel_command_b,vel_command_w,heading_target,heading_error;is_heading_env,is_standing_env,is_world_env,is_forward_envmasks;_host_uniformreset sampling;G1GaitPhase.resetreset_root_state_uniformandreset_scene_to_defaultusage with their tensor event peers;origins_tensor.abs().max().item()hot-path synchronization by making origin applicability a cold-path owner fact or otherwise device-resident branch.G1PenaltyCurriculum/EpisodeLengthTrackerB. MJWarp motion owner residual cleanup
The FlashSAC motion owner is already largely tensor-native, but closeout must verify and, where needed, remove:
MotionCommandbase-class NumPy methods still reachable from an MJWarp owner;C. Legacy path removal
No dormant MJWarp NumPy fallback should remain:
MuJoCo HOST_BRIDGE remains a production backend, so its public NumPy boundary is not a fallback and is not deleted.
Benchmark governance
scripts/benchmarkmust be updated to remain useful but subordinate to the end-to-end gate:Any benchmark change must not lengthen the canonical loop merely to produce a more impressive number.
Required validations
Each child PR must run the smallest relevant checks, and the final head must run the full gate.
Minimum final gates:
Mandatory end-to-end runs:
Record for each run:
completed_iterationstotal_env_stepstail_env_steps_per_sectail_iteration_counttraining_wall_time_secGovernance
dev/issue-1811-tensor-manager.main.develop/tensor-runtime.unisimunilab_rlSuccess criteria
g1_walk_flat/ MJWarp reaches ≥100,000 tail steps/s.make check,make test, andmake test-allpass on the final head.Part of #1811 closeout.