Skip to content

perf: end-to-end MJWarp training throughput gate and tensor-native closeout #2023

Description

@TATP-233

perf: end-to-end MJWarp training throughput gate and tensor-native closeout

Objective

Close the remaining #1811 runtime work using end-to-end training throughput as the decisive performance gate. The canonical production behavior is the real uv run train path—not a collector-only microbenchmark and not a final single iteration.

Canonical workloads

Both workloads are mandatory and run on MJWarp:

  1. FlashSAC / g1_motion_tracking
    uv run train \
      --algo flashsac \
      --task g1_motion_tracking \
      --sim mjwarp
  2. SAC / g1_walk_flat
    uv run --extra mjwarp train \
      --algo sac \
      --task g1_walk_flat \
      --sim mjwarp

Measurement contract

The acceptance metric is the new runner-owned:

run_summary.json:tail_env_steps_per_sec

from unilab-rl #74.

Definition:

sum(env steps represented by the final 20 completed training iterations)
/
sum(wall time of those iterations)

Rules:

  • It is an aggregate-rate over one bounded window, not the mean of per-iteration rates.
  • The window is 20 iterations: large enough to suppress single-iteration jitter, small enough not to extend development loops.
  • Startup, backend materialization, Torch compilation, graph capture, replay fill, and first-update warm-up are excluded.
  • Existing final_env_steps_per_sec is diagnostic only and is not acceptance evidence.
  • Collector scripts under scripts/benchmark are diagnostic attribution tools only. They must never replace the end-to-end train gate.

Current local feasibility evidence

Local RTX 4090, 160 iterations, 2048 envs, no playback/ONNX, CUDA_VISIBLE_DEVICES unset:

Workload tail steps/s final-iter steps/s
FlashSAC / g1_motion_tracking / MJWarp 104,652 104,797
SAC / g1_walk_flat / MJWarp 96,383 82,203

This proves the metric is feasible and demonstrates why the final iteration is unsuitable: SAC's final iteration reports 82k while its bounded tail is 96k.

Artifacts: /tmp/flashsac-tail-contract/run_summary.json and /tmp/sac-tail-contract/run_summary.json.

Acceptance targets

Local workstation

  • Local results establish feasibility and guide optimization.
  • FlashSAC motion already exceeds 100k steps/s.
  • SAC walk must reach ≥100,000 tail steps/s.

592 / RTX 5090 server

Both canonical workloads must reach:

tail_env_steps_per_sec >= 100,000

A performance claim requires repeated interleaved end-to-end training runs under the same contract. A single run is not acceptance evidence.

Backend and architecture scope

Active runtime scope:

  • mjwarp
  • mujoco

Shelved for this effort:

  • Genesis
  • all other adapters

Non-negotiable runtime invariants

The following must remain true after every change:

  1. Manager-Based API is the sole training entry point.
    • No direct environment mode.
    • No task-owned direct runtime registration.
  2. No rollback to a generic NumPy Manager.
  3. No rollback to a NumPy RNG owner.
    • The Manager-owned Torch RNG remains authoritative.
  4. No re-introduction of env.tensor_runtime / env.tensor_runtime_device.
    • Placement is derived from actual public env carrier placement and backend capabilities.
  5. No backend-private probing from task/env/training code.
    • Backend behavior must remain behind declared SimBackend public capabilities/APIs.
  6. No compatibility layer or legacy NumPy owner path for MJWarp.
    • Tensor-native is the only MJWarp execution path.
  7. NumPy is permitted only at the explicit MuJoCo HOST_BRIDGE D2H/H2D boundary.
    • MJWarp owners must not use NumPy state carriers, NumPy reset composers, or host RNG reset paths.

Required tensor-native migration

A. g1_walk_flat host-composer removal

This is the primary architectural blocker and the nearest local performance gap.

Remove the NumPy host owner paths from the canonical MJWarp SAC owner:

  • UniformVelocityCommand / G1VelocityCommand
    • remove NumPy command mirrors such as vel_command_b, vel_command_w, heading_target, heading_error;
    • remove NumPy is_heading_env, is_standing_env, is_world_env, is_forward_env masks;
    • remove selected-row D2H and _host_uniform reset sampling;
    • execute resampling, dead-zone filtering, and update using Torch rows and Manager Torch RNG.
  • G1GaitPhase.reset
    • remove selected-row D2H and NumPy RNG sampling;
    • sample phases on the public device using Manager Torch RNG.
  • Reset events
    • replace canonical reset_root_state_uniform and reset_scene_to_default usage with their tensor event peers;
    • remove the origins_tensor.abs().max().item() hot-path synchronization by making origin applicability a cold-path owner fact or otherwise device-resident branch.
  • G1PenaltyCurriculum / EpisodeLengthTracker
    • keep episode-length statistics and scaling device-resident during steady training;
    • convert to host scalar only at the unavoidable reward-weight mutation boundary, and avoid doing so on every reset.
  • G1 walk sensor/reward owners
    • remove capability fallbacks that silently select a NumPy host reader on MJWarp;
    • fail closed if a required sensor is absent from the public tensor read plan.

B. MJWarp motion owner residual cleanup

The FlashSAC motion owner is already largely tensor-native, but closeout must verify and, where needed, remove:

  • legacy MotionCommand base-class NumPy methods still reachable from an MJWarp owner;
  • capability fallbacks that select host readers on a DEVICE_RESIDENT backend;
  • remaining selected-reset NumPy row mirrors;
  • reward/log host publication outside the interval-controlled collector contract.

C. Legacy path removal

No dormant MJWarp NumPy fallback should remain:

  • remove unreachable NumPy execution branches for MJWarp owners rather than preserving them as compatibility;
  • do not add an owner-level tensor/NumPy execution switch;
  • do not preserve “legacy behavior” in config comments or runtime dispatch.

MuJoCo HOST_BRIDGE remains a production backend, so its public NumPy boundary is not a fallback and is not deleted.

Benchmark governance

scripts/benchmark must be updated to remain useful but subordinate to the end-to-end gate:

  • mirror the canonical two-workload set;
  • exclude warmup/compile;
  • report a bounded tail aggregate rate;
  • preserve detailed phase attribution for optimization decisions;
  • clearly label results as diagnostic, not acceptance evidence.

Any benchmark change must not lengthen the canonical loop merely to produce a more impressive number.

Required validations

Each child PR must run the smallest relevant checks, and the final head must run the full gate.

Minimum final gates:

make check
make test
make test-all

Mandatory end-to-end runs:

uv run train \
  --algo flashsac \
  --task g1_motion_tracking \
  --sim mjwarp \
  algo.max_iterations=160 \
  algo.num_envs=2048 \
  training.no_play=true \
  training.export_onnx=false

uv run --extra mjwarp train \
  --algo sac \
  --task g1_walk_flat \
  --sim mjwarp \
  algo.max_iterations=160 \
  algo.num_envs=2048 \
  training.no_play=true \
  training.export_onnx=false

Record for each run:

  • UniLab commit
  • unisim commit
  • unilab-rl commit
  • GPU and CUDA visibility
  • completed_iterations
  • total_env_steps
  • tail_env_steps_per_sec
  • tail_iteration_count
  • training_wall_time_sec
  • run artifact path

Governance

  • Integration branch: dev/issue-1811-tensor-manager.
  • Do not modify main.
  • Final integration target: develop/tensor-runtime.
  • No PyPI release.
  • Direct changes to local sibling repositories are authorized:
    • unisim
    • unilab_rl
  • Sibling changes require their own reviewable PR, tests, and linked evidence; they must not become hidden dependencies of a UniLab PR.
  • Maintain ADR-0012 and the support matrix as scope and behavior change.
  • Do not close this issue merely because local throughput improves. Server repeated end-to-end evidence and the tensor-native architecture boundary are both mandatory.

Success criteria

  • MJWarp tensor-native is the sole execution path for both canonical tasks.
  • No canonical MJWarp owner uses a NumPy host composer, host RNG reset path, or hidden NumPy state mirror.
  • Local SAC / g1_walk_flat / MJWarp reaches ≥100,000 tail steps/s.
  • 592 / RTX 5090 repeated runs show both canonical tasks at ≥100,000 tail steps/s.
  • Diagnostic benchmarks mirror the end-to-end contract and clearly remain non-acceptance tools.
  • make check, make test, and make test-all pass on the final head.
  • Both canonical 160-iteration training runs complete with normal shutdown and no residual collector process.
  • Manager-only factory, scoped backends, tensor-switch removal, Torch RNG owner, and no-direct-mode invariants remain intact.
  • Documentation, ADR-0012, and generated support matrix reflect the final runtime.

Part of #1811 closeout.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions