Skip to content

feat(rl): add WarpSAC G1 owners for unilab-rl 1.4.0 - #1645

Merged
TATP-233 merged 5 commits into
mainfrom
feat/unilab-rl-1.4.0-g1-warpsac
Sep 26, 2026
Merged

TATP-233 merged 5 commits into
mainfrom
feat/unilab-rl-1.4.0-g1-warpsac

Conversation

@TATP-233

@TATP-233 TATP-233 commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Pinned standard and ROCm dependencies/locks to unilab-rl==1.4.0; no release or PyPI publication is included.
  • Added the WarpSAC entrypoint/config tree, structured config, runner dispatch, CLI/completion/docs/support-matrix integration, and 1.4.0 field cleanup.
  • Added and tuned WarpSAC owners for g1_walk_flat on MuJoCo/mjwarp and g1_motion_tracking on MuJoCo/mjwarp.
  • Fixed explicit post-training playback run selection so ONNX/video are generated from the final checkpoint rather than a stale algo.load_run selector.

Linked Work

Tuning and full-run evidence

Tuning decisions:

  • Walk: decay_step=2048, replay_min_weight=0.10, actor/critic parameter normalization disabled.
  • Motion: uniform replay (decay_step=0), replay_min_weight=0.05, actor/critic parameter normalization enabled.

All four final WarpSAC runs completed with final checkpoints, verified ONNX, and video:

Task Backend Env steps Final reward Best reward Episode length Training wall
Walk MuJoCo 20,680,704 296.136 298.545 993.12 167.77 s
Walk mjwarp 20,680,704 299.591 300.279 998.94 158.62 s
Motion MuJoCo 51,400,704 26.847 32.470 290.49 486.92 s
Motion mjwarp 51,400,704 28.915 32.868 278.36 582.10 s

Existing FastSAC/FlashSAC logs were analyzed without rerunning them. Full tables and comparability caveats are in logs/warpsac_comparison/REPORT.md; curves use cumulative environment steps and reward/mean_ep100:

  • logs/warpsac_comparison/g1_walk_mujoco_warpsac_comparison.png
  • logs/warpsac_comparison/g1_walk_mjwarp_warpsac_comparison.png
  • logs/warpsac_comparison/g1_motion_mujoco_warpsac_comparison.png
  • logs/warpsac_comparison/g1_motion_mjwarp_warpsac_comparison.png

Walk caveat: FastSAC uses 2,048 envs, 10.260M steps, and a different reward-owner profile. Its curve is only a common-prefix reference; FlashSAC/WarpSAC are directly comparable at 20.681M steps. Motion runs all use 2,048 envs and approximately 51.2–51.4M steps.

Final artifacts:

  • Walk MuJoCo: logs/warp_sac/G1WalkFlat/final_fba92b8f_mujoco/{policy.onnx,play_video.mp4}
  • Walk mjwarp: logs/warp_sac/G1WalkFlat/final_fba92b8f_mjwarp/{policy.onnx,play_video.mp4}
  • Motion MuJoCo: logs/warp_sac/G1MotionTrackingSAC/final_ddb2908b_mujoco/{policy.onnx,play_video.mp4}
  • Motion mjwarp: logs/warp_sac/G1MotionTrackingSAC/final_ddb2908b_mjwarp/{policy.onnx,play_video.mp4}

ONNX checker results: walk [1,98] -> [1,29]; motion [1,160] -> [1,29]. ffprobe verified every video as H.264, 1280x720, 800 frames at 50 fps (16 s).

Validation

  • make test-all passed on the final local head before this PR was created
  • Additional task-specific validation listed below

Commands and outcomes:

uv run pytest --override-ini='addopts=--tb=short' \
  tests/visualization/test_interactive_playback.py::test_sac_playback_session_runs_sim2sim_preflight \
  tests/config/test_locomotion_params.py::test_offpolicy_warpsac_g1_task_overrides
# 2 passed

make check
# ruff, mypy, pyright, and focused test lint checks passed

make test
# 1548 passed, 26 skipped, 604 deselected, 12 warnings

make test-all
# check passed; 1548 passed, 26 skipped, 604 deselected;
# benchmark module smoke 35/35 and script smoke 36/36 passed

uv run scripts/audit_sim2sim_contracts.py --trees warpsac
# 0 transferable, 0 blocked, 0 guard-blind-spot fields

uv lock --check
# standard lock resolved and valid

# ROCm files were temporarily swapped into pyproject.toml/uv.lock, checked, then restored:
uv lock --check
# ROCm lock resolved and valid

uv tree --package unilab-rl
# unilab-rl v1.4.0

The unrelated ignored scripts/benchmark/outputs/ tree was temporarily moved outside the repo for make test/make test-all and restored afterward; its docs-check contents are unrelated local benchmark output.

Experimental validation used local, pc823, and pc825 in parallel. Remote pc825 was updated to the final playback-fix head before regenerating its final motion ONNX/video from model_25000.pt.

Remote CI route:

Impact

  • Backend impact: MuJoCo and mjwarp training/playback paths; other backends unchanged.
  • Platform impact: Linux (training validation used CUDA); configuration/docs are cross-platform.
  • Training effect expected: yes, adds WarpSAC and tuned G1 owners.

Artifacts

  • W&B: none (TensorBoard only)
  • Benchmark result: logs/warpsac_comparison/summary_metrics.json
  • Video / screenshot: four final play_video.mp4 files listed above
  • ONNX / checkpoint: four final policy.onnx files and run checkpoints listed above

Checklist

  • Added or updated tests where needed
  • Updated docs if behavior or workflow changed
  • Linked the driving issue
  • Noted any follow-up work explicitly

@TATP-233
TATP-233 requested a review from caozx1110 as a code owner September 25, 2026 16:19
@TATP-233
TATP-233 merged commit d72f7f0 into main Sep 26, 2026
8 checks passed
@TATP-233
TATP-233 deleted the feat/unilab-rl-1.4.0-g1-warpsac branch September 26, 2026 08:31
@TATP-233 TATP-233 mentioned this pull request Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add validated WarpSAC G1 owners for unilab-rl 1.4.0

1 participant