fix(training): export ONNX at opset 18 and decouple play artifact failures - #1656
Merged
Merged
Conversation
…lures (#1654) torch 2.14's dynamo exporter emits the onnxscript function library (e.g. aten_isnan from torch.nan_to_num in the SAC actor) at opset 18; requesting opset 17 crashed the optimizer's InlinePass with an opset mismatch after version conversion. Bump the shared export default to 18. Also wrap ONNX export and video rendering in a shared nonfatal_play_step guard across the off-policy, APPO, and rsl-rl play paths so one artifact failure no longer blocks the other.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
export_policy_onnxdefault opset 17 → 18. With torch 2.14 the dynamo exporter emits the onnxscript function library (e.g.aten_isnanfromtorch.nan_to_numin the uni_rl SAC actor wrapper) at opset 18; requesting 17 made torch version-convert the model back while the function stayed at 18, and the optimizerInlinePasscrashed with an opset mismatch — killing the play phase right after training finished.nonfatal_play_step(unilab.training.run) and wrapped both post-training artifacts — ONNX/JIT policy export and video rendering — in it across all three play paths (train_offpolicy.py,train_appo.py,train_rsl_rl.py), so one artifact's failure no longer blocks the other or crashes the play command; failures print a warning plus traceback and the run continues.Linked Work
mainValidation
make test-allpassed on the final local head before this PR was created or updatedCommands actually run:
Remote CI route:
main: current-head CI all pass — https://github.com/Motphys/UniLab/actions/runs/36266050382 (ruff-format, ruff-lint, mypy, pyright, test (ubuntu-slim), benchmark-smoke) and https://github.com/Motphys/UniLab/actions/runs/36266050384 (Build Sphinx).Impact
mujoco)Artifacts
logs/fast_sac/G1WalkFlat/2026-09-27_03-24-29_mujoco/play_video.mp4(local)logs/fast_sac/G1WalkFlat/2026-09-27_03-24-29_mujoco/policy.onnx(opset 18, ORT-verified vs PyTorch, max_diff 8.94e-08)Checklist