Skip to content

perf(training): training.log_interval 配置项 + rsl_rl PPO TensorBoard 批量写 patch - #1648

Merged
TATP-233 merged 2 commits into
mainfrom
perf/tb-batched-scalar-logging
Sep 26, 2026
Merged

TATP-233 merged 2 commits into
mainfrom
perf/tb-batched-scalar-logging

Conversation

@TATP-233

@TATP-233 TATP-233 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

概要

UniLab 侧配套改动,落地 #1646 的优化方案:training.log_interval 配置项 + PPO(rsl_rl)TensorBoard 批量写 patch。

uni_rl 侧的批量写/降频实现在 unilabsim/unilab_rl#44,并由 unilab-rl==1.4.0 发布。

改动

  • conf(sac / flashsac / warp_sac / appo / ppo 五个 owner config.yaml):新增 training.log_interval: 1。默认 1、行为不变;off-policy builder 在 unilab-rl 侧直接读取该键。
  • src/unilab/scripts/train_appo.py:透传 log_interval 给 APPORunner。
  • src/unilab/training/experiment.py:新增 patch_rsl_rl_tensorboard_logging——包装 rsl_rl logger 的 SummaryWriter,把每 iteration 约 15–30 次 add_scalar 合并为单条 event record,并按 log_interval 节流后端写;控制台输出与 episode 簿记不受影响,最后一个 iteration 始终记录。
  • src/unilab/scripts/train_rsl_rl.py:在既有 patch_rsl_rl_action_std_logging 旁挂接。
  • tests/config/test_config_system.py:断言五个算法 owner config 均暴露有效 training.log_interval。

依赖与合入顺序

验证

Rebase 后最终本地 head(4bed1f18):

uv run --no-sync ruff format --check .   # 383 files already formatted
uv run --no-sync ruff check .            # All checks passed
uv lock --check                          # resolved 240 packages
uv run --no-sync pytest -q tests/scripts/test_torch_cuda_source.py  # 2 passed(#1647 契约)
uv run --no-sync pytest -q -m slow tests/config/test_config_system.py -k 'algo_config_composes'  # 5 passed
uv run --no-sync pytest -q tests/utils/test_experiment_tracking.py  # 23 passed
uv run --no-sync mypy src/unilab          # Success: 134 source files
uv run --no-sync pyright                  # 0 errors, 1 warning
uv run --no-sync python scripts/benchmark/smoke_test.py  # module-mode 35/35;script-mode 36/36
uv run --no-sync pytest -q -m "not slow"  # 1549 passed, 27 skipped, 1 failed
  • 远端 CI 已在 rebase 后重新触发并全部通过(7 pass / 1 skip:ruff lint/format、mypy、pyright、test、benchmark-smoke、Build Sphinx;gh-pages deploy 按规则 skip)。
  • pyright 的唯一 warning 是可选依赖 drake_uni.runtime import 无法解析,非本 PR 引入。
  • 非 slow 全量测试中的 tests/scripts/test_check_docs.py::test_documentation_files_match_current_repo_contracts 失败,已在干净 origin/main@d72f7f07 复现,属 main 既有问题;本 PR 不顺手修改以保持范围。
  • 原优化验证(RTX 5090,g1_walk_flat/mjwarp,fast_sac 5000 iterations,日志目录在 GFS/FUSE 上):修复前 cycle 墙钟约 206 ms,其中约 165 ms 被 TensorBoard 写盘阻塞;修复后墙钟与埋点时间一致,collector/wait_for_learner_action 从约 156 ms 降到 33 ms。

Closes #1646

@TATP-233
TATP-233 requested a review from caozx1110 as a code owner September 26, 2026 13:28
- pyproject: collapse the platform-split torch pins into torch>=2.9,<2.15
- uv.sources: route linux/win torch to the cu130 index (cu128 tops out at
  torch 2.11); drop the pytorch-cu128 index
- uv.lock: torch 2.8.0/2.8.0+cu128/2.9.0+cu130 -> 2.14.0/2.14.0+cu130,
  triton 3.8.0, nvidia deps cu12 -> cu13
- rocm: torch==2.14.0 + triton-rocm==3.8.0 (torch 2.14 pairs with
  triton~=3.8); relax the setuptools<70 typo-era pin which conflicts with
  torch 2.14+rocm7.2 (requires setuptools>=77); regenerate uv.rocm.lock
- update the torch CUDA source contract tests and cu128 doc references

Measured on Apple Silicon (M5 Max, g1_walk_flat, mujoco): FlashSAC
end-to-end 16.1k -> 21.7k steps/s (+35%), learner 213ms -> 147ms/iter;
FastSAC unaffected (GEMM-bound).
…ging for rsl_rl PPO

- conf (sac/flashsac/appo/ppo): new training.log_interval key (default 1,
  no behavior change); off-policy builders in unilab-rl read it directly,
  train_appo forwards it to APPORunner.
- experiment.patch_rsl_rl_tensorboard_logging: wrap the rsl_rl logger's
  SummaryWriter so each iteration's ~15-30 add_scalar calls become a
  single event record, and gate backend writes to every log_interval
  iterations (console output unaffected, final iteration always logged);
  wired up in train_rsl_rl alongside the existing rsl_rl patches.

On network filesystems (FUSE/GFS) per-record writes in TensorBoard's
writer thread saturate the async queue and block the training loop
(~160 ms/iteration for fast_sac g1_walk_flat/mjwarp on RTX 5090).

Requires unilab-rl with log_interval support (unilabsim/unilab_rl#44);
bump the unilab-rl pin when that release lands.

Closes #1646
@TATP-233
TATP-233 force-pushed the perf/tb-batched-scalar-logging branch from 162dd45 to 4bed1f1 Compare September 26, 2026 18:00
@TATP-233
TATP-233 merged commit 8ff9341 into main Sep 26, 2026
8 checks passed
@TATP-233
TATP-233 deleted the perf/tb-batched-scalar-logging branch September 26, 2026 18:10
@TATP-233 TATP-233 mentioned this pull request Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

TensorBoard 标量逐条写盘在网络文件系统(FUSE)上阻塞训练主线程:offpolicy/appo 每 iteration ~50 条 record,GPU 利用率降至 0-20%

1 participant