Skip to content

[c10d][xccl2] Add backend-agnostic TorchComms helpers for xccl2 - #5

Open
frost-intel wants to merge 1 commit into
masterfrom
xccl2/01-foundation
Open

frost-intel wants to merge 1 commit into
masterfrom
xccl2/01-foundation

Conversation

@frost-intel

@frost-intel frost-intel commented Aug 25, 2026 •

Copy link
Copy Markdown
Owner

Stack (bottom → top)

PR Branch
→ #5 xccl2/01-foundation
#6 xccl2/02-api
#7 xccl2/03-work
#8 xccl2/04-engine-decl
#9 xccl2/05-engine-impl
#10 xccl2/06-backend-surface
#11 xccl2/07-build-wiring
#12 xccl2/08-tests
#3 xccl2/09-integration-tests

Each PR is exactly one commit and targets the one below it. Review bottom-up.


xccl2 needs three of the backend-agnostic helpers that nccl2 uses:
TracingGuard (profiler and flight-recorder records), Logging (TC_LOG) and
Batch (BatchSendRecv). All three were upstreamed from torchcomms'
comms/torchcomms/utils/ and contain no backend-specific code, but they
currently live under c10d/nccl2/ behind USE_C10D_NCCL and compile into
libtorch_cuda. An XPU-only build never compiles them, so xccl2 cannot link
against them.

Vendor a copy under c10d/xccl2/ in namespace c10d::xccl2. The files are
identical to the nccl2 originals apart from the namespace, the compile guard
(USE_C10D_XCCL) and the include paths. nccl2's Utils.{hpp,cpp} is deliberately
not copied: xccl2 takes its rank and size from the c10d Store rather than from
query_ranksize(), so none of its entry points are reachable here.

This duplicates the helpers rather than relocating the originals, so that
landing xccl2 requires no changes to nccl2. Consolidating the copies into a
shared directory is left to a follow-up that can be reviewed on its own merits,
once both backends are in tree.

No functional change to any existing backend.

xccl2 needs three of the backend-agnostic helpers that nccl2 uses:
TracingGuard (profiler and flight-recorder records), Logging (TC_LOG) and
Batch (BatchSendRecv). All three were upstreamed from torchcomms'
comms/torchcomms/utils/ and contain no backend-specific code, but they
currently live under c10d/nccl2/ behind USE_C10D_NCCL and compile into
libtorch_cuda. An XPU-only build never compiles them, so xccl2 cannot link
against them.

Vendor a copy under c10d/xccl2/ in namespace c10d::xccl2. The files are
identical to the nccl2 originals apart from the namespace, the compile guard
(USE_C10D_XCCL) and the include paths. nccl2's Utils.{hpp,cpp} is deliberately
not copied: xccl2 takes its rank and size from the c10d Store rather than from
query_ranksize(), so none of its entry points are reachable here.

This duplicates the helpers rather than relocating the originals, so that
landing xccl2 requires no changes to nccl2. Consolidating the copies into a
shared directory is left to a follow-up that can be reviewed on its own merits,
once both backends are in tree.

No functional change to any existing backend.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant