Skip to content

[c10d][xccl2] Port the XCCL engine implementation - #9

Open
frost-intel wants to merge 1 commit into
xccl2/04-engine-declfrom
xccl2/05-engine-impl
Open

frost-intel wants to merge 1 commit into
xccl2/04-engine-declfrom
xccl2/05-engine-impl

Conversation

@frost-intel

@frost-intel frost-intel commented Aug 25, 2026 •

Copy link
Copy Markdown
Owner

Stack (bottom → top)

PR Branch
#5 xccl2/01-foundation
#6 xccl2/02-api
#7 xccl2/03-work
#8 xccl2/04-engine-decl
→ #9 xccl2/05-engine-impl
#10 xccl2/06-backend-surface
#11 xccl2/07-build-wiring
#12 xccl2/08-tests
#3 xccl2/09-integration-tests

Each PR is exactly one commit and targets the one below it. Review bottom-up.


Adds ProcessGroupXCCL.cpp, the port of torchcomms' TorchCommXCCL.cpp. This is
the backend-internal engine: the communicator lifecycle (init / finalize /
abort), the XCCLException definition, and the *Impl collective methods that the
c10d virtual overrides will call in the next commit.

The oneCCL-specific workarounds are carried over verbatim, since they are the
substance of what makes this an XCCL backend rather than a retargeted NCCL one:

Three mechanical substitutions were needed because XpuApi -- torchcomms'
driver-mock indirection layer -- is not part of this port:

  • memcpyAsync for the local (i == rank) chunk of all_to_all, all_to_all_v,
    scatter and gather becomes a stream-ordered at::Tensor::copy_ under a
    c10::StreamGuard. all_to_all_v_single slices along dim 0 to build the views.
  • stream/event creation becomes c10::xpu::getStreamFromPool and at::xpu::
    XPUEvent, and the barrier scratch buffer becomes an at::DataPtr from the XPU
    caching allocator.
  • The device-property and free-memory probes in init() are dropped; they only
    existed to be intercepted by the mock.

One behavioural narrowing, following nccl2's precedent: tensors must live on the
device this backend is bound to. torchcomms accepted CPU tensors as well
(checkAllTensorsOnXPUorCPU) so its mock tests could run without XPU hardware,
which also required the MAYBE_STREAM_GUARD dance. c10d dispatches XPU tensors
here, so checkTensorDevice/checkTensorsDevice/checkSameDtype are used instead
and the stream guards are unconditional. checkTensorDevice and
checkTensorsDevice are defined here rather than in ProcessGroupXCCLUtils.cpp
(where nccl2 keeps them) simply because this is the commit that introduces
their callers.

abortXcclComm() cannot actually abort: onecclCommAbort is declared
CCL_C_NOT_IMPLEMENTED, so XcclApi::commAbort returns onecclNotImplemented and
the communicator is dropped without teardown. The destructor therefore falls
back to onecclCommDestroy. Both are torchcomms' behaviour and both are
documented at the call sites.

Still to come so that each commit stays link-clean: the constructor,
ensureInitialized(), operationTimeout(), runAbortHooks() and the c10d virtual
overrides land with ProcessGroupXCCLBackend.cpp.

all_gather handles the uneven case, which is torchcomms' separate
all_gather_v. c10d has one all_gather entry point, so an uneven gather is just
an output list whose per-rank sizes differ; only this rank's slot has to match
what it contributes. oneCCL has no all_gatherv, so that path is issued as one
grouped broadcast per rank, mirroring how reduce_scatter already covers
reduce_scatter_v. The even path still stages through a flattened buffer.

split() carves a sub-communicator out of this one with onecclCommSplit, and
initFromComm() adopts the result instead of bootstrapping through the Store.
This follows nccl2's ProcessGroupNCCL::split closely, including the rule that
an empty global_ranks_in_group -- or one shorter than the group, as a merged
group inherits -- means "spans the world in rank order" and so is not indexed.
onecclCommSplit is collective over the parent, so the parent is initialized
first if a collective has not already done it; ranks outside the new group
still call it, with ONECCL_SPLIT_NOCOLOR, and get back no child backend.

The caching-allocator hook is attached once XCCL resources exist and detached
in finalize(), abortXcclComm() and the destructor. Detaching first in each of
those is deliberate: the hook holds a raw pointer to this backend and fires
from allocator threads, so it has to be dropped while the communicator is
still valid and before anything else is torn down.

Adds ProcessGroupXCCL.cpp, the port of torchcomms' TorchCommXCCL.cpp. This is
the backend-internal engine: the communicator lifecycle (init / finalize /
abort), the XCCLException definition, and the *Impl collective methods that the
c10d virtual overrides will call in the next commit.

The oneCCL-specific workarounds are carried over verbatim, since they are the
substance of what makes this an XCCL backend rather than a retargeted NCCL one:

  - PREMUL_SUM and AVG are lowered to SUM. oneCCL implements neither natively
    (uxlfoundation/oneCCL#196 and pytorch#195), so PREMUL_SUM pre-scales the input and
    AVG divides the result by comm_size afterwards. all_reduce scales the
    caller's buffer in place because the collective is already in place;
    reduce and the reduce_scatter variants scale a clone.
  - batch_op_issue issues every send before any recv inside the group. oneCCL
    replays grouped P2P in enqueue order and a recv blocks in a host-side IPC
    handle exchange until its peer's send is reached, so recv-first on both
    peers deadlocks.
  - scatter and gather group the sends *and* the receives, unlike nccl2 which
    groups only the sends, to avoid a hang (uxlfoundation/oneCCL#193).
  - Every collective keeps its zero-element early-out; oneCCL does not accept
    zero-sized buffers on these paths.
  - all_gather stages through a flattened buffer rather than nccl2's N-way
    broadcast, matching torchcomms.

Three mechanical substitutions were needed because XpuApi -- torchcomms'
driver-mock indirection layer -- is not part of this port:

  - memcpyAsync for the local (i == rank) chunk of all_to_all, all_to_all_v,
    scatter and gather becomes a stream-ordered at::Tensor::copy_ under a
    c10::StreamGuard. all_to_all_v_single slices along dim 0 to build the views.
  - stream/event creation becomes c10::xpu::getStreamFromPool and at::xpu::
    XPUEvent, and the barrier scratch buffer becomes an at::DataPtr from the XPU
    caching allocator.
  - The device-property and free-memory probes in init() are dropped; they only
    existed to be intercepted by the mock.

One behavioural narrowing, following nccl2's precedent: tensors must live on the
device this backend is bound to. torchcomms accepted CPU tensors as well
(checkAllTensorsOnXPUorCPU) so its mock tests could run without XPU hardware,
which also required the MAYBE_STREAM_GUARD dance. c10d dispatches XPU tensors
here, so checkTensorDevice/checkTensorsDevice/checkSameDtype are used instead
and the stream guards are unconditional. checkTensorDevice and
checkTensorsDevice are defined here rather than in ProcessGroupXCCLUtils.cpp
(where nccl2 keeps them) simply because this is the commit that introduces
their callers.

abortXcclComm() cannot actually abort: onecclCommAbort is declared
CCL_C_NOT_IMPLEMENTED, so XcclApi::commAbort returns onecclNotImplemented and
the communicator is dropped without teardown. The destructor therefore falls
back to onecclCommDestroy. Both are torchcomms' behaviour and both are
documented at the call sites.

Still to come so that each commit stays link-clean: the constructor,
ensureInitialized(), operationTimeout(), runAbortHooks() and the c10d virtual
overrides land with ProcessGroupXCCLBackend.cpp.

all_gather handles the uneven case, which is torchcomms' separate
all_gather_v. c10d has one all_gather entry point, so an uneven gather is just
an output list whose per-rank sizes differ; only this rank's slot has to match
what it contributes. oneCCL has no all_gatherv, so that path is issued as one
grouped broadcast per rank, mirroring how reduce_scatter already covers
reduce_scatter_v. The even path still stages through a flattened buffer.

split() carves a sub-communicator out of this one with onecclCommSplit, and
initFromComm() adopts the result instead of bootstrapping through the Store.
This follows nccl2's ProcessGroupNCCL::split closely, including the rule that
an empty global_ranks_in_group -- or one shorter than the group, as a merged
group inherits -- means "spans the world in rank order" and so is not indexed.
onecclCommSplit is collective over the parent, so the parent is initialized
first if a collective has not already done it; ranks outside the new group
still call it, with ONECCL_SPLIT_NOCOLOR, and get back no child backend.

The caching-allocator hook is attached once XCCL resources exist and detached
in finalize(), abortXcclComm() and the destructor. Detaching first in each of
those is deliberate: the hook holds a raw pointer to this backend and fires
from allocator threads, so it has to be dropped while the communicator is
still valid and before anything else is torn down.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant