[c10d][xccl2] Port the XCCL engine implementation - #9
Open
frost-intel wants to merge 1 commit into
Open
frost-intel wants to merge 1 commit into
frost-intel wants to merge 1 commit into
Conversation
Adds ProcessGroupXCCL.cpp, the port of torchcomms' TorchCommXCCL.cpp. This is
the backend-internal engine: the communicator lifecycle (init / finalize /
abort), the XCCLException definition, and the *Impl collective methods that the
c10d virtual overrides will call in the next commit.
The oneCCL-specific workarounds are carried over verbatim, since they are the
substance of what makes this an XCCL backend rather than a retargeted NCCL one:
- PREMUL_SUM and AVG are lowered to SUM. oneCCL implements neither natively
(uxlfoundation/oneCCL#196 and pytorch#195), so PREMUL_SUM pre-scales the input and
AVG divides the result by comm_size afterwards. all_reduce scales the
caller's buffer in place because the collective is already in place;
reduce and the reduce_scatter variants scale a clone.
- batch_op_issue issues every send before any recv inside the group. oneCCL
replays grouped P2P in enqueue order and a recv blocks in a host-side IPC
handle exchange until its peer's send is reached, so recv-first on both
peers deadlocks.
- scatter and gather group the sends *and* the receives, unlike nccl2 which
groups only the sends, to avoid a hang (uxlfoundation/oneCCL#193).
- Every collective keeps its zero-element early-out; oneCCL does not accept
zero-sized buffers on these paths.
- all_gather stages through a flattened buffer rather than nccl2's N-way
broadcast, matching torchcomms.
Three mechanical substitutions were needed because XpuApi -- torchcomms'
driver-mock indirection layer -- is not part of this port:
- memcpyAsync for the local (i == rank) chunk of all_to_all, all_to_all_v,
scatter and gather becomes a stream-ordered at::Tensor::copy_ under a
c10::StreamGuard. all_to_all_v_single slices along dim 0 to build the views.
- stream/event creation becomes c10::xpu::getStreamFromPool and at::xpu::
XPUEvent, and the barrier scratch buffer becomes an at::DataPtr from the XPU
caching allocator.
- The device-property and free-memory probes in init() are dropped; they only
existed to be intercepted by the mock.
One behavioural narrowing, following nccl2's precedent: tensors must live on the
device this backend is bound to. torchcomms accepted CPU tensors as well
(checkAllTensorsOnXPUorCPU) so its mock tests could run without XPU hardware,
which also required the MAYBE_STREAM_GUARD dance. c10d dispatches XPU tensors
here, so checkTensorDevice/checkTensorsDevice/checkSameDtype are used instead
and the stream guards are unconditional. checkTensorDevice and
checkTensorsDevice are defined here rather than in ProcessGroupXCCLUtils.cpp
(where nccl2 keeps them) simply because this is the commit that introduces
their callers.
abortXcclComm() cannot actually abort: onecclCommAbort is declared
CCL_C_NOT_IMPLEMENTED, so XcclApi::commAbort returns onecclNotImplemented and
the communicator is dropped without teardown. The destructor therefore falls
back to onecclCommDestroy. Both are torchcomms' behaviour and both are
documented at the call sites.
Still to come so that each commit stays link-clean: the constructor,
ensureInitialized(), operationTimeout(), runAbortHooks() and the c10d virtual
overrides land with ProcessGroupXCCLBackend.cpp.
all_gather handles the uneven case, which is torchcomms' separate
all_gather_v. c10d has one all_gather entry point, so an uneven gather is just
an output list whose per-rank sizes differ; only this rank's slot has to match
what it contributes. oneCCL has no all_gatherv, so that path is issued as one
grouped broadcast per rank, mirroring how reduce_scatter already covers
reduce_scatter_v. The even path still stages through a flattened buffer.
split() carves a sub-communicator out of this one with onecclCommSplit, and
initFromComm() adopts the result instead of bootstrapping through the Store.
This follows nccl2's ProcessGroupNCCL::split closely, including the rule that
an empty global_ranks_in_group -- or one shorter than the group, as a merged
group inherits -- means "spans the world in rank order" and so is not indexed.
onecclCommSplit is collective over the parent, so the parent is initialized
first if a collective has not already done it; ranks outside the new group
still call it, with ONECCL_SPLIT_NOCOLOR, and get back no child backend.
The caching-allocator hook is attached once XCCL resources exist and detached
in finalize(), abortXcclComm() and the destructor. Detaching first in each of
those is deliberate: the hook holds a raw pointer to this backend and fires
from allocator threads, so it has to be dropped while the communicator is
still valid and before anything else is torn down.
This was referenced Aug 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack (bottom → top)
xccl2/01-foundationxccl2/02-apixccl2/03-workxccl2/04-engine-declxccl2/05-engine-implxccl2/06-backend-surfacexccl2/07-build-wiringxccl2/08-testsxccl2/09-integration-testsEach PR is exactly one commit and targets the one below it. Review bottom-up.
Adds ProcessGroupXCCL.cpp, the port of torchcomms' TorchCommXCCL.cpp. This is
the backend-internal engine: the communicator lifecycle (init / finalize /
abort), the XCCLException definition, and the *Impl collective methods that the
c10d virtual overrides will call in the next commit.
The oneCCL-specific workarounds are carried over verbatim, since they are the
substance of what makes this an XCCL backend rather than a retargeted NCCL one:
(Enable
PreMulSumUser-Defined Reduction Across All Reduce-Type Collectives and Scale-Up/Scale-Out Configurations uxlfoundation/oneCCL#196 and Look for libcudart in default CUDA installation paths pytorch/pytorch#195), so PREMUL_SUM pre-scales the input andAVG divides the result by comm_size afterwards. all_reduce scales the
caller's buffer in place because the collective is already in place;
reduce and the reduce_scatter variants scale a clone.
replays grouped P2P in enqueue order and a recv blocks in a host-side IPC
handle exchange until its peer's send is reached, so recv-first on both
peers deadlocks.
groups only the sends, to avoid a hang (Support group API with only SEND or RECV operations uxlfoundation/oneCCL#193).
zero-sized buffers on these paths.
broadcast, matching torchcomms.
Three mechanical substitutions were needed because XpuApi -- torchcomms'
driver-mock indirection layer -- is not part of this port:
scatter and gather becomes a stream-ordered at::Tensor::copy_ under a
c10::StreamGuard. all_to_all_v_single slices along dim 0 to build the views.
XPUEvent, and the barrier scratch buffer becomes an at::DataPtr from the XPU
caching allocator.
existed to be intercepted by the mock.
One behavioural narrowing, following nccl2's precedent: tensors must live on the
device this backend is bound to. torchcomms accepted CPU tensors as well
(checkAllTensorsOnXPUorCPU) so its mock tests could run without XPU hardware,
which also required the MAYBE_STREAM_GUARD dance. c10d dispatches XPU tensors
here, so checkTensorDevice/checkTensorsDevice/checkSameDtype are used instead
and the stream guards are unconditional. checkTensorDevice and
checkTensorsDevice are defined here rather than in ProcessGroupXCCLUtils.cpp
(where nccl2 keeps them) simply because this is the commit that introduces
their callers.
abortXcclComm() cannot actually abort: onecclCommAbort is declared
CCL_C_NOT_IMPLEMENTED, so XcclApi::commAbort returns onecclNotImplemented and
the communicator is dropped without teardown. The destructor therefore falls
back to onecclCommDestroy. Both are torchcomms' behaviour and both are
documented at the call sites.
Still to come so that each commit stays link-clean: the constructor,
ensureInitialized(), operationTimeout(), runAbortHooks() and the c10d virtual
overrides land with ProcessGroupXCCLBackend.cpp.
all_gather handles the uneven case, which is torchcomms' separate
all_gather_v. c10d has one all_gather entry point, so an uneven gather is just
an output list whose per-rank sizes differ; only this rank's slot has to match
what it contributes. oneCCL has no all_gatherv, so that path is issued as one
grouped broadcast per rank, mirroring how reduce_scatter already covers
reduce_scatter_v. The even path still stages through a flattened buffer.
split() carves a sub-communicator out of this one with onecclCommSplit, and
initFromComm() adopts the result instead of bootstrapping through the Store.
This follows nccl2's ProcessGroupNCCL::split closely, including the rule that
an empty global_ranks_in_group -- or one shorter than the group, as a merged
group inherits -- means "spans the world in rank order" and so is not indexed.
onecclCommSplit is collective over the parent, so the parent is initialized
first if a collective has not already done it; ranks outside the new group
still call it, with ONECCL_SPLIT_NOCOLOR, and get back no child backend.
The caching-allocator hook is attached once XCCL resources exist and detached
in finalize(), abortXcclComm() and the destructor. Detaching first in each of
those is deliberate: the hook holds a raw pointer to this backend and fires
from allocator threads, so it has to be dropped while the communicator is
still valid and before anything else is torn down.