Repository navigation
[c10d][xccl2] Port the ProcessGroupXCCL declaration and engine helpers - #8
Open
frost-intel wants to merge 1 commit into
Open
frost-intel wants to merge 1 commit into
frost-intel wants to merge 1 commit into
Conversation
Adds the ProcessGroupXCCL class declaration and the backend-internal helper
definitions, ported from torchcomms' TorchCommXCCL.hpp / TorchCommXCCLUtils.cpp
and adapted to c10d the same way nccl2 adapted TorchCommNCCL.
ProcessGroupXCCL.hpp declares `c10d::xccl2::ProcessGroupXCCL : c10d::Backend`.
The TorchComm / TorchCommBackend / BackendWrapper layers are collapsed away and
the collective surface takes the c10d option objects directly. The namespace is
what keeps this distinct from the existing c10d::ProcessGroupXCCL in
third_party/torch-xpu-ops, which drives the oneCCL C++ (ccl::) API rather than
the v2 C API this backend uses.
Unlike nccl2 -- which reuses ::c10d::ProcessGroupNCCL::Options because that type
already exists in core -- there is no core-side XCCL options type to reuse, so
the backend defines its own nested Options deriving from c10d::Backend::Options.
ProcessGroupXCCLUtils.cpp carries the RedOpRAII reduction-op wrapper, the
dtype/reduce-op mapping, the work factory/queue plumbing, the
operation-stream selection and the timeout watchdog. The one c10d-specific
change is the PREMUL_SUM factor: torchcomms carried it on its own ReduceOp
(op.factor()), while c10d hangs it off the ReduceOp supplement, so
getPreMulSumFactor() unpacks NCCLPreMulSumSupplement exactly as nccl2 does.
The event pool is not a member of the backend, unlike torchcomms and unlike
nccl2. It is an XcclEventPool held by shared_ptr and shared with every work
the backend creates, because ~WorkXCCL returns its events to the pool and a
work can outlive the backend; see the previous commit.
Behaviour is otherwise unchanged from torchcomms. In particular the watchdog
does not poll the communicator for async errors and does not abort the
communicator on timeout: oneCCL declares both onecclCommGetAsyncError and
onecclCommAbort with CCL_C_NOT_IMPLEMENTED, which expands to
__attribute__((error(...))), so referencing them does not compile. The
divergence from nccl2 here is deliberate and documented at the call sites.
Definitions deferred to later commits in this stack so that every commit stays
link-clean once the build is wired up: XCCLException, throwAsyncError(),
abortXcclComm(), runAbortHooks(), the init/finalize path and the *Impl engine
methods land with ProcessGroupXCCL.cpp; the c10d virtual overrides land with
ProcessGroupXCCLBackend.cpp.
This also lets WorkXCCL.cpp and XCCLBootstrap.cpp from the previous commit
resolve against the real header.
The event pool is sized from the max_event_pool_size hint rather than the
compile-time default, matching TorchCommXCCL. XCCLBootstrap already accepted
that key as a TorchComm-layer hint, so before this it was validated and then
silently ignored. The pool is const and declared ahead of options_c10d_, so
the constructor parses the hint off its own parameter (see
ProcessGroupXCCLBackend.cpp) rather than the member.
XCCLCachingAllocatorHook.{hpp,cpp} port TorchCommXCCLCCA. It is a process-wide
singleton that watches XPU caching-allocator SEGMENT_ALLOC/SEGMENT_FREE trace
events and mirrors them into every live ProcessGroupXCCL as
register_address()/deregister_address(), which wrap onecclCommRegister and
onecclCommDeregister. Registering a segment once lets oneCCL skip its
per-operation buffer setup, so this is purely a throughput optimization --
collectives are correct without it.
Two details are taken from nccl2's NCCLCachingAllocatorHook rather than from
torchcomms. The singleton is deliberately leaked, because allocator trace
trackers cannot be detached and the hook must outlive every allocator event.
And default-pool segments are skipped when the allocator is in expandable
segments mode, which is incompatible with registration. Comms are walked in
group-name order rather than set (pointer) order so that every rank issues
these calls in the same sequence.
split() and its helper initFromComm() are declared here and defined with the
engine implementation.
This was referenced Aug 25, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack (bottom → top)
xccl2/01-foundationxccl2/02-apixccl2/03-workxccl2/04-engine-declxccl2/05-engine-implxccl2/06-backend-surfacexccl2/07-build-wiringxccl2/08-testsxccl2/09-integration-testsEach PR is exactly one commit and targets the one below it. Review bottom-up.
Adds the ProcessGroupXCCL class declaration and the backend-internal helper
definitions, ported from torchcomms' TorchCommXCCL.hpp / TorchCommXCCLUtils.cpp
and adapted to c10d the same way nccl2 adapted TorchCommNCCL.
ProcessGroupXCCL.hpp declares
c10d::xccl2::ProcessGroupXCCL : c10d::Backend.The TorchComm / TorchCommBackend / BackendWrapper layers are collapsed away and
the collective surface takes the c10d option objects directly. The namespace is
what keeps this distinct from the existing c10d::ProcessGroupXCCL in
third_party/torch-xpu-ops, which drives the oneCCL C++ (ccl::) API rather than
the v2 C API this backend uses.
Unlike nccl2 -- which reuses ::c10d::ProcessGroupNCCL::Options because that type
already exists in core -- there is no core-side XCCL options type to reuse, so
the backend defines its own nested Options deriving from c10d::Backend::Options.
ProcessGroupXCCLUtils.cpp carries the RedOpRAII reduction-op wrapper, the
dtype/reduce-op mapping, the work factory/queue plumbing, the
operation-stream selection and the timeout watchdog. The one c10d-specific
change is the PREMUL_SUM factor: torchcomms carried it on its own ReduceOp
(op.factor()), while c10d hangs it off the ReduceOp supplement, so
getPreMulSumFactor() unpacks NCCLPreMulSumSupplement exactly as nccl2 does.
The event pool is not a member of the backend, unlike torchcomms and unlike
nccl2. It is an XcclEventPool held by shared_ptr and shared with every work
the backend creates, because ~WorkXCCL returns its events to the pool and a
work can outlive the backend; see the previous commit.
Behaviour is otherwise unchanged from torchcomms. In particular the watchdog
does not poll the communicator for async errors and does not abort the
communicator on timeout: oneCCL declares both onecclCommGetAsyncError and
onecclCommAbort with CCL_C_NOT_IMPLEMENTED, which expands to
attribute((error(...))), so referencing them does not compile. The
divergence from nccl2 here is deliberate and documented at the call sites.
Definitions deferred to later commits in this stack so that every commit stays
link-clean once the build is wired up: XCCLException, throwAsyncError(),
abortXcclComm(), runAbortHooks(), the init/finalize path and the *Impl engine
methods land with ProcessGroupXCCL.cpp; the c10d virtual overrides land with
ProcessGroupXCCLBackend.cpp.
This also lets WorkXCCL.cpp and XCCLBootstrap.cpp from the previous commit
resolve against the real header.
The event pool is sized from the max_event_pool_size hint rather than the
compile-time default, matching TorchCommXCCL. XCCLBootstrap already accepted
that key as a TorchComm-layer hint, so before this it was validated and then
silently ignored. The pool is const and declared ahead of options_c10d_, so
the constructor parses the hint off its own parameter (see
ProcessGroupXCCLBackend.cpp) rather than the member.
XCCLCachingAllocatorHook.{hpp,cpp} port TorchCommXCCLCCA. It is a process-wide
singleton that watches XPU caching-allocator SEGMENT_ALLOC/SEGMENT_FREE trace
events and mirrors them into every live ProcessGroupXCCL as
register_address()/deregister_address(), which wrap onecclCommRegister and
onecclCommDeregister. Registering a segment once lets oneCCL skip its
per-operation buffer setup, so this is purely a throughput optimization --
collectives are correct without it.
Two details are taken from nccl2's NCCLCachingAllocatorHook rather than from
torchcomms. The singleton is deliberately leaked, because allocator trace
trackers cannot be detached and the hook must outlive every allocator event.
And default-pool segments are skipped when the allocator is in expandable
segments mode, which is incompatible with registration. Comms are walked in
group-name order rather than set (pointer) order so that every rank issues
these calls in the same sequence.
split() and its helper initFromComm() are declared here and defined with the
engine implementation.