test(otel): add GPU, Neuron, and EFA DRA-path integration tests - #753
samehkhalil wants to merge 1 commit into
Conversation
077c30e to
f803c04
Compare
f803c04 to
d43ae5d
Compare
d43ae5d to
3073b10
Compare
Add integration coverage for the awsdevicepodcorrelation processor's DRA (Dynamic Resource Allocation) path, mirroring the device-plugin GPU/Neuron/EFA correlation tests. Each package exposes its devices via a DRA driver (through a ResourceClaimTemplate) instead of the device-plugin resource, and asserts per-device pod correlation. - test/otel/multi_efa_dra: EFA via dranet (driver dra.net); efaburn claims one of two devices, the other stays unclaimed. Guards the per-device correlation collapse (ResourceSlice keying via dra.net/rdmaDevice plus the groupbyattrs split before the resource-level promote). - test/otel/neuron_dra: Neuron via the AWS Neuron DRA driver (DeviceClass neuron.aws.com). Single Trainium device (trn1.2xlarge); the claimed device's two cores attribute to the burn pod and to no other pod. The Neuron DRA driver supports Trainium only, so this targets trn1.2xlarge. - test/otel/gpu_dra: GPU via the NVIDIA DRA driver (DeviceClass gpu.nvidia.com). g4dn.12xlarge (4 GPUs); one claimed GPU correlates to the burn pod, the other three stay uncorrelated. Asserts device count, consecutive indices, and all DCGM metrics per device. New terraform modules under terraform/eks/daemon (otel-multi-efa-dra, otel-neuron-dra, otel-gpu-dra) install the DRA driver in place of the device plugin and apply a ResourceClaimTemplate burn workload. The processor uses the GA resource.k8s.io/v1 DRA API (available since Kubernetes 1.34), so the clusters run k8s 1.35 like the rest of the suite. Wired into the test case generator. Requires a chart carrying the DRA correlation config and resource.k8s.io RBAC.
3073b10 to
b0171bf
Compare
| } | ||
| } | ||
|
|
||
| require.Len(t, claimed, expectedClaimedGPUs, |
There was a problem hiding this comment.
This checks how many GPUs are correlated, not which one. The DRA keying (the gpu-(\d+) pattern) is what's under test here. If it pointed the pod at the wrong GPU (e.g. the DRA device index not matching DCGM's gpu label), we'd still get "1 claimed, 3 unclaimed" and the test would pass.
Cheap way to pin it: under DRA, gpu_burn only sees its claimed GPU, so the correlated GPU should be the only busy one:
busy := map[string]bool{}
for _, r := range multi {
if r.Value > 0 {
busy[r.Labels.Datapoint["gpu"]] = true
}
}
require.True(t, busy[claimed[0]], "GPU %s is correlated to %s* but idle; the burn pod is on a different GPU", claimed[0], burnPodPrefix)For an exact check, read the ground truth from the API the way test/otel/gpu/k8s_helpers_test.go already does: burn pod → status.resourceClaimStatuses → ResourceClaim status.allocation.devices.results[].device → ResourceSlice uuid attribute. Then assert that the aDCGM series with that UUID carries the burn pod.
| } | ||
|
|
||
| // Claimed side: exactly the number of EFAs efaburn requested map to a pod. | ||
| require.Len(t, claimed, expectedClaimedEFACount, |
There was a problem hiding this comment.
Same gap as in gpu_dra: with 2 devices, a keying bug that attaches efaburn to the other EFA still passes "1 claimed / 1 unclaimed". Since the dra.net/rdmaDevice attribute lookup is the main thing this test guards, could we assert the correlated aws.efa.device equals the attribute of the device actually allocated? That's the efaburn ResourceClaim's status.allocation.devices.results[].device, then the matching ResourceSlice device's dra.net/rdmaDevice.
Description of the issue
The
awsdevicepodcorrelationprocessor's Dynamic Resource Allocation (DRA) path had nointegration coverage - only the device-plugin GPU/Neuron/EFA correlation tests existed. There
was nothing exercising devices allocated via a
ResourceClaimTemplateend to end (driver ->ResourceSlice/ResourceClaim -> metric correlation), so regressions on the DRA path could ship
undetected.
Depends on aws-observability/helm-charts#356 (DRA correlation config +
resource.k8s.ioRBAC) and the processor change in amazon-contributing/opentelemetry-collector-contrib#631.
These go green in upstream CI once those land; until then they run against a chart/agent image
carrying those changes.
Description of changes
Add integration coverage mirroring the device-plugin tests, each exposing devices via a DRA
driver instead of the device-plugin resource and asserting per-device pod correlation:
test/otel/multi_efa_dra: EFA via dranet (driver dra.net);efaburnclaims one of twodevices, the other stays unclaimed. Guards the per-device correlation collapse
(ResourceSlice keying via
dra.net/rdmaDeviceplus the groupbyattrs split before theresource-level promote).
test/otel/neuron_dra: Neuron via the AWS Neuron DRA driver (DeviceClass neuron.aws.com).Single Trainium device (trn1.2xlarge); the claimed device's two cores attribute to the burn
pod and to no other pod. The Neuron DRA driver supports Trainium only, so this targets
trn1.2xlarge.
test/otel/gpu_dra: GPU via the NVIDIA DRA driver (DeviceClass gpu.nvidia.com).g4dn.12xlarge (4 GPUs); one claimed GPU correlates to the burn pod, the other three stay
uncorrelated. Asserts device count, consecutive indices, and all DCGM metrics per device.
New terraform modules under
terraform/eks/daemon(otel-multi-efa-dra, otel-neuron-dra,otel-gpu-dra) install the DRA driver in place of the device plugin and apply a
ResourceClaimTemplate burn workload. All run k8s 1.35 (the processor uses the GA
resource.k8s.io/v1DRA API, available since 1.34), and are wired into the test casegenerator.
License
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.
Tests
go vet -tags integrationis clean for all three packages and the generator. Validatedend-to-end on live EKS 1.35 clusters (all tests pass):
and correlated, the other 3 uncorrelated.
Test output from the live runs (
count=N= correlated-series counts each assertion queriedfrom CloudWatch):
PR checklist
makepasses locally —make simple-lint(checklicense + impi, the Go checks build-check.yml runs) passes on all tracked files;make compilepasses. Note these targets do not pass-tags integration, so the added test code (all//go:build integration) is compile-checked withgo vet -tags integrationon all three DRA packages + the generator (clean).