diff --git a/README.md b/README.md index bf710b7..1b3906f 100644 --- a/README.md +++ b/README.md @@ -6,6 +6,10 @@ CAUTION: This is an beta / non-production software, do not use on production clu Intel GPU Base operator allows automatic deployment of GPU related components to enable use of Intel GPU hardware within the Kubernetes cluster. +For the production security baseline, image guidance, RBAC, webhook TLS, +metrics, firmware updates, and optional integration risks, see the +[Security Configuration Guide](SECURITY-CONFIGURATION.md). + ## Description ![Architecture diagram](docs/architecture.svg) diff --git a/SECURITY-CONFIGURATION.md b/SECURITY-CONFIGURATION.md new file mode 100644 index 0000000..c88bf47 --- /dev/null +++ b/SECURITY-CONFIGURATION.md @@ -0,0 +1,67 @@ +# Security Configuration Guide + +This guide describes the recommended security baseline for deploying the Intel +GPU Base Operator. It supplements the installation and feature documentation +in `README.md`, the Helm chart READMEs, and `FWUPDATE.md`. + +## Production baseline + +- Install a released chart version, not a development checkout or `:devel` + image. Prefer images referenced by immutable digests. Review every + customer-supplied image before allowing it to run on GPU nodes. +- Install cert-manager and verify that the webhook certificate is issued correctly. +- Install only the integrations required by the cluster. See [optional integrations](#optional-integrations-and-their-security-impact). +- Restrict `create` and `update` access to `ClusterPolicy`, + `GPUFirmwareUpdate`, and `GPURecoveryPlan` resources to trusted cluster + administrators. These resources control privileged workloads and can + modify nodes, drain workloads, reset GPUs, or write firmware. +- Use dedicated namespaces and review the generated RBAC, ServiceAccount, + SCC, and NetworkPolicy resources before installation. +- Use the default secure metrics configuration. If metrics are enabled, keep + HTTPS and authentication/authorization enabled and grant scrape access only + to the intended Prometheus service account. + +## Images and registries + +Use a released operator image and pin component images to digests whenever +possible. Mutable tags such as `latest` and `devel` are unsuitable for +production because the container executed on nodes can change without a chart +change. If a private registry is used, configure the chart's private registry +settings and protect the resulting Kubernetes Secret. + +For firmware updates, use separate, reviewed images for `spec.updaterImage` +and `spec.content.containerImage`. The firmware content image should be +digest-pinned whenever checksums are specified; this binds pre-flight +verification to the image pulled by the update job. + +For GPU recovery, review and preferably digest-pin both +`spec.xpuSmi.image` and `spec.firmware.source.containerSource.name`. Recovery +Jobs run the supplied image with a shell as root, privileged access, and +writable host `/sys`, pinned to the affected node. The reset command or +firmware filename is generated by the operator, but the image is +administrator-controlled and must be treated as trusted host-level code. +Keep image-pull credentials scoped to the operator namespace and do not allow +untrusted users to edit recovery plans. + +## Webhooks, TLS, and validation + +The GPUFirmwareUpdate, GPURecoveryPlan, and ClusterPolicy webhooks validate +CR inputs and protect update state transitions. Webhook failure policy is +fail-closed for GPU firmware updates. Treat an unavailable webhook as an +installation or certificate problem to resolve, not as a reason to disable +validation. + +## Optional integrations and their security impact + +Enable only the following integrations that are needed: + +| Integration | Default | Security considerations | +|---|---|---| +| NFD | Disabled | Creates node-feature rules and labels. Review label consumers and remove generated rules during uninstall if no longer needed. | +| Kueue | Disabled | Creates cluster-scoped queues and controls workload admission. Review queue administrators, local-queue namespaces, and resource allocation policy. | +| Prometheus/ServiceMonitor | Disabled | Exposes GPU and operator metrics to the configured monitoring stack. Restrict scrape access and review TLS and selector settings. | +| Xpumd monitoring | Policy-dependent; enabled by the policy chart sample defaults | Collects GPU health and telemetry and requires host/device access. Configure the monitoring backend's access and retention controls. | +| DRA | Selected by `spec.resourceRegistration` | Uses ResourceClaims and ResourceSlices and affects workload resource allocation. Review namespace access and DRA administrators. | +| OpenShift support | Disabled | Creates SCC, Role, RoleBinding, and ServiceAccount resources. Review privileged SCC permissions and bindings before enabling. | +| Firmware updates | Explicit `GPUFirmwareUpdate` operation | Requires trusted images and elevated access. It taints and drains nodes, runs update Jobs, and should use checksum verification, digest-pinned content, and canary mode where appropriate. | +| GPU recovery plans | Explicit `GPURecoveryPlan` operation | Detects DRA device taints and waits for approval before draining nodes and running privileged, node-pinned reset or reflash Jobs. Restrict plan and approval RBAC, review images, and monitor event state and cleanup. | diff --git a/charts/gpu-base-operator-policy/README.md b/charts/gpu-base-operator-policy/README.md index 610ee69..3f89aa0 100644 --- a/charts/gpu-base-operator-policy/README.md +++ b/charts/gpu-base-operator-policy/README.md @@ -2,6 +2,9 @@ Helm chart is for installing the Intel GPU base operator policy. The operator has to be installed before the policy. See [the operator chart](../gpu-base-operator/README.md). +Review the [Security Configuration Guide](../../SECURITY-CONFIGURATION.md) +before enabling privileged GPU components or optional integrations. + ## Helm install ``` diff --git a/charts/gpu-base-operator/README.md b/charts/gpu-base-operator/README.md index a989ef7..7a3b885 100644 --- a/charts/gpu-base-operator/README.md +++ b/charts/gpu-base-operator/README.md @@ -2,6 +2,10 @@ Helm chart is for installing the Intel GPU base operator. Operator installation is a dependency for the [policy chart](../gpu-base-operator-policy/README.md). Once the operator is installed, the policy chart can configure the cluster in a certain way. +For the recommended production baseline, image pinning, webhook TLS, RBAC, +metrics, and optional integration security considerations, see the +[Security Configuration Guide](../../SECURITY-CONFIGURATION.md). + ## Prerequisites - [cert-manager](https://cert-manager.io/docs/installation/) [required — provisions TLS certificates for the admission webhook] - [Node Feature Discovery NFD](https://kubernetes-sigs.github.io/node-feature-discovery/master/get-started/deployment-and-usage.html) [recommended, optional] diff --git a/docs/architecture.svg b/docs/architecture.svg index 4637183..a3a46ba 100644 --- a/docs/architecture.svg +++ b/docs/architecture.svg @@ -30,28 +30,36 @@ GPUFirmwareUpdate - nodeSelector, fwImageTag - retryLimit, timeout + nodeSelector, updaterImage + holdAfterCanary, pciDeviceID + + + + GPURecoveryPlan + deviceId, approvals, defaultResetType + reset / reflash, drain, timeouts + - - 🔒 Admission Webhook - Validates ClusterPolicy & GPUFirmwareUpdate on create / update — rejects malformed CRs + 🔒 Admission Webhook + Validates ClusterPolicy, GPUFirmwareUpdate & GPURecoveryPlan on create / update — rejects malformed CRs - + - - - - Intel GPU Base Operator (controller-manager) + + + Intel GPU Base Operator (controller-manager) • Device Plugin Reconciler • DRA Reconciler • XPU-Manager Reconciler - • Misc / NFD Rule Reconciler - • Kueue Reconciler + • Misc / Kueue / NFD Reconciler + • KMM Reconciler - Detects at startup: DRA · OpenShift · Prometheus · Kueue + Detects at startup: DRA · OpenShift · Prometheus Enables / disables reconcilers accordingly Watches one ClusterPolicy per cluster Reconciles on CR create / update / delete @@ -83,34 +91,48 @@ • Node selection • Status / conditions reporting + + + + + RecoveryPlan Controller + • Approval, drain & recovery Jobs + • Reset / reflash status + Managed K8s Resources - - + + - - - - + + + + + - - DaemonSet: Intel GPU Device Plugin + DS: GPU Device Plugin - - DaemonSet: DRA Plugin + DS: GPU DRA Plugin - - DaemonSet: Intel XPU-Manager + DS: Intel Xpumd - - Job: GPU Firmware Update + Job: GPU Firmware Update + + + Jobs: GPU Recovery Cluster dependencies & integrations: @@ -133,6 +155,12 @@ workload queueing & resource reservation + + KMM + out-of-tree Xe driver + kernel module builds + @@ -140,18 +168,18 @@ Device Plugin - gpu.intel.com/i915 extended resource - LevelZero health checks + gpu.intel.com/xe extended resource + Xpumd health data scheduling via kubelet DRA Plugin ResourceSlice · driver: gpu.intel.com - health via XPU-Manager + health via Xpumd scheduling via ResourceClaim - XPU-Manager + Xpumd per-GPU telemetry & health metrics endpoint (REST/gRPC) scraped by Prometheus @@ -163,7 +191,7 @@ nodeSelector + Kueue queuing - deployed to / runs on