Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,10 @@ CAUTION: This is an beta / non-production software, do not use on production clu

Intel GPU Base operator allows automatic deployment of GPU related components to enable use of Intel GPU hardware within the Kubernetes cluster.

For the production security baseline, image guidance, RBAC, webhook TLS,
metrics, firmware updates, and optional integration risks, see the
[Security Configuration Guide](SECURITY-CONFIGURATION.md).

## Description

![Architecture diagram](docs/architecture.svg)
Expand Down
67 changes: 67 additions & 0 deletions SECURITY-CONFIGURATION.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
# Security Configuration Guide

This guide describes the recommended security baseline for deploying the Intel
GPU Base Operator. It supplements the installation and feature documentation
in `README.md`, the Helm chart READMEs, and `FWUPDATE.md`.

## Production baseline

- Install a released chart version, not a development checkout or `:devel`
image. Prefer images referenced by immutable digests. Review every
customer-supplied image before allowing it to run on GPU nodes.
- Install cert-manager and verify that the webhook certificate is issued correctly.
- Install only the integrations required by the cluster. See [optional integrations](#optional-integrations-and-their-security-impact).
- Restrict `create` and `update` access to `ClusterPolicy`,
`GPUFirmwareUpdate`, and `GPURecoveryPlan` resources to trusted cluster
administrators. These resources control privileged workloads and can
modify nodes, drain workloads, reset GPUs, or write firmware.
- Use dedicated namespaces and review the generated RBAC, ServiceAccount,
SCC, and NetworkPolicy resources before installation.
- Use the default secure metrics configuration. If metrics are enabled, keep
HTTPS and authentication/authorization enabled and grant scrape access only
to the intended Prometheus service account.

## Images and registries

Use a released operator image and pin component images to digests whenever
possible. Mutable tags such as `latest` and `devel` are unsuitable for
production because the container executed on nodes can change without a chart
change. If a private registry is used, configure the chart's private registry
settings and protect the resulting Kubernetes Secret.

For firmware updates, use separate, reviewed images for `spec.updaterImage`
and `spec.content.containerImage`. The firmware content image should be
digest-pinned whenever checksums are specified; this binds pre-flight
verification to the image pulled by the update job.

For GPU recovery, review and preferably digest-pin both
`spec.xpuSmi.image` and `spec.firmware.source.containerSource.name`. Recovery
Jobs run the supplied image with a shell as root, privileged access, and
writable host `/sys`, pinned to the affected node. The reset command or
firmware filename is generated by the operator, but the image is
administrator-controlled and must be treated as trusted host-level code.
Keep image-pull credentials scoped to the operator namespace and do not allow
untrusted users to edit recovery plans.

## Webhooks, TLS, and validation

The GPUFirmwareUpdate, GPURecoveryPlan, and ClusterPolicy webhooks validate
CR inputs and protect update state transitions. Webhook failure policy is
fail-closed for GPU firmware updates. Treat an unavailable webhook as an
installation or certificate problem to resolve, not as a reason to disable
validation.

## Optional integrations and their security impact

Enable only the following integrations that are needed:

| Integration | Default | Security considerations |
|---|---|---|
| NFD | Disabled | Creates node-feature rules and labels. Review label consumers and remove generated rules during uninstall if no longer needed. |
| Kueue | Disabled | Creates cluster-scoped queues and controls workload admission. Review queue administrators, local-queue namespaces, and resource allocation policy. |
| Prometheus/ServiceMonitor | Disabled | Exposes GPU and operator metrics to the configured monitoring stack. Restrict scrape access and review TLS and selector settings. |
| Xpumd monitoring | Policy-dependent; enabled by the policy chart sample defaults | Collects GPU health and telemetry and requires host/device access. Configure the monitoring backend's access and retention controls. |
| DRA | Selected by `spec.resourceRegistration` | Uses ResourceClaims and ResourceSlices and affects workload resource allocation. Review namespace access and DRA administrators. |
| OpenShift support | Disabled | Creates SCC, Role, RoleBinding, and ServiceAccount resources. Review privileged SCC permissions and bindings before enabling. |
| Firmware updates | Explicit `GPUFirmwareUpdate` operation | Requires trusted images and elevated access. It taints and drains nodes, runs update Jobs, and should use checksum verification, digest-pinned content, and canary mode where appropriate. |
| GPU recovery plans | Explicit `GPURecoveryPlan` operation | Detects DRA device taints and waits for approval before draining nodes and running privileged, node-pinned reset or reflash Jobs. Restrict plan and approval RBAC, review images, and monitor event state and cleanup. |
3 changes: 3 additions & 0 deletions charts/gpu-base-operator-policy/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,9 @@

Helm chart is for installing the Intel GPU base operator policy. The operator has to be installed before the policy. See [the operator chart](../gpu-base-operator/README.md).

Review the [Security Configuration Guide](../../SECURITY-CONFIGURATION.md)
before enabling privileged GPU components or optional integrations.


## Helm install
```
Expand Down
4 changes: 4 additions & 0 deletions charts/gpu-base-operator/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,10 @@

Helm chart is for installing the Intel GPU base operator. Operator installation is a dependency for the [policy chart](../gpu-base-operator-policy/README.md). Once the operator is installed, the policy chart can configure the cluster in a certain way.

For the recommended production baseline, image pinning, webhook TLS, RBAC,
metrics, and optional integration security considerations, see the
[Security Configuration Guide](../../SECURITY-CONFIGURATION.md).

## Prerequisites
- [cert-manager](https://cert-manager.io/docs/installation/) [required — provisions TLS certificates for the admission webhook]
- [Node Feature Discovery NFD](https://kubernetes-sigs.github.io/node-feature-discovery/master/get-started/deployment-and-usage.html) [recommended, optional]
Expand Down
92 changes: 60 additions & 32 deletions docs/architecture.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.