diff --git a/README.md b/README.md
index bf710b7..1b3906f 100644
--- a/README.md
+++ b/README.md
@@ -6,6 +6,10 @@ CAUTION: This is an beta / non-production software, do not use on production clu
Intel GPU Base operator allows automatic deployment of GPU related components to enable use of Intel GPU hardware within the Kubernetes cluster.
+For the production security baseline, image guidance, RBAC, webhook TLS,
+metrics, firmware updates, and optional integration risks, see the
+[Security Configuration Guide](SECURITY-CONFIGURATION.md).
+
## Description

diff --git a/SECURITY-CONFIGURATION.md b/SECURITY-CONFIGURATION.md
new file mode 100644
index 0000000..c88bf47
--- /dev/null
+++ b/SECURITY-CONFIGURATION.md
@@ -0,0 +1,67 @@
+# Security Configuration Guide
+
+This guide describes the recommended security baseline for deploying the Intel
+GPU Base Operator. It supplements the installation and feature documentation
+in `README.md`, the Helm chart READMEs, and `FWUPDATE.md`.
+
+## Production baseline
+
+- Install a released chart version, not a development checkout or `:devel`
+ image. Prefer images referenced by immutable digests. Review every
+ customer-supplied image before allowing it to run on GPU nodes.
+- Install cert-manager and verify that the webhook certificate is issued correctly.
+- Install only the integrations required by the cluster. See [optional integrations](#optional-integrations-and-their-security-impact).
+- Restrict `create` and `update` access to `ClusterPolicy`,
+ `GPUFirmwareUpdate`, and `GPURecoveryPlan` resources to trusted cluster
+ administrators. These resources control privileged workloads and can
+ modify nodes, drain workloads, reset GPUs, or write firmware.
+- Use dedicated namespaces and review the generated RBAC, ServiceAccount,
+ SCC, and NetworkPolicy resources before installation.
+- Use the default secure metrics configuration. If metrics are enabled, keep
+ HTTPS and authentication/authorization enabled and grant scrape access only
+ to the intended Prometheus service account.
+
+## Images and registries
+
+Use a released operator image and pin component images to digests whenever
+possible. Mutable tags such as `latest` and `devel` are unsuitable for
+production because the container executed on nodes can change without a chart
+change. If a private registry is used, configure the chart's private registry
+settings and protect the resulting Kubernetes Secret.
+
+For firmware updates, use separate, reviewed images for `spec.updaterImage`
+and `spec.content.containerImage`. The firmware content image should be
+digest-pinned whenever checksums are specified; this binds pre-flight
+verification to the image pulled by the update job.
+
+For GPU recovery, review and preferably digest-pin both
+`spec.xpuSmi.image` and `spec.firmware.source.containerSource.name`. Recovery
+Jobs run the supplied image with a shell as root, privileged access, and
+writable host `/sys`, pinned to the affected node. The reset command or
+firmware filename is generated by the operator, but the image is
+administrator-controlled and must be treated as trusted host-level code.
+Keep image-pull credentials scoped to the operator namespace and do not allow
+untrusted users to edit recovery plans.
+
+## Webhooks, TLS, and validation
+
+The GPUFirmwareUpdate, GPURecoveryPlan, and ClusterPolicy webhooks validate
+CR inputs and protect update state transitions. Webhook failure policy is
+fail-closed for GPU firmware updates. Treat an unavailable webhook as an
+installation or certificate problem to resolve, not as a reason to disable
+validation.
+
+## Optional integrations and their security impact
+
+Enable only the following integrations that are needed:
+
+| Integration | Default | Security considerations |
+|---|---|---|
+| NFD | Disabled | Creates node-feature rules and labels. Review label consumers and remove generated rules during uninstall if no longer needed. |
+| Kueue | Disabled | Creates cluster-scoped queues and controls workload admission. Review queue administrators, local-queue namespaces, and resource allocation policy. |
+| Prometheus/ServiceMonitor | Disabled | Exposes GPU and operator metrics to the configured monitoring stack. Restrict scrape access and review TLS and selector settings. |
+| Xpumd monitoring | Policy-dependent; enabled by the policy chart sample defaults | Collects GPU health and telemetry and requires host/device access. Configure the monitoring backend's access and retention controls. |
+| DRA | Selected by `spec.resourceRegistration` | Uses ResourceClaims and ResourceSlices and affects workload resource allocation. Review namespace access and DRA administrators. |
+| OpenShift support | Disabled | Creates SCC, Role, RoleBinding, and ServiceAccount resources. Review privileged SCC permissions and bindings before enabling. |
+| Firmware updates | Explicit `GPUFirmwareUpdate` operation | Requires trusted images and elevated access. It taints and drains nodes, runs update Jobs, and should use checksum verification, digest-pinned content, and canary mode where appropriate. |
+| GPU recovery plans | Explicit `GPURecoveryPlan` operation | Detects DRA device taints and waits for approval before draining nodes and running privileged, node-pinned reset or reflash Jobs. Restrict plan and approval RBAC, review images, and monitor event state and cleanup. |
diff --git a/charts/gpu-base-operator-policy/README.md b/charts/gpu-base-operator-policy/README.md
index 610ee69..3f89aa0 100644
--- a/charts/gpu-base-operator-policy/README.md
+++ b/charts/gpu-base-operator-policy/README.md
@@ -2,6 +2,9 @@
Helm chart is for installing the Intel GPU base operator policy. The operator has to be installed before the policy. See [the operator chart](../gpu-base-operator/README.md).
+Review the [Security Configuration Guide](../../SECURITY-CONFIGURATION.md)
+before enabling privileged GPU components or optional integrations.
+
## Helm install
```
diff --git a/charts/gpu-base-operator/README.md b/charts/gpu-base-operator/README.md
index a989ef7..7a3b885 100644
--- a/charts/gpu-base-operator/README.md
+++ b/charts/gpu-base-operator/README.md
@@ -2,6 +2,10 @@
Helm chart is for installing the Intel GPU base operator. Operator installation is a dependency for the [policy chart](../gpu-base-operator-policy/README.md). Once the operator is installed, the policy chart can configure the cluster in a certain way.
+For the recommended production baseline, image pinning, webhook TLS, RBAC,
+metrics, and optional integration security considerations, see the
+[Security Configuration Guide](../../SECURITY-CONFIGURATION.md).
+
## Prerequisites
- [cert-manager](https://cert-manager.io/docs/installation/) [required — provisions TLS certificates for the admission webhook]
- [Node Feature Discovery NFD](https://kubernetes-sigs.github.io/node-feature-discovery/master/get-started/deployment-and-usage.html) [recommended, optional]
diff --git a/docs/architecture.svg b/docs/architecture.svg
index 4637183..a3a46ba 100644
--- a/docs/architecture.svg
+++ b/docs/architecture.svg
@@ -30,28 +30,36 @@
GPUFirmwareUpdate
- nodeSelector, fwImageTag
- retryLimit, timeout
+ nodeSelector, updaterImage
+ holdAfterCanary, pciDeviceID
+
+
+
+ GPURecoveryPlan
+ deviceId, approvals, defaultResetType
+ reset / reflash, drain, timeouts
+
-
- 🔒 Admission Webhook
- Validates ClusterPolicy & GPUFirmwareUpdate on create / update — rejects malformed CRs
+ 🔒 Admission Webhook
+ Validates ClusterPolicy, GPUFirmwareUpdate & GPURecoveryPlan on create / update — rejects malformed CRs
-
+
-
-
-
- Intel GPU Base Operator (controller-manager)
+
+
+ Intel GPU Base Operator (controller-manager)
• Device Plugin Reconciler
• DRA Reconciler
• XPU-Manager Reconciler
- • Misc / NFD Rule Reconciler
- • Kueue Reconciler
+ • Misc / Kueue / NFD Reconciler
+ • KMM Reconciler
- Detects at startup: DRA · OpenShift · Prometheus · Kueue
+ Detects at startup: DRA · OpenShift · Prometheus
Enables / disables reconcilers accordingly
Watches one ClusterPolicy per cluster
Reconciles on CR create / update / delete
@@ -83,34 +91,48 @@
• Node selection
• Status / conditions reporting
+
+
+
+
+ RecoveryPlan Controller
+ • Approval, drain & recovery Jobs
+ • Reset / reflash status
+
Managed K8s Resources
-
-
+
+
-
-
-
-
+
+
+
+
+
-
- DaemonSet: Intel GPU Device Plugin
+ DS: GPU Device Plugin
-
- DaemonSet: DRA Plugin
+ DS: GPU DRA Plugin
-
- DaemonSet: Intel XPU-Manager
+ DS: Intel Xpumd
-
- Job: GPU Firmware Update
+ Job: GPU Firmware Update
+
+
+ Jobs: GPU Recovery
Cluster dependencies & integrations:
@@ -133,6 +155,12 @@
workload queueing
& resource reservation
+
+ KMM
+ out-of-tree Xe driver
+ kernel module builds
+
@@ -140,18 +168,18 @@
Device Plugin
- gpu.intel.com/i915 extended resource
- LevelZero health checks
+ gpu.intel.com/xe extended resource
+ Xpumd health data
scheduling via kubelet
DRA Plugin
ResourceSlice · driver: gpu.intel.com
- health via XPU-Manager
+ health via Xpumd
scheduling via ResourceClaim
- XPU-Manager
+ Xpumd
per-GPU telemetry & health
metrics endpoint (REST/gRPC)
scraped by Prometheus
@@ -163,7 +191,7 @@
nodeSelector + Kueue queuing
-
deployed to
/ runs on