NVIDIA GPU Operator: GPU Nodes as Ordinary Kubernetes Nodes
NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes
At a glance
- What is it?
- The GPU Operator packages the driver, container runtime, device plugin, DCGM monitoring and node labelling into one Helm-installed controller. This review covers what it manages, how to install it, and where the containerised driver model becomes a constraint.
- Who is it for?
- Adopt the GPU Operator if your Kubernetes nodes are provisioned from a standard OS image and you need the driver, device plugin, container runtime and DCGM monitoring managed as cluster workloads rather than baked into an image. Do not adopt it if you cannot grant the DaemonSets privileged access to the host, or if your nodes already carry a validated driver stack you are unwilling to hand to a controller.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem: a GPU node is a stack, not a device
Kubernetes reaches special hardware through the device plugin framework. That framework only exposes the resource. It does not install the kernel driver that makes the device usable, does not configure the container runtime to inject the driver libraries into pods, and does not label the node so the scheduler knows what it holds. On a CPU-only cluster none of this matters. On a GPU cluster each of those steps is a separate piece of software with its own version and its own failure mode.
The README states the case plainly: configuring nodes with these hardware resources requires configuration of multiple software components such as drivers, container runtimes or other libraries, and this is difficult and prone to errors. The GPU Operator exists to move that work out of the machine image and into the cluster. The stated audience is cluster administrators who want to manage GPU nodes the way they manage CPU nodes, using one standard OS image for both and letting the operator provision the GPU software afterwards. The README also names the scenario where this pays off most: clusters that need to scale quickly, provisioning additional GPU nodes on cloud or on-prem and managing the lifecycle of the components underneath. Because every component runs as a container, swapping a version is a matter of starting or stopping containers rather than rebuilding images.
What the operator actually manages on each node
The component list in the README is the architecture. The NVIDIA drivers enable CUDA. The Kubernetes device plugin advertises GPUs as schedulable resources. The NVIDIA Container Runtime wires the driver into containers. Automatic node labelling tells the scheduler which nodes carry GPUs. DCGM based monitoring reports GPU health and utilisation. The operator drives all of these as operands.
The repository layout matches that description. There is a controllers directory for the reconciliation logic, an api directory for the custom resource types, manifests and deployments for the rendered objects, and a bundle directory, which is the Operator Lifecycle Manager packaging format used by OpenShift. The go.mod file shows the project is built on sigs.k8s.io/controller-runtime and k8s.io/client-go, the standard Go controller stack, with github.com/NVIDIA/nvidia-container-toolkit as a direct dependency. The Makefile confirms the shipped image coordinates: the default registry is nvcr.io/nvidia/cloud-native and the image name is gpu-operator. So the operator is itself a Kubernetes controller, running in-cluster, reconciling node state toward a declared configuration.
One design note worth stating directly: because the driver is delivered as a container, the operator needs privileged access to the host to load kernel modules. That is the price of the containerised model, and it is why the prerequisite checks matter more than they would for an ordinary workload.
Installing the GPU Operator with Helm
The README gives a two-step quick start for the data center driver. First add the NVIDIA Helm repository and refresh the index:
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia \
&& helm repo updateThen install the chart into its own namespace, creating the namespace if it does not exist. The --wait flag makes Helm block until the release resources are ready, and --generate-name lets Helm pick the release name instead of you supplying one:
helm install --wait --generate-name \
-n gpu-operator --create-namespace \
nvidia/gpu-operatorAfter installation, the README says the GPU Operator and its operands should be up and running. In practice that means a controller pod plus the operand DaemonSets it creates. The README does not document a rollback procedure, so treat the release name that --generate-name produced as something to record before you need it. Before either command, the README points to two checks: the prerequisites page and the platform support page, which lists supported operating systems and Kubernetes platforms. OpenShift users are told not to follow this path at all but to use the OpenShift-specific instructions in the official documentation.
Where the containerised driver model bites
The operator's main selling point is also its main constraint. Running the driver as a container means the node's kernel module state is managed by a controller, not by the image you validated. If the operator's driver container does not match the running kernel, the node does not silently degrade, it fails to expose GPUs. That makes kernel and OS version drift a first-class operational concern rather than a background detail.
Privilege is the second constraint. Loading kernel modules and configuring the container runtime are host-level operations, so the operand pods need elevated permissions. Clusters with restrictive pod security policy, or environments where a platform team will not approve privileged DaemonSets, are the wrong fit. In those cases a pre-baked GPU node image with the driver installed at provisioning time is the more honest choice, even though it costs you the fast-swap property the README highlights.
Scale is the third. The README frames the operator around clusters that need to scale quickly. For a single long-lived GPU node with a hand-validated driver, the operator adds a controller, a set of DaemonSets and a CRD surface in exchange for automating something you already did once. That trade is not obviously worth it.
Alternatives and how they differ in approach
The clearest alternative is the one the README implicitly argues against: bake the driver, container runtime and device plugin into a GPU-specific machine image, and let the cluster treat those nodes as pre-provisioned. The difference is where the state lives. With the operator, the node starts as a standard image and the software arrives as cluster workloads, so upgrading a component is a container swap. With a baked image, the node starts complete and the software is versioned with the image, so upgrading means rolling the node pool. The baked approach removes the privileged DaemonSets and the kernel-matching problem, at the cost of a second image to maintain and a slower path to adding capacity.
A second point of comparison sits inside the repository itself. The roadmap lists promoting the NVIDIADriver CRD to General Availability and integrating NVIDIA's DRA Driver for GPUs as a managed component. The DRA Driver is a separate project, referenced in the README, that takes a different route to exposing accelerators than the classic device plugin. The roadmap entry means the operator is not the only NVIDIA-managed path to GPU scheduling in Kubernetes, and the two are converging rather than competing. If your cluster already standardises on DRA, that roadmap item is the one to watch.
Maintenance, releases and licence
The project is not archived, and the last push was on 2026-09-24. Releases are dated and frequent: v26.7.1 on 2026-09-23, v26.7.0 on 2026-08-21, and v26.3.3 on 2026-06-25. That cadence matters because the operator tracks driver and Kubernetes versions, and a stale operator is a compatibility liability rather than a stable baseline. The README's roadmap is concrete about what is coming: the latest NVIDIA data center GPUs, systems and drivers, RHEL 10 support, KubeVirt with Ubuntu 24.04, the NVIDIADriver CRD promotion, and DRA Driver integration. Each of those is a reason an upgrade may be needed rather than optional.
The licence is Apache-2.0. For most adopters that is a permissive licence with the usual obligations around notices and attribution. The repository carries a THIRD_PARTY_NOTICES.md file, which is where the bundled dependencies and their terms are recorded; if your organisation runs licence review, that file plus the go.mod dependency list is the material to hand over. Nothing here is legal advice, and the operator ships NVIDIA-built container images from nvcr.io, whose terms are separate from the source licence.
What to verify before you commit
Check three things in order. First, the platform support page, because the README makes it a prerequisite rather than a suggestion: your OS and Kubernetes distribution need to be listed. Second, whether your cluster policy permits privileged DaemonSets that touch the host kernel and container runtime. Third, whether your nodes already run a driver you depend on, because the operator will manage that layer and you need to know how the two interact before an upgrade, not after.
The README does not document rollback, does not describe a dry-run mode, and does not cover air-gapped installation. If any of those matter to you, the official documentation is the place to look, and their absence from the README is worth noting rather than assuming away.
Editorial conclusion
Adopt the GPU Operator if your Kubernetes nodes are provisioned from a standard OS image and you need the driver, device plugin, container runtime and DCGM monitoring managed as cluster workloads rather than baked into an image. Do not adopt it if you cannot grant the DaemonSets privileged access to the host, or if your nodes already carry a validated driver stack you are unwilling to hand to a controller. Before rolling it out, verify your OS and Kubernetes version against the platform support page, and confirm the Helm release lands in the gpu-operator namespace with its operands running.
Frequently asked questions
What is the NVIDIA GPU Operator?
It is a Kubernetes operator that automates the management of the NVIDIA software components needed to provision GPUs, including the drivers, the Kubernetes device plugin, the NVIDIA Container Runtime, automatic node labelling and DCGM based monitoring.
How do I install the NVIDIA GPU Operator?
Add the NVIDIA Helm repository with helm repo add nvidia https://helm.ngc.nvidia.com/nvidia and then install the chart into the gpu-operator namespace with helm install --wait --generate-name -n gpu-operator --create-namespace nvidia/gpu-operator. The README notes that OpenShift users should follow the separate OpenShift instructions instead.
What does the NVIDIA GPU Operator manage on a node?
According to the README, it manages the NVIDIA drivers, the Kubernetes device plugin for GPUs, the NVIDIA Container Runtime, automatic node labelling and DCGM based monitoring.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-gpu-operator)