Self-hosted service
NVIDIA/k8s-device-plugin avatar
NVIDIA/k8s-device-plugin

NVIDIA/k8s-device-plugin: exposing GPUs to Kubernetes without touching kubelet

NVIDIA device plugin for Kubernetes

3,885 stars871 forksGoApache-2.0

At a glance

What is it?
The official NVIDIA device plugin is a DaemonSet that advertises node GPUs to kubelet and keeps their health visible. It is the plumbing layer, not the whole GPU platform, and the README is explicit about what it does not do.
Who is it for?
Adopt it if you already run the NVIDIA driver and the NVIDIA Container Toolkit on your GPU nodes and you want kubelet to see GPUs as schedulable resources; the plugin is the official implementation and the README states support is only provided for it, not for forks. Do not adopt it expecting health checking or GPU cleanup, because the README lists both as lacking.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the NVIDIA device plugin actually solves

Kubernetes has no native concept of a GPU. The device plugin framework exists so a vendor can register a resource name with kubelet and report how many of that resource a node has. NVIDIA's implementation is the DaemonSet that does this for GPUs: the README says it exposes the number of GPUs on each node, keeps track of their health, and lets you run GPU enabled containers in the cluster.

The audience is narrow and specific. You are running Kubernetes on nodes with NVIDIA hardware, you have already installed the driver and the NVIDIA Container Toolkit, and you want a Pod to request a GPU the way it requests memory. If you are still choosing a driver stack or a container runtime, this project is downstream of those decisions and cannot make them for you.

One detail worth reading twice: the README says the NVIDIA device plugin API is beta as of Kubernetes v1.10. That is a statement about the API surface the plugin implements, not about the plugin's release cadence.

How the plugin registers GPUs with kubelet

The repository is a Go module, github.com/NVIDIA/k8s-device-plugin, and go.mod pins the pieces that matter: go-nvml for talking to the driver, go-gpuallocator for choosing which device to hand out, go-nvlib for device discovery, nvidia-container-toolkit, and the Kubernetes client libraries. The plugin does not shell out to nvidia-smi to build its inventory; NVML is the interface.

As of v0.15.0 the repository also holds the GPU Feature Discovery implementation, which the README points to separately. That matters for deployment topology: you can run the plugin alone and get schedulable GPUs, or add gpu-feature-discovery to get node labels describing the hardware. The helm chart documents both, including a standalone mode for gpu-feature-discovery.

Sharing is configured, not automatic. The README covers CUDA time-slicing and CUDA MPS as two ways to give more than one workload access to the same physical GPU. Time-slicing is the coarser option; MPS is the one that changes how concurrent CUDA contexts behave. Both are opt-in through the plugin's configuration, and both change the arithmetic of what a node can schedule.

Prerequisites before you install anything

The README puts the prerequisites first, and they are the part people skip. You need NVIDIA drivers around 384.81 or newer, nvidia-docker 2.0 or newer or nvidia-container-toolkit 1.7.0 or newer (1.11.0 or newer if you are using integrated GPUs on Tegra-based systems), nvidia-container-runtime configured as the default low-level runtime, and Kubernetes 1.10 or newer.

The README assumes the driver and the toolkit are pre-installed and states that the nvidia-container-runtime must already be the default low-level runtime. The plugin does not install any of that for you. If it is missing, the DaemonSet will come up and find nothing to advertise.

If the nvidia runtime is not set as the default, the README says a RuntimeClass needs to be defined instead. That YAML is short, and it is the difference between a Pod that starts and one that fails to find the driver.

Installing the plugin and running a first GPU job

The README's quick start begins on the GPU nodes. The CRI-O path is the most concrete example it gives: a config file under /etc/crio/crio.conf.d/99-nvidia.conf sets nvidia-container-runtime as the default low-level OCI runtime, taking priority over the default crun config at /etc/crio/crio.conf.d/10-crun.conf. The README notes this file can be generated rather than hand-written.

shell
sudo nvidia-ctk runtime configure --runtime=crio --set-as-default --config=/etc/crio/crio.conf.d/99-nvidia.conf

The README also notes that CRI-O uses crun as its default low-level OCI runtime, so crun needs to be added to the runtimes of the nvidia-container-runtime in /etc/nvidia-container-runtime/config.toml. After applying runtime configuration changes, the README says to restart each runtime.

If the nvidia runtime is not the cluster default, define a RuntimeClass and reference it from the Pod:

yaml
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
  name: nvidia
handler: nvidia

The README's deployment section covers helm. The chart is documented with a ConfigMap for passing plugin configuration, and the README describes three patterns: a single config file, multiple config files, and per-node configuration updated with a node label. The chart also has a value for deploying gpu-feature-discovery alongside the plugin, and a standalone mode for that component.

The README also documents a helm install that takes a direct URL to the helm package, for readers who do not want to add a chart repository first. Once the DaemonSet is running and the node has been prepared, the README's final quick start step is running GPU jobs: a Pod requests the NVIDIA GPU resource and the plugin's allocator picks a device. What you should see is the Pod scheduled onto a GPU node with a device assigned, rather than sitting Pending because the resource is unknown to the scheduler.

Health checking and cleanup are explicitly missing

The README is unusually direct about the gaps. It states that the plugin is currently lacking comprehensive GPU health checking features and GPU cleanup features. Read that as a boundary on what you can build on top of it. If a GPU fails in a way the plugin does not detect, the node may keep advertising a device that workloads cannot use, and the plugin is not the component that will quarantine it.

There is a second limitation in the same list: support will only be provided for the official NVIDIA device plugin, and not for forks or other variants. That is a support policy rather than a technical constraint, but it changes the calculus if your organization patches the plugin locally.

The API status is the third thing to internalize. The README says the device plugin API is beta as of Kubernetes v1.10. For a component that sits directly under your scheduler, that is a reason to pin versions and read the changelog before upgrading the cluster, not a reason to avoid the project.

Where this fits against the GPU Operator and other device plugins

The most common alternative in practice is the NVIDIA GPU Operator, which is a different level of abstraction: it manages the driver, the container toolkit, the plugin, and the feature discovery components as one lifecycle. If your nodes are provisioned fresh and you want the whole stack installed and upgraded together, the Operator covers more ground. If your driver and toolkit are already managed by your image pipeline or your node provisioning system, adding the Operator duplicates that management, and the standalone plugin is the smaller thing to run.

The other category is device plugins for other hardware. ROCm has its own Kubernetes device plugin for AMD GPUs, and there are generic device plugins that expose arbitrary devices through a declarative configuration. Those solve the same kubelet registration problem for different hardware, so they are alternatives in the sense of occupying the same slot, not in the sense of being drop-in replacements. An NVIDIA plugin and a ROCm plugin on the same cluster would advertise different resource names, and a workload written for one will not schedule against the other.

One clarification the naming invites: this is not the same as the generic device plugin project, nor is it an RDMA device plugin. It advertises NVIDIA GPUs and, through the included GPU Feature Discovery component, the labels that describe them.

Licence, upgrades and what maintenance looks like

The repository is Apache-2.0. That is a permissive licence, and the repository carries the usual governance and contribution files alongside it, including GOVERNANCE.md, CONTRIBUTING.md and SECURITY.md. Apache-2.0 includes an express patent grant, which is a detail worth confirming with your own counsel rather than taking from an article.

The README has a versioning section and a section on upgrading Kubernetes with the device plugin, which is the practical signal that cluster upgrades and plugin upgrades are treated as a coupled operation. Release history runs from v0.19.3 in June 2026 through v0.20.0 in August 2026 to v0.20.1 in September 2026, and the repository's last push was on 2026-09-23. The upgrade cost is not the plugin binary; it is the coordination between the plugin version, the Kubernetes version, and the node-level driver and toolkit versions, which the README's prerequisites tie together.

Editorial conclusion

Adopt it if you already run the NVIDIA driver and the NVIDIA Container Toolkit on your GPU nodes and you want kubelet to see GPUs as schedulable resources; the plugin is the official implementation and the README states support is only provided for it, not for forks. Do not adopt it expecting health checking or GPU cleanup, because the README lists both as lacking. Before rollout, verify that your driver is at least 384.81, that your Kubernetes version is at least 1.10, and that nvidia-container-runtime is configured as the default low-level runtime on every GPU node, since the plugin assumes that work is already done.

Frequently asked questions

What is the NVIDIA device plugin for Kubernetes?

It is a DaemonSet that exposes the number of GPUs on each node, tracks their health, and lets GPU enabled containers run in the cluster. The README describes it as NVIDIA's official implementation of the Kubernetes device plugin interface, and as of v0.15.0 the repository also contains the GPU Feature Discovery implementation.

How do I install the NVIDIA device plugin with helm?

The README documents deployment via helm, including passing configuration to the plugin through a ConfigMap with single, multiple, or per-node config file patterns. It also documents a helm install that takes a direct URL to the helm package, and chart values for deploying gpu-feature-discovery alongside the plugin or in standalone mode.

What are the prerequisites for running the NVIDIA device plugin?

The README lists NVIDIA drivers around 384.81 or newer, nvidia-docker 2.0 or newer or nvidia-container-toolkit 1.7.0 or newer (1.11.0 or newer for integrated GPUs on Tegra-based systems), nvidia-container-runtime configured as the default low-level runtime, and Kubernetes 1.10 or newer.

Does the NVIDIA device plugin do GPU health checking?

The README states that the plugin is currently lacking comprehensive GPU health checking features, and lists GPU cleanup features as missing as well. Health tracking is described as one of the things it does, but the README does not claim comprehensive checking.

Can multiple workloads share one GPU with the NVIDIA device plugin?

Yes, through configuration. The README documents shared access to GPUs with CUDA time-slicing and with CUDA MPS as two separate approaches.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. NVIDIA/k8s-device-plugin on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-k8s-device-plugin.svg)](https://hysenlabs.com/projects/nvidia-k8s-device-plugin)