nvidia_gpu_exporter: Prometheus GPU metrics without the DCGM stack
Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML
At a glance
- What is it?
- A Go exporter that parses nvidia-smi output (or reads NVML directly) to expose NVIDIA GPU metrics to Prometheus. It targets consumer cards, homelabs and restricted setups where DCGM-exporter is a poor fit.
- Who is it for?
- Adopt nvidia_gpu_exporter if you run GeForce or RTX cards, a small Kubernetes cluster, a homelab, or a vGPU or MIG guest where nvidia-smi answers but deeper counters do not. Skip it if you already run datacenter cards under the NVIDIA GPU Operator, where the README points to DCGM-exporter instead.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What nvidia_gpu_exporter solves, and for whom
Prometheus needs a scrape endpoint. NVIDIA GPUs do not have one. The gap is normally filled by DCGM-exporter, which links against NVIDIA's datacenter libraries. That works well on data center cards with the GPU Operator installed, and the README says so directly: if you run datacenter cards on Kubernetes with the GPU Operator already installed, DCGM-exporter is probably the better fit.
Everything else is left out. GeForce and RTX cards expose little through the datacenter tooling. Homelabs, edge boxes and small Kubernetes clusters do not want to install the whole GPU Operator stack just to see utilization. Virtualized guests, MIG slices and locked-down containers often hide the deeper counters entirely, while nvidia-smi still answers. Mixed fleets of old and new cards need one exporter that behaves the same on every machine.
That is the audience this project addresses. It is a side project, and the README says as much in a warning near the top: issues and pull requests may take a long time or never be handled. Weigh that against the alternative before adopting it in a production path.
How the nvidia-smi backend collects and exports metrics
The default backend shells out to the nvidia-smi binary, parses its output, and converts the parsed fields into Prometheus metric families. The README describes this as collecting, parsing and exporting. Because the only dependency is the binary, the exporter runs on Linux, Windows and macOS, on bare metal, in Docker or on Kubernetes, with no C bindings.
Field discovery is automatic. The README lists auto-discovery of the metric fields nvidia-smi can expose as a highlight, which is how the exporter stays compatible with future driver output rather than hard-coding a fixed column set. Two optional behaviors change the data flow. Per-process GPU metrics show which process uses how much GPU memory. Background collection runs nvidia-smi on a timer instead of on every scrape, which matters when the binary is slow or the scrape interval is short.
The exporter does not have to run on the monitored machine. The README states it can be configured to execute the nvidia-smi command remotely. That is a real architectural difference from exporters that must live beside the device, and it is also the point where you take on the cost of making a remote command execution reliable.
Installing nvidia_gpu_exporter and scraping it for the first time
The quick start assumes a Linux machine with the NVIDIA driver and the NVIDIA Container Toolkit already present. The container runs with all GPUs attached, the utility driver capability enabled, and port 9835 published. The image tag is utkuozdemir/nvidia_gpu_exporter:latest.
docker run -d --name nvidia_gpu_exporter --restart unless-stopped \
--gpus all -e NVIDIA_DRIVER_CAPABILITIES=utility -p 9835:9835 \
utkuozdemir/nvidia_gpu_exporter:latest
curl http://localhost:9835/metricsThe curl call is the check that matters. If the container started but the driver capability was not passed through, the metrics endpoint will not carry GPU families. The Dockerfile explains why the image is not built on a static base: the NVIDIA container runtime injects nvidia-smi from the host, and that binary needs a libc and its loader inside the container. It also runs as numeric user 65534:65534, because the kubelet cannot verify runAsNonRoot against a named image user.
If you have no GPU on hand, demo mode serves synthetic metrics on any machine. The README states it simulates two H200 GPUs with fluctuating values, a MIG topology and an XID error history by default.
nvidia_gpu_exporter --collect.backend demoFor Windows, macOS, packages, Kubernetes and running without Docker, the README points to docs/INSTALL.md, which is where the winget instructions live as well.
The experimental NVML backend and what it adds
On Linux the exporter can skip nvidia-smi and read metrics directly from the NVIDIA driver library through NVML. The README makes a specific compatibility promise: every metric the default backend serves stays identical in name, labels and value, so existing dashboards and alerts keep working. That is the right design choice, because it means switching backends is not a dashboard migration.
The NVML backend adds families nvidia-smi cannot provide. The README names per-MIG-instance metrics, XID error counters, a total energy counter and PCIe throughput. The official Grafana dashboards have panels for all of these. On the default backend those panels sit empty; on the NVML backend they fill up.
It ships as its own release flavor. You can take a -nvml archive from the releases page, or use a -nvml image tag.
docker run -d \
--name nvidia_gpu_exporter \
--restart unless-stopped \
--gpus all \
-e NVIDIA_DRIVER_CAPABILITIES=utility \
-p 9835:9835 \
utkuozdemir/nvidia_gpu_exporter:latest-nvmlThe README marks it experimental mainly because it needs more testing across driver versions and GPU generations, and asks users to open an issue about how it went, good or bad. The full backend comparison and current limits live in docs/CONFIGURE.md. Treat the experimental label as a statement about test coverage, not about the metric contract, which the README says is unchanged.
Where nvidia_gpu_exporter is the wrong tool
The clearest failure case is the one the README states itself. If you run datacenter cards on Kubernetes and the GPU Operator is already installed, DCGM-exporter is probably the better fit. Choosing this exporter there means giving up the datacenter metric surface for no gain in simplicity, since the operator is already deployed.
The second constraint is the nvidia-smi dependency on the default backend. Anything that removes or hides the binary breaks collection, and the Dockerfile notes that the binary is injected from the host by the container runtime, so a misconfigured runtime is a silent failure rather than a startup error. Restricted environments are listed as a use case precisely because nvidia-smi still answers there, but that is a property of the environment, not something the exporter can guarantee.
The third is maintenance. The README carries an explicit warning that this is a side project maintained in spare time, and that issues or pull requests may take a long time or not be handled at all. If your deployment needs a vendor-backed support path, this is not it. The NVML backend is additionally marked experimental, so the extra metric families it adds should not be the reason you commit to the project.
DCGM-exporter versus nvidia_gpu_exporter: the actual difference
DCGM-exporter is the alternative the README names, and the difference is in where the metrics come from. DCGM-exporter uses NVIDIA's datacenter GPU management stack. nvidia_gpu_exporter parses the output of the nvidia-smi command line tool, or reads NVML directly on Linux with the experimental backend. One depends on a management layer being present; the other depends on a binary being callable.
That difference drives everything else. DCGM-exporter assumes a datacenter deployment, typically on Kubernetes with the GPU Operator. nvidia_gpu_exporter runs on Windows and macOS as well, needs no C bindings, and can execute nvidia-smi on a remote machine instead of the monitored one. The README frames the split as consumer and prosumer GPUs, small clusters, edge and homelab boxes, virtualized or restricted setups, and mixed fleets on one side, with DCGM-exporter on the other for datacenter cards under the GPU Operator.
The two are not mutually exclusive across a fleet. If you have RTX workstations and a datacenter cluster, the per-host choice can differ. What you should not do is run both on the same host and expect the metric names to line up, since they come from different sources.
Maintenance cost, packaging and licence
The release cadence is visible in the repository: v1.15.1 and v1.15.0 both landed on 2026-09-02, and v1.14.0 on 2026-08-12. The last push to the default branch was on 2026-09-10. That is a recent cadence, but it sits alongside the README's own warning about spare-time maintenance, so plan for slow responses on issues rather than for the absence of releases.
The packaging surface is broad and each channel is a separate upgrade path: a Docker image on Docker Hub, a Helm chart under charts/ published to Artifact Hub, release archives including the -nvml flavor, and a winget package. The go.mod pins Go 1.27.1 and a set of Prometheus client libraries, so building from source tracks those versions. The Dockerfile pins its distroless base image by digest, which means base image updates arrive only when the project rebuilds.
The licence is MIT, as stated in the repository and shown in the README badge. MIT is permissive: it allows use, modification and redistribution with the licence and copyright notice retained. It provides no patent grant and no warranty. That is a description of the licence text, not legal advice; if your organisation has specific obligations around bundled NVIDIA components or container base images, check those separately, since MIT covers this project's code and not the driver or the nvidia-smi binary it calls.
Editorial conclusion
Adopt nvidia_gpu_exporter if you run GeForce or RTX cards, a small Kubernetes cluster, a homelab, or a vGPU or MIG guest where nvidia-smi answers but deeper counters do not. Skip it if you already run datacenter cards under the NVIDIA GPU Operator, where the README points to DCGM-exporter instead. Before rolling it out, verify that nvidia-smi exists at the path the exporter expects on each host, confirm the container runtime passes the utility driver capability, and decide whether the experimental -nvml flavor is worth its extra testing burden.
Frequently asked questions
What does a DCGM exporter do, and how does nvidia_gpu_exporter differ?
DCGM-exporter uses NVIDIA's datacenter GPU management stack, while nvidia_gpu_exporter parses nvidia-smi output or reads NVML directly on Linux. The README says DCGM-exporter is probably the better fit for datacenter cards on Kubernetes with the GPU Operator already installed, and positions nvidia_gpu_exporter for consumer GPUs, small clusters, homelabs and restricted setups.
What is Prometheus' exporter, and is nvidia_gpu_exporter one?
nvidia_gpu_exporter is a Prometheus exporter: the README describes it as a simple exporter that uses the nvidia-smi(.exe) binary to collect, parse and export metrics. It serves those metrics on port 9835 at the /metrics endpoint for Prometheus to scrape.
Is there a shortage of Nvidia GPUs that affects running nvidia_gpu_exporter?
The README does not discuss GPU supply or availability. It does note that you can run the exporter without any GPU by starting it with --collect.backend demo, which serves synthetic metrics including the NVML-only families on any machine.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/utkuozdemir-nvidia-gpu-exporter)