Self-hosted service
kubernetes/node-problem-detector avatar
kubernetes/node-problem-detector

node-problem-detector: turning node faults into Kubernetes events and conditions

This is a place for various problem detectors running on the Kubernetes nodes.

3,469 stars704 forksGoApache-2.0

At a glance

What is it?
node-problem-detector is a per-node daemon that watches kernel logs, systemd units, kubelet health and user scripts, then reports what it finds as Node conditions and Events. This is how its monitors, exporters and config files fit together, and where the limits are.
Who is it for?
Adopt node-problem-detector if you run Kubernetes on hardware or kernels you do not fully control and want node faults visible as Node conditions and Events rather than only in local logs. Skip it if you only need metric scraping of kubelet endpoints, since its SystemStatsMonitor is explicitly a metrics collector and its condition output has no remedy loop of its own.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Go, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap node-problem-detector fills on a Kubernetes node

Kubernetes schedules pods based on what the control plane knows about a node. A kernel deadlock, a read-only filesystem, a dead NTP daemon or an unresponsive container runtime are not part of that picture. The scheduler keeps placing pods on a machine that is degrading, and the first signal anyone gets is usually a workload failing. The README lists this exact set of cases: infrastructure daemon issues, hardware faults, kernel problems and container runtime problems, and states that these are invisible to the upstream layers of the cluster management stack.

node-problem-detector is the daemon that closes that loop. It runs on each node, either as a DaemonSet or standalone, and reports what it finds to the API server. The README notes it ships as a Kubernetes addon enabled by default in GKE and is enabled by default in AKS as part of the AKS Linux Extension. That distribution detail matters for evaluation: on those two clouds you are likely tuning an existing install, not introducing a new component.

The intended audience is platform and cluster operators who own the node layer. If your nodes are fully managed and you have no ability to change what runs on them, the detector's value to you is mostly in reading the conditions it already emits.

How NodeCondition and Event map to problem severity

The reporting model is deliberately split in two. The README states that a permanent problem which makes the node unavailable for pods should be reported as a NodeCondition, while a temporary problem with limited impact on pods but informative value should be reported as an Event. That distinction is the design decision everything else hangs off.

A NodeCondition is durable state attached to the Node object. Tools that already watch node status, such as autoscalers or custom controllers, can react to it without any new integration. An Event is a timestamped record that ages out. A flapping NTP check is informative but not a reason to cordon a machine, so it belongs in Events. A read-only root filesystem is a reason to stop scheduling, so it belongs in a condition. Getting this classification wrong in your own config is the most common way to make the detector either noisy or useless, and the choice is yours to make per rule, not something the daemon infers.

Problem daemons: log rules, stats, plugins and health checks

The detector is a host process with several sub-daemons, called problem daemons, running as goroutines inside the same binary. The README says the plan is to separate them into different containers composed through pod specification, but that is not the current state.

SystemLogMonitor is the core. It reads system logs and applies predefined rules, and it is the source of the KernelDeadlock, ReadonlyFilesystem, FrequentKubeletRestart, FrequentDockerRestart and FrequentContainerdRestart conditions. Its configs are split by input path: kernel-monitor-filelog.json, kernel-monitor.json for kmsg, kernel-monitor-counter.json, abrt-adaptor.json and systemd-monitor-counter.json. That split is worth reading before deployment, because a filelog config and a kmsg config are not interchangeable inputs.

SystemStatsMonitor collects health-related system stats as metrics and, per the README, produces no NodeCondition today, with the note that conditions could be added later. CustomPluginMonitor runs user-defined check scripts, with NTPProblem given as an existing example and custom-plugin-monitor.json as the example config. HealthChecker covers kubelet and container runtime health, producing KubeletUnhealthy and ContainerRuntimeUnhealthy, with separate config files for kubelet, docker and containerd.

Each category can be removed at compile time with a build tag: disable_system_log_monitor, disable_system_stats_monitor, disable_custom_plugin_monitor. The README is explicit that this trims build dependencies, global variables and background goroutines out of the binary. For an image you build yourself, that is a real size and attack-surface lever. For the stock image, it is not available to you.

Exporters decide where a detected problem actually lands

Detection and reporting are separate components. The Kubernetes exporter writes temporary problems as Events and permanent ones as Node Conditions on the API server. The Prometheus exporter exposes node problems and metrics locally as Prometheus metrics instead of, or in addition to, writing to the API server. The Stackdriver exporter sends problems and metrics to the Stackdriver Monitoring API and can be compiled out with disable_stackdriver_exporter.

This is where deployments diverge. A cluster that only runs the Kubernetes exporter has no local metric endpoint to scrape, so a Grafana dashboard has nothing to read. A cluster that only runs the Prometheus exporter leaves the control plane blind, and the scheduler keeps placing pods on the failing node, which defeats the stated purpose of the project. The practical configuration is both, with the Kubernetes exporter carrying the conditions that controllers act on and the Prometheus exporter carrying the counters that humans graph.

Installing node-problem-detector and a first CustomPluginMonitor check

The repository ships a Dockerfile whose entrypoint already runs the detector against two config files, kernel-monitor.json and readonly-monitor.json. The builder stage runs make bin/node-problem-detector bin/health-checker bin/log-counter, and the runtime image is debian-base with util-linux, bash and libsystemd-dev installed. The Makefile's default REGISTRY is gcr.io/k8s-staging-npd, and PLATFORMS covers linux_amd64, linux_arm64 and windows_amd64.

To build the binaries locally, the Makefile targets are the entry point:

bash
make build-binaries

Running the detector directly on a node requires at least one monitor config. The flag takes a comma-separated list of paths, and the README gives this example shape:

bash
./node-problem-detector --config.system-log-monitor=/config/kernel-monitor.json,/config/readonly-monitor.json

Node name resolution is worth knowing before you debug anything. The README states the detector takes the node name first from --hostname-override, then from the NODE_NAME environment variable, and finally falls back to os.Hostname. A DaemonSet that does not set NODE_NAME correctly will report conditions against the wrong Node object.

For a custom check, the CustomPluginMonitor invokes a script you supply and maps its result to a condition. The example config to copy from is config/custom-plugin-monitor.json, and the NTPProblem condition is the worked example in the README's table. The detector's own version flag is the fastest sanity check after install:

bash
./node-problem-detector --version

Where node-problem-detector is the wrong tool

The detector reports. It does not remediate. The README describes the reporting path and points at a remedy system as a separate discussion, which means nothing in this daemon drains a node, restarts a service or evicts a pod. If your requirement is automatic recovery, you are building that layer yourself on top of the conditions.

Second, its coverage is bounded by the monitors that exist. SystemStatsMonitor emits no NodeCondition at all today. If your problem class is not in the SystemLogMonitor rule set, the HealthChecker set, or a script you write for CustomPluginMonitor, the detector will not see it. The README's own list of problem categories is the ceiling of what is addressed out of the box.

Third, log-based detection is inherently a parsing exercise. Rules match against log text, so a kernel or systemd version that changes its message format can silently stop matching. That is a maintenance obligation, not a one-time install.

Finally, if you run nodes you cannot modify, or a managed platform that already installs the detector for you, adding another copy is redundant and risks duplicate Events and conflicting condition writes.

node-problem-detector compared with a generic log shipper

The obvious alternative is shipping node logs to a central store with a general-purpose collector and alerting there. The difference is structural, not a matter of feature checklists. A log shipper moves text off the node and leaves interpretation to a query engine downstream; the node stays schedulable the whole time. node-problem-detector interprets on the node and writes the verdict back into the Kubernetes API as a condition or an event, so the control plane itself can see it.

That difference sets the cost profile. A log pipeline scales with log volume and needs a backend. node-problem-detector scales with node count and needs no storage tier, but it only answers the questions its rules ask, and the answers live in the API server rather than in a searchable index. Teams that need forensic search across months of kernel output still need the log pipeline. Teams that need the scheduler and their controllers to know a node is unhealthy need this daemon.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived and the last push was on 2026-09-21. Three releases landed on 2026-07-11: v1.36.0, v1.35.3 and v1.34.4. The parallel patch releases on older minors suggest backports are maintained rather than only the newest line, which matters if you pin versions across a fleet.

The build toolchain is pinned hard. go.mod declares go 1.26.7 and k8s.io/api, k8s.io/apimachinery and k8s.io/client-go at v0.37.0, and the Dockerfile builder stage uses golang:1.26.7-bookworm. Upgrading the detector therefore tends to track Kubernetes client library versions, not just the detector's own changelog. The Dockerfile also documents that the builder-base stage can be overridden with docker buildx's --build-context flag for users who need an older OS to avoid depending on recent glibc versions, which is a hint that glibc compatibility has been a real constraint for some builders.

The licence is Apache-2.0. For most users that is a permissive licence with a patent grant and a requirement to preserve notices. If you modify and redistribute the detector, or embed it in a product, read the LICENSE and NOTICE handling yourself; this is not legal advice and the repository's own terms govern.

Editorial conclusion

Adopt node-problem-detector if you run Kubernetes on hardware or kernels you do not fully control and want node faults visible as Node conditions and Events rather than only in local logs. Skip it if you only need metric scraping of kubelet endpoints, since its SystemStatsMonitor is explicitly a metrics collector and its condition output has no remedy loop of its own. Before rolling it out, verify three things: that your nodes expose the log paths your chosen config files expect, that the NodeCondition names it emits do not collide with conditions your controllers already consume, and that the CustomPluginMonitor scripts you write return the exit codes the plugin monitor expects.

Frequently asked questions

What is node-problem-detector in Kubernetes?

It is a daemon that runs on each node, detects node problems and reports them to the API server, either as a DaemonSet or standalone. It uses NodeCondition for permanent problems that make a node unavailable for pods and Event for temporary, informative problems.

How does node-problem-detector check node health?

Several problem daemons run inside the binary. SystemLogMonitor applies rules to system logs, HealthChecker checks kubelet and container runtime health, SystemStatsMonitor collects health-related system stats as metrics, and CustomPluginMonitor runs user-defined check scripts.

Which NodeConditions does node-problem-detector produce by default?

The README lists KernelDeadlock, ReadonlyFilesystem, FrequentKubeletRestart, FrequentDockerRestart and FrequentContainerdRestart from SystemLogMonitor, plus KubeletUnhealthy and ContainerRuntimeUnhealthy from HealthChecker. CustomPluginMonitor conditions depend on the user's configuration.

Can node-problem-detector run checks I write myself?

Yes. CustomPluginMonitor invokes user-defined check scripts, with NTPProblem given as an existing example and config/custom-plugin-monitor.json as the example configuration file to start from.

Does node-problem-detector fix the problems it finds?

No. It detects problems and reports them as Events and NodeConditions. The README treats a remedy system as a separate discussion, so any eviction, drain or restart logic has to be built on top of the conditions it emits.

How do I build node-problem-detector from source?

The Makefile provides the build targets, and its default container registry is gcr.io/k8s-staging-npd. The Dockerfile builds bin/node-problem-detector, bin/health-checker and bin/log-counter in a golang:1.26.7-bookworm builder stage.

Official sources

  1. Issues
  2. kubernetes/node-problem-detector on GitHub
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/kubernetes-node-problem-detector.svg)](https://hysenlabs.com/projects/kubernetes-node-problem-detector)