# Kuberhealthy: Synthetic Checks as Kubernetes Pods, With Prometheus Metrics

> Kuberhealthy is a Kubernetes operator that schedules short-lived checker pods from HealthCheck custom resources and reports results to a status UI, a JSON API and a Prometheus endpoint. It fits teams that want multi-step validation written in any language and shipped as manifests.

**kuberhealthy/kuberhealthy** — A Kubernetes operator for running synthetic checks as pods. Works great with Prometheus!

- Repository: https://github.com/kuberhealthy/kuberhealthy
- Website: https://kuberhealthy.github.io/kuberhealthy/
- Stars: 2,269 · Forks: 295
- Language: Go
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/kuberhealthy-kuberhealthy

## What Kuberhealthy solves that probes do not

A liveness probe answers one question: is this container still working? It cannot log in, create a record, verify the record, and delete it. Kuberhealthy exists for that second class of test. The README describes it as an operator for synthetic monitoring and continuous validation, and the checks it runs are ordinary pods, so the validation can be a multi-step workflow in Go, Python, Rust, bash, or anything that fits in a container.

The audience is narrow and specific. If your team already writes smoke tests and runs them from CI on a schedule, Kuberhealthy moves that logic inside the cluster and gives it a Kubernetes-native lifecycle: a HealthCheck object, a run interval, a timeout, and a result. The README states that checks are Kubernetes manifests, which is the whole pitch. You ship the test beside the application it exercises, in the same repository, with the same review process.

## The HealthCheck CRD, the controller and the reporting path

The mechanism is visible in the README's architecture diagram. Kuberhealthy provides the HealthCheck custom resource definition. The controller watches those objects and schedules a short-lived checker pod for each one. When the pod finishes its validation, it reports back to the controller with a POST to /check, and the result then flows to three consumers: the built-in status UI, a JSON API, and the Prometheus metrics endpoint.

That reporting step is why the check clients exist. The README says the check client handles KH_REPORTING_URL, KH_RUN_UUID, and deadline enforcement automatically. In other words, the controller injects the address to report to and an identifier for the run, and the client library wraps the call. A Go check calls checkclient.ReportSuccess() or checkclient.ReportFailure() with a slice of error strings. If you write a check in a language without a client, you are responsible for making the same HTTP call yourself before the deadline.

The scheduling model is deliberately simple. Each HealthCheck carries a runInterval and a timeout in its spec, plus a podSpec describing the container to run. The README's deployment example uses runInterval: 10m and timeout: 5m. Two numbers, one container, one result. There is no dependency graph between checks and no notion of a check that triggers another check.

## Installing Kuberhealthy and applying a first check

The README gives three install paths and calls Helm the recommended one. All three assume you have a cluster and kubectl pointed at it. The Helm command creates the kuberhealthy namespace if it does not exist.

```bash
helm install kuberhealthy deploy/helm/kuberhealthy -n kuberhealthy --create-namespace
```

The Kustomize path pulls the base manifests straight from the repository at the main ref, which means you are tracking the branch rather than a tagged release. Pin a ref if that matters to you.

```bash
kubectl apply -k github.com/kuberhealthy/kuberhealthy/deploy/kustomize/base?ref=main
```

Once installed, the README suggests port-forwarding the service to reach the status UI. The service listens on port 80 inside the cluster, and the example maps it to 8080 locally.

```bash
kubectl -n kuberhealthy port-forward svc/kuberhealthy 8080:80
```

Open http://localhost:8080 and the status UI should render. From there the README points you at docs/CHECKS_REGISTRY.md to apply an existing check, or docs/CHECK_CREATION.md to build your own. A first check is a single manifest. This one is the built-in deployment check, which creates a test deployment, rolls it out, and tears it down on a schedule.

```yaml
apiVersion: kuberhealthy.github.io/v2
kind: HealthCheck
metadata:
  name: deployment
  namespace: kuberhealthy
spec:
  runInterval: 10m
  timeout: 5m
  podSpec:
    spec:
      containers:
        - name: deployment
          image: docker.io/kuberhealthy/deployment-check:v0.1.1
          env:
            - name: CHECK_DEPLOYMENT_REPLICAS
              value: "4"
```

Apply it and the README shows what to expect: kubectl apply prints that the healthcheck was created, and kubectl get healthcheck lists the check with its namespace, last run, age and an OK column. You can shorten the resource name to hc. From that point the check appears in the UI, in the JSON response at /json, and as metrics.

## Reading results from Prometheus and the JSON API

The metrics endpoint is where Kuberhealthy earns its keep if you already run Prometheus. The README lists four series: kuberhealthy_check with a status label, kuberhealthy_check_duration_seconds, and kuberhealthy_check_success_total, all labelled by check and namespace. That label pair is the join key back to the HealthCheck object, so an alert rule can name a specific check without guessing.

The JSON API is a different shape. A GET to /json returns a top-level ok boolean and a checks map keyed by namespace and check name, with per-check fields for ok, errors, lastRun and runDuration. That structure is convenient for a script that needs the failure text, which the metrics endpoint does not carry. Note that the README does not document authentication on either endpoint, so treat the service as internal until you have confirmed otherwise.

## Where Kuberhealthy is the wrong tool

The cost model is the first limitation. Every check is a pod. A check with runInterval: 10m produces 144 pod starts a day, and the README's deployment example asks for 4 replicas and a rolling update inside that pod. On a small cluster, a handful of frequent checks is real API server and scheduler load, and the checker pods compete for the same capacity as your workloads. Kuberhealthy is a poor fit for a cluster that is already tight on resources or on a control plane with strict rate limits.

The second limitation is that Kuberhealthy observes, it does not repair. Nothing in the README describes remediation, retries beyond the run interval, or rollback. A failing check produces a failed status and a metric, and the response is whatever your alerting does with it.

The third is the reporting contract. A check that crashes, hangs, or exits without calling the reporting endpoint is handled by the timeout, not by a partial result. If your validation logic has a long tail, the timeout is the only knob, and the README does not document what happens to a check that reports success and then keeps running. Finally, if all you need is to know whether a pod is alive, this is a large amount of machinery for a question a readiness probe already answers.

## How it compares to Prometheus Blackbox Exporter and CI smoke tests

The closest alternative for synthetic monitoring in a Kubernetes cluster is the Prometheus Blackbox Exporter. It probes endpoints from outside, over HTTP, TCP, ICMP or gRPC, and it needs no operator, no CRD and no checker pod per run. The difference in approach is fundamental: Blackbox Exporter measures reachability and response from a prober's point of view, while Kuberhealthy runs your code inside the cluster with its own service account, its own environment variables and its own deadline. If your test is a login, a write, a read-back and a cleanup, Blackbox Exporter cannot express it. If your test is an HTTP GET against /health and nothing more, Blackbox Exporter does it with far less to operate.

The other alternative is running the same smoke tests from CI on a schedule. That keeps the test out of the cluster, which is simpler to reason about, but it means the test's network path and credentials differ from what runs in production, and a failure appears in a pipeline rather than next to the workload.

## Upgrades, the v2 to v3 question, and the Apache-2.0 licence

Maintenance is current: the last push to the default branch was on 2026-08-26, the same day as the v3.0.14 release, and the repository is not archived. Releases arrive frequently, with v3.0.12, v3.0.13 and v3.0.14 all dated 2026-08-20 or later.

One detail deserves attention before you install. The README's HealthCheck example declares apiVersion: kuberhealthy.github.io/v2, while the Go module path in go.mod is github.com/kuberhealthy/kuberhealthy/v3 and the repository root carries a MIGRATING_TO_V3.md file. The README does not explain the mismatch, so read that migration document and confirm the correct apiVersion for your installed version rather than copying the example blindly. The rest of the dependency set in go.mod is unremarkable: controller-runtime v0.25.1 and the k8s.io libraries at v0.37.0, on Go 1.26.0. That pins your upgrade cadence loosely to Kubernetes client library releases, which is the usual cost of building on controller-runtime.

The licence is Apache-2.0, and the repository includes a NOTICE file, which Apache-2.0 expects you to preserve when redistributing. If you fork the operator or vendor the check clients into your own image, keep the LICENSE and NOTICE intact and check how your organisation handles the patent grant. That is a description of the licence terms, not legal advice.

## Conclusion

Adopt Kuberhealthy if your validation logic is a multi-step workflow that needs a real container, and you want it versioned as a HealthCheck manifest next to the app it tests. Do not adopt it if you only need liveness and readiness probes, or if you want a managed service with no operator to run. Before applying anything, read MIGRATING_TO_V3.md, because the CRD in the README uses apiVersion kuberhealthy.github.io/v2 while the Go module path is v3, and confirm which apiVersion your cluster should use. Then check docs/CHECK_CREATION.md for the environment variables injected into every check pod, since that list is what your checker must respect.

## FAQ

### What exactly is Kubernetes used for?

The repository does not answer this. It assumes you already run a Kubernetes cluster and describes Kuberhealthy as an operator that schedules, tracks, monitors and manages check pods inside it.

### How does Kubernetes do a health check, and how is Kuberhealthy different?

Kubernetes itself uses liveness and readiness probes on a container. Kuberhealthy adds a HealthCheck custom resource that schedules a separate short-lived checker pod, which runs your validation logic and reports the result back to the controller.

### What is the best Kubernetes monitoring tool?

The repository does not make that comparison. It states that Kuberhealthy ships metrics to Prometheus and is meant to be used alongside it, so the question is whether your checks need to run as containers rather than as endpoint probes.

### How can I check if a pod is healthy with Kuberhealthy?

You apply a HealthCheck manifest that describes the checker container to run, then read the result from the status UI, the /json endpoint, or the kuberhealthy_check metric. The README's example uses kubectl get healthcheck, which lists each check with a last run time and an OK column.

## Sources

- [Official documentation](https://kuberhealthy.github.io/kuberhealthy/)
- [Official README](https://github.com/kuberhealthy/kuberhealthy#readme)
- [Project repository](https://github.com/kuberhealthy/kuberhealthy)
- [Release notes](https://github.com/kuberhealthy/kuberhealthy/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kuberhealthy-kuberhealthy
