# k8sgpt-operator: running K8sGPT analysis as a Kubernetes controller

> The k8sgpt-operator turns K8sGPT from a CLI you run by hand into a K8sGPT custom resource that scans your cluster and writes Result objects. The install is short, the auto-remediation feature is alpha and opt-in, and the README is thinner than the CRD surface suggests.

**k8sgpt-ai/k8sgpt-operator** — Automatic SRE Superpowers within your Kubernetes cluster

- Repository: https://github.com/k8sgpt-ai/k8sgpt-operator
- Website: https://k8sgpt.ai
- Stars: 485 · Forks: 145
- Language: Go
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/k8sgpt-ai-k8sgpt-operator

## The gap between running K8sGPT once and watching a cluster

K8sGPT is a scanner. You point it at a cluster, it inspects workloads, and it asks a model to explain what it finds. Run as a CLI, that is a snapshot: the analysis happens when a human types the command, and the output lands on that human's terminal. Nothing about the cluster changes when the terminal closes.

The k8sgpt-operator exists to close that gap. It installs a controller that reconciles a custom resource of kind K8sGPT, and that resource defines the behaviour and scope of a managed K8sGPT workload. Instead of a person deciding when to scan, the operator owns a Deployment that performs the analysis, and the findings are written back into the cluster as Result objects. The audience is platform and SRE teams who already accept K8sGPT's explanations as useful and want them to accumulate somewhere queryable, not scroll past in a shell.

The multi-cluster story is the second half of that audience. The README describes running the operator in a Cluster API management cluster and pointing it at provisioned clusters through a kubeconfig Secret, so one operator instance can watch a fleet without installing anything on the seed clusters.

## How the K8sGPT custom resource drives a managed Deployment

The operator is a Go controller built on sigs.k8s.io/controller-runtime, which the go.mod confirms, and it follows the usual pattern: a CRD in api/, a reconciler in internal/ or pkg/, and a manager binary in cmd/main.go. The Dockerfile builds exactly that binary and ships it on gcr.io/distroless/static:nonroot, running as user 65532, so the controller image carries no shell.

What makes the design interesting is where the analysis actually runs. The operator does not scan workloads itself. It creates a Deployment from the image named in spec.repository and spec.version, and that Deployment runs K8sGPT. The custom resource is therefore a description of a workload, not the workload. Settings like the AI backend, model, language, anonymisation flag and cache behaviour are passed down to that managed process.

Findings come back as Result objects in the operator namespace. The README's example output shows a Result whose spec.details contains a model-written explanation of missing service endpoints, which tells you the shape: the cluster stores prose, not just a status code. Sinks extend that, with a Slack webhook type documented in the sample spec, and the sample also shows an optional Trivy integration and a filter list for narrowing which resource kinds get analysed.

## Installing the k8sgpt-operator Helm chart and reading your first Result

The README gives a three-command Helm install. The chart lives in the k8sgpt repository on charts.k8sgpt.ai, and the install creates the k8sgpt-operator-system namespace if it does not exist.

```bash
helm repo add k8sgpt https://charts.k8sgpt.ai/
helm repo update
helm install release k8sgpt/k8sgpt-operator -n k8sgpt-operator-system --create-namespace
```

After that, the operator needs an AI backend. The sample uses OpenAI, and the API key goes into a Secret in the same namespace. The variable name matches what the Makefile expects for its own test flow.

```bash
kubectl create secret generic k8sgpt-sample-secret --from-literal=openai-api-key=$OPENAI_TOKEN -n k8sgpt-operator-system
```

Then you apply a K8sGPT object. This is the minimal shape from the README, with the optional keys commented out in the original. Note that spec.version pins the K8sGPT image tag separately from the operator's own version, so upgrading the operator does not silently move the scanner.

```yaml
apiVersion: core.k8sgpt.ai/v1alpha1
kind: K8sGPT
metadata:
  name: k8sgpt-sample
  namespace: k8sgpt-operator-system
spec:
  ai:
    enabled: true
    model: gpt-4o-mini
    backend: openai
    secret:
      name: k8sgpt-sample-secret
      key: openai-api-key
  noCache: false
  repository: ghcr.io/k8sgpt-ai/k8sgpt
  version: v0.4.32
```

Once the object is accepted, the operator creates the K8sGPT Deployment, and after some minutes the README says Results appear if the cluster has problems worth reporting. The command it shows for reading them is `kubectl get results -n k8sgpt-operator-system -o json | jq .`. An empty list is a legitimate outcome: a healthy cluster produces nothing to explain.

## Watching remote clusters through a kubeconfig Secret

For the fleet case, the K8sGPT spec accepts a kubeconfig block with a Secret name and key. The README's Cluster API example uses a Secret named capi-quickstart-kubeconfig with data key value, which is the naming convention CAPI produces. The operator then builds the K8sGPT Deployment using that remote kubeconfig, so the scanner talks to the seed cluster rather than the cluster it runs in.

The README flags its own security problem here, and it is worth repeating plainly: the kubeconfig Cluster API generates is bound to the admin user with cluster-admin permissions. The project's own advice is that a least-privilege setup needs a different kubeconfig. If you follow the example literally, you are handing a workload that calls a third-party AI API a cluster-admin credential. That is a real design tension, not a documentation nit, and the README does not describe a narrower role to use instead.

The rest of the multi-cluster flow is left implicit. The README stops after stating that the operator will create the Deployment using the seed cluster kubeconfig, so questions about what happens when that kubeconfig rotates, or how Results are namespaced across many watched clusters, are not answered in the documentation available.

## Auto-remediation is alpha, opt-in, and narrower than the headline

The most prominent part of the README is the auto-remediation notice. The operator can repair selected Kubernetes workload image failures automatically. The described flow is: an AI provider analyses an ImagePullBackOff, the operator generates a one-image repair, Kubernetes rolls out the corrected Deployment, and the Mutation object reaches Successful.

The safety model is the interesting part, and it is stated clearly. The model never receives write authority. The operator re-fetches the object, computes the semantic JSON patch itself, permits exactly one approved image path, dry-runs the change against the API server, and rejects proposals that are stale, broad or unsafe. Findings on owned Pods are repaired through the owning workload, so the change lands in desired state rather than on an ephemeral Pod. That is a sensible answer to the obvious objection about letting a language model mutate a cluster.

Two constraints matter for adoption. First, the README calls this deliberately alpha and opt-in, and points to AUTO_REMEDIATION.md for the policy model, supported workload types, audit fields and operational limits. Second, the scope is image failures on selected workload types, not general remediation. If your incidents are OOM kills, misconfigured probes or DNS failures, this feature does not address them. The README also says provider setup is independent of the remediation policy, so the same safeguards apply regardless of backend.

## Where k8sgpt-operator is the wrong choice

If you want an explanation of a cluster only when you ask for one, the operator is overhead. It runs a Deployment continuously, and every analysis cycle that finds a problem may call your model provider. A developer running the K8sGPT CLI against a staging cluster pays for exactly the tokens they request. The operator's model is ambient: the cluster is watched, and findings accumulate whether or not anyone is looking.

Cost control is the operator's weak point in the documentation. The spec exposes a backOff block with maxRetries, and noCache can be set to false, but the README does not document a scheduling interval, a budget cap, or a way to stop analysis during known maintenance windows. You are trusting the managed K8sGPT workload's own behaviour, configured through a CRD whose fields are mostly shown as commented-out examples rather than described.

There is also a dependency question. The operator is a thin controller around an image tag you choose. If the upstream K8sGPT image changes its flags or output format, the operator's assumptions may not hold, and the README does not describe a compatibility matrix between operator releases and K8sGPT versions. Pinning spec.version is the only lever the documentation gives you.

## How it differs from GitOps health tools and from K8sGPT in CI

Argo CD is the comparison people reach for, and the difference is what each one considers a problem. Argo CD reconciles declared manifests against live state and reports drift or sync failure. It answers whether the cluster matches the repository. The k8sgpt-operator asks a model to read the cluster and explain why something is unhealthy, which covers ground that has no manifest to compare against: a Service with no endpoints, an image that will not pull, a workload stuck in a state nobody declared.

The other alternative is running K8sGPT in a pipeline. A CI job can run the CLI on a schedule, capture the output, and fail a build or post to Slack. That keeps the analysis outside the cluster and inside tooling your team already audits. What it cannot do is see a live cluster continuously, and it cannot write Result objects that other controllers could consume. The operator's value is that findings become Kubernetes API objects with a known kind and API group, which is a different integration surface from a log line.

Choosing between them is mostly about whether you want the scanner inside the blast radius of the cluster it inspects. The operator says yes, and accepts the credential exposure that follows.

## Licence, release cadence and the cost of upgrading

The repository is Apache-2.0, and the Dockerfile labels the published image with the same identifier. For most users that means permissive use with an explicit patent grant and the usual notice obligations. It says nothing about the licence of the AI provider you configure, and the operator ships no model. Your provider's terms govern the data you send it, which is why the spec exposes an anonymized flag and a language setting. If your cluster data cannot leave your network, the README's example configuration is not for you, and the documentation does not describe a local backend option.

On cadence, the last push to the repository was on 2026-09-14, and the most recent tagged release in the list is v0.2.29 from 2026-08-13. The gap between v0.2.27 in March 2026 and v0.2.28 and v0.2.29 on the same day in August suggests releases cluster around feature merges rather than a fixed schedule. The repository uses release-please, visible in release-please-config.json and the manifest file, so version bumps are automated from commit messages.

Upgrade cost has two layers. The operator itself upgrades through Helm, and the CRD may change between minor versions, which is the usual source of friction. The K8sGPT image is pinned separately in spec.version, so you can move the operator without moving the scanner, and vice versa. The README does not document rollback behaviour for either, so a failed upgrade path is something you would have to work out from the changelog.

## Conclusion

Adopt it if you already want K8sGPT findings to land as Kubernetes objects, either on the cluster it runs in or on a remote cluster whose kubeconfig you supply in the spec. Do not adopt it for auto-remediation yet: the README calls that path alpha and opt-in, and it is limited to selected workload image failures. Before enabling anything, verify your AI backend and Secret are wired the way the sample manifest expects, confirm the K8sGPT image version in spec.version is the one you intend to run, and read AUTO_REMEDIATION.md for the policy model and operational limits, since the README itself tells you to do that before turning it on.

## FAQ

### What is K8sGPT?

K8sGPT is the scanner the operator manages. It inspects a Kubernetes cluster and asks an AI backend to explain the problems it finds, and the operator runs it as a Deployment driven by a K8sGPT custom resource.

### What does a Kubernetes operator do?

It reconciles a custom resource into the workloads and objects that resource describes. The k8sgpt-operator watches K8sGPT objects and creates the K8sGPT Deployment plus the Result objects that hold the findings.

### Which AI is best for Kubernetes?

The documentation does not rank providers. It shows openai with model gpt-4o-mini in the sample and states that provider setup is independent of the remediation policy, so the same safeguards apply to every backend.

### Is Argo CD a Kubernetes operator?

Argo CD reconciles declared manifests against live state and reports drift or sync failure. The k8sgpt-operator instead asks a model to explain why a resource is unhealthy, which covers cases with no manifest to compare against.

## Sources

- [k8sgpt-ai/k8sgpt-operator on GitHub](https://github.com/k8sgpt-ai/k8sgpt-operator)
- [License: Apache-2.0](https://github.com/k8sgpt-ai/k8sgpt-operator/blob/main/LICENSE)
- [Project website](https://k8sgpt.ai)
- [README](https://github.com/k8sgpt-ai/k8sgpt-operator/blob/main/README.md)
- [Releases](https://github.com/k8sgpt-ai/k8sgpt-operator/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/k8sgpt-ai-k8sgpt-operator
