Kubernaut: An AIOps Agent That Investigates Kubernetes Alerts Before It Fixes Them
An AIOps platform that closes the loop from Kubernetes alert to automated remediation: An AI Agent investigates via MCP tools, selects a fix from a pre-seeded workflow catalog and delegates execution (K8s Job, Tekton, Ansible) — or escalates with a full RCA. Approval gates, OPA policies, and audit trails keep humans in control.
At a glance
- What is it?
- Kubernaut is an open-source AIOps platform that turns Kubernetes alerts into investigated, approved, and executed remediations. It uses an LLM agent with native Go bindings to inspect clusters, picks a fix from a workflow catalog, and escalates with a root cause analysis when it cannot act.
- Who is it for?
- Adopt Kubernaut if you operate Kubernetes clusters with a heavy alert load and have a catalog of runbooks that can be turned into Tekton, Job, or Ansible workflows. Its LLM-based investigation adds value where rule-based tools fail on multi-root-cause incidents.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem: Alerts Don't Tell You the Root Cause
Kubernetes operators spend hours triaging the same alerts at 3am, pulling logs and metrics from scattered sources, and following runbooks that drift out of date. Rule-based remediation tools handle deterministic cases: if X, then Y. But a single symptom like a crash loop can have multiple root causes, and the right fix depends on context the rule cannot see. Kubernaut targets exactly that gap. It is for platform teams that have a large fleet, a noisy alert stream, and a catalog of known remediation steps, but still need a human to decide when a fix is safe. The README frames it as a diagnostician that also adjusts the thermostat, not just a thermostat. That positioning is the core of the project's value proposition, and it is worth keeping in mind when evaluating whether the added complexity is justified.
Architecture: Twelve Go Services Coordinating a Pipeline
Kubernaut is not a single binary. The repository lists twelve services under cmd/, each with a distinct responsibility. The Gateway ingests AlertManager webhooks and Kubernetes Events. Signal Processing enriches and deduplicates by fingerprint. The Remediation Orchestrator manages CRD lifecycles. AI Analysis dispatches to the Kubernaut Agent, which does the LLM-powered root cause analysis and MCP tool execution. The API Frontend exposes MCP and A2A protocols with OIDC auth. Workflow Execution runs Tekton Pipelines, Kubernetes Jobs, or Ansible. Data Storage keeps the workflow catalog and audit trail in PostgreSQL. Notification sends Slack or webhook messages. Effectiveness Monitor runs post-remediation health checks and scores how well a fix worked. Auth Webhook provides service identity, and Fleet Metadata Cache handles multi-cluster targeting. That is a lot of moving parts. The design separates concerns cleanly, but it also means a production deployment involves many replicas, a database, and careful configuration.
How the Agent Investigates and Chooses a Fix
The investigation phase is where Kubernaut differs from rule-based tools. The Kubernaut Agent uses native Go client-go bindings against the Kubernetes API, Prometheus, and log endpoints. It does not just match an alert name to a runbook. It inspects live cluster state, traces cross-namespace dependencies, and filters noise. It can run autonomously end to end, or interactively, where an operator joins via MCP or A2A and guides the investigation in real time. After investigation, the agent selects a remediation workflow from a searchable catalog. The catalog is pre-seeded, meaning the team defines the workflows in advance. Execution goes through Tekton, Kubernetes Jobs, or Ansible, with optional human approval gates. The agent can escalate to a human with a full root cause analysis when it cannot determine a safe fix. This mechanism is the heart of the project, and it is more ambitious than a simple alert-to-action mapper.
Getting It Running: Commands and Configuration
The README does not include a quickstart command, but it points to full documentation at jordigilh.github.io/kubernaut-docs for installation and usage. The repository layout shows a Go project with a main branch and recent release candidates like v1.6.0-rc7. You would clone the repo, build the services, and deploy them into a Kubernetes cluster. The services use CRDs, so you need to install those first. The API Frontend exposes MCP and A2A endpoints, which means you configure OIDC authentication for external agents. The Workflow Execution service needs access to Tekton, Kubernetes Jobs, or Ansible AWX/AAP. The Data Storage service requires PostgreSQL. The README mentions a demo scenarios repository, kubernaut-demo-scenarios, with 37 runnable scenarios. That is the fastest way to see the system in action: clone that repo and run scenarios against your own cluster. There is no single-install command visible, so expect a multi-step deployment.
Approval Gates, OPA Policies, and Audit Trails
Kubernaut keeps humans in control through three mechanisms. First, approval gates can pause execution before a workflow runs. Second, OPA policies constrain which workflows are allowed. Third, audit trails record every action. The README lists these as part of the platform's safety features. The interactive mode also lets operators take over an autonomous session mid-flight, review findings, and approve next steps. That is a sensible design for production use, because an LLM agent making changes to a cluster without oversight is a hard sell for most platform teams. The effectiveness monitor closes the loop by evaluating whether the fix worked via health checks, alert resolution, and spec hash drift detection. It feeds effectiveness scores back into future investigations. This feedback loop is a genuine differentiator, but it also means the system needs enough history to be useful.
Limitations and Failure Modes
Kubernaut is not the right tool for every situation. The README acknowledges that rule-based tools are adequate for known, deterministic problems. If your alerts are simple and your runbooks are stable, Kubernaut's complexity is overkill. The platform depends on an LLM agent, which introduces non-determinism. Even with prompt-injection detection and a shadow agent, an LLM can misjudge a root cause. The README lists safety and reasoning robustness as a validated scenario category, but that does not guarantee correctness in all cases. Another limitation is the operational overhead: twelve services, a PostgreSQL database, and multiple execution backends. A small team may not have the capacity to run and maintain that. Also, the remediation catalog must be pre-seeded and kept current. If your runbooks are outdated or incomplete, the agent will select a fix from a stale catalog. The effectiveness monitor can detect failures, but it cannot fix a missing workflow.
Alternatives: Rule-Based Tools and the Thermostat Comparison
The README explicitly compares Kubernaut to rule-based remediation tools, calling them thermostats. These tools, such as Keptn or Flux's notification controllers, trigger predefined actions on alert conditions. They are deterministic, lightweight, and easy to audit. The difference is in the investigation step. A rule-based tool sees an alert and immediately executes a mapping, with no context beyond the alert labels. Kubernaut investigates first, using live cluster state and observability data, then selects a workflow. That means Kubernaut can handle multi-root-cause incidents where the same symptom has different fixes. The trade-off is latency: an LLM investigation takes longer than a rule match. It also costs more in compute and operational complexity. If you have a small cluster with a handful of alerts, a rule-based tool is faster and cheaper. If you have a large fleet with complex failure chains, Kubernaut's investigation phase is worth the overhead.
Maintenance, License, and Upgrade Cost
Kubernaut is licensed under Apache-2.0, which permits commercial use, modification, and distribution with attribution. The project is actively developed, with release candidates v1.6.0-rc5 through rc7 pushed within days of each other in August 2026. That cadence indicates active maintenance, but it also means you should expect frequent updates and potential breaking changes between release candidates. The documentation site is separate, and the README links to it for architecture and installation guides. The multi-service architecture means upgrades are not a single binary swap; you need to coordinate updates across twelve services. The fleet operations feature in v1.6 adds multi-cluster targeting, which increases complexity further. Before adopting, verify that the project's release stability matches your tolerance for change. The Apache-2.0 license gives you freedom, but you are responsible for maintaining your fork if the project changes direction.
Editorial conclusion
Adopt Kubernaut if you operate Kubernetes clusters with a heavy alert load and have a catalog of runbooks that can be turned into Tekton, Job, or Ansible workflows. Its LLM-based investigation adds value where rule-based tools fail on multi-root-cause incidents. Do not adopt it if you cannot accept an AI agent touching production without human approval gates, or if your remediation steps are still ad hoc and undocumented. Before deploying, verify that your cluster can run the required CRDs and services, that your observability stack exposes the metrics and logs the agent needs, and that your OPA policies cover the workflows you plan to allow. Also verify that your team can handle the operational overhead of running a multi-service Go platform with PostgreSQL, because this is not a single-binary tool.
Frequently asked questions
What is a kubestronaut?
That term does not appear in this repository. The project is Kubernaut, a Go AIOps platform that investigates Kubernetes alerts with an LLM agent and executes remediations from a workflow catalog through Tekton, Jobs or Ansible.
How much does it cost to become a Kubestronaut?
No cost or pricing information is recorded here. What the repository states is that the latest published tags are release candidates of the 1.6.0 line, including v1.6.0-rc19 on 2026-09-29.
How can I become a Kubestronaut?
Nothing in this repository describes that. It covers how an operator joins an investigation, since in interactive mode you can take over an autonomous session mid-flight, review findings and approve the next step through MCP or A2A sessions authenticated with OIDC.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jordigilh-kubernaut)