Model or dataset
HolmesGPT/holmesgpt avatar
HolmesGPT/holmesgpt

HolmesGPT: the CNCF SRE agent that queries your live observability stack

SRE Agent - CNCF Sandbox Project

3,460 stars503 forksPythonApache-2.0

At a glance

What is it?
HolmesGPT is an Apache-2.0 Python agent that investigates production incidents by running an agentic loop against Prometheus, Kubernetes, Datadog and other data sources. Here is how it installs, what the architecture implies, and where it stops being the right tool.
Who is it for?
Adopt HolmesGPT if you already run Prometheus, Kubernetes or a comparable observability stack and want an agent that reads those sources during an incident rather than a human pasting screenshots into a chat window. Do not adopt it if you have no LLM API budget, no observability data for it to query, or a compliance rule against shipping cluster credentials and log contents to a hosted model.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap HolmesGPT fills between alert and root cause

An alert fires. Someone opens a dashboard, then a second dashboard, then kubectl, then the log query UI, and twenty minutes later has a hypothesis. HolmesGPT targets that interval. It is an AI agent that takes an investigation request and queries live observability data across multiple sources to produce a root cause, according to the README. The intended user is an SRE or on-call engineer who already has telemetry but not the time to correlate it by hand.

The project is a Cloud Native Computing Foundation sandbox project, originally created by Robusta.Dev with contributions from Microsoft, per the README. That provenance matters for procurement: sandbox status is an early CNCF stage, not a graduation, so it signals a project under foundation governance rather than a mature, widely deployed standard.

The README is explicit that Kubernetes is not required. HolmesGPT works against VMs, bare metal, cloud services and containers, which separates it from tools that assume a cluster. The Operator mode is the exception: the README states the operator itself runs in Kubernetes, while the health checks it runs can query any connected data source.

How the agentic loop reaches your Prometheus and Kubernetes data

The mechanism the README describes is an agentic loop. The agent does not hold your telemetry; it holds tools. Each integration is a toolset, and the agent decides which tool to call, reads the result, and decides what to call next until it can state a root cause. The README lists Prometheus, Grafana, Datadog, Kubernetes, Elasticsearch/OpenSearch, Coralogix, ArgoCD, Crossplane, Docker, AWS, Azure, GCP, GitHub, GitLab, Jira and Confluence among the built-in toolsets, with a documented path for adding your own or wrapping any REST API.

Two design choices in that loop are worth naming. The first is context management. The README claims petabyte-scale data handling through server-side filtering, JSON tree traversal and tool output transformers, with the stated goal of keeping large payloads out of context windows. The second is memory safety: per-tool memory limits, streaming large results to disk, and automatic output budgeting to prevent OOM kills when querying large observability datasets. Both are admissions that the naive version of this architecture breaks. An agent that pulls a day of logs into a prompt will fail, and the README is describing the mitigations rather than the happy path.

Alert handling is bidirectional. The README says HolmesGPT can fetch alerts from AlertManager, PagerDuty, OpsGenie or Jira and write findings back. That write-back direction is where the operator mode gets its leverage: the README states that with the GitHub integration connected, the operator can open pull requests to fix what it finds.

Installing HolmesGPT and running a first investigation

The repository ships a docker-compose.yaml, which is the shortest path to a running server. It publishes port 5050 bound to 127.0.0.1, mounts ~/.holmes for state, and mounts your kubeconfig and cloud credential directories read-only. The compose file expects an API key from the environment; OPENAI_API_KEY is active and the Anthropic, Gemini, AWS, and Azure variables are present but commented out.

bash
export OPENAI_API_KEY=sk-...
docker compose up

After the container starts, the healthcheck defined in the compose file curls http://localhost:5050/healthz every 30 seconds, with a 15 second start period. A passing healthcheck means the server is up; it does not mean your LLM credentials are valid.

There is also a CLI. The pyproject.toml declares a Poetry script named holmes mapped to holmes.main:run, so an installed package exposes a holmes command. The project requires Python >=3.10,<3.14. The upper bound is deliberate: a comment in pyproject.toml states that litellm (>=1.84.0) declares Requires-Python <3.14, and that the cap should be lifted once the referenced litellm pull request is merged and released.

bash
poetry install
poetry run holmes --help

For the Kubernetes operator, the repository contains a helm/ directory and a Dockerfile.operator, which indicates a Helm chart is the intended deployment path. The README does not give the chart name or the values schema, so read the chart before installing rather than guessing at flags.

The Makefile shows the project's own test split. Unit-style tests run without an LLM, while the LLM-backed tests are separated:

bash
poetry run pytest tests -m "not llm"
poetry run pytest tests/llm/test_ask_holmes.py -n 6 -vv

That separation is a useful signal about where the non-determinism lives. Anything in tests/llm depends on a model's output and is run in parallel across six workers.

Where the agentic loop becomes the wrong tool

The README does not document rollback for Operator mode. It describes the operator running in the background 24/7, spotting problems and messaging Slack, and optionally opening pull requests through the GitHub integration. Nothing in the supplied documentation describes how to undo an action the operator took, or how to bound what it is allowed to change. If you are considering letting it open PRs against production repositories, that is the first thing to resolve, and the README is silent on it.

Cost and latency are structural, not incidental. Every investigation is a sequence of LLM calls interleaved with tool calls against live systems. That means an incident investigation consumes tokens, and the token count scales with how many tools the agent decides to invoke. The README's context-management features exist precisely because unbounded tool output is expensive. A team with no LLM budget line, or one whose provider rate limits are shared with production traffic, will feel this immediately.

Data egress is the other hard edge. The compose file mounts your kubeconfig, ~/.aws and ~/.config/gcloud into the container, and the agent queries live observability data. Those results are sent to whichever LLM provider you configure, and pyproject.toml lists OpenAI, Bedrock and Azure SDK dependencies alongside litellm. If your logs or cluster metadata cannot leave your network, you need a provider path that keeps inference inside it, and the README does not describe a local-model setup.

Finally, this is an investigation tool, not a mitigation tool. It finds root causes. It does not roll back a deployment, drain a node, or page a human, unless the operator's GitHub integration is configured to open a fix.

HolmesGPT compared with kagent and the alert-only path

The most direct comparison people search for is HolmesGPT versus kagent. Both are Kubernetes-adjacent agents, but the repository layout points to a different centre of gravity. HolmesGPT's pyproject.toml carries a broad dependency set spanning AWS, Azure, GCP, Kafka and OpenSearch clients, and the README's toolset table leads with SaaS and cloud data sources such as Datadog, Coralogix, Jira and Confluence. The README also states plainly that Kubernetes is not required. That is a tool built to correlate across whatever telemetry you have, with Kubernetes as one source among many.

A narrower alternative is the alert-only path: keep AlertManager, PagerDuty or OpsGenie doing what they do, and treat root-cause analysis as a human task with runbooks in Confluence. HolmesGPT can read those same runbooks through its Confluence toolset, so the difference is not the data. It is who does the correlation. The alert-only path costs nothing per incident and is fully deterministic. HolmesGPT trades determinism and per-incident cost for an automated first hypothesis.

A third point of comparison is a general-purpose LLM chat session where an engineer pastes log excerpts. That approach has no tool credentials, so it cannot query Prometheus for the metric that would confirm or kill the hypothesis. HolmesGPT's value is concentrated in that difference: the agent can go look, rather than reason about what you pasted.

Licence, maintenance and the upgrade cost you inherit

HolmesGPT is Apache-2.0, and the repository carries a LICENSE file at the top level. Apache-2.0 permits commercial use and modification and includes a patent grant, which is why it is a common choice for infrastructure tooling. It is not a copyleft licence, so you are not obliged to publish modifications. This is a description of the licence text, not legal advice; if you are embedding HolmesGPT in a product, have counsel read the NOTICE and attribution requirements.

The repository is not archived. The last push was on 2026-09-09, and releases 0.41.0, 0.41.0-alpha and 0.40.0 landed in the weeks before that. The version field in pyproject.toml is 0.0.0, which is typical of projects that set the version at release time rather than in the source tree.

The upgrade cost is visible in the dependency pins. pyproject.toml pins litellm to exactly 1.89.0 and mcp to 1.28.1, with comments tying those floors to specific CVEs, including CVE-2026-49468 in litellm and CVE-2026-52869, CVE-2026-52870 and CVE-2026-59950 in mcp. The Dockerfile pins minimum versions for pip, wheel and setuptools for the same reason, and pulls librdkafka from Alpine edge because Alpine 3.23 ships 2.12.1 while confluent-kafka 2.14.0 needs 2.14.0 or newer. This is a project that tracks its supply chain closely, and the flip side is that upgrading means re-checking those pins rather than running a blanket update. The Python cap at <3.14 is the constraint most likely to bite you first, because it is waiting on an upstream fix rather than on HolmesGPT itself.

Editorial conclusion

Adopt HolmesGPT if you already run Prometheus, Kubernetes or a comparable observability stack and want an agent that reads those sources during an incident rather than a human pasting screenshots into a chat window. Do not adopt it if you have no LLM API budget, no observability data for it to query, or a compliance rule against shipping cluster credentials and log contents to a hosted model. Before rolling it out, verify three things: which LLM provider you will point it at and at whose expense, that the toolset credentials you configure are read-only, and whether you want the CLI path or the Kubernetes operator, since the README presents them as separate deployment shapes. The repository is not archived and the last push was on 2026-09-09, so the codebase is moving, but the README does not document a rollback procedure for the operator, and that gap is worth resolving before you let it open pull requests.

Frequently asked questions

What is HolmesGPT?

It is an open-source AI agent for investigating production incidents and finding root causes, described in the README as a CNCF sandbox project originally created by Robusta.Dev with contributions from Microsoft. It runs an agentic loop against live observability data from toolsets such as Prometheus, Kubernetes, Datadog, Elasticsearch and Jira. The README states that Kubernetes is not required and that it works with VMs, bare metal, cloud services and containers.

what is holmesgpt

The same project under its lowercase repository name HolmesGPT/holmesgpt. It is written in Python, licensed Apache-2.0, and installs either as a Docker Compose service on port 5050 or as a Poetry package exposing a holmes command. Its distinguishing feature is the agentic loop, which queries multiple data sources rather than reasoning over pasted context.

How does HolmesGPT compare with kagent?

The search question is incomplete, so this answer is limited to what the README and repository files support. HolmesGPT's dependency set in pyproject.toml spans AWS, Azure, GCP, Kafka and OpenSearch clients, and its README toolset table leads with SaaS sources such as Datadog, Coralogix, Jira and Confluence, while stating that Kubernetes is not required. No information about kagent appears in the README, so a direct comparison cannot be made here.

Official sources

  1. HolmesGPT/holmesgpt on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/holmesgpt-holmesgpt.svg)](https://hysenlabs.com/projects/holmesgpt-holmesgpt)