# OpenSRE: an Apache-2.0 toolkit for building AI SRE agents on your own infrastructure

> OpenSRE is a Python framework from Tracer-Cloud that connects the observability tools you already run to an LLM agent, with an interactive shell, a headless CLI and a Docker deployment. It is a public alpha, and the training-and-benchmark ambition is bigger than what the current install path delivers.

**Tracer-Cloud/opensre** — Build your own AI SRE agents. The open source toolkit for the AI era.

- Repository: https://github.com/Tracer-Cloud/opensre
- Website: https://discord.com/invite/opensre
- Stars: 11,288 · Forks: 1,649
- Language: Python
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/tracer-cloud-opensre

## The scattered-evidence problem OpenSRE targets

The README frames the problem in one sentence: when something breaks in production, the evidence is scattered across logs, metrics, traces, runbooks and Slack threads. That is a description of the work, not of a missing tool. Most teams already have dashboards and a pager. What they do not have is a single place where an agent can pull a metric spike, a log line, a deploy timestamp and a runbook step into one reasoning pass.

OpenSRE is aimed at the SRE or platform engineer who is willing to run that reasoning loop on their own infrastructure rather than inside a vendor's black box. The pyproject description is blunt about the scope: "Open-source SRE agent for automated incident investigation and root cause analysis. Automatically analyzes alerts from Slack, Grafana, Datadog, and other tooling." The topics list on the repository adds alerting, incident management and remediation to that set.

The stated motivation is worth reading closely, because it explains why the project exists at all. The README argues that SWE-bench gave coding agents scalable training data and clear feedback, while production incident response has no equivalent, since distributed failures are slower, noisier and harder to simulate than local code tasks. OpenSRE is presented as an attempt to build that missing layer: an open reinforcement learning environment for agentic infrastructure incident response. That is a research ambition attached to a working CLI, and the two should be judged separately.

## How the agent, the integrations and the gateway fit together

The repository layout tells you more about the architecture than the README prose does. Top-level directories include core, gateway, integrations, surfaces, tools, infrastructure, config and bootstrap, with main.py as the process entrypoint. That split is meaningful: integrations hold the connectors to external services, gateway holds the two-way messaging transport, surfaces hold the user-facing entry points, and core holds the agent harness.

The README points Python users at `core.agent_harness.AgentSession`, showing that the agent can be driven in-process from your own code, with a source checkout required. That is the lowest-level surface. Above it sit the interactive shell, the headless CLI, and the Docker deployment.

The Dockerfile is the clearest statement of the runtime model. One image supports three modes selected by the MODE environment variable: MODE=web runs a FastAPI web API with health, alerts and async investigations on port 8000; MODE=gateway runs a two-way messaging gateway using Slack Socket Mode plus Telegram; MODE=scheduler runs a cron or loop scheduler service with no gateway or web component. The container runs as a non-root user with uid and gid 1000, with /workspace as the writable runtime area, and installs the postgresql extra so psycopg2 can back an investigations store through DATABASE_URL.

One detail in the Dockerfile deserves attention because it contradicts a common assumption about agent frameworks: gateway mode is described as outbound-only long-polling, so the comment notes that EXPOSE and HEALTHCHECK apply only to web mode. If you are deploying the Slack or Telegram transport, you are not opening an inbound port for it.

## Installing OpenSRE and running a first investigation

The README's install path is a single curl pipe into bash, which fetches the latest build from main and does not require sudo.

```bash
curl -fsSL https://install.opensre.com | bash
```

After that, running `opensre` with no arguments starts the CLI. The README notes that if the command is not found, you should follow the PATH instructions the installer printed or open a new terminal, and that macOS and Linux terminals are the supported path, with WSL recommended on Windows.

```bash
opensre
```

The first launch activates the hosted model and asks you to sign in. With no subcommand, `opensre` validates your account and opens a REPL, which requires a TTY. Session control lives in slash commands: /help, /status, /cost, /sessions, /resume, /compact, /new and /exit. Integration management uses /integrations list and /integrations verify, and /agents monitors a local agent fleet. Ctrl+C cancels an in-flight turn without losing session state.

For scripted use, the headless CLI runs one agent turn non-interactively. This is the form to put in a CI job or a terminal alias.

```bash
opensre ask "why is checkout-api slow?"
```

The environment file describes a more deliberate setup path than the one-liner suggests. It recommends setting LLM_PROVIDER and the matching API key, then running `opensre onboard` for guided local setup, configuring one integration with `opensre integrations setup <service>`, and verifying with `opensre health` and `opensre integrations verify <service>`. The preferred local store for integration credentials is `~/.opensre/integrations.json`, with environment variables still supported as a fallback. LLM_PROVIDER accepts a long list of values, including anthropic, openai, openrouter, deepseek, gemini, bedrock and ollama, and several CLI-backed options such as claude-code, codex and gemini-cli that work after you authenticate the underlying CLI.

If you prefer containers, the web mode path is three commands, and the health endpoint is the thing to check first.

```bash
docker build -t opensre:latest .
docker run -p 8000:8000 --env-file .env opensre:latest
curl http://localhost:8000/health
```

The gateway mode container needs SLACK_BOT_TOKEN and SLACK_APP_TOKEN, or TELEGRAM_BOT_TOKEN and TELEGRAM_ALLOWED_USERS, plus LLM_PROVIDER and the matching API keys.

## Where OpenSRE will disappoint you

The README labels the project a public alpha in a banner and in the badge, with the warning that core workflows are usable for early exploration but not yet fully stable, and that APIs and integrations may evolve. Take that literally when you plan. Anything you build against `core.agent_harness.AgentSession` is built against an interface the maintainers have reserved the right to change, and the README explicitly requires a source checkout to use it.

The account requirement is the second constraint, and it is easy to miss. The interactive shell opens only for an active account, and the first launch activates the hosted model. If your organisation cannot send production incident context through a hosted model, the default path is the wrong one, even though the environment file lists local and self-hosted provider options such as ollama. The README does not document what the hosted activation does with your data; it links to a separate security page at trust.tracer.cloud for that.

The third gap is that the README does not document rollback. There is no described procedure for reverting an agent action that made an incident worse, and no described approval gate for remediation steps. The pyproject topics include remediation, so the capability is in scope, but the safety story around it is not spelled out in the README. If you need an agent that can act on production, verify that boundary yourself before you wire it to anything that can change state.

Finally, OpenSRE is not an alerting or paging system. It analyzes alerts that arrive from Slack, Grafana, Datadog and similar tools. Teams looking for a pager replacement, an uptime monitor or a status page are looking at the wrong project.

## OpenSRE compared with HolmesGPT and IncidentFox

HolmesGPT and IncidentFox appear alongside OpenSRE in the searches people run about this project, so the comparison is worth making concrete rather than gestural. The difference visible in the repository is the shape of the deliverable. OpenSRE ships a framework plus a training and evaluation environment: the README describes an open reinforcement learning environment for agentic infrastructure incident response, with end-to-end tests under tests/e2e and a semantic test-catalog naming convention documented in tests/README.md so that e2e versus unit and local versus cloud boundaries stay obvious.

That is a different bet from a tool whose primary output is an investigation report. OpenSRE's stated mission is to scale to thousands of realistic infrastructure failure scenarios and to serve as a benchmark and training ground for AI SRE. If you want an agent that answers a question about a slow service today, both approaches can serve you. If you want to evaluate how an agent behaves across many failure scenarios, or to contribute scenarios yourself, the test catalog and e2e layout are the part of OpenSRE that distinguishes it.

The cost of that orientation is surface area. A framework with a gateway, a web API, a scheduler, a CLI and an in-process harness has more moving parts to operate than a single-purpose investigation tool, and the Dockerfile's three MODE values are the evidence. Choosing OpenSRE means accepting that you will configure and run more than one of those surfaces.

## Licence, telemetry and the real upgrade cost

OpenSRE is licensed under Apache-2.0, stated in the README badge and in the pyproject license field. For most teams that means permissive use with an attribution and notice obligation, and no copyleft reach into your own code. This is not legal advice; check the LICENSE file and your own counsel for anything that matters.

The dependency list is where the operational cost shows up. The project requires Python 3.12 or newer and pins a wide set of libraries, including anthropic, openai, litellm, mcp, kubernetes, fastapi, pydantic, boto3, slack-sdk, discord.py and cryptography. Two entries deserve a second look. LiteLLM is constrained to a range with both a floor and a ceiling, so a major upgrade there is gated by the project rather than by you. Cryptography is pinned to a narrow band, and the pyproject comment explains why: version 50.0.0 is described as the first version patched for a specific CVE, and the comment notes that cryptography 49 ships no x86-64 macOS wheel, so the frozen darwin-x64 release binary builds it from source against OpenSSL 3.2 and the release workflow re-copies those dylibs into the PyInstaller bundle. That is a real build constraint, not a stylistic pin.

Upgrade cadence is visible in the release list: v0.1.2026.9.20 and v0.1.2026.9.19 were published one day apart. The repository's last push was on 2026-09-20, and it is not archived. A daily versioning scheme on a project that calls itself a public alpha is a signal to read release notes before upgrading rather than to upgrade automatically. The README also references a telemetry section in its table of contents, so if your environment has data-handling rules, read that section before you install, not after.

## Conclusion

Adopt OpenSRE if you already run Slack, Grafana or Datadog and want an agent that reads those signals on infrastructure you control, and if you are comfortable tracking a public alpha where, per the README, APIs and integrations may evolve. Do not adopt it as a replacement for your alerting or on-call stack; the README positions it as investigation and root cause analysis, not paging. Before committing, verify three things: that `opensre health` reports your LLM provider as reachable, that `opensre integrations verify` passes for the one service you actually care about, and that the hosted model activated on first launch is the one you want carrying production log content.

## FAQ

### What is OpenSRE?

OpenSRE is an open-source framework for building AI SRE agents, distributed under Apache-2.0 and written in Python. Its pyproject description says it performs automated incident investigation and root cause analysis by analyzing alerts from Slack, Grafana, Datadog and other tooling.

### What are AI SRE agents used for in OpenSRE?

In this project they are used for production incident investigation and root cause analysis, pulling evidence that is otherwise scattered across logs, metrics, traces, runbooks and Slack threads. The repository topics also list alerting, incident management and remediation.

### Is OpenSRE itself open source?

Yes. The README carries an Apache 2.0 licence badge, and the pyproject license field reads Apache-2.0. The repository is not archived.

### What are the alternatives to OpenSRE?

HolmesGPT and IncidentFox are the names that come up alongside it in searches. The visible difference is that OpenSRE ships a training and evaluation environment as well as agents, with end-to-end tests under tests/e2e and a documented test-catalog naming convention.

## Sources

- [License: Apache-2.0](https://github.com/Tracer-Cloud/opensre/blob/main/LICENSE)
- [Project website](https://discord.com/invite/opensre)
- [README](https://github.com/Tracer-Cloud/opensre/blob/main/README.md)
- [Releases](https://github.com/Tracer-Cloud/opensre/releases)
- [Tracer-Cloud/opensre on GitHub](https://github.com/Tracer-Cloud/opensre)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tracer-cloud-opensre
