# RunbookHermes: An AIOps Incident Response Agent with Evidence-First RCA and Approval-Gated Remediation

> RunbookHermes extends the Hermes Agent runtime with an AIOps domain layer for production incident response. It enforces evidence collection before root cause analysis, gates dangerous remediation actions behind approval and dry-run checks, and writes each resolved incident into a self-improving runbook memory.

**Tommy-yw/RunbookHermes** — Hermes-native AIOps agent for evidence-driven incident response, approval-gated remediation, and runbook learning.

- Repository: https://github.com/Tommy-yw/RunbookHermes
- Stars: 547 · Forks: 41
- Language: Python
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/tommy-yw-runbookhermes

## What RunbookHermes adds to the Hermes Agent runtime

RunbookHermes is built by layering an AIOps domain layer on top of the Hermes Agent framework from NousResearch. The architecture is explicit in the README:

```text
Hermes Agent Runtime
+ RunbookHermes AIOps Domain Layer
= RunbookHermes
```

The Hermes Agent provides the underlying capabilities: an agent main loop, a tool registry, multi-model LLM provider support (via OpenRouter), CLI and Gateway entry points, a memory provider, skills, context compression, session state management, and reinforcement learning trajectory collection.

RunbookHermes adds the SRE-specific layer: alert ingestion from Alertmanager, Feishu, WeCom, and API endpoints; structured evidence collection from Prometheus, Loki, distributed traces, and deployment history; an EvidenceStack context type that groups all evidence for an incident; root cause analysis guards that require evidence before generating a hypothesis; an action policy system; approval, checkpoint, and dry-run gates for dangerous operations; controlled execution with recovery verification; and post-incident memory persistence including RAG knowledge base updates, skill generation, and evaluation case recording.

The result is an agent that cannot skip evidence collection and cannot execute a dangerous action without going through a defined safety chain.

## Evidence-first: why guessing root cause is explicitly blocked

The README's first design principle is stated plainly: incident handling cannot start with a root cause guess. Before generating any RCA hypothesis, RunbookHermes must collect from at least these sources: metrics, logs, distributed traces, deployment history, service profile, existing memory context, RAG citations, and any multimodal evidence available.

The README articulates the distinction between memory and evidence. Historical incident memory is a weak prior: it can suggest that a similar fault pattern appeared in the past, but it cannot substitute for the current incident's metrics and logs. The README gives this example: memory might note that a service has shown similar failures before, but the conclusion must come from the current evidence, not from that historical note.

This design prevents a class of agent failure common in LLM-based diagnostic tools, where the model confidently states a root cause based on pattern-matching against training data rather than actual signal from the monitored system. By requiring evidence collection as a precondition for the RCA step, RunbookHermes makes it harder for the model to shortcut the diagnostic process.

## The approval chain for dangerous actions

Dangerous actions in RunbookHermes cannot be executed directly by the agent. The README lists the categories: rollback, restart, scaling mutation, traffic switching, configuration mutation, database-affecting operations, cache flush, and dependency failover. Any action in these categories goes through a defined safety chain:

```text
action policy
→ approval
→ checkpoint
→ dry-run
→ controlled execution
→ recovery verification
→ audit timeline
```

The Approval Center in the web console is the human-in-the-loop gate. Operators see the action, its risk level, the checkpoint state, and the full payload before deciding to approve or reject. The audit timeline records every step of the incident lifecycle: evidence collection, hypothesis generation, action planning, checkpoint creation, approval request, approval decision, and execution result.

The recovery verification step runs after execution to confirm the service has returned to a healthy state, rather than assuming success from a successful API call.

This chain makes the pattern auditable. Every incident produces a timeline that shows exactly what evidence was collected, what the agent concluded, what action was proposed, who approved it, and what the post-execution state was.

## Setting up the environment and connecting observability

RunbookHermes uses environment variables for all service credentials. The .env.runbook.example file (in the repository root alongside .env.example) covers the runbook-specific configuration. The base .env.example covers the LLM provider, which defaults to OpenRouter:

```bash
# OpenRouter provides access to many models through one API
# All LLM calls go through OpenRouter - no direct provider keys needed
# Get your key at: https://openrouter.ai/keys
# OPENROUTER_API_KEY=
```

The model is configured in ~/.hermes/config.yaml rather than as an environment variable. The README notes that LLM_MODEL is no longer read from .env.

Connecting observability tools (Prometheus, Loki, distributed tracing systems) and alerting integrations (Alertmanager, Feishu, WeCom) is configured through the settings page in the web console, which shows whether each integration interface is connected or missing. The Dockerfile builds the full stack from source using a Debian base with uv for Python dependency management, npm for Node dependencies, and Playwright for browser tools. The container runs as a non-root user (UID 10000) for runtime isolation.

## Self-evolving memory and the RAG knowledge base

Each resolved incident updates RunbookHermes's memory layer with a set of structured artifacts: an incident summary, identified fault patterns, inferred service governance rules, observed team troubleshooting habits, a generated SKILL.md runbook file, an updated RAG document, an evaluation case, a training trajectory, and a reward label.

The memory console exposes this layer to operators. It includes a local memory store backed by SQLite FTS5 for full-text search, HRR-based (Holographic Reduced Representation) offline semantic recall, and skill indexing with trust scores. Operators can search historical incident summaries and fault patterns, and can write stable operational knowledge directly into memory: service profiles, rollback rules, approval preferences, and recurring failure patterns.

The RAG knowledge base accepts SRE manuals, service troubleshooting documents, architecture notes, and runbook content. RunbookHermes uses the RAG layer as background knowledge during incident response, with citations showing which document supported each recommendation. This means the agent can reference the specific runbook section that informed an action suggestion.

This combination of per-incident learning and operator-curated knowledge is what the README means by the self-improving property: each incident leaves traces that influence how similar incidents are handled next time.

## Limitations and what the project does not cover

RunbookHermes is a reference implementation, not a production-hardened platform. The README describes the project as a transformation of Hermes Agent for AIOps use cases, and the Hermes Agent itself is version 0.11.0 per the pyproject.toml. The system is complex: it requires a configured LLM provider, at least one connected observability source, alert routing from Alertmanager or Feishu/WeCom, and a running web console to be useful.

The README does not document rollback procedures for the RunbookHermes deployment itself. The web console screenshots in the README describe the interface but there is no getting-started guide for connecting a live Prometheus or Loki instance.

The benchmark and evaluation console is described as measuring whether the agent improves over time, but the README does not describe the specific metrics used or the baselines. The RL training dataset pipeline is mentioned as a feature but its practical use requires configuring reward labels and a compatible training setup.

For teams looking to adopt AIOps tooling today, mature commercial platforms such as PagerDuty or Datadog provide alert correlation and incident response with existing production integrations. RunbookHermes is most useful as a research reference for teams building custom agent-based incident response workflows.

## Maintenance, licence, and dependencies

The last push to the repository was on 2026-05-18. The repository is not archived. There are no GitHub releases.

The project is MIT-licensed. The pyproject.toml identifies the project as hermes-agent version 0.11.0, based on the NousResearch Hermes Agent framework. Key runtime dependencies include openai>=2.21.0, anthropic>=0.39.0, httpx with SOCKS support, pydantic>=2.12.5, and exa-py and firecrawl-py for web search and crawling.

The messaging integrations (Telegram, Discord, Slack, WhatsApp via Matrix) are optional extras in the pyproject.toml. The cron extra adds croniter for scheduled tasks. Node.js 20 or later is required for the web dashboard component.

## Conclusion

Site reliability engineers and platform teams who want an agent that collects metrics, logs, and traces before diagnosing an incident and blocks dangerous remediation commands behind human approval will find RunbookHermes a concrete reference implementation. Teams looking for a production-deployed, vendor-supported AIOps product should note that the last push was on 2026-05-18 and the project carries no support SLA. Before deploying, configure the OPENROUTER_API_KEY in .env and connect at least one observability source (Prometheus or Loki) to verify that the evidence-collection pipeline works for your environment before relying on the RCA output.

## FAQ

### How does RunbookHermes differ from a standard AI chatbot for SRE?

RunbookHermes enforces two constraints that a general chatbot lacks: it requires evidence collection (metrics, logs, traces, deploy history) before generating a root cause hypothesis, and it gates all dangerous remediation actions behind an approval chain with checkpoint, dry-run, controlled execution, and recovery verification steps.

### What observability sources does RunbookHermes connect to?

The README describes evidence collection from Prometheus metrics, Loki logs, distributed traces, and deployment history as the primary observability sources. Alert ingestion is supported from Alertmanager, Feishu, WeCom, the web console, and a direct API endpoint.

### What does RunbookHermes do with incident data after resolution?

After each incident, RunbookHermes persists an incident summary, identified fault patterns, service governance rules, a generated SKILL.md runbook skill, an evaluation case, and a training trajectory into its memory layer. This data is indexed in SQLite FTS5 for text search and through HRR-based recall for semantic search.

## Sources

- [Issues](https://github.com/Tommy-yw/RunbookHermes/issues)
- [License: MIT](https://github.com/Tommy-yw/RunbookHermes/blob/main/LICENSE)
- [README](https://github.com/Tommy-yw/RunbookHermes/blob/main/README.md)
- [Tommy-yw/RunbookHermes on GitHub](https://github.com/Tommy-yw/RunbookHermes)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tommy-yw-runbookhermes
