Open-source project
derisk-ai/OpenDerisk avatar
derisk-ai/OpenDerisk

OpenDeRisk: a multi-agent SRE framework built around the OpenRCA dataset

AI-Native Risk Intelligence Systems, OpenDeRisk——Your application system risk intelligent manager provides 7* 24-hour comprehensive and in-depth protection.

971 stars130 forksPythonMIT

At a glance

What is it?
OpenDeRisk is an MIT-licensed Python system that runs SRE-Agent, Code-Agent, ReportAgent, Vis-Agent and Data-Agent together to perform root cause analysis. It ships as a uv workspace, and its reference workload is Microsoft's OpenRCA dataset.
Who is it for?
Adopt OpenDeRisk if you already run a multi-agent Python stack and want a working reference for agent-driven root cause analysis over the OpenRCA dataset, or if you want to study how SRE-Agent and Code-Agent divide diagnostic work. Do not adopt it as a drop-in production monitoring platform: the README does not describe alert ingestion from your own systems, and the default path assumes a 26GB external dataset.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What OpenDeRisk actually diagnoses

The README frames OpenDeRisk as an AI-Native Risk Intelligence System and an application system's intelligent manager, offering what it calls 7x24 hour comprehensive and in-depth protection. Strip the marketing framing and the concrete deliverable is narrower and more interesting: a multi-agent pipeline that takes logs, traces and code and produces a root cause hypothesis with a visible evidence chain.

The intended audience is SRE and platform engineering teams who already accept that an LLM-driven agent can read telemetry, and who want to see the orchestration layer rather than a black-box alerting product. The README's own citation describes the work as an industrial framework for AI-driven SRE with design, implementation and case studies, which sets expectations at the level of a reference architecture you can extend. If you want a hosted incident-response tool with an on-call schedule and paging integrations, this is not that, and the README never claims it is.

Five agents, a data layer, and a rendering protocol

The architecture section splits the system into three layers. The data layer pulls the OpenRCA dataset from GitHub, decompresses it locally and processes it for analysis. The logic layer runs the multi-agent collaboration. The visualization layer uses what the README calls the Vis protocol to render the processing flow, the evidence chain, and the switching between roles.

The agents named in the README are SRE-Agent, Code-Agent, ReportAgent, Vis-Agent and Data-Agent. Code-Agent is the one worth paying attention to: the README states that on the OpenRCA dataset, root cause analysis runs through multi-agent collaboration with Code-Agent dynamically writing code for the final analysis. That is a different design from a fixed tool-calling loop. Instead of predefining every diagnostic query as a tool, the system generates analysis code at runtime. It gives the agents open-ended reach into the data, and it also means the quality of a diagnosis depends on generated code being correct, which is a much harder thing to verify than a fixed query.

The README is explicit that the code currently implements the highlighted components of the architecture diagram, so the diagram is a target, not a description of the repository. Treat the five-agent list as the design intent and check packages/derisk-ext for what is actually wired up.

Installing OpenDeRisk and running the quickstart server

The README offers two install paths. The recommended one is a curl-piped shell script that downloads and installs the latest version, then initializes a configuration file at ~/.openderisk/configs/derisk-proxy-aliyun.toml. If you prefer to read the installer before running it, the repository keeps install.sh at the top level.

bash
curl -fsSL https://raw.githubusercontent.com/derisk-ai/OpenDerisk/main/install.sh | bash

After installation you edit the generated TOML file and set your API keys, then start the server with the openderisk-server command. The README does not document which keys are required beyond the instruction to set your API keys, so read the file before starting.

bash
vi ~/.openderisk/configs/derisk-proxy-aliyun.toml
openderisk-server

For development, the README requires uv and installs from a cloned repository. The sync command enables a specific set of extras, and the README notes that channel_dingtalk is optional and can be skipped if you do not need DingTalk channel support.

bash
git clone https://github.com/derisk-ai/OpenDerisk.git
cd OpenDerisk
uv sync --all-packages --frozen \
    --extra "base" \
    --extra "proxy_openai" \
    --extra "rag" \
    --extra "storage_chromadb" \
    --extra "derisks" \
    --extra "storage_oss2" \
    --extra "client" \
    --extra "ext_base" \
    --extra "channel_dingtalk"

The quickest way to see the system is the zero-configuration start, which the README lists first. It runs without any configuration file, and you then configure models and settings through the web UI at http://localhost:7777.

bash
uv run derisk quickstart

The README also documents a port flag and a configuration-file variant, which is the path to use once you have filled in the TOML file.

bash
uv run derisk quickstart -p 8888
uv run derisk quickstart -c configs/derisk-proxy-aliyun.toml

Beyond the AI-SRE mode, the README lists two lighter usage modes that need no dataset: a flame graph assistant that accepts Java or Python flame graphs from your local application, and DataExpert, which takes metrics, logs, traces or Excel data for conversational analysis. If you want to evaluate the agents without committing to a large download, start there.

The 26GB dataset is the real entry cost

The single biggest constraint is the data. The README states that alert awareness is based on Microsoft's OpenRCA dataset, that the decompressed dataset is approximately 26GB, and elsewhere that the data layer pulls a 20GB dataset from GitHub. Whatever the exact figure, this is a multi-tens-of-gigabytes download and decompression before the flagship AI-SRE mode does anything. The README gives a gdown command for the Telecom dataset and a separate one for the Bank dataset, and instructs you to move the results into pilot/datasets/. There is no documented path for pointing the AI-SRE mode at your own production telemetry, and no schema description for what a local dataset would have to look like.

That has a direct consequence for evaluation. If you want to know how OpenDeRisk performs on your incidents, the README does not tell you how to get there. You would be reading the agent and data-layer code to build the adapter yourself.

The second limitation is scope. The README describes an architecture with five agents, then notes that the code currently implements the highlighted components. Those two statements sit in the same document, and the gap between them is where most of the integration work lives. A team that adopts OpenDeRisk expecting a finished five-agent product will spend its first weeks reconciling the diagram with the repository.

OpenDeRisk versus a scripted diagnostic runbook

The obvious alternative is the runbook you already have: a set of dashboards, saved queries and shell scripts that an engineer runs when an alert fires. The difference in approach is where the branching happens. A runbook encodes the branch points in advance, so a human decides which path to take and the tooling executes it. OpenDeRisk moves that decision into the agents, with Code-Agent writing analysis code on the fly for the final step.

The trade-off is legibility. A runbook can be reviewed, version-controlled and tested against a known incident. Generated analysis code cannot be reviewed before it runs, and the README's answer to that is the visualization layer: the Vis protocol renders the processing flow and the evidence chain so a human can inspect the reasoning after the fact. That is a real design response to a real problem, and it is also post-hoc. If your organization requires pre-approved diagnostic steps for production systems, the runbook wins on that axis regardless of how well the agents perform.

A second alternative, closer in spirit, is to build the agent loop yourself on top of a general agent framework. OpenDeRisk's value there is that the SRE-specific division of labor (SRE-Agent, Code-Agent, ReportAgent, Vis-Agent, Data-Agent) and the OpenRCA integration are already written down and, per the README, partially implemented. You are trading control of the orchestration for a head start on the domain wiring.

Maintenance, upgrades and the MIT licence

OpenDeRisk is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are preserved. The README does not discuss trademark use, and the repository carries a LEGAL.md and a MODEL_DELETION_FIX_REPORT.md at the top level, both of which are worth reading before you redistribute anything. Nothing here is legal advice; if you plan to ship OpenDeRisk inside a product, have counsel review those two files.

The last push to the default branch was on 2026-09-03, and the most recent release is v0.3.0 from 2026-03-31, following v0.2.1-archive and v0.2.0. The repository is not archived. The release cadence visible in that list is roughly one significant version every few months, with the archive suffix on v0.2.1 suggesting the project has already retired at least one release line. For an upgrade budget, plan on reading release notes rather than assuming API stability, and note that the workspace pins dependencies through uv.lock, so a sync against a newer lock file is the upgrade mechanism the README implies.

The install script writes to ~/.openderisk/configs/, which means an upgrade can touch a file you have edited. The README does not document rollback or configuration migration, so keep your TOML under version control outside that directory.

Editorial conclusion

Adopt OpenDeRisk if you already run a multi-agent Python stack and want a working reference for agent-driven root cause analysis over the OpenRCA dataset, or if you want to study how SRE-Agent and Code-Agent divide diagnostic work. Do not adopt it as a drop-in production monitoring platform: the README does not describe alert ingestion from your own systems, and the default path assumes a 26GB external dataset. Before committing, verify two things in the repository itself. First, open packages/derisk-ext to see which agents are actually implemented, since the README says only the highlighted architecture components are present in code. Second, read configs/derisk-proxy-aliyun.toml to confirm which model provider keys the proxy expects, because the default configuration file is named after a specific cloud proxy.

Frequently asked questions

What is OpenDeRisk?

It is an AI-Native Risk Intelligence System that performs root cause analysis on logs, traces and code using five collaborating agents: SRE-Agent, Code-Agent, ReportAgent, Vis-Agent and Data-Agent. The README describes it as an industrial framework for AI-driven SRE, with a visualization layer that renders the processing flow and evidence chain.

How do I install OpenDeRisk?

The README recommends a curl-piped install script that writes a configuration file to ~/.openderisk/configs/derisk-proxy-aliyun.toml. From source, you clone the repository and run uv sync with the extras listed in the README, then start the server with uv run derisk quickstart or the openderisk-server command.

Does OpenDeRisk need the OpenRCA dataset to run?

Only the AI-SRE mode does. The README states that alert awareness is based on Microsoft's OpenRCA dataset, which is roughly 26GB decompressed. The flame graph assistant and DataExpert modes accept local uploads instead, so you can start there without the download.

What licence does OpenDeRisk use?

MIT. The repository also carries LEGAL.md and MODEL_DELETION_FIX_REPORT.md at the top level, and the README does not address trademark use, so read those files if you intend to redistribute.

Official sources

  1. derisk-ai/OpenDerisk on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/derisk-ai-openderisk.svg)](https://hysenlabs.com/projects/derisk-ai-openderisk)