# Ongrid: a self-hosted ops AI agent that reads metrics, logs and traces from chat

> Ongrid is a Go-based, self-hosted AI agent for infrastructure operations. It answers operator questions in Slack, Telegram, Lark or DingTalk, walks the topology to find a root cause, and gates any write behind an approval step.

**ongridio/ongrid** — An ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.

- Repository: https://github.com/ongridio/ongrid
- Website: https://ongrid.cloud
- Stars: 1,103 · Forks: 238
- Language: Go
- License: AGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ongridio-ongrid

## What Ongrid actually does, and who it is for

Ongrid is an ops AI agent. The README describes it as something that "understands your infrastructure, finds the root cause, and fixes it, right from Slack or Telegram." The practical shape of that claim is narrower and more useful than it sounds: an operator asks a question in a chat channel, or an alert fires, and the agent assembles evidence from metrics, logs, traces and a topology graph before answering.

The target user is an SRE or platform engineer who already has observability data but not the time to pivot between four dashboards during an incident. The repository topics list aiops, chatops, incident-response, root-cause-analysis, kubernetes and opentelemetry, which maps to that audience. It is not a general-purpose chatbot and not a monitoring system of record. It sits on top of data you already collect.

The feature list is broad: a coordinator agent that dispatches to SRE, network and DB sub-agents; alert-driven auto-investigation; a root-cause walk over topology; browser SSH through a reverse tunnel; a knowledge vault that indexes runbooks and repositories; workflow automation; and Kubernetes lifecycle management. That breadth is the first thing to judge. A tool that claims to do all of it in one self-hosted stack is either genuinely integrated or thinly spread, and the README alone will not tell you which.

## The coordinator, specialist agents and the edge dial-out model

The architecture visible in the repository has two halves. A manager component runs the agents, the chat integrations and the database. An edge component runs on the hosts you want to inspect. The .env.example points at a frontier broker: the manager dials ONGRID_FRONTIER_ADDR, which defaults to frontier:40011 inside Docker Compose, and registers under ONGRID_FRONTIER_SERVICE_NAME=ongrid-manager. That is the transport that lets the manager reach edge hosts without the hosts accepting inbound connections.

The README states the consequence plainly: "Zero inbound ports, edge dials out; no port 22 / 80 / 443 on hosts." That is the single most consequential design decision in the project. It means the agent can reach hosts inside a private subnet without a jumpbox and without opening a firewall rule per host. Browser SSH is built on the same reverse tunnel, so a shell session is audited by construction rather than by a separate bastion.

Agent behaviour is layered. A coordinator receives the request and dispatches to specialist sub-agents for SRE, network or database work. Tool calls run against a bash sandbox plus what the README calls "26+ inspection tools," and every call is audited. The go.mod confirms the machinery: casbin for policy, eino for the agent framework, gosnmp for network device polling, and the OpenTelemetry SDK plus Prometheus client for instrumentation. A separate investigator path handles alerts by spawning an RCA worker that writes its conclusion back into chat.

## Installing Ongrid on a server and running a first investigation

The README gives a binary install rather than a package manager route. You pick the release archive for your server architecture, extract it, and run install.sh as root. The supported platforms listed are Ubuntu 22.04+, Debian 12+, and RHEL/Rocky 9. The commands below are the AMD64 path from v0.15.2; the ARM64 archive has the same name with linux-arm64 substituted.

```bash
wget https://github.com/ongridio/ongrid/releases/download/v0.15.2/ongrid-v0.15.2-linux-amd64.tar.xz
tar -xf ongrid-v0.15.2-linux-amd64.tar.xz && cd ongrid-v0.15.2-linux-amd64
sudo ./install.sh
```

The installer brings up the full stack, which the README summarises as "Self-host in one command." For mainland China, the README offers a CDN mirror at ongrid.cloud/dl/ with the same filenames, because GitHub downloads are slow there.

Before the first real use you need a model provider and a database. The .env.example defaults to MySQL with ONGRID_DB_DIALECT=mysql and an ONGRID_DB_DSN pointing at 127.0.0.1:3306; SQLite is available as an opt-in for single-user local development via ONGRID_DB_DIALECT=sqlite and ONGRID_DB_PATH. Model keys are set per provider, for example ONGRID_OPENAI_API_KEY with ONGRID_OPENAI_MODEL=gpt-4o. The README lists Anthropic, OpenAI, GLM, DeepSeek, Gemini and Kimi as supported, with hot routing between them.

```bash
ONGRID_DB_DIALECT=mysql
ONGRID_DB_DSN=ongrid:ongrid@tcp(127.0.0.1:3306)/ongrid?parseTime=true&charset=utf8mb4&loc=Local
ONGRID_OPENAI_API_KEY=
ONGRID_OPENAI_MODEL=gpt-4o
```

Two settings deserve attention before you expose anything. ONGRID_JWT_SECRET ships as change-me-to-a-long-random-string and must be replaced. ONGRID_ADMIN_EMAIL and ONGRID_ADMIN_PASSWORD are seeded on first startup if the email is unused; the .env.example warns that leaving both empty starts with an empty users table, so nobody can log in until you set them and restart. The host ports are ONGRID_HTTP_PORT=443 and ONGRID_HTTP_REDIRECT_PORT=80, fronted by nginx per ADR-008, with 0 disabling the redirect.

## Where Ongrid stops being the right tool

The install path is a single server, and the .env.example confirms it: MySQL on 127.0.0.1, one manager, one frontier broker. There is no documented multi-region or high-availability manager topology in the repository. If your incident response depends on the tool being available when a region is degraded, that is a gap you have to solve yourself, and the README does not document rollback for a failed upgrade.

The agent also depends on data you already have. The README says built-in observability wires Prometheus, Loki, Tempo and Grafana together, but the .env.example is explicit that the Prometheus integration is off until you deploy it: "Leave false until you actually deploy a Prom container." If you have no metrics, logs or traces for the affected service, the root-cause walk has nothing to correlate, and you are paying for an LLM call to be told that.

Write actions are the other boundary. The product tour shows an approval and write gate, and the README describes read-only host tools with a bash sandbox. That is a sensible default, but it means Ongrid does not autonomously remediate anything until a human approves. Teams expecting unattended auto-remediation will find the gate is the feature, not a limitation to route around.

Finally, the licence. AGPL-3.0 with a network copyleft clause is a real constraint for anyone embedding this in a commercial service. There is also a TRADEMARK.md in the repository root, which suggests the project separates code rights from name rights.

## How Ongrid differs from wiring your own agent to Grafana MCP

The obvious alternative is assembling the same thing yourself: a Grafana MCP server plus an LLM client, or a chat bot that shells out to kubectl and promtool. The difference is in what comes pre-integrated. Ongrid ships the topology graph, the edge dial-out transport, the chat channel adapters for Slack, Telegram, Larksuite, DingTalk and WeCom, the Casbin policy layer, and the approval gate as one stack. Building that from parts means writing the topology correlation and the audit trail yourself.

The cost of that integration is coupling. A hand-rolled setup lets you swap the agent framework, keep your existing bastion, and avoid running MySQL. Ongrid's go.mod shows a fairly opinionated set of dependencies, including the eino agent framework, fastembed-go for local embeddings, and geminio for the frontier transport. If you want to replace the agent loop, you are working inside someone else's abstractions.

A second comparison is against the observability vendors' own AI assistants. Those are hosted, priced per seat, and read only their own telemetry. Ongrid is self-hosted, reads across Prometheus, Loki and Tempo, and can reach hosts over SSH. The trade is operational burden: you run the database, the TLS terminator, the broker and the model keys.

## Maintenance, upgrade cost and the licence you are accepting

The project is not archived, and the last push was on 2026-09-10. Releases v0.15.0, v0.15.1 and v0.15.2 all landed within the same week in early September 2026, which tells you the release cadence is currently tight but says nothing about long-term stability. The Makefile describes itself as the single build, test and deployment entry point and says CI, Dockerfiles and the README should only call make targets rather than bare go build or docker build. That is a good sign for reproducible builds and a warning that building outside the Makefile is unsupported.

Upgrade cost is where you should look hardest. The Makefile notes that every production package caches both architectures so mixed fleets can be upgraded, and that edge binaries remain external CNB assets. The download URLs are architecture-specific, so an upgrade is a re-download and a re-run of install.sh rather than a package manager transaction. The README does not document a rollback procedure, and the .env.example does not mention schema migration behaviour, so verify how your database is migrated before upgrading a production instance.

The licence is AGPL-3.0. If you run a modified Ongrid as a network service for third parties, the copyleft terms likely reach your modifications; that is a question for your own counsel, not something this article can settle. TRADEMARK.md exists separately, which usually means the name is governed apart from the code.

## Conclusion

Ongrid fits teams that already run Prometheus, Loki, Tempo or Grafana and want an agent that reads the same data from chat, with write actions behind an approval gate. It is the wrong tool for anyone unwilling to run a MySQL-backed stack, expose a TLS port, or hold their own model keys. Before adopting, verify that your server matches the supported Ubuntu, Debian or RHEL/Rocky releases, that your chosen model provider is reachable from the host, and that your security team accepts the AGPL-3.0 network copyleft and the TRADEMARK.md terms.

## FAQ

### What is Ongrid used for?

Ongrid is an ops AI agent that answers infrastructure questions from chat channels such as Slack, Telegram, Lark or DingTalk. It correlates metrics, logs, traces and topology to find a root cause, and can run remote inspection on enrolled hosts.

### How do I install Ongrid on a Linux server?

Download the release archive for linux-amd64 or linux-arm64, extract it, and run sudo ./install.sh as root. The README lists Ubuntu 22.04+, Debian 12+ and RHEL/Rocky 9 as supported.

### Which chat platforms and model providers does Ongrid support?

The README lists Slack, Telegram, Larksuite, DingTalk and WeCom as chat channels, each with a per-channel locale. Model providers include Anthropic, OpenAI, GLM, DeepSeek, Gemini and Kimi, with hot routing between them.

### Does Ongrid need inbound ports opened on my hosts?

No. The README states that edge agents dial out, so there are no inbound ports such as 22, 80 or 443 on the monitored hosts. The manager connects through the frontier broker instead.

### Can Ongrid change my infrastructure on its own?

Not without approval. The product tour shows an approval and write gate, and the README describes read-only host tools running in a bash sandbox with every call audited.

### What licence does Ongrid use?

The repository is licensed under AGPL-3.0, and there is a separate TRADEMARK.md at the repository root. The network copyleft clause is worth reviewing with your own counsel before embedding it in a commercial service.

## Sources

- [License: AGPL-3.0](https://github.com/ongridio/ongrid/blob/main/LICENSE)
- [ongridio/ongrid on GitHub](https://github.com/ongridio/ongrid)
- [Project website](https://ongrid.cloud)
- [README](https://github.com/ongridio/ongrid/blob/main/README.md)
- [Releases](https://github.com/ongridio/ongrid/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ongridio-ongrid
