Superlog: self-hosted OpenTelemetry observability with agent runners
Open-source observability tool that uses AI agents to self-heal your software
At a glance
- What is it?
- Superlog is an open-core, local-first workspace that ingests OTLP traces, logs and metrics, groups noisy signals into incidents, and ships a default community agent runner. Here is what the repository actually contains, how to start the local stack, and where the design still leaves gaps.
- Who is it for?
- Adopt Superlog if you already emit OTLP and want the ingest, storage and incident-grouping path to live on your own machines, and you accept that the community agent runner only records a local incident summary. Do not adopt it if you need a documented upgrade or rollback procedure, or if you expect the open-source edition to investigate incidents on its own; the README names a hosted Cloud edition for that.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Superlog targets: OTLP data that nobody groups
Most teams that adopt OpenTelemetry end up with the same shape of problem. The collector works. Traces, logs and metrics land somewhere. What is missing is the layer that turns a thousand near-identical error spans into one thing a human can look at. Superlog positions itself exactly there. The README describes it as an "open-source agentic telemetry system" that "ingests traces, logs, and metrics, groups noisy signals into incidents, and watches your infra while you sleep." The audience is engineers running production systems who already speak OTLP and do not want to hand their telemetry to a vendor. The repository is explicit that this is the community edition: the web app and API, the OTLP ingest proxy, the worker processes, the Postgres schema, ClickHouse-backed telemetry queries, and agent runner interfaces all live here. A hosted Superlog Cloud edition with a free tier, a pay-to-go plan and monthly credit packs exists separately. That split matters more than the feature list, because it tells you which parts of the product you are actually self-hosting.
How the ingest and incident-grouping pipeline is put together
The architecture is visible in the repository layout rather than in prose. Five apps and two packages carry the work: apps/web is a Vite and React frontend, apps/api is the HTTP API, apps/proxy is the OTLP intake proxy, apps/worker holds background workers and agent orchestration, and packages/db holds the Drizzle schema and migrations. packages/fingerprint contains telemetry fingerprinting helpers, which is the piece that makes grouping possible at all. Two datastores are used for different jobs. Postgres holds the application schema that Drizzle migrates. ClickHouse holds telemetry and answers the queries, and the bundled collector writes into it directly: the docker-compose file sets CLICKHOUSE_ENDPOINT to tcp://clickhouse:9000 and CLICKHOUSE_DB to superlog on the otel/opentelemetry-collector-contrib service. That is a deliberate split. Incident records, projects and configuration are relational; span and log volume is not. The fingerprint package plus the worker is the grouping mechanism the README refers to, and the agent runner interfaces are the extension point where a pluggable investigation runtime would attach. The default community agent runner, per the README, "records a local incident summary." Read that sentence twice. The open-source edition groups and summarises; it does not claim to remediate anything on its own.
Installing Superlog and starting the local stack
The README gives two entry points. The first is through a coding agent, using the project's skills repository:
Run npx skills add superloglabs/skills --all and use the skills to install Superlog in this projectThat line is quoted as it appears; note that it is written as an instruction to an agent rather than as a shell command you would paste. The second entry point is the manual quick start, and it is the one to follow if you want to see what is happening. Prerequisites are Node.js 20 or newer, pnpm 9 or newer, and Docker.
pnpm install
docker compose up -d
pnpm --filter @superlog/db db:migrate
pnpm devThe compose file brings up three services: postgres:16 on host port 5434 by default, clickhouse/clickhouse-server:26.1 on 8123 and 9000, and the OpenTelemetry collector on 4317 (gRPC) and 4318 (HTTP). Ports are parameterised, so POSTGRES_HOST_PORT, CLICKHOUSE_HTTP_HOST_PORT, CLICKHOUSE_TCP_HOST_PORT, COLLECTOR_GRPC_HOST_PORT and COLLECTOR_HTTP_HOST_PORT can all be overridden. The collector mounts ./infra/collector/config.yaml read-only and waits for ClickHouse to report healthy before starting. After pnpm dev, the README lists the default local services as Web on http://localhost:5173, API on http://localhost:4100, and OTLP intake on http://localhost:4101. A first real use is to point an existing OTLP exporter at the intake on 4101, or at the bundled collector on 4317 or 4318, and watch incidents appear in the web app. The repository also ships demo scripts, including demo:bootstrap:acme, demo:seed:acme-telemetry and demo:seed:everything, which generate telemetry without needing your own services.
Where the community edition stops short
The most important limitation is stated plainly and then softened by the marketing around it. The default community agent runner records a local incident summary. There is no claim in the README that the open-source edition diagnoses root causes, proposes fixes or takes action. The agent runner interfaces are described as interfaces for pluggable investigation runtimes, which means the investigation capability is something you supply or something the hosted edition supplies. If your reason for looking at Superlog is the phrase "self-heal your software" in the repository description, the community edition is not that. A second constraint is operational weight. This is not a single binary. You are running Postgres, ClickHouse and an OpenTelemetry collector before the application services start, and ClickHouse is configured with a nofile soft and hard limit of 262144 in the compose file. That is a real footprint for a small team. Third, the README documents no upgrade path, no migration rollback and no version compatibility matrix between the app and the collector config. The CHANGELOG.md and ROADMAP.md files exist at the top level, but the README does not tell you what an upgrade involves. Finally, the repository contains a server.json and a skills-lock.json, which suggests the MCP surface is part of the product, but the README does not document how to use it.
Superlog against Grafana and a plain collector pipeline
The honest comparison is not against another agentic tool. It is against the pipeline most OTLP users already have: an OpenTelemetry collector writing to ClickHouse or Prometheus, with Grafana reading from it. That stack is mature, heavily documented, and you configure alert rules by hand. Its weakness is exactly Superlog's premise. Alert rules fire on thresholds you wrote in advance, and they do not group a burst of related failures into one incident with a summary. Superlog's fingerprinting package and worker exist to do that grouping, and the incident is the unit the product is built around rather than the individual alert. The trade is control for opinion. With a collector-plus-Grafana setup you choose every storage detail and every query. With Superlog you accept its Postgres schema, its ClickHouse layout and its incident model, and you get the grouping for free. If your team already has dashboards and runbooks built on raw queries, Superlog will sit beside that stack rather than replace it, and you should treat the OTLP intake proxy as an additional consumer rather than a migration target.
Maintenance, licence and what an upgrade actually costs
The repository is not archived, and the last push was on 2026-09-04, which puts it within recent weeks. That is the only maintenance signal available here; there are no retrieved releases, so there is no published version history to reason about. The root package.json pins packageManager to [email protected] and engines.node to >=20.0.0, and the workspace is a Turbo monorepo with build, dev, lint, format and typecheck scripts. Lint and format run through Biome. On the licence side, the repository is Apache-2.0, stated in both the README badge and the package.json license field, with the full text at LICENSE.md. Apache-2.0 permits commercial use and modification and includes a patent grant, but it also requires that you preserve notices and state significant changes. That is a general property of the licence, not advice about your situation. The practical upgrade cost is the part the documentation does not cover. Because the collector config is mounted from infra/collector/config.yaml and the schema is managed by Drizzle migrations under packages/db, an upgrade touches three moving parts at once: the application, the database schema and the collector. Budget for reading CHANGELOG.md yourself before each bump, because the README will not warn you.
Editorial conclusion
Adopt Superlog if you already emit OTLP and want the ingest, storage and incident-grouping path to live on your own machines, and you accept that the community agent runner only records a local incident summary. Do not adopt it if you need a documented upgrade or rollback procedure, or if you expect the open-source edition to investigate incidents on its own; the README names a hosted Cloud edition for that. Before committing, run the local stack, open http://localhost:5173, and confirm that your collector can reach the OTLP intake on port 4101.
Frequently asked questions
How does Superlog work?
It ingests traces, logs and metrics over OTLP, uses fingerprinting helpers to group noisy signals into incidents, and stores telemetry in ClickHouse while the application schema lives in Postgres. The bundled OpenTelemetry collector writes directly into ClickHouse at tcp://clickhouse:9000, and worker processes handle incident grouping and background jobs.
What are the prerequisites for running Superlog locally?
The README lists Node.js 20 or newer, pnpm 9 or newer, and Docker. The root package.json pins [email protected] and requires node >=20.0.0, and docker compose brings up Postgres, ClickHouse and the OpenTelemetry collector.
Does the open-source Superlog edition fix incidents automatically?
No. The README states that the default community agent runner records a local incident summary, and describes the agent runner interfaces as a way to plug in investigation runtimes. Remediation is not claimed for the community edition.
Which ports does Superlog use by default?
The README lists the web app on http://localhost:5173, the API on http://localhost:4100 and OTLP intake on http://localhost:4101. The compose file additionally exposes the collector on 4317 and 4318, ClickHouse on 8123 and 9000, and Postgres on host port 5434.
What licence is Superlog released under?
Apache-2.0, declared in the package.json license field and shown in the README badge, with the text at LICENSE.md. The README also notes a separate hosted Superlog Cloud edition with a free tier and paid plans.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/superloglabs-superlog)