ccfos/nightingale: an open source alerting engine that sits on top of your existing metrics
Nightingale is to monitoring and alerting what Grafana is to visualization.
At a glance
- What is it?
- Nightingale is a Go alerting server that connects to VictoriaMetrics, ElasticSearch and similar stores, then handles rule evaluation, alarm generation and notification. It is not a collector, and it is not an on-call product.
- Who is it for?
- Adopt Nightingale if you already have metrics and logs landing in a time-series store or ElasticSearch and the missing piece is rule evaluation and alarm distribution, including the n9e-edge mode for data centers with poor connectivity to the center. Do not adopt it as a collector (the README points to Categraf) or as a replacement for on-call scheduling, escalation and collaborative handling, where the README itself recommends PagerDuty and FlashDuty.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Nightingale fills: rule evaluation, not collection or visualization
Most observability stacks end up with two separate halves. One half collects and stores data. The other half draws it. The README describes Nightingale as an open-source monitoring project that focuses on alerting, and draws the comparison directly: like Grafana it connects to various existing data sources, but where Grafana emphasizes visualization, Nightingale emphasizes the alerting engine plus the processing and distribution of alarms.
That framing tells you who this is for. The intended user already has metrics and logs stored somewhere, and the missing component is the thing that evaluates conditions on a schedule, turns them into events, and pushes those events to humans. Nightingale is explicit that it does not provide monitoring data collection. The README recommends Categraf as the collector, which gathers data from operating systems, network devices, middleware and databases and pushes it via the Prometheus Remote Write protocol.
So the boundary is deliberate. If you have nothing collecting data yet, Nightingale is the wrong first install. If you have data and no alerting discipline, it is aimed at you.
Data flow: remote write in, time-series store behind, alarms out
The path described in the README is short. A collector pushes monitoring data to Nightingale over Prometheus Remote Write. Nightingale stores that data in a time-series database such as Prometheus or VictoriaMetrics. Alert rules and notification rules are configured inside Nightingale against connected data sources, which the README lists as VictoriaMetrics, ElasticSearch and similar stores.
The repository layout matches that split. Top-level directories include pushgw for the write path, alert for the alerting engine, center for the main server, tsdb and storage for data handling, and notify-related code under the alert tree. The Makefile builds these as separate binaries: n9e from cmd/center, n9e-alert from cmd/alert, and n9e-pushgw from cmd/pushgw. That means the write gateway can be scaled or restarted independently of the alerting engine.
There is also a distributed mode worth knowing about before you design a topology. For edge data centers with poor network connectivity to the central server, Nightingale offers a distributed deployment mode for the alerting engine. The README's diagram has Data Center A using the central Nightingale process as its alerting engine, while Data Center B deploys n9e-edge to handle alerting for its own data sources. The claim is that even if the network is disconnected, alerting remains unaffected. That is a meaningful architectural commitment: rule evaluation moves next to the data rather than the data moving to the rules.
Installing Nightingale and getting to a first alert rule
The README points at the project documentation site for setup, and the repository ships a Makefile that builds the binaries from source. The Makefile's all target runs prebuild then build. prebuild calls fe.sh, which the Makefile describes as downloading and embedding the front-end file, so the UI is baked into the binary rather than served from a separate directory.
Building the main server produces a binary named n9e:
make buildAfter that command you should see an n9e binary in the repository root. The Makefile also defines build-edge, build-alert, build-pushgw and build-cli, producing n9e-edge, n9e-alert, n9e-pushgw and n9e-cli respectively. For a single-node evaluation, the main server is the one you want.
Running it in the background is what the Makefile's run target does:
make runThat target executes nohup ./n9e > n9e.log 2>&1 &, so output lands in n9e.log in the working directory. The configuration file referenced throughout the README is etc/config.toml, which is where the HTTP and MCP settings live.
The MCP endpoint is enabled without any configuration. The README states that the n9e process serves the Model Context Protocol at /mcp over Streamable HTTP, at http(s)://<nightingale>:17000/mcp, and that this is a root path rather than something under /api/n9e. Authentication uses a personal access token created in the web UI under Profile, then Token Management, sent in the X-User-Token header. A client configuration looks like this:
{
"mcpServers": {
"nightingale": {
"type": "http",
"url": "http://127.0.0.1:17000/mcp",
"headers": { "X-User-Token": "<your-token>" }
}
}
}One thing to check early: the README says /mcp reuses the [HTTP.TokenAuth] section for authentication, so keep that enabled. The endpoint is read-only by default, and write tools are an explicit opt-in.
The MCP endpoint is the most opinionated thing in this release
Nightingale now exposes an MCP server from the same process, and the README is unusually direct about the permission model. Every tool call is dispatched onto Nightingale's own HTTP API inside the process, carrying your token, so RBAC and business-group permissions apply exactly as they do for that user in the UI. The stated consequence is that a client can never reach anything its token's owner cannot. That is the right design choice, and it is also the reason the endpoint is read-only by default: registering the write tools is a deliberate configuration step, not something you get by accident.
The tool surface is large. The README counts 74 fine-grained tools, 42 read and 32 write, across 13 toolsets including alerts, targets, datasource, mutes, busi_groups, notify_rules, alert_subscribes, event_pipelines, users, metrics, logs, dashboards and roles. If you want to narrow that, MCPToolsets restricts the exposed set, and an empty value means all of them.
Authentication has two paths beyond the personal token. Nightingale can act as the authorization server itself, with RFC 7591 dynamic client registration and PKCE, which the README says lets hosted clients such as Claude or ChatGPT connect with zero pre-registration. Alternatively it can act as a resource server for an existing enterprise IdP (Keycloak, Entra ID, Okta, Auth0), mapping each token to its local user so permissions and audit stay per-person. The same process also exposes an A2A endpoint at /a2a wrapping a built-in AI assistant for agent-to-agent integration. If you would rather not run this inside the main process, the README names a standalone n9e-mcp-server that provides the same tools against a remote Nightingale.
Where Nightingale stops, in the project's own words
The README contains a section that most projects would not write. It lists advanced requirements, then says Nightingale is not suitable for them. The two examples given are consolidating events from multiple monitoring systems into one platform for unified noise reduction, response handling and data analysis, and supporting personnel scheduling, practicing on-call culture, alert escalation to avoid missing alerts, and collaborative handling. For those, the README recommends on-call products such as PagerDuty and FlashDuty, describing them as simple and easy to use.
That is an honest boundary and it should shape your evaluation. Nightingale generates alarms and distributes them by rule across 20 built-in notification media, which the README lists as including phone calls, SMS, email, DingTalk and Slack. It does not do rotation schedules, escalation policies or incident collaboration. If your problem is that alerts vanish into a shared channel and nobody owns them, Nightingale's notification layer will not fix that by itself.
A second limitation is structural rather than functional: Nightingale does not collect data. Everything downstream depends on a collector you supply. The README recommends Categraf, and the integration is described as pushing over Prometheus Remote Write. If your existing collection path does not speak that protocol or cannot be pointed at the push gateway, you are adding a component before you get any alerting value. The README does not document a rollback path for a data source or a rule set, so plan configuration changes as forward-only and keep your own copy of rule definitions.
Compared with running alerts inside Prometheus or Grafana
The obvious alternative is to keep alerting where the data already lives. Prometheus has its own rule evaluation and Alertmanager handles routing and grouping, and Grafana can evaluate alerts against the same sources Nightingale connects to. The difference in approach is where the state lives. With Prometheus rules, the rule file sits next to the scrape configuration and the alerting pipeline is coupled to that Prometheus instance. Nightingale separates the concern: the data source is a connection, the rules are objects managed through the UI and the API, and the alerting engine is its own process that can be deployed separately, including as n9e-edge next to a remote data center.
That separation is the whole argument. It buys you one place to manage rules across heterogeneous sources, which matters when some metrics live in VictoriaMetrics and some logs live in ElasticSearch, since the README lists both as connectable data sources. It costs you a component. You now run a server, a database for its own state, and a push gateway, and you have to keep their versions aligned.
The choice is not about which evaluates expressions better. It is about whether rule management should be a property of each storage system or a shared service on top of them. Nightingale bets on the second.
Licence, release cadence and what upgrades actually cost
Nightingale is Apache-2.0, which permits commercial use and modification and includes an explicit patent grant. The practical implication for adopters is that embedding it in a product or running it internally does not require publishing your changes. It also means the project carries no copyleft obligation back to you. This is a description of the licence text, not legal advice; if you are redistributing it inside a commercial product, have your own counsel read the NOTICE and attribution requirements.
The repository is not archived, and the last push was on 2026-08-18. The release history shows v9.0.0 on 2026-07-25, v9.1.0 on 2026-08-06 and v9.1.1 on 2026-08-18, so the 9.x line has been moving in roughly two-week steps. The go.mod declares module github.com/ccfos/nightingale/v6 and go 1.25.0, which means the module path version and the release tag version are not the same number. Do not assume a v9 tag corresponds to a v9 import path.
The upgrade cost that matters is the configuration surface. etc/config.toml is where HTTP and MCP settings live, and the README shows the MCP options grouped under an [HTTP.A2A] section, with DisableMCP noted as also turning off /a2a when set. Because /mcp is on by default and reuses [HTTP.TokenAuth], an upgrade that changes token handling changes both the UI and any MCP client you have connected. Pin a version, read the release notes for the two or three releases you skipped, and check that section of the config after each bump.
Editorial conclusion
Adopt Nightingale if you already have metrics and logs landing in a time-series store or ElasticSearch and the missing piece is rule evaluation and alarm distribution, including the n9e-edge mode for data centers with poor connectivity to the center. Do not adopt it as a collector (the README points to Categraf) or as a replacement for on-call scheduling, escalation and collaborative handling, where the README itself recommends PagerDuty and FlashDuty. Before committing, verify three things: that your store is reachable as a data source, whether you want the built-in /mcp endpoint left read-only or opted into write tools, and how the 20 notification media map onto the channels your team actually answers.
Frequently asked questions
How do I set up a Nightingale server?
Build the main server with make build, which produces an n9e binary, then start it with make run, which runs it under nohup and writes to n9e.log. Configuration lives in etc/config.toml, and the README points at the project documentation site for the full setup procedure.
How do I use the Nightingale app?
The workflow described in the README is to connect an existing storage repository such as VictoriaMetrics or ElasticSearch as a data source, then configure alert rules and notification rules inside Nightingale so alarms are generated and distributed. Nightingale itself does not collect data; the README recommends Categraf as the collector, pushing via Prometheus Remote Write.
Can an AI assistant manage Nightingale alerting?
Yes. The n9e process serves an MCP endpoint at /mcp over Streamable HTTP, and the README says any MCP client such as Claude Code, Claude Desktop, Cursor or ChatGPT connectors can query and manage Nightingale in natural language. Authentication uses a personal access token sent in the X-User-Token header, and the endpoint is read-only by default.
Does Nightingale collect monitoring data itself?
No. The README states plainly that Nightingale does not provide monitoring data collection capabilities, and recommends Categraf as the collector for operating systems, network devices, middleware and databases. Categraf pushes data to Nightingale over the Prometheus Remote Write protocol, and Nightingale stores it in a time-series database such as Prometheus or VictoriaMetrics.
Is Nightingale a replacement for PagerDuty?
No, and the README says so directly. It lists personnel scheduling, on-call culture, alert escalation and collaborative handling as requirements Nightingale is not suitable for, and recommends on-call products such as PagerDuty and FlashDuty instead. Nightingale covers rule evaluation and distribution across 20 built-in notification media.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ccfos-nightingale)