Open-source project
keephq/keep avatar
keephq/keep

Keep (keephq/keep): an open source AIOps and alert management platform you can run with Docker

The open-source AIOps and alert management platform

12,365 stars1,519 forksPythonNOASSERTION

At a glance

What is it?
Keep is a Python alert management layer that ingests alerts from monitoring tools, deduplicates and enriches them, and triggers workflows. It is aimed at teams whose alert sources have outgrown a single inbox, and the docker-compose.yml in the repository is the fastest way to see whether that trade is worth making.
Who is it for?
Adopt Keep if you already run several monitoring tools and want one place where alerts are deduplicated, enriched and routed, and if you are willing to operate a backend, a UI and a websocket server yourself. Do not adopt it if a single alert source and a chat channel already cover your needs, or if you need a vendor to carry the pager.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Keep solves, and who ends up running it

Alert fatigue in a multi-tool estate is a routing problem before it is a machine learning problem. Five monitoring systems each fire their own version of the same incident, each with its own severity scale, and a human reconciles them by hand. Keep positions itself as the layer between those tools and the responder: one pane for alerts and incidents, with deduplication, correlation, filtering and enrichment applied on the way in. The README describes the combination as a Swiss Army Knife for alerts, which is marketing language, but the underlying claim is concrete enough: bi-directional integrations with monitoring tools plus workflows that act on what arrives.

The intended user is a platform or SRE team that already has observability coverage and now needs a control point. A single-service shop running one Prometheus and one Slack channel is not the audience; there is nothing to deduplicate and nothing to correlate. The audience is the team whose on-call rotation is spending its first twenty minutes of every incident working out which of four alerts is the real one. Keep is Python, and the repository carries a keep-ui directory alongside the keep package, so you are adopting a backend service and a separate frontend rather than a library you import.

How alerts flow through Keep: providers, deduplication, workflows

The repository layout makes the architecture legible without reading source. Providers live under examples/providers, workflows under examples/workflows, and the README links a provider documentation page for each integration, from AppDynamics and Azure Monitoring to CloudWatch and Checkmk. A provider is the adapter: it defines how a monitoring tool's alerts enter Keep and, where the integration is bi-directional, how state changes travel back. That bi-directional claim is the design decision worth noticing. A one-way alert forwarder is simple; pushing acknowledgement or resolution back to the source system means Keep holds credentials for the tools it talks to, which widens the blast radius of a misconfiguration.

Once an alert is in, the pipeline applies deduplication, correlation, filtering and enrichment. Enrichment is where the AI backends attach: the README lists Anthropic, OpenAI, DeepSeek, Ollama, LlamaCPP, Grok and Gemini as options for enrichment, correlation and incident context gathering. Listing Ollama and LlamaCPP next to hosted APIs matters, because it means the enrichment step can run against a local model, which is the difference between usable and unusable for teams that cannot send alert payloads to a third party.

Workflows are the action layer, described in the README as GitHub Actions for your monitoring tools. The comparison is apt in one respect: a workflow is a triggered, declarative unit rather than a script you cron. It is a weaker comparison in another, because workflows here run inside a service you operate, not on managed runners.

Installing Keep with docker-compose and seeing the first alert

The repository ships several compose files, and the default docker-compose.yml is the shortest path to a running stack. It defines four services: keep-frontend, keep-backend, keep-websocket-server, and a grafana service behind a profile. Both frontend and backend are pulled as prebuilt images from a Google Artifact Registry path, so the default path does not build from source. The compose file sets AUTH_TYPE=NO_AUTH on both, which is appropriate for a local look and wrong for anything reachable from a network.

Start the stack from the repository root. The compose file does not name a command, so the standard Compose invocation applies:

yaml
services:
  keep-frontend:
    extends:
      file: docker-compose.common.yml
      service: keep-frontend-common
    image: us-central1-docker.pkg.dev/keephq/keep/keep-ui
    environment:
      - AUTH_TYPE=NO_AUTH
      - API_URL=http://keep-backend:8080

That excerpt is the frontend service as it appears in docker-compose.yml. The frontend talks to the backend over the compose network at http://keep-backend:8080, which is the API_URL value. The backend service in the same file sets AUTH_TYPE=NO_AUTH, PROMETHEUS_MULTIPROC_DIR=/tmp/prometheus and KEEP_METRICS=true. Both the frontend and the backend mount ./state, so alert state survives a container restart as long as you keep that directory.

The grafana service also sits in that file behind a profile, with ports 3001:3000, GF_SECURITY_ADMIN_USER=admin, GF_SECURITY_ADMIN_PASSWORD=admin and GF_USERS_ALLOW_SIGN_UP=false:

yaml
  grafana:
    image: grafana/grafana:latest
    profiles:
      - grafana
    ports:
      - "3001:3000"
    environment:
      - GF_SECURITY_ADMIN_USER=admin
      - GF_SECURITY_ADMIN_PASSWORD=admin
      - GF_USERS_ALLOW_SIGN_UP=false

The prometheus service in the same file exposes 9090:9090, reads ./prometheus/prometheus.yml and passes --config.file=/etc/prometheus/prometheus.yml. The README does not give a first-alert walkthrough, so for a real integration the reference points are the provider and workflow examples under examples/providers and examples/workflows, plus the platform documentation at docs.keephq.dev. Two compose variants are worth knowing about before you commit: docker-compose-with-auth.yml for deployments that need authentication, and docker-compose-with-otel.yaml for OpenTelemetry wiring. The backend dependency list includes opentelemetry-instrumentation-fastapi and the OTLP exporters, so tracing is a supported path rather than an afterthought.

Where Keep is the wrong tool

Keep is a service you run, and that is the limitation that decides most evaluations. The default compose file runs three containers before you add Grafana and Prometheus, plus a state directory you are responsible for. If your team has no one who wants to own a backend upgrade path, adding Keep converts an alerting problem into an operations problem. The README does not document a rollback procedure for a failed upgrade, and the repository's release cadence is brisk, with v0.54.1, v0.54.2 and v0.54.3 landing within roughly three months. Frequent releases are not a defect, but they do mean the upgrade path gets exercised often and should be tested on your own state directory before it touches production.

The Python version constraint is narrower than it looks. pyproject.toml requires >=3.11 and <3.14, so a deployment pinned to an older interpreter cannot install the package, and the upper bound means a jump to a newer Python release needs a dependency review first.

Licensing is the other unresolved item. The repository reports the licence as NOASSERTION, which means the automated detection could not classify the LICENSE file. That is not the same as saying there is no licence, and it is not something to resolve by guessing. Read LICENSE yourself before you build a commercial plan on top of Keep, and treat the question as open until you have.

Keep against a general-purpose workflow engine

The obvious alternative for a team that mostly wants automation is a general-purpose workflow engine such as n8n, which connects services and runs steps in response to events. The difference is where the alert semantics live. In n8n you model deduplication and correlation yourself, as nodes and branches over whatever payload arrives, and you own the schema of an alert. Keep treats the alert as a first-class object with providers that already know how to parse a given monitoring tool's payload, and it ships the deduplication, correlation and enrichment stages as part of the pipeline rather than as something you assemble.

That cuts both ways. A workflow engine is more general: it will automate tasks that have nothing to do with alerts, and it does not ask you to adopt an alert data model. Keep is narrower and, for the specific job of consolidating noisy monitoring output, arrives with more of the work already done. If your problem is genuinely alert-shaped, the narrower tool is the better fit. If you are looking for one automation layer for everything, Keep will be a second system to run alongside it.

The same reasoning applies to the AI enrichment features. Keep gives you a choice of backends, including local ones, but the value depends on your alerts being in Keep in the first place.

Maintenance cost, releases and the licence question

The last push to the default branch was on 2026-09-19, two days before this was written, and the most recent release is v0.54.3 from 2026-09-09. The repository is not archived. On the evidence of the push dates and release tags, this is a project under current development, and the practical consequence for an adopter is that you should expect to move. The gap between v0.54.1 in late June and v0.54.2 in mid July, then v0.54.3 in early September, suggests a steady rather than frantic cadence.

Upgrade cost is concentrated in two places. The first is the state directory mounted at ./state by both the frontend and the backend; any migration that touches the alert schema runs against that data, and the compose file gives no backup step. The second is the image tags. The default compose file references images without a version tag, so pulling again moves you to whatever is current. Pinning is a change you make, not one the repository makes for you.

On the licence, the repository reports NOASSERTION. The only responsible statement is that the LICENSE file exists at the top level and you should read it before deciding how Keep fits a commercial deployment. Nothing in the README or the package metadata resolves the question, and the pyproject.toml authors field names Keep Alerting LTD without stating terms.

What to check before you commit to Keep

Three checks decide most adoptions. First, open the provider documentation for the monitoring tool you actually run and confirm it is listed; the README links a page per provider, and the list in the README itself is truncated, so absence from the README is not absence from the docs. Second, confirm your deployment target satisfies the Python range of >=3.11 and <3.14 if you plan to install the package rather than run the images. Third, decide whether NO_AUTH is acceptable anywhere beyond your laptop; docker-compose-with-auth.yml exists precisely because it is not.

If those three pass, the compose stack is a cheap experiment. The state directory is local, the images are prebuilt, and the Grafana profile gives you a metrics view of the backend without extra configuration. What you are really evaluating is whether the deduplication and enrichment stages reduce the number of alerts a human has to read, and that answer only shows up once real alerts are flowing through a real provider.

Editorial conclusion

Adopt Keep if you already run several monitoring tools and want one place where alerts are deduplicated, enriched and routed, and if you are willing to operate a backend, a UI and a websocket server yourself. Do not adopt it if a single alert source and a chat channel already cover your needs, or if you need a vendor to carry the pager. Verify first that the provider you depend on exists in the documentation, that your Python version falls inside the pyproject.toml range, and that you have read the LICENSE file, since the repository reports the licence as NOASSERTION rather than naming one.

Frequently asked questions

What are some open source AIOps tools?

Keep describes itself as an open-source AIOps and alert management platform, offering a single pane of glass, alert deduplication, enrichment, filtering and correlation, bi-directional integrations and workflows. The repository is public on GitHub, and the README points to docs.keephq.dev for the full provider list.

How do I install Keep with Docker?

The repository root contains docker-compose.yml, which defines keep-frontend, keep-backend and keep-websocket-server services using prebuilt images. The README does not give an install command, so follow the platform documentation at docs.keephq.dev for the deployment steps.

Which Python version does Keep require?

The pyproject.toml file declares python = ">=3.11,<3.14". A deployment pinned to an older interpreter cannot install the package, and the upper bound means a newer Python release needs a dependency review first.

Which AI backends can Keep use for enrichment and correlation?

The README lists Anthropic, OpenAI, DeepSeek, Ollama, LlamaCPP, Grok and Gemini as AI backends for enrichments, correlations and incident context gathering. Ollama and LlamaCPP allow the enrichment step to run against a local model.

What licence does Keep use?

The repository reports the licence as NOASSERTION, meaning automated detection could not classify the LICENSE file. The file exists at the top level and should be read directly before making a commercial decision.

Official sources

  1. Issues
  2. keephq/keep on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/keephq-keep.svg)](https://hysenlabs.com/projects/keephq-keep)