slok/sloth: a Prometheus SLO generator with a CLI and a Kubernetes operator
🦥 Easy and simple Prometheus SLO (service level objectives) generator
At a glance
- What is it?
- sloth turns a short YAML SLO spec into Prometheus recording rules and multi window multi burn alerts, either from a single binary or as a Kubernetes controller. Here is how it installs, what it generates, and where it stops being the right tool.
- Who is it for?
- Adopt sloth if you already run Prometheus with Prometheus-operator and want one uniform SLO spec across teams, and start with the CLI plus the validate command before touching the Kubernetes controller. Skip it if your SLOs need to be computed outside Prometheus or to drive paging through something other than Prometheus alert rules.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem sloth solves is spec drift, not alerting
Writing one Prometheus alert for one service is easy. Writing the same alert for forty services across six teams is not, because every team picks its own window, its own burn rate thresholds and its own label names, and six months later nobody can compare the numbers. sloth's answer is to make the SLO spec the only artifact a team writes. The README describes the goal as generating "understandable, uniform and reliable Prometheus SLOs for any kind of service" from "a simple SLO spec" that results in multiple metrics and multi window multi burn alerts.
The audience is platform and observability engineers who already run Prometheus and want SLOs expressed as rule files rather than as a dashboard someone maintains by hand. The spec is deliberately small: a version string, a service name, service level labels, and a list of SLOs, each with a name, an objective, an SLI and an alerting block. The multi window multi burn alerting method comes from the Google SRE workbook, which the README links as the basis for both the SLO implementation and the alert framework. That link matters: sloth does not invent an alerting theory, it automates one that was already written down.
How the generator turns one spec into rule groups
The input is a YAML document with version "prometheus/v1". Each SLO carries an sli block, and the SLI is where the queries live. In the events form there are two PromQL expressions, error_query and total_query, and both use a {{.window}} placeholder. sloth substitutes that placeholder for each time window it needs, which is how one pair of queries becomes a set of recording rules covering several ranges. The repository also ships a windows/ directory under examples, which matches the README's mention of customizable SLO period windows alongside the default 30 and 28 day periods.
From there the generator emits three kinds of output: SLI recording rules per window, SLO metadata rules, and the multi window multi burn alert rules split into a page alert and a ticket alert. The alerting block in the spec lets a team override the alert name, add labels such as severity and routing_key, and set annotations, and the feature list says labels can be customized and different alert types disabled. The CLI is a single binary, and the repository layout shows cmd/, internal/ and pkg/ alongside deploy/, docker/ and docs/, with a Kubernetes controller mode backed by CRDs. go.mod confirms the shape of the thing: prometheus-operator API and client packages for the Kubernetes side, the OpenSLO library, and kooper as the controller framework.
Installing sloth and generating your first rules
The README's getting started section is one command against a bundled example. It assumes sloth is on your PATH; the project points to sloth.dev for documentation, and the repository ships a Makefile whose build targets produce the binary and the container images, with ghcr.io/slok/sloth as the production image name.
sloth generate -i ./examples/getting-started.ymlThe input file is examples/getting-started.yml, a service named myservice with a single SLO named requests-availability at 99.9 percent. Its SLI is the events form: an error_query matching code=~"(5..|429)" and a total_query, both wrapped in sum(rate(...[{{.window}}])). The alerting block names the alert MyServiceHighErrorRate and gives the page alert severity pageteam and the ticket alert severity slack with a slack_channel label.
version: "prometheus/v1"
service: "myservice"
slos:
- name: "requests-availability"
objective: 99.9
sli:
events:
error_query: sum(rate(http_request_duration_seconds_count{job="myservice",code=~"(5..|429)"}[{{.window}}]))
total_query: sum(rate(http_request_duration_seconds_count{job="myservice"}[{{.window}}]))What you should see is a Prometheus rule file. The README points at examples/_gen/getting-started.yml as the exact result of that spec, so diffing your output against it is the fastest way to confirm the version you installed behaves as documented. For GitOps and CI, the feature list names a validate command, which is the piece worth wiring into a pipeline: it checks the spec before the generated rules ever reach a Prometheus instance.
sloth validate -i ./examples/getting-started.ymlKubernetes users have a second path. The repository contains deploy/ and examples/k8s-getting-started.yml, and the README lists a Kubernetes Controller/operator mode with CRDs plus Prometheus-operator support, so the same spec can be applied as a custom resource instead of piped through the CLI.
Where sloth is the wrong tool
sloth generates Prometheus rules. That is the whole output surface, and it constrains everything else. If your alerting path is not Prometheus rule files, if paging runs through a system that evaluates its own conditions rather than consuming alert rules, sloth has nothing to hand it. The alerting block can attach labels and annotations, but the evaluation still happens inside Prometheus.
The SLI queries are the second boundary. Both error_query and total_query are PromQL, and the {{.window}} placeholder is a string substitution, so the query has to be valid for every window sloth generates. A query that is only correct over a 5 minute range, or one that references a recording rule that does not exist yet, will produce rules that load and then evaluate to nothing. Nothing in the generator can tell you that the metric you named is the metric you meant; that check is yours.
There is also a real question about how much of the alerting decision is now hidden. The multi window multi burn thresholds, the window set and the page versus ticket split are produced by sloth from a specification the user does not write. That is the point of the tool, and it is also the cost: tuning the burn rate behaviour means going through sloth's own configuration surface rather than editing the rules directly, and any hand edit to the generated file is lost on the next generate. The README does not document a rollback path for generated rules, so recovery from a bad spec means regenerating and reapplying, not reverting a stored artifact.
sloth against hand written rules and against OpenSLO tooling
The obvious alternative is writing the recording and alerting rules yourself. That gives full control over every threshold and window, and it needs no extra binary in the pipeline. The difference is volume: one SLO in sloth's format expands into multiple recording rules across windows plus metadata rules plus two alert tiers, and doing that by hand for every service is where inconsistency creeps in. If you have three services and one team, hand written rules are defensible. At twenty services, the uniformity argument wins.
The second comparison is OpenSLO. sloth supports OpenSLO, and go.mod pulls in github.com/OpenSLO/oslo v0.12.0, with examples/openslo-getting-started.yml and examples/openslo-kubernetes-apiserver.yml in the tree. The approaches differ in what the spec is for. OpenSLO is a vendor neutral specification for describing SLOs; sloth is a generator that consumes a spec and emits Prometheus artifacts. Using OpenSLO input with sloth means your SLO definitions are portable even though the generated rules are Prometheus specific. Using sloth's own prometheus/v1 format is more direct and exposes sloth specific knobs such as alert type disabling, but ties the spec to this tool. Which one is right depends on whether anything other than Prometheus will ever read these definitions.
Maintenance, licence and the cost of upgrading
The repository is not archived, and the last push was on 2026-09-04. The most recent release listed is v0.16.0 from 2026-04-04, with a matching Helm chart release sloth-helm-chart-0.16.0 the same day, and v0.15.0 before it on 2025-10-31. Two releases roughly six months apart is a slow but real cadence, and the README's project status section states that work is underway on internal improvements to make the project more extensible, with more flexibility planned for how SLOs are generated. That statement is a plan, not a shipped feature.
Upgrade cost concentrates in the generated output. Because sloth owns the rule content, a change in how windows or burn rates are emitted shows up as a diff across every generated rule file in your repository, and that diff has to be reviewed and applied. The validate command is the cheap half of that work; the expensive half is re-generating and re-applying rules across every service. Pinning the version and regenerating in a branch before merging is the practical pattern.
The licence is Apache-2.0, and the README carries the standard Apache 2 badge pointing at the LICENSE file. That is a permissive licence, but the repository also depends on prometheus-operator packages and the OpenSLO library, so if you vendor or redistribute anything built from this tree, read the licence files of those dependencies rather than assuming the top level licence covers the whole distribution. This is not legal advice; the LICENSE file and the dependency licences are the sources that matter.
Editorial conclusion
Adopt sloth if you already run Prometheus with Prometheus-operator and want one uniform SLO spec across teams, and start with the CLI plus the validate command before touching the Kubernetes controller. Skip it if your SLOs need to be computed outside Prometheus or to drive paging through something other than Prometheus alert rules. Before committing, run sloth generate against examples/getting-started.yml and diff the output against examples/_gen/getting-started.yml, then check that the generated rule groups land in the namespace or file your Prometheus instance actually loads.
Frequently asked questions
How do I install sloth?
The README's getting started section shows the CLI being used directly, and the documentation is at sloth.dev. The repository ships a Makefile whose build targets produce the binary and the container images, with ghcr.io/slok/sloth as the production image name, and a Helm chart release exists alongside v0.16.0.
How do I use sloth to generate SLO rules?
Run sloth generate with an input spec, as in the README's example against examples/getting-started.yml. The spec declares a version of prometheus/v1, a service name and a list of SLOs with an objective, an SLI and an alerting block, and the output is a Prometheus rule file.
Does sloth work with Kubernetes and Prometheus-operator?
Yes. The README lists Kubernetes Controller/operator mode with CRDs and Prometheus-operator support, and go.mod depends on the prometheus-operator API and client packages. The repository includes deploy/ and Kubernetes example files such as examples/k8s-getting-started.yml.
Can sloth validate an SLO spec before it reaches Prometheus?
The feature list names SLO spec validation including a validate command intended for GitOps and CI. Running it against a spec file checks the spec before the generated rules are applied.
What SLI formats does sloth support?
The README lists different SLI types, SLI plugins, a library of common SLI plugins, and OpenSLO support. The getting started example uses the events SLI form with error_query and total_query PromQL expressions containing a {{.window}} placeholder.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/slok-sloth)