linkedin/Burrow: Kafka consumer lag checking without thresholds
Kafka Consumer Lag Checking
At a glance
- What is it?
- Burrow watches committed offsets across Kafka clusters and scores consumer groups on a sliding window instead of a fixed lag number. It is a Go service with an HTTP endpoint, email and HTTP notifiers, and a Docker Compose stack for local evaluation.
- Who is it for?
- Adopt Burrow if you run Kafka and want a single service that scores every committed-offset consumer group without you inventing lag thresholds per topic. Do not adopt it if your consumers commit offsets outside Kafka and outside the Zookeeper or Storm paths the README lists, or if you need a per-partition lag dashboard rather than a group status endpoint.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 39 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The threshold problem Burrow was built to remove
Most Kafka lag alerting starts with a number someone picked: alert when lag exceeds ten thousand messages. That number is wrong in both directions. A group processing a low-volume topic pages nobody at ten thousand, and a high-throughput group crosses ten thousand during a routine rebalance and pages everyone. The number also has to be maintained per topic, and topics multiply.
Burrow's README states the design goal directly: it provides consumer lag checking "without the need for specifying thresholds." Instead of comparing lag to a constant, it monitors committed offsets for all consumers and calculates the status of those consumers on demand. Groups are evaluated over a sliding window, per the features list. The audience is the team that already runs Kafka and wants one service covering every consumer group rather than a growing set of per-topic alert rules. It is a monitoring companion, not a broker-side component: it reads offsets and reports, it does not sit in the consumption path.
How the sliding window evaluation works
The mechanism visible in the repository is offset collection plus windowed evaluation. Burrow reads committed offsets, and the README lists three sources it can be configured for: Kafka-committed offsets, which it monitors automatically for all consumers, plus configurable support for Zookeeper-committed offsets and Storm-committed offsets. That distinction matters because a consumer group that commits to Zookeeper is invisible to a tool that only reads the Kafka offset topic.
The evaluation is a sliding window rather than a threshold check. Burrow keeps a history of offset positions and judges the group's status from how the offsets move across that window, which is why the README can promise no thresholds. A group that is behind but steadily advancing is a different case from a group that has stopped advancing, and the window is what lets the service tell them apart without a configured limit.
The service is a single Go binary. The dependency list in go.mod shows the pieces: IBM/sarama for the Kafka protocol, linkedin/go-zk for Zookeeper, httprouter for the HTTP endpoint, zap for logging, viper for configuration, and prometheus/client_golang for metrics. Notifiers are wired in as libraries: gomail.v2 for the emailer and the HTTP client path for sending alerts to another system. Storage of the offset history is not described in the README, which points to the wiki for configuration detail. If you need to know where that window data lives and how long it is retained, the wiki is the place to look, not the README.
Installing Burrow and reading your first consumer group
The README gives two paths. The source path requires Go: the minimum supported version is 1.24 per go.mod, and CI builds with Go 1.25. Clone the repository outside GOPATH, then build and install with the module commands the README lists.
go mod tidy
go installAfter that, run the binary with a config directory. The README's invocation is:
$GOPATH/bin/Burrow --config-dir /path/containing/configThe config directory is where the wiki-documented configuration files live. The repository ships a config/ directory and a docker-config/ directory with a burrow.toml, and the Dockerfile copies docker-config/burrow.toml to /etc/burrow/ and starts the binary with --config-dir /etc/burrow. That is the fastest way to see the service running without writing configuration first.
The Docker Compose file brings up Burrow alongside zookeeper and kafka, with the Burrow container on port 8000 and the config mounted from docker-config. The README's steps are to build and then run the stack:
docker-compose build
docker-compose down; docker-compose upThe compose file pre-creates three topics (test-topic, test-topic2, test-topic3) through KAFKA_CREATE_TOPICS, so there is something for the service to observe. Once the stack is up, the README says Burrow can be accessed at http://localhost:8000/v3/kafka. That endpoint is the cluster listing under the v3 API; the README describes the HTTP endpoint as serving consumer group status along with broker and consumer information, so expect to walk from the cluster path to a consumer group path to get a status verdict. The README does not spell out the full route list, so check the wiki or the running service for the exact group endpoint before you build a dashboard against it.
Where Burrow is the wrong tool
The offset-source constraint is the real boundary. Burrow reads Kafka-committed offsets automatically, and Zookeeper and Storm committed offsets only when configured. A consumer that tracks its position somewhere else, in an application database or a custom checkpoint store, produces no offset for Burrow to read, and no amount of configuration fixes that. The README is silent on any other offset source.
There is also a scope question. Burrow reports group status, not per-partition lag charts. If your operational question is "which partition of this topic is the straggler," this is not the layer that answers it; the HTTP endpoint returns consumer group status plus broker and consumer information. Teams that want time-series graphs of lag per partition usually pair a metrics pipeline with a broker-side exporter instead.
Operationally, Burrow is another service to run, and its alerting depends on the notifiers you configure. The README lists a configurable emailer for specific groups and a configurable HTTP client for all groups. The asymmetry is deliberate but easy to miss: the emailer is scoped per group, while the HTTP client sends for all groups. If you want per-group routing to a paging system over HTTP, that routing logic lives downstream of Burrow, not inside it. The README does not document an alert deduplication or suppression mechanism, so repeated status sends are a downstream concern.
Burrow against a broker-side lag exporter
The common alternative approach is a Prometheus exporter that scrapes Kafka broker metrics and exposes per-partition consumer lag as a gauge, with alert rules written in PromQL. The difference in approach is where the judgement happens. An exporter hands you raw lag numbers and leaves the threshold to your alert rules; you still choose the number, and you still maintain it per topic or per group.
Burrow moves that judgement into the service. It evaluates groups over a sliding window and returns a status, so the alerting layer consumes a verdict rather than a measurement. It also covers Zookeeper-committed and Storm-committed offsets, which a broker-metrics exporter generally does not, because those offsets are not in the broker's metric surface.
The trade-off runs the other way for graphing. A metrics exporter fits an existing Prometheus and Grafana stack with no new service and gives you the raw series to chart however you like. Burrow is a separate process with its own configuration directory, its own HTTP API, and its own notifier configuration. If your team already lives in PromQL and wants lag as a time series, adding Burrow means operating a second system to answer a question your existing one partly answers. go.mod does include prometheus/client_golang, so Burrow can expose metrics of its own, but the README does not describe that surface, so treat it as unverified until you read the wiki.
Maintenance, releases and the Apache-2.0 terms
The repository is not archived, and the last push was on 2026-08-21. The release history shows v1.9.6 on 2026-05-11, v1.9.5 on 2025-10-03, and v1.9.4 on 2025-05-15, so the cadence is roughly a minor release every five to eight months, with the most recent one a few months before the last push. The presence of .goreleaser.yml and Dockerfile.gorelease indicates releases are built through GoReleaser, and the Dockerfile pins golang:1.25.1-alpine for the builder stage and alpine:3.22 for the runner, so container builds are reproducible against those tags.
Upgrade cost is low on the surface: the artifact is a single static binary (the Makefile builds with CGO_ENABLED=0), and the Docker image is a two-stage build that copies the binary and a config file. The friction is configuration compatibility. The README defers all configuration questions to the wiki, and the repository carries a CHANGELOG.md, so the practical upgrade step is to read the changelog for the version you are moving to and diff your config directory against the shipped config/ and docker-config/ examples. Nothing in the README describes a config migration tool or a schema version field.
Licensing is Apache-2.0, per the LICENSE file and the README's licence section, with copyright held by LinkedIn Corp. The README reproduces the standard grant and the disclaimer that the software is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND. There is also a NOTICE file at the repository root, which Apache-2.0 expects redistributors to carry. If you vendor or redistribute Burrow, that NOTICE file and the attribution requirements are the parts to hand to whoever handles compliance; this is a description of what the repository contains, not legal advice.
Editorial conclusion
Adopt Burrow if you run Kafka and want a single service that scores every committed-offset consumer group without you inventing lag thresholds per topic. Do not adopt it if your consumers commit offsets outside Kafka and outside the Zookeeper or Storm paths the README lists, or if you need a per-partition lag dashboard rather than a group status endpoint. Before rollout, verify the v3 HTTP responses against your own cluster, confirm which offset storage your consumers actually use, and read the wiki for the notifier and storage settings, since the README defers configuration to it.
Frequently asked questions
What is linkedin/Burrow?
It is a monitoring companion for Apache Kafka that provides consumer lag checking as a service without requiring thresholds. It monitors committed offsets for all consumers and calculates group status on demand, exposing it over an HTTP endpoint.
How do I install linkedin/Burrow?
Install Go 1.24 or later, clone the repository outside GOPATH, then run go mod tidy and go install. The README also provides a Docker path: build the container, mount your configuration into /etc/burrow, and run it.
Does linkedin/Burrow need me to set lag thresholds?
No. The README states there are no thresholds and that groups are evaluated over a sliding window, so the service derives status from how offsets move rather than comparing lag to a configured number.
Which offset sources can linkedin/Burrow read?
It automatically monitors all consumers using Kafka-committed offsets, and the README lists configurable support for Zookeeper-committed and Storm-committed offsets. No other offset source is documented.
What port does linkedin/Burrow listen on in the Docker Compose setup?
The docker-compose.yml maps port 8000 on the Burrow service, and the README says Burrow can be accessed at http://localhost:8000/v3/kafka once the stack is running.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/linkedin-burrow)