# cruise-control: automated rebalance and self-healing for large Kafka clusters

> The LinkedIn project that watches broker and partition metrics, generates multi-goal rebalance proposals, and heals a Kafka cluster while you sleep. What the branch matrix and release cadence say about running it.

**cruise-control-for-kafka/cruise-control** — Cruise-control is the first of its kind to fully automate the dynamic workload rebalance and self-healing of a Kafka cluster. It provides great value to Kafka users by simplifying the operation of Kafka clusters.

- Repository: https://github.com/cruise-control-for-kafka/cruise-control
- Website: https://github.com/linkedin/cruise-control/tags
- Stars: 3,045 · Forks: 653
- Language: Java
- License: Apache-2.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/cruise-control-for-kafka-cruise-control

## The problem is stated in terms of broker deaths

The introduction is short and it is worth quoting the reasoning rather than the marketing. Cruise Control is a product that helps run Apache Kafka clusters at large scale. Because Kafka is popular, many companies have increasingly large clusters with hundreds of brokers. At LinkedIn there are more than ten thousand Kafka brokers, which means broker deaths are an almost daily occurrence and balancing the workload becomes a big overhead.

That is a much more specific framing than a typical platform tool, and it explains the entire feature list. If broker death is normal, then failover, rebalance and repair cannot be incident response, they have to be a continuous background loop that nobody has to be awake for. The project describes itself as the first of its kind to fully automate the dynamic workload rebalance and self-healing of a Kafka cluster.

The scale assumption is the first filter to apply. Ten thousand brokers is not a typical deployment, and the design inherits its priorities from that number. On a cluster with a dozen brokers, the built-in Kafka tooling plus a human is usually cheaper than installing an additional control plane. The value curve is steep somewhere past the point where a single rebalance takes longer than your change window.

The repository is a LinkedIn project under the cruise-control-for-kafka organisation, Java, Apache 2.0 licensed, with roughly three thousand stars, six hundred and fifty forks and two hundred and seventy five open issues. That issue count is normal for a project this size and is a reminder that this is a live codebase where issues are how work arrives, not a sign of neglect.

## Goals are the configuration surface, and there are a lot of them

The feature list is organised around what the system observes and what it will try to fix, and the shape of it is genuinely useful for sizing up the project.

Observation comes first: resource utilization tracking for brokers, topics and partitions, plus a query interface over current cluster state. That query returns online and offline partitions, in-sync and out-of-sync replicas, replicas under min.insync.replicas, online and offline logDirs, and the distribution of replicas across the cluster. Reading that list tells you what the tool considers a fault rather than a preference, and replicas falling under min.insync.replicas being first-class is a durability concern ranked alongside capacity.

The second piece is multi-goal rebalance proposal generation, and this is the part that makes the tool different from a partition reassigner. The goals include rack-awareness, resource capacity violation checks across CPU, disk and network I/O, per-broker replica count violation checks, resource utilization balance across the same three resources, leader traffic distribution, replica distribution for topics, global replica distribution, global leader replica distribution, and custom goals that you wrote and plugged in.

Read that list as the constraint system. A rebalance that only moves replicas to fix distribution will happily push a broker past its disk capacity, and a rebalance that only respects capacity will leave you with a lopsided cluster. You are expressing both as goals and letting the optimizer find a proposal that satisfies all of them at once, which is why a goals list matters more here than in most Kafka tooling. The custom goal extension point is the part that tends to matter in practice, because disk headroom policy and traffic weighting are exactly the things that differ between teams.

The third piece is anomaly detection, alerting and self-healing: goal violation detection, broker failure detection, metric anomaly detection, disk failure detection and slow broker detection. The last two are explicitly marked as not available on the older kafka_0_11_and_1_0 branch, which is a useful hint about how the older branches have aged.

## The metrics reporter is the load-bearing piece

Cruise Control does not observe Kafka directly. A separate component does, and it has to be built and installed on your brokers before anything else works, which is the single most common reason a first attempt stalls.

The metrics reporter periodically samples Kafka's raw broker metrics and sends them to a Kafka topic. That means Kafka observing Kafka, which is an elegant self-contained design and also a hard dependency: if the metrics topic is missing, has the wrong cleanup policy, or is not reachable, the control plane has nothing to act on and will not be obviously broken, it will simply be idle.

The build and install sequence in the README is worth walking through because the order matters. Build the jar with `./gradlew jar`, which requires Java 17. Copy the resulting jar from the build output into your Kafka server dependency jar folder, which is `core/build/dependant-libs-SCALA_VERSION/` for a Kafka source checkout and `libs/` for a release download. Then set `metric.reporters` in your Kafka server properties to the reporter class, and the server properties file for an Apache Kafka release lives at `./config/server.properties`.

Two configuration traps are called out explicitly. If SSL is enabled, the relevant client configurations have to be set for every broker, because the reporter takes all its configuration from vanilla KafkaProducer settings under a `cruise.control.metrics.reporter.` prefix, so a truststore password is spelled `cruise.control.metrics.reporter.ssl.truststore.password`. And if the broker's default cleanup policy is compact, the metrics topic must be created with the delete cleanup policy instead, because the default metrics topic name is `__CruiseControlMetrics` and compaction would drop the samples the optimizer needs to reason about time series.

There is also a shell pair at the repository root, `kafka-cruise-control-start.sh` and `kafka-cruise-control-stop.sh`, which tells you the service is run as a process rather than deployed as a Kubernetes workload. The control plane itself then reads `config/cruisecontrol.properties`, where `bootstrap.servers` and `capacity.config.file` are marked required.

## Capacity is a file you own, and that is where the operational risk lives

Rebalancing without capacity information is guessing, so Cruise Control requires a capacity file, a JSON document describing the capacity of the brokers. The README marks `capacity.config.file` as required alongside `bootstrap.servers`.

This is the design decision with the most operational consequence. Capacity is not something the control plane discovers; it is an input you maintain, and every goal that references CPU, disk or network I/O is evaluated against those numbers. If a broker is recorded as having more headroom than it has, the optimizer will keep proposing rebalances that push it past its real limit. The capacity file is therefore a piece of infrastructure state that lives outside both Kafka and the control plane, and it has no owner by default.

The README also mentions that you can start from a capacity recommendation produced by GitHub Copilot, and it is worth being blunt about what that is. It is a convenient bootstrap for a file you must otherwise write by hand for every broker, and it is a starting point rather than an authority. Disk capacities are easy to read from a filesystem, CPU is not, and network I/O capacity is the hardest of the three to state honestly because it depends on the shape of your traffic rather than a fixed link speed.

The other required piece of configuration is the broker admin surface. The README lists admin operations including adding brokers, removing brokers, demoting brokers, rebalancing the cluster, fixing offline replicas, performing preferred leader election, and adjusting replication factor. Demote and remove in particular are the operations where Cruise Control's autonomy should make you cautious, because both move partitions away from a machine on its own initiative.

If you are deploying this, treat capacity file maintenance as a runbook item with a review step, not as one-time setup. Every hardware change, every storage tier change and every broker replacement invalidates part of it, and the failure mode is silent rather than loud.

## Branch names encode the Kafka version compatibility matrix

The environment requirements section is the most dated part of the README and the most likely to confuse a new evaluator, because branch names no longer match what they once meant.

The main branch, previously called migrate_to_kafka_2_5, is compatible with Apache Kafka 2.5 and up, which the README spells out as a specific list of Cruise Control releases: 2.5.*, 2.5.11 and up for Kafka 2.6, 2.5.36 and up for 2.7, 2.5.66 and up for 2.8, 2.5.85 and up for both 3.0 and 3.1, 2.5.142 and up for 3.8, 2.5.143 and up for 3.9, and 2.5.144 and up for 4.0 and 4.3.

Below that there are separate maintenance branches, which is the part people miss. There is a migrate_to_kafka_2_4 branch for Kafka 2.4, a deprecated kafka_2_0_to_2_3 branch covering Kafka 2.0 through 2.3, and a deprecated kafka_0_11_and_1_0 branch covering Kafka 0.11, 1.0 and 1.1. A known compatibility issue is documented for the older versions: support for Kafka 2.0 through 2.3 requires the KAFKA-8875 hotfix.

The language requirements follow the same pattern. The two deprecated branches compile with Scala 2.11, migrate_to_kafka_2_4 compiles with Scala 2.12, and migrate_to_kafka_2_5 compiles with Scala 2.13. This matters because the Kafka build you link against has to agree with the Scala version Cruise Control was compiled with. The project requires Java 17, and message.format.version 0.10.0 or above is needed.

So the practical reading is that main tracks modern Kafka and the older branches exist for clusters that have not moved. Before you plan an upgrade, check which branch your Kafka version maps to, because moving Kafka and moving Cruise Control are separate decisions with separate compatibility notes.

## A Gradle multi-module build with a release process you can read

The repository layout is a Gradle multi-module build with four published modules and one that is easy to miss. cruise-control-core holds the model and optimization logic, cruise-control-client is the HTTP client library, cruise-control is the service, and cruise-control-metrics-reporter is the on-broker component described earlier. The wrapper scripts gradlew and gradlew.bat are present, along with gradle.properties and settings.gradle.

Two files describe how releases work. semantic-build-versioning.gradle implies versions are derived from commit messages rather than chosen by hand, and build_api_wiki.sh generates an API wiki, which is a hint that the HTTP API surface is treated as a published contract. There is also a checkstyle directory for style enforcement and a buildSrc directory for custom build logic.

The recent release history shows how the version scheme is used in practice. There is 3.0.4 published in July 2026, preceded by 3.0.4-RC.1 in July and 3.0.4-RC.0 in May. The final 3.0.4 turned out to be very small: an Artifactory publish workflow fix that set the Artifactory key from secrets, and a version bump on the migrate_to_kafka_3_0 branch. The two release candidates carried the substantive work, which is the expected shape for a project with a strict release process rather than continuous deployment.

The RCs show where the maintenance effort goes. 3.0.4-RC.1 was entirely dependency work under a single ticket: Netty from 4.1.100 to 4.1.118.Final, Vert.x from 4.5.8 to 4.5.11, jackson-databind from 2.15.2 to 2.18.4 with a jackson-core pin, a global Logback exclude with a resolutionStrategy force block, and Kafka itself from 3.5.1 to 3.5.2. Note that Kafka 3.5.x is what the tool is developed against even though it supports Kafka 4.x, which is a good detail to remember when you wonder how far ahead of the broker it can be.

3.0.4-RC.0 is more interesting for anyone planning an upgrade. Most of it is a feature: broker connection count and connection capacity metric types added, socket-server-metrics reporting wired in, and unknown metric ids dropped during deserialization instead of throwing, which is the kind of robustness fix that matters when a broker reports a metric the control plane has never heard of. The rest is a testing migration off PowerMock, onto Mockito-inline and EasyMock partial mocks, with EasyMock bumped from 4.3 to 5.4.0 for JDK 17 compatibility and java.base opened to the test JVM for reflection. A project removing a mocking framework usually does it because the framework blocked the language version it needed.

## Conclusion

Cruise Control makes the most sense once broker count is high enough that a single broker death becomes a routine event rather than an incident, because that is the threshold its design assumes. The cost is a real dependency chain: a metrics reporter running on every broker, a capacity file you have to keep honest, and a GitHub Copilot-generated capacity recommendation you are expected to review rather than trust. Read the branch compatibility list against your actual Kafka version before anything else, since the project maintains parallel branches per Kafka generation and picking the wrong one wastes an afternoon. Then treat the goal list as configuration surface you own, because a rebalance is only as good as the constraints you hand it.

## FAQ

### What does cruise control do for Kafka?

It tracks resource utilization for brokers, topics and partitions, generates rebalance proposals that satisfy multiple goals at once such as rack-awareness, capacity limits and leader traffic distribution, and performs self-healing for goal violations, broker failures, metric anomalies, disk failures and slow brokers. It also exposes admin operations including adding, removing and demoting brokers, preferred leader election and replication factor changes.

### When should cruise control not be used?

It is built for clusters with hundreds or thousands of brokers where broker deaths are routine, as at LinkedIn with more than ten thousand. On a small cluster Kafka's own tooling plus an operator is usually simpler, and you would also be adding a metrics reporter on every broker, a capacity file to keep accurate, and a second service to run.

### What does it take to run Cruise Control for Kafka?

Java 17 and a Gradle build. You build the metrics reporter jar with `./gradlew jar`, copy it into your Kafka dependency jar directory, set `metric.reporters` in server.properties, and ensure the `__CruiseControlMetrics` topic uses the delete cleanup policy. On the control plane side, set `bootstrap.servers` and `capacity.config.file` in cruisecontrol.properties.

### Which Kafka versions does Cruise Control support?

The main branch, previously migrate_to_kafka_2_5, supports Kafka 2.5 and later, mapped to specific Cruise Control releases per Kafka version, through Kafka 4.0 and 4.3. Older clusters use separate branches: migrate_to_kafka_2_4 for Kafka 2.4, and deprecated kafka_2_0_to_2_3 and kafka_0_11_and_1_0 branches for older versions. Kafka 2.0 through 2.3 need the KAFKA-8875 hotfix.

## Sources

- [cruise-control-for-kafka/cruise-control on GitHub](https://github.com/cruise-control-for-kafka/cruise-control)
- [License: Apache-2.0](https://github.com/cruise-control-for-kafka/cruise-control/blob/main/LICENSE)
- [Project website](https://github.com/linkedin/cruise-control/tags)
- [README](https://github.com/cruise-control-for-kafka/cruise-control/blob/main/README.md)
- [Releases](https://github.com/cruise-control-for-kafka/cruise-control/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cruise-control-for-kafka-cruise-control
