Apache Kafka: building and running the broker from the apache/kafka repository
Apache Kafka - A distributed event streaming platform. We build and test Apache Kafka with Java versions 17 and 25.
At a glance
- What is it?
- Apache Kafka is a distributed event streaming platform, and the apache/kafka repository is the source tree you build the broker, clients and Streams libraries from. The build is Gradle-based and targets Java 17 and 25, and the README's own quickstart points at the website rather than the repo.
- Who is it for?
- Adopt Apache Kafka if you need a durable, partitioned event log that several independent consumers can read at their own pace, and you are prepared to run Java 17 or 25 and a Gradle build. Do not adopt it as a drop-in task queue for a single small service, and do not treat the repository README as the operator manual: it defers to kafka.apache.org/quickstart for running a cluster.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What problem apache/kafka solves, and for whom
Apache Kafka is a distributed event streaming platform. The README describes it as being used for high-performance data pipelines, streaming analytics, data integration, and mission-critical applications. The problem it addresses is moving events between systems without coupling the producer to each consumer: a producer writes a record once, and any number of consumers read that record independently and at their own pace.
The repository is not only the broker. The top-level entries include clients/, streams/, connect/, core/, server/, storage/, metadata/, raft/, tools/, shell/, and several coordinator modules (coordinator-common/, group-coordinator/, share-coordinator/, transaction-coordinator/). That layout tells you what you are actually adopting: a broker plus client libraries plus a stream processing library plus a connector framework, all released together from one tree. If you only want a client library, you are still working inside a repository that builds the whole platform.
The audience implied by the build instructions is people who compile and test Kafka itself, or who need to patch and run it: contributors, integrators embedding the clients, and operators building a release tarball. If you just want a running broker, the README's own path is the Docker image or the website quickstart, not a source build.
How the repository is put together and how a build flows
The build is Gradle, driven by build.gradle, settings.gradle, gradle.properties and the checked-in wrapper (gradlew, gradlewAll, plus gradle/ and wrapper.gradle). The README states that Kafka is built and tested with Java versions 17 and 25, and that the release parameter in javac is set to 11 for the clients and streams modules and 17 for the rest. The same split applies to scalac: 11 for the streams modules, 17 for the rest. Scala 2.13 is the only supported version in Apache Kafka.
That split is the single most useful fact in the README for anyone embedding Kafka. The client and Streams artifacts are compiled against Java 11 bytecode, so they can be consumed by applications running on an older JVM than the one needed to build the broker. The rest of the platform requires 17. You can therefore run a Java 11 application against the clients while the broker side needs 17 or newer.
The module boundaries are real build targets, not just directories. The README gives ./gradlew core:jar and ./gradlew core:test for the core, examples and clients projects, and notes that Streams has multiple sub-projects so :streams:testAll runs all of their tests. The protocol layer has its own generation step: clients/src/main/resources/common/message/ holds the message definitions, and ./gradlew processMessages processTestMessages regenerates the RPC message data when switching branches.
Installing from source and running a first broker
The README requires Java to be installed but does not pin a single JDK for building; it says Kafka is built and tested with Java 17 and 25, and a note under the IDE section says to ensure JDK 17 is used when developing Kafka. Start by building the JARs from the repository root:
./gradlew jarThe README then says to follow the instructions at https://kafka.apache.org/quickstart, so the repository itself does not carry the full first-run walkthrough. What it does give is the sequence for starting a broker from the compiled files, which needs a cluster ID and a formatted storage directory before the server will start:
KAFKA_CLUSTER_ID="$(./bin/kafka-storage.sh random-uuid)"
./bin/kafka-storage.sh format --standalone -t $KAFKA_CLUSTER_ID -c config/server.properties
./bin/kafka-server-start.sh config/server.propertiesThe first command generates a random cluster ID, the second formats local storage in standalone mode against config/server.properties, and the third starts the broker. If you would rather not build at all, the README gives the container route instead:
docker run -p 9092:9092 apache/kafka:latestThat publishes the broker's default client port 9092 on the host, and the README points at docker/README.md for details. Building a distributable release is a separate target: ./gradlew clean releaseTarGz, with the tarball landing in ./core/build/distributions/. Note that the README's own examples for using the platform, once you have a broker, live in examples/, and the Streams quickstart archetype is published with mvn deploy from streams/quickstart rather than Gradle.
Where the source build is the wrong tool
Building Kafka from source is a heavy way to get a broker. The README's own default answer for running one is the Docker image, and its answer for learning the API is the website quickstart. A source build pulls the entire platform, including the connector framework, the Raft module and every coordinator, and the repository's test surface is correspondingly large: ./gradlew test runs both unit and integration tests, with separate unitTest and integrationTest tasks, and the README documents retry flags for flaky tests because failures happen.
There is also a version boundary to respect. The IDE note says JDK 17 must be used when developing Kafka, and the build is tested on 17 and 25. If your build environment is pinned to an older JDK, the client and Streams modules may compile to Java 11 bytecode, but that does not make the whole build runnable on 11.
Finally, consider whether you need Kafka at all. Kafka is a partitioned, replayable log with consumer groups, and that model imposes operational surface: brokers, storage formatting, cluster IDs, and the coordinators visible in the repository layout. For a single service that needs to hand off background jobs, a simpler queue is less machinery. Kafka earns its place when multiple independent consumers must read the same ordered stream, or when you need to replay history.
Apache Kafka compared with RabbitMQ
The most common comparison people search for is Kafka versus RabbitMQ, and the difference is architectural rather than a feature checklist. Kafka's unit is an append-only partition in a log; consumers track their position and can re-read. RabbitMQ's unit is a message routed through an exchange to a queue, and once a consumer acknowledges it, the broker discards it. If your requirement is "three services must each see every order event, and one of them may need to reprocess last week's events," the log model is the natural fit. If your requirement is "dispatch this job to exactly one available worker," the queue model is simpler, and Kafka's consumer-group mechanics are extra work for no gain.
The repository reflects the log-centric design: metadata/, raft/, storage/ and the coordinator modules exist because the broker must agree on partition leadership and consumer group state across nodes. RabbitMQ's design does not need those components in the same form. Neither is universally better; the choice follows from whether you need replay and multiple independent readers. Note that Kafka Connect, present here as connect/, is for moving data between Kafka and external systems, which is a different job from the broker-to-worker dispatch that a task queue performs.
Maintenance, licence and upgrade considerations
The repository is not archived, and the default branch is trunk. No last-push date was retrieved for this review, so no claim can be made here about how frequently it is updated; check the commit history on trunk directly if that matters to you. The README shows an active CI setup, with workflow badges for ci.yml on pushes to trunk and a scheduled generate-reports.yml, and it documents a contributor path through CONTRIBUTING.md and AGENTS.md at the top level.
Apache Kafka is licensed under Apache-2.0. The repository carries LICENSE and NOTICE alongside LICENSE-binary and NOTICE-binary, which is the standard Apache Software Foundation split between source and binary distributions. If you redistribute a built binary, the binary licence and notice files are the ones that travel with it. That is a factual note about the repository contents, not legal advice; consult your own counsel for compliance questions.
Upgrade cost is shaped by the module split. Because the clients and streams modules target Java 11 bytecode while the rest targets 17, upgrading your application's Kafka client and upgrading your broker are separable decisions. The README also warns that auto-generated RPC message data can fail to build when switching between branches, which is why ./gradlew processMessages processTestMessages exists. If you build from source and switch branches often, expect to run that.
Editorial conclusion
Adopt Apache Kafka if you need a durable, partitioned event log that several independent consumers can read at their own pace, and you are prepared to run Java 17 or 25 and a Gradle build. Do not adopt it as a drop-in task queue for a single small service, and do not treat the repository README as the operator manual: it defers to kafka.apache.org/quickstart for running a cluster. Before committing, verify the Java version on your build machines, confirm the release parameter targets (11 for clients and streams, 17 for the rest) match your deployment JDK, and read the docker/README.md if you intend to run the container image instead of a local build.
Frequently asked questions
How do I install Apache Kafka from the apache/kafka repository?
Install Java, then run ./gradlew jar from the repository root to build the JARs. The README then directs you to the quickstart at kafka.apache.org for running it. If you prefer not to build, the README gives the container command docker run -p 9092:9092 apache/kafka:latest.
How do I use Apache Kafka once it is running?
The repository README does not carry a usage walkthrough; it points to https://kafka.apache.org/quickstart for that. What it does show is starting a broker from compiled files: generate a cluster ID with ./bin/kafka-storage.sh random-uuid, format storage with --standalone against config/server.properties, then start it with ./bin/kafka-server-start.sh.
How does Kafka messaging work?
The README describes Apache Kafka as a distributed event streaming platform used for data pipelines, streaming analytics, data integration and mission-critical applications. The repository layout, with separate storage, metadata, raft and coordinator modules, reflects a broker that partitions and replicates an event log across nodes.
How do I install Apache Kafka on Windows?
The README does not give Windows-specific steps. It requires Java to be installed, then ./gradlew jar to build, and it offers the Docker route with docker run -p 9092:9092 apache/kafka:latest, which works wherever Docker runs. For running a broker it defers to the quickstart at kafka.apache.org.
How do I use Apache Kafka in Spring Boot?
The repository README does not describe Spring Boot integration. It documents the clients module and states that the javac release parameter is set to 11 for the clients and streams modules, so those artifacts can be consumed by an application running on Java 11. Any Spring-specific setup would come from outside this repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-kafka)