CLI tool
dpkp/kafka-python avatar
dpkp/kafka-python

kafka-python: a pure-Python Kafka client, and what the 3.0 rewrite changed

Python client for Apache Kafka

5,906 stars1,481 forksPythonApache-2.0

At a glance

What is it?
kafka-python is a pure-Python client for Apache Kafka with no C or Rust core. The 3.0 line regenerates the protocol stack from Kafka's JSON schemas and adds an admin CLI, but the README is thin on transactions, error handling and upgrade paths.
Who is it for?
Adopt kafka-python when you want a Kafka client that installs with pip on Python 3.8 or newer and needs no JVM: the admin CLI alone justifies it on hosts where the Apache Kafka bin/ scripts are awkward. Do not adopt it if you need a documented migration path from 2.x, or if you depend on transaction and error-handling behaviour the README does not spell out.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem kafka-python solves: Kafka access without a JVM or a compiled extension

The Apache Kafka distribution ships Java command line tools under bin/, and using them assumes a compatible JVM on the machine you are working from. kafka-python exists to remove that assumption on the Python side. The README describes it as "a pure-python client library for Apache Kafka" with "no external dependencies and no Cython/C/rust core", which is the whole pitch: `pip install kafka-python` and you have a producer, a consumer and an admin client. The admin CLI is positioned explicitly as "a simple alternative to the apache kafka bin/ scripts, particularly if/when you do not have easy access to an installed/compatible jvm".

The audience follows from that. If you write Python services that publish or read Kafka topics, this is the library you reach for when you do not want to build a C extension as part of your deployment. The pyproject.toml classifiers list CPython and PyPy, and requires-python is ">=3.8", so the supported surface is broad rather than narrow. The trade-off is stated in the same paragraph that sells the package: pure Python means no compiled core, and the README points at `pip install crc32c` as an optional dependency to offload "one of the most CPU intensive subsystems" to a C library. That is an admission that the default path spends CPU in Python.

How the 3.0 protocol stack works: generated from Kafka JSON schemas

The most consequential change listed under "What's New in 3.0" is that the protocol stack is "dynamically generated from Apache Kafka json message schemas". Instead of hand-maintaining request and response encoders for every Kafka API version, the project derives them from the schemas Kafka itself publishes. The pyproject.toml package-data entry confirms the shape of this: `kafka = ["protocol/schemas/resources/*.json", ...]`, so the schema resources ship inside the installed package.

The second half of that mechanism is encode/decode performance work, described as "compiled/cached python bytecode". Generated code that is compiled once and cached avoids re-deriving encoders at runtime. Networking was also reworked: the release notes say the changes "leverage kafka.net event-loop and async/await syntax". So the data flow is a select-style event loop in `kafka.net`, with the higher-level KafkaConsumer and KafkaProducer classes sitting on top of it, and a generated protocol layer underneath both.

Feature coverage in 3.0 is enumerated by KIP: Cooperative Rebalance (KIP-429), Rack-aware Fetch (KIP-392), Log-Truncation detection (KIP-320), Transactional Producer improvements (KIP-360, KIP-447, KIP-654), Sticky Partitioner (KIP-480) and splitting of oversized producer batches (KIP-126). KafkaAdminClient was refactored and expanded. The README states the compatibility badge covers Kafka 4.3 down to 0.8, and the Makefile defaults `KAFKA_VERSION ?= 4.3.0` for integration testing. That is a wide version range to hold with one generated stack, and the Makefile carries a warning that older brokers need specific Java handling, which tells you the compatibility claim is maintained through testing rather than assumed.

Installing kafka-python and running a first produce/consume round trip

Installation is a single pip command. The README gives it directly, and no build toolchain is involved because there is no compiled core.

bash
pip install kafka-python

After that, the package is callable both as a module and as a CLI script. The README shows the admin CLI against a broker on localhost:9092, which is the default Kafka port and the one the examples use throughout.

bash
kafka-python admin -b localhost:9092 cluster describe

Creating a topic uses the same binary through the module entry point, with `topics create` and a `-t` flag for the topic name.

bash
python -m kafka.admin -b localhost:9092 topics create -t foo-topic

Producing from the shell is a pipe into the producer module. The message body goes to stdout of the echo and into the topic named by `-t`.

bash
echo "foo message" | python -m kafka.producer -b localhost:9092 -t foo-topic

Consuming uses the consumer module with a group id and two configuration overrides passed as `-C key=value`. Here `auto_offset_reset=earliest` makes a fresh group read from the beginning, and `consumer_timeout_ms=1000` makes the process exit after a second of no messages rather than blocking forever.

bash
python -m kafka.consumer -b localhost:9092 -C auto_offset_reset=earliest -C consumer_timeout_ms=1000 -g foo-group -t foo-topic

If the round trip worked you should see the string you piped in. In application code the same consumer is a for loop, and the README notes that ConsumerRecords are namedtuples exposing topic, partition, offset, key and value.

python
from kafka import KafkaConsumer
consumer = KafkaConsumer('my_favorite_topic', group_id='my_favorite_group')
for msg in consumer:
    print(msg)

Keys and values arrive as raw bytes unless you configure a deserializer. The README offers `JsonSerializer` and `DefaultSerializer` for that, and shows `value_deserializer=JsonSerializer()` producing dict values. Note that this is the reverse of the naming you may expect if you have used other clients: the same class name is used on both sides of the wire.

Where kafka-python stops: transactions, throughput and thin documentation

The README ends mid-sentence. Its last line is "KafkaProducer also supports transactions and message headers when", and the text available stops there. That is a truncation of the excerpt rather than necessarily of the file, but it is a fair illustration of the documentation posture: the consumer and producer sections are example-driven and the prose around them is sparse. The consumer side does document `isolation_level=IsolationLevel.READ_COMMITTED` for reading only committed messages from a transactional topic, and `consumer.metrics()` for metrics. The producer side documents `send()` as asynchronous, `future.get(timeout=60)` as the way to learn whether a delivery succeeded, and `flush()` with an explicit warning that it "does not guarantee delivery or success" and is "really only useful if you configure internal batching using linger_ms". Those are the right warnings to have, but they are scattered across code comments rather than collected into a delivery-guarantee discussion.

The throughput trade-off is the other hard boundary. A pure-Python client does its record encoding and CRC work in Python. The README's answer is `pip install crc32c`, which moves checksumming to C, and optional compression dependencies for lz4, snappy and zstd (the pyproject optional-dependencies group lists `lz4`, `python-snappy` and `zstandard`). Gzip works from the standard library with no extra install. If your workload is CPU-bound on serialization rather than network-bound, this is the wrong tool and a client with a compiled core is the right one.

There is also an upgrade cost the README does not address. It documents 3.0 as a protocol stack rewrite and a KafkaAdminClient refactor, and requires Python 3.8 or newer, but it contains no migration guide from 2.x. The repository has a CHANGES.md at the top level and the pyproject.toml points the Changelog URL at it, so that file is where the upgrade detail lives, not the README. Anyone moving a production service from 2.x to 3.0.11 should read CHANGES.md before touching a requirements file.

kafka-python compared with confluent-kafka-python and aiokafka

The alternative people most often weigh against this one is confluent-kafka-python, which wraps librdkafka, a C library. The difference is not cosmetic. librdkafka does its protocol work and checksumming in C, and it is maintained by the company behind the commercial Kafka offering, so its feature surface tends to track Kafka releases closely. kafka-python takes the opposite bet: everything in Python, generated from Kafka's own JSON schemas, with no build step. That makes it easier to install on constrained or unusual platforms and easier to read when you need to understand why a request looks the way it does. It also means the CPU cost is yours. The README's own recommendation to add `crc32c` is the clearest statement of that trade.

aiokafka is the other common comparison, and the split is about concurrency model rather than language. aiokafka is built for asyncio applications from the ground up. kafka-python's 3.0 networking changes brought in "async/await syntax" and a `kafka.net` event loop, but the high-level KafkaConsumer and KafkaProducer classes shown in the README are the synchronous, iterator-based interface. If your application is already an asyncio service, the shape of the API you want is probably the async-first one; if it is threads and blocking loops, the classes here fit without an adapter. Neither comparison is settled by a benchmark in the available documentation, so the deciding factor should be your deployment constraints: no compiler and no JVM points at kafka-python, maximum throughput points elsewhere.

Maintenance, licensing and what a 3.0 upgrade actually costs

The repository is not archived, and the last push was on 2026-09-21. Releases are frequent: 3.0.9 on 2026-07-21, 3.0.10 on 2026-08-04 and 3.0.11 on 2026-08-16. That cadence matters more than usual here, because a client library has to track broker protocol changes, and a generated protocol stack only stays correct if the schemas it is generated from are refreshed.

Licensing is Apache-2.0, declared in pyproject.toml as `license = "Apache-2.0"` with `license-files = ["LICENSE"]`, and the README carries the Apache 2 badge. Apache-2.0 is permissive and includes an explicit patent grant, which is usually what corporate policy wants from a client library; the usual obligations around notices and attribution still apply, and how they land in your distribution is a question for your own legal review rather than something this article can settle. The optional dependencies are separate packages with their own licences: `crc32c`, `lz4`, `python-snappy` and `zstandard`. Adding one of them to get compression or faster checksums also adds a dependency to audit.

The upgrade cost is the part to plan for. Python 3.8 is the floor, so anything older is blocked. The 3.0 release notes describe a protocol stack rewrite, an expanded KafkaAdminClient and networking changes, which is a large enough surface that a version bump in a requirements file is not a sufficient plan. The repository keeps CHANGES.md for exactly this, and the Makefile shows how the project itself validates against brokers: `make test` builds an integration fixture and runs pytest, with `KAFKA_VERSION` defaulting to 4.3.0. Mirroring that against your own broker version before upgrading is the concrete step, and the consumer and producer API pages under kafka-python.readthedocs.io are where the configuration keys you rely on are listed.

Editorial conclusion

Adopt kafka-python when you want a Kafka client that installs with pip on Python 3.8 or newer and needs no JVM: the admin CLI alone justifies it on hosts where the Apache Kafka bin/ scripts are awkward. Do not adopt it if you need a documented migration path from 2.x, or if you depend on transaction and error-handling behaviour the README does not spell out. Before committing, run `kafka-python admin -b localhost:9092 cluster describe` against your own broker, then read the KafkaConsumer and KafkaProducer API pages for the configuration keys your application sets.

Frequently asked questions

How do I install kafka-python?

Install it with pip, as shown in the README: `pip install kafka-python`. There is no compiled core and no external dependency, so no build toolchain is required.

What is kafka-python?

It is a pure-Python client library for Apache Kafka. The README describes high-level classes for consumer, producer and admin clients, plus CLI scripts for interactive tasks, with no Cython, C or Rust core.

Which is better, confluent-kafka-python or kafka-python?

confluent-kafka-python wraps librdkafka, a C library, while kafka-python keeps everything in Python and generates its protocol stack from Kafka's JSON schemas. The README itself points at installing the optional `crc32c` package to offload a CPU-intensive subsystem, which is the clearest signal of where the pure-Python approach costs you.

How do I use kafka-python to consume messages?

Construct a KafkaConsumer and iterate it, as the README shows: `KafkaConsumer('my_favorite_topic', group_id='my_favorite_group')` and then a for loop over the consumer. Keys and values come back as raw bytes unless you set a deserializer such as JsonSerializer.

Official sources

  1. dpkp/kafka-python on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dpkp-kafka-python.svg)](https://hysenlabs.com/projects/dpkp-kafka-python)