# python-mysql-replication: reading MySQL binlogs from Python

> A pure Python client for the MySQL replication protocol, built on PyMySQL. It hands you row events and the raw SQL behind them, which makes it a CDC primitive rather than a finished pipeline.

**julien-duponchelle/python-mysql-replication** — Pure Python Implementation of MySQL replication protocol build on top of PyMYSQL

- Repository: https://github.com/julien-duponchelle/python-mysql-replication
- Stars: 2,414 · Forks: 691
- Language: Python
- License: not declared
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/julien-duponchelle-python-mysql-replication

## The gap it fills: raw binlog events in Python

MySQL already writes every committed change to its binary log, and the replication protocol is how a replica asks for that stream. Most teams consume it through a server-side replica, a Debezium connector or a hosted CDC product. python-mysql-replication takes the low-level route: it speaks the replication protocol directly from your Python process, using PyMySQL for the connection, and yields parsed events. The README lists the intended uses as MySQL to NoSQL or search-engine replication, cache invalidation, audit, and real-time analytics. The audience is therefore engineers who want the event stream inside their own application, where they can route it to their own sink without operating a Kafka Connect cluster or a second MySQL server. The project describes itself as a pure Python implementation built on top of PyMySQL, and setup.py confirms the only runtime dependencies are packaging and pymysql>=1.1.0. There is no C extension and no separate service to deploy.

## How BinLogStreamReader turns the binlog into Python objects

The public entry point is BinLogStreamReader, which opens a replication connection and reads the binlog stream. The README frames the output as events like insert, update and delete with their data and raw SQL queries. The repository layout backs that up: the pymysqlreplication package holds the parser, and it ships constants and util subpackages, which is where event type codes and helpers such as column-schema lookups live. The server side has to cooperate. The README's MySQL configuration block sets server-id, log_bin, binlog_expire_logs_seconds, max_binlog_size and binlog-format=ROW, with a comment that ROW is very important if you want write, update and delete row events. For MySQL 8.0.14 and later the README also requires binlog_row_metadata='FULL' and binlog_row_image='FULL'. That second pair matters more than it looks: without full row metadata the stream does not carry enough column information for the parser to reconstruct rows reliably. The data flow is one-directional. Your process connects as a replica, the server pushes events, the library decodes them, and your callback decides what to do. There is no built-in checkpoint store, no offset management and no schema registry in the dependency list.

## Install and read your first binlog event

The README gives a single install command, and the distribution name differs from the import name, which trips people up. The package on PyPI is mysql-replication; the module you import is pymysqlreplication.

```bash
pip install mysql-replication
```

On the server, the README's configuration block is the prerequisite. At minimum the binlog must be in ROW format, and on MySQL 8.0.14 or later the row metadata and row image must be FULL.

```toml
[mysqld]
server-id                  = 1
log_bin                    = /var/log/mysql/mysql-bin.log
binlog_expire_logs_seconds = 864000
max_binlog_size            = 100M
binlog-format              = ROW
binlog_row_metadata        = FULL
binlog_row_image           = FULL
```

The repository ships runnable scripts under examples/, including dump_events.py, redis_cache.py, mysql_to_kafka.py and mysql_to_rabbitmq.py. The README points at that directory rather than inlining a program, so the fastest first contact is to read dump_events.py and run it against a server with the settings above. Expect a stream of event objects as you write rows in another session; if nothing arrives, the binlog format is the first thing to check.

## What the project does not promise

The README links to a limitations page in the ReadTheDocs documentation instead of listing constraints inline, which is a signal worth taking seriously: the sharp edges are known, documented separately, and not summarised where most readers will look. Two concrete constraints come from the README itself. First, it states that MySQL version 8.0.14 and later need binlog_row_metadata='FULL' and binlog_row_image='FULL'. Changing those settings on an existing server is a server-level operation, not a library flag, and it affects every consumer of that binlog. Second, the README's own project-status section says the project is used in production for critical work at some medium internet corporations, then adds that all use cases have not been perfectly tested in the real world. That is an unusually candid sentence and it should shape expectations. The version support list is also split: MySQL 5.5, 5.6 and 5.7 map to releases v0.1 through v0.45, MySQL 8.0.14 maps to v1.0 and later, and MariaDB 10.6 is listed separately. If you are on an older MySQL, you are on an older branch of the library. And because the library gives you events rather than a pipeline, anything you need around them (deduplication, ordering guarantees across restarts, dead-letter handling) is your code. The wrong tool here is a team that wants a supported connector with an operations story; a library that hands you a stream is a different product.

## Where it sits next to Debezium and pg_chameleon

The comparison that matters is against Debezium, the Kafka Connect based CDC family. Debezium runs as a connector inside Kafka Connect, manages offsets in Kafka, and emits change events to topics; you operate Kafka and Connect. python-mysql-replication is a library inside your process, so there is no broker, no Connect cluster and no offset store supplied for you, but also no infrastructure to run. The difference is where state lives: with Debezium the offset is a Kafka concern, with this library it is yours. The README also lists projects built on top of it, and two are instructive. pg_chameleon uses it for migration and replication from MySQL to PostgreSQL, and the Singer.io tap for MySQL is listed as well, so if you want a ready-made tap rather than a parser, that route exists. binlog2sql is listed too, and it is a binlog parser that converts raw binlog to SQL and can generate flashback SQL, a different output shape aimed at recovery rather than streaming. aiomysql_replication is a fork supporting asyncio, which tells you the core project is synchronous; if your application is asyncio-based, that fork is the relevant starting point rather than this package. The library also appears in the stacks of Localstack, PGSync and django-mysql-replication among others, so the integration pattern is well trodden even though the library itself stays minimal.

## Maintenance, licence and upgrade cost

The repository is not archived and the last push was on 2026-09-25, with releases 1.0.17 on 2026-08-06, 1.0.16 on 2026-07-12 and 1.0.15 on 2026-02-07. That is a steady release cadence, and version 1.0.17 in setup.py matches the most recent release. Upgrades are cheap by construction: the dependency list is packaging and pymysql>=1.1.0, there is no compiled component, and the package is pure Python, so a version bump is a pip install away. The version number in setup.py is the single source of truth for the release, which means a fork or a pinned install needs to track that file. The README does not document a rollback procedure, and it does not describe a deprecation policy, so pinning a version in your own requirements file is the only rollback story the README supports. On licensing, setup.py declares license="Apache 2" while the repository metadata supplied here lists the licence as unknown; the two disagree, so read the LICENSE file in the checkout before you rely on either. Nothing in the README discusses what Apache 2 means for redistribution or for a hosted service, and that is a question for your own counsel rather than for this article.

## Conclusion

Adopt it if you are building a CDC consumer in Python and your MySQL or MariaDB server can be configured with binlog-format=ROW, binlog_row_metadata=FULL and binlog_row_image=FULL. Do not adopt it if you need a managed service, a GUI, or a tool that survives a schema change without code changes on your side; the library gives you events, not a pipeline. Before writing code, verify the binlog settings on the target server, check how GTID mode is configured, and read the limitations page in the ReadTheDocs documentation, which the README links to instead of summarising.

## FAQ

### How do I install python-mysql-replication?

The README gives one command, pip install mysql-replication. The distribution name is mysql-replication while the import name is pymysqlreplication, so import from the latter in your code.

### Can python-mysql-replication connect MySQL to Python?

Yes. It implements the MySQL replication protocol on top of PyMySQL, so a Python process connects as a replica and receives insert, update and delete events with their data. The only runtime dependencies are packaging and pymysql>=1.1.0.

### Which MySQL settings does python-mysql-replication require?

The README configures server-id, log_bin, binlog_expire_logs_seconds, max_binlog_size and binlog-format=ROW, noting that ROW is very important for row events. MySQL 8.0.14 and later additionally need binlog_row_metadata='FULL' and binlog_row_image='FULL'.

### Does python-mysql-replication handle replication lag?

The README does not describe any lag handling. It links to a limitations page in the ReadTheDocs documentation rather than summarising constraints, so that page is where to check before assuming the library manages lag for you.

## Sources

- [Issues](https://github.com/julien-duponchelle/python-mysql-replication/issues)
- [julien-duponchelle/python-mysql-replication on GitHub](https://github.com/julien-duponchelle/python-mysql-replication)
- [README](https://github.com/julien-duponchelle/python-mysql-replication/blob/main/README.md)
- [Releases](https://github.com/julien-duponchelle/python-mysql-replication/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/julien-duponchelle-python-mysql-replication
