Debezium: Change Data Capture for Production Databases
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
At a glance
- What is it?
- Debezium is an open-source CDC platform that reads row-level changes directly from database transaction logs and streams those events to Kafka, removing the need for dual writes and polling across MySQL, PostgreSQL, Oracle, SQL Server, MariaDB, and MongoDB.
- Who is it for?
- Engineering teams running MySQL, PostgreSQL, Oracle, SQL Server, MariaDB, or MongoDB in environments that already operate Kafka and Kafka Connect will find Debezium a direct fit for cache invalidation, CQRS read-side rebuilds, and event-driven microservice coordination. Teams without a Kafka cluster should evaluate Debezium Server first to understand the additional sink configuration it requires.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Dual-Write Problem Debezium Solves
Most application architectures write to a primary database and then immediately propagate that change to secondary systems: an in-memory cache, a full-text search index, a read model for CQRS, or an analytics pipeline. The standard approach is to write to both within the same application code path. When the application crashes between the database commit and the secondary write, the downstream systems fall out of sync with the primary. This is the dual-write problem, and it is common enough that it has produced several generations of workarounds.
Database triggers can catch changes inside the engine but are tightly coupled to the database and cannot reach external services in a reliable, ordered way. Polling the database for recently modified rows introduces latency and forces the application to maintain a watermark column or a change log table. Neither approach guarantees ordering across concurrent transactions.
Debezium takes a different approach. It reads from the database's own transaction log, which the database already maintains for its own replication. Every committed row change is recorded there in order. Debezium reads that log and publishes each change as a structured event to a Kafka topic. Because Kafka stores those events in durable, replicated partitions, any downstream consumer can subscribe and process changes at its own pace, independent of the others. If a consumer is offline for hours, it will receive all missed events in order when it reconnects. The application itself only writes to one place: the primary database.
Architecture: Kafka Connect at the Core
Debezium is built on top of Kafka Connect, the distributed worker framework that ships as part of the standard Apache Kafka distribution. Each Debezium connector is a Kafka Connect source connector plugin: a JVM component that connects to one upstream database server, reads from its replication log, and writes structured change events to one or more Kafka topics. The standard deployment model runs connectors in Kafka Connect distributed mode, which provides automatic offset tracking, task rebalancing across Connect workers, and restart recovery.
Each database table typically maps to a single Kafka topic. Change events include both a before and after image for update operations, so consumers can compute the delta without querying the source database again. Each event also carries a source block that records the connector type, the database name, the table name, and the transaction ID of the originating commit.
Kafka handles durability. Because events are stored in replicated topic partitions, a consumer that was offline will receive every missed event when it restarts, in the order they occurred in the source database. Debezium itself does not deduplicate: the delivery guarantee at the Kafka Connect level is at-least-once, so consumers should be prepared to handle occasional duplicates during connector restart scenarios.
For teams that do not operate a full Kafka cluster, the repository includes a Debezium Server module. This standalone process wraps a connector with an event router that can forward change events directly to Amazon Kinesis, Google Pub/Sub, Apache Pulsar, or other sinks, without Kafka Connect in the path.
Deploying Debezium: Installation Path and Configuration
Debezium does not ship as a standalone binary with a single install command. It is distributed as a set of Kafka Connect connector JARs, one per supported database. The deployment sequence is: obtain the connector archive from debezium.io, extract the JARs into the Kafka Connect plugin path, restart the Connect workers, and then POST a connector configuration document to the Kafka Connect REST API. That configuration document specifies the connector class, the database host and credentials, the list of tables to capture, and the snapshot mode to use on first start.
The project's documentation at debezium.io provides Docker Compose examples that bring up a Kafka cluster, a Kafka Connect worker with Debezium installed, and a sample source database. These examples are the fastest route to a working environment without manually assembling the stack.
For the Debezium Server deployment, the debezium-server module in the repository packages a connector with a self-contained runner. Configuration is done through an application.properties file that specifies the source connector type and the target sink.
Supported Databases and the Differences Between Connectors
The repository contains separate connector modules for MySQL, PostgreSQL, Oracle, SQL Server, MariaDB, and MongoDB, each reflecting meaningful differences in how those databases expose their change logs.
The MySQL connector reads from MySQL's binary log (binlog). It can also handle DDL changes, because MySQL writes schema modification statements into the binlog alongside row changes. This means column additions and table renames propagate automatically through the event stream.
The PostgreSQL connector uses PostgreSQL's logical replication protocol, which requires setting the wal_level parameter to logical in postgresql.conf. Unlike MySQL, PostgreSQL logical replication does not include DDL events, so adding a column requires additional handling outside of Debezium.
The Oracle connector interfaces with Oracle LogMiner. It has additional prerequisites around database edition, supplemental logging configuration, and permissions that MySQL and PostgreSQL connectors do not require. The README does not document the Oracle connector setup steps; those are covered in the project's website documentation at debezium.io.
The MariaDB connector exists as a separate module because MariaDB's binlog format diverged from MySQL's in ways that prevent the MySQL connector from handling it correctly. The MongoDB connector is different in kind: it reads from the MongoDB replica set oplog and emits documents rather than relational before/after row pairs.
The debezium-ai module visible in the repository structure suggests active development in directions the current README does not describe fully.
Snapshot Mode and the Initial Capture Constraint
When a Debezium connector starts for the first time against a populated database, it must establish a baseline before it can stream incremental changes. This is the snapshot phase. By default, Debezium performs an initial snapshot of each tracked table, reading the full table contents and emitting a change event for every existing row before switching to log streaming.
For tables with tens or hundreds of millions of rows, this initial snapshot can take hours and can increase I/O load on the source database. The snapshot.mode configuration key controls this behavior. The initial mode performs the full snapshot; schema_only captures only the schema without existing row data; never skips the snapshot entirely and starts reading from the current log position. Each mode involves a trade-off between completeness and time to operational streaming, and the right choice depends on whether the downstream system needs the full historical dataset or only changes going forward.
The snapshot phase is also a point of failure: if the connector restarts mid-snapshot, it will attempt the snapshot again from the beginning by default. Teams with very large tables should review the connector's offset tracking and snapshot lock behavior before deploying to production.
Maxwell and the Single-Database Alternative
Maxwell is a MySQL-only change data capture daemon, written in Java, that reads the MySQL binlog and publishes change events to Kafka, a file, or standard output. Its deployment model is simpler than Debezium's: the daemon connects directly to MySQL and Kafka without requiring Kafka Connect. For teams using only MySQL and unwilling to operate the full Kafka Connect infrastructure, Maxwell covers the same core use case with a smaller operational footprint.
Debezium's advantage is coverage. It handles PostgreSQL, Oracle, SQL Server, MariaDB, and MongoDB alongside MySQL, and it uses a consistent change event format across all of them. Its integration with Kafka Connect enables fanout to many consumers, restart from arbitrary offsets, and the full suite of Kafka topic configuration options. Maxwell does not offer a standalone server mode for non-Kafka sinks the way Debezium Server does.
Licence, Maintenance, and Project Structure
The project is released under the Apache License, Version 2.0. One internal module, debezium-ddl-parser, contains Antlr grammars licensed separately under the MIT License. The repository accepted its last push on 2026-09-26, indicating active development at the time of writing.
Debezium does not publish tagged GitHub releases in the main repository. Versioned releases are managed separately and announced through debezium.io. The project uses Zulip for community discussion, with separate streams for users and developers, and also maintains a Google Group. Issues are filed in a separate repository at github.com/debezium/dbz rather than in the main codebase repository.
The repository structure is organized by concern rather than by deployment mode. The debezium-core module provides the shared framework, debezium-api defines the connector SPI, each connector lives in its own module, debezium-server is the standalone runner, and debezium-embedded contains the engine for running connectors inside an application process. The debezium-sink module handles the outbound side for Kafka and other sinks.
Editorial conclusion
Engineering teams running MySQL, PostgreSQL, Oracle, SQL Server, MariaDB, or MongoDB in environments that already operate Kafka and Kafka Connect will find Debezium a direct fit for cache invalidation, CQRS read-side rebuilds, and event-driven microservice coordination. Teams without a Kafka cluster should evaluate Debezium Server first to understand the additional sink configuration it requires. Before committing, verify that your specific database version supports the replication mechanism Debezium depends on: PostgreSQL needs wal_level=logical, MySQL needs binlog enabled, and Oracle requires LogMiner access with the appropriate privileges.
Frequently asked questions
What is Debezium used for?
Debezium captures row-level changes from databases such as MySQL, PostgreSQL, Oracle, and MongoDB, then streams those changes as events to Kafka topics. Teams use it for cache invalidation, replacing dual writes in microservices, keeping search indexes in sync, and building CQRS read models. The project's README also describes data integration and multi-application database sharing as common scenarios.
Is Debezium part of Kafka?
Debezium is not part of the Apache Kafka project. It is a separate open-source CDC platform deployed as Kafka Connect source connector plugins. It depends on Kafka Connect for durability and scaling, but the Debezium Server variant can route events to non-Kafka sinks such as Amazon Kinesis and Google Pub/Sub without running Kafka Connect.
Can Debezium be used without Kafka?
The README describes a Debezium Server deployment that routes change events directly to sinks including Amazon Kinesis, Google Pub/Sub, and Apache Pulsar, making Kafka optional. An embedded connector engine is also available for running a connector inside an application process, which sends events directly to application code rather than to a broker.
What is a Debezium connector?
A Debezium connector is a Kafka Connect source connector plugin that monitors one upstream database, reads from its transaction log, and publishes change events to Kafka topics. The repository includes distinct connectors for MySQL, PostgreSQL, Oracle, SQL Server, MariaDB, and MongoDB, each packaged as a separate Maven module.
What is Debezium Server?
Debezium Server is a standalone runtime in the repository that packages a connector with a built-in event router. It allows change events to be forwarded to sinks like Amazon Kinesis, Google Pub/Sub, or Apache Pulsar without running the full Kafka Connect distributed service.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/debezium-debezium)