# alibaba/canal: MySQL binlog subscription and consumption for change data capture

> Canal disguises itself as a MySQL replica, pulls the binary log, and hands parsed row changes to Java clients, Kafka, or RocketMQ. It is a data pipeline component with a real operational surface, not a drop-in library.

**alibaba/canal** — 阿里巴巴 MySQL binlog 增量订阅&消费组件 

- Repository: https://github.com/alibaba/canal
- Stars: 29,737 · Forks: 7,625
- Language: Java
- License: Apache-2.0
- Published: 2026-09-11 · Updated: 2026-09-11 · Language: en
- Canonical page: https://hysenlabs.com/projects/alibaba-canal

## The problem canal solves: getting row changes out of MySQL without touching application code

The README frames canal as a response to a specific historical problem at Alibaba: dual data center deployments in Hangzhou and the United States needed cross-datacenter synchronization, and the first implementation used business triggers to capture changes. Triggers put load on the write path and tie capture logic to the application schema. Starting in 2010 the team moved to parsing the database log instead, and canal is the general-purpose result of that move.

The README lists the workloads it was built for: database mirroring, real-time backup, index construction and maintenance (split heterogeneous indexes, inverted indexes), business cache refresh, and incremental processing with business logic attached. Those have one thing in common. The source of truth is a MySQL table, and something downstream needs to know about each committed change as it happens, not on a nightly schedule.

The audience is infrastructure and data platform engineers, not application developers. Canal is a server process with its own deployment, its own account on the source database, and its own failure modes. If you are writing a Spring service and want a library that emits events on save, this is the wrong shape of tool.

## How canal works: pretending to be a MySQL slave

The README describes MySQL replication in three steps. The master writes data changes into the binary log, where each entry is a binary log event visible through show binlog events. A slave copies those events into its relay log. The slave then replays the relay log to apply the changes to its own data.

Canal inserts itself at the second step. It speaks the MySQL slave protocol, identifies itself to the master as a slave, and sends the dump protocol request. The master accepts the request and starts pushing binary log events to canal the same way it would push them to a real replica. Canal then parses the raw byte stream into structured objects.

The README states that the client-server interaction protocol uses protobuf 3.0, and that clients can be implemented in different languages. That is the design decision worth noting: the parsing and the consumption are separated by a network protocol, so the Java server does the binlog work and any consumer that can speak protobuf can read the stream. The README lists community clients for C#, Go, PHP, Python, Rust, and Node.js, none of which are part of the main repository.

There is a second delivery path. Canal can push change records into a message queue, specifically Kafka or RocketMQ, and the README points to that as the way to get multi-language consumption without writing a protocol client at all. The repository layout reflects this split: parse/, store/, sink/, connector/, client-adapter/, and server/ are separate top-level modules, with admin/ and deployer/ handling management and packaging.

## Installing canal and consuming your first binlog event

The README does not include install commands. It points to a QuickStart wiki page and to a Download page under the GitHub releases, with a separate Docker QuickStart page for container deployment. The repository does contain a docker/ directory and a deployer/ module, which is consistent with that documentation but is not itself an install procedure.

What the README does state is the supported source range: MySQL 5.1.x, 5.5.x, 5.6.x, 5.7.x, and 8.0.x. Before canal can see anything, the source database has to be configured for row-based replication. The README's own explanation of the mechanism implies the requirement: canal reads binary log events and parses them, and statement-based logging does not carry the row images that parsing depends on.

For the client side, the README links a ClientExample wiki page and a ClientAPI wiki page rather than embedding code. Those pages carry the class names and method signatures for the Java client, and they are the source to copy from. The README does not reproduce a connection snippet, so the exact constructor arguments, the server port, and the subscribe filter string should be taken from the ClientExample page for your version.

For a queue-based setup, the README directs you to the Canal Kafka/RocketMQ QuickStart page rather than describing the configuration inline. Canal 1.1.x added native Kafka delivery, per the release notes in the README, and the client-adapter module is the piece that maps parsed rows onto downstream targets.

For operators, version 1.1.4 introduced canal-admin, a Web UI for dynamic management of canal instances, including configuration, tasks, and logs. The README links a Canal Admin Guide, a ServerGuide, and a Docker page for it. This is where the deployment stops looking like a single jar.

## Where canal gets awkward: state, accounts, and the boundary of its job

Canal is not a library you add to a build file. It is a long-running server that holds a replication position and must keep that position consistent across restarts. The repository has a store/ module and a meta/ module, which matches the README's description of a system that maintains subscription state rather than a stateless parser. If the position is lost or reset, the practical recovery is re-subscribing from a point you choose, and the README does not document rollback or replay semantics on its wiki index. Treat position management as your responsibility until you find that documented.

The source database needs an account with replication privileges. That is a real organizational constraint: in many managed environments, granting REPLICATION SLAVE to an application-adjacent process requires a review, and some providers restrict it entirely. The README addresses one such case directly by listing native support for Aliyun RDS binlog subscription, including automatic primary-standby switchover and offline OSS binlog parsing. That support is specific to Aliyun RDS. If you run MySQL elsewhere, the dump connection is yours to manage.

Canal is also the wrong tool for bulk extraction. It reads the incremental log. It does not scan tables, and it has no concept of a full snapshot unless you build one. The README's related projects list DataX as Alibaba's offline synchronization project, and that is the correct category for periodic full or range extracts.

Finally, the README lists supported MySQL versions through 8.0.x and does not mention MariaDB support in the version list, although the wiki index includes a BinlogChange(MariaDB) page. The README itself is silent on MariaDB in the compatibility statement, so verify against that page rather than assuming parity.

## Canal versus Debezium: same idea, different center of gravity

Debezium is the obvious comparison for anyone doing change data capture on MySQL, and the difference is architectural rather than a feature checklist. Debezium is a set of Kafka Connect connectors. It runs inside a Connect cluster, and Kafka is not optional: the connector's output is a Kafka topic, and the Connect framework supplies offset storage, restart behavior, and scaling. Canal is a standalone server with its own client protocol, and Kafka is one of several sinks rather than the substrate it lives on.

That difference decides deployments. If you already operate Kafka Connect, Debezium's offsets live in Connect's internal topics and you get its rebalancing for free. If you do not run Kafka, canal can still deliver to a Java client over its own protobuf protocol, which the README presents as a first-class path with community clients in six other languages. Canal also ships canal-admin for Web UI management from version 1.1.4 onward, which is a different operational model from editing Connect worker configs.

The comparison is not a verdict. Canal's README documents Alibaba's own production lineage and an Aliyun RDS integration that Debezium does not have. Debezium's documentation covers connector-level offset and snapshot semantics that canal's README does not address at all. Pick based on what you already run, not on which project is better in the abstract.

## Maintenance, licensing, and what the release history tells you

The repository is not archived. The last push was on 2026-07-30, which is recent enough that the project cannot be described as dormant. The most recent release listed is canal-1.1.8 from 2025-01-16, preceded by two alpha builds of the same version line in 2024. There is a gap between the release cadence and the commit activity, and the README does not explain it. Check the release notes page linked from the README before planning an upgrade, because the README's own version summary stops at the 1.1.4 feature description and does not cover 1.1.8.

Upgrade cost is dominated by the client protocol and the admin layer, not by the parser. The README states the protocol is protobuf 3.0 and that clients exist in multiple languages outside the main repository. Those clients are community projects, so a canal server upgrade can move ahead of the client you depend on. If you use the Java client from the main repository, that risk is smaller. If you use one of the listed third-party clients, verify its compatibility before upgrading the server.

The licence is Apache-2.0, per the repository's LICENSE.txt and the license badge in the README. Apache-2.0 permits commercial use and modification and includes a patent grant. It also requires that you preserve the licence and notice files and state significant changes. This is a factual description of the licence, not legal advice; get your own counsel for anything that matters.

## Conclusion

Adopt canal when you need row-level change events from MySQL and you are willing to run a stateful server with a ZooKeeper or admin dependency, or when you already run Kafka or RocketMQ and want binlog changes in that pipeline. Do not adopt it if you only need periodic batch extraction (use DataX) or if you cannot grant a replication account on the source. Verify first: that your MySQL version is in the supported range, that binlog_format is set to ROW, that the canal account has REPLICATION SLAVE and REPLICATION CLIENT, and that your network path can hold a long-lived dump connection from canal to the master.

## FAQ

### What is alibaba/canal and what does it do?

It is a Java server that parses MySQL binary logs to provide incremental data subscription and consumption. It connects to a MySQL master as if it were a slave, receives binary log events, and parses them into structured change records for downstream consumers.

### How do I install alibaba/canal?

The README does not give install commands. It points to a QuickStart wiki page and a Download page under GitHub releases, with a separate Docker QuickStart page for container deployment. The repository contains a docker/ directory and a deployer/ module consistent with those pages.

### Which MySQL versions does alibaba/canal support?

The README states support for source MySQL versions 5.1.x, 5.5.x, 5.6.x, 5.7.x, and 8.0.x. MariaDB is not mentioned in that compatibility list, though the wiki index includes a BinlogChange(MariaDB) page.

## Sources

- [alibaba/canal on GitHub](https://github.com/alibaba/canal)
- [Issues](https://github.com/alibaba/canal/issues)
- [License: Apache-2.0](https://github.com/alibaba/canal/blob/master/LICENSE)
- [README](https://github.com/alibaba/canal/blob/master/README.md)
- [Releases](https://github.com/alibaba/canal/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alibaba-canal
