Open-source project
apache/flink-cdc avatar
apache/flink-cdc

Flink CDC: a YAML, SQL and DataStream route to change data capture

Flink CDC is a streaming data integration tool

6,481 stars2,196 forksJavaApache-2.0

At a glance

What is it?
Flink CDC is Apache's streaming data integration tool built on Flink, with a zero-code YAML pipeline API, Flink SQL connectors and a DataStream API. It is strongest when you already run Flink and need schema evolution and sharded-table sync, and weakest when a lightweight standalone capture agent would do.
Who is it for?
Adopt Flink CDC when you already operate a Flink cluster and need full-database sync, sharded-table merging or schema evolution into a lakehouse sink such as Paimon or Iceberg; skip it when a single-table capture agent with no Flink dependency would be enough.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Flink CDC solves, and who ends up using it

The project describes itself as a distributed data integration tool for real-time data and batch data, built on top of Apache Flink. The problem it targets is the gap between an operational database and everything downstream of it: analytics stores, search indexes, message buses and lakehouse tables. Rather than writing per-table extraction jobs, you declare a source, a sink, optional routing and transformations, and let the pipeline run continuously.

The audience is narrower than the description suggests. Because everything runs as a Flink job, the person operating it is usually already responsible for a Flink cluster. The README positions three API layers explicitly, which maps onto three kinds of user: a data engineer who wants a YAML file and the flink-cdc.sh CLI, an analyst or SQL-first developer who wants Flink SQL DDL, and a Java developer who wants the DataStream API as a Maven dependency. If none of those three describes your team, the tool is probably heavier than you need.

The feature list in the README is where the real differentiation sits: full database synchronization, sharding table synchronization, schema evolution and data transformation. Full-database sync and sharded-table merging are the two that are awkward to reproduce with hand-written capture jobs, because they require coordinating many tables and many physical shards under one logical identity.

Three API layers and the pipeline model underneath them

The repository layout shows the machinery behind the APIs: flink-cdc-pipeline-model, flink-cdc-composer, flink-cdc-runtime, flink-cdc-cli, flink-cdc-common and a connect directory holding the connectors. The pipeline model is the shared representation, the composer assembles a pipeline from it, and the runtime executes it as a Flink job. The SQL and DataStream APIs sit beside that model rather than on top of it, which is why the SQL connectors are distributed as separate artifacts such as flink-sql-connector-mysql-cdc.

In the YAML API, a pipeline has four parts. source declares a connector type and its connection properties plus a table pattern. sink declares the destination. transform applies projections and filters per source table. route maps source tables onto sink tables, and the README's example shows a catch-all pattern routing app_db.\.* to ods_db.others after two explicit routes. Order matters here in the sense that the specific routes are written first; the pattern is the fallback.

pipeline carries the execution settings: name, parallelism and schema.change.behavior, which the README sets to evolve. That single key is the clearest statement of the project's ambition. A schema change on the source is not treated as an error to be handled out of band, it is a pipeline-level policy. The README does not document what happens for values other than evolve, and the docs site is the place to check the full set.

The connector lists are split by API, and the split is not cosmetic. Pipeline connectors include Doris, Elasticsearch, Fluss, Hudi, Iceberg, Kafka, MaxCompute, MySQL, OceanBase, Oracle, Paimon, PostgreSQL and StarRocks. SQL source connectors include MySQL, PostgreSQL, Oracle, SQL Server, MongoDB, OceanBase, TiDB, Db2 and Vitess. If your source is MongoDB and your sink is Iceberg, no single API layer covers both ends, and you will be composing the pipeline yourself.

Installing Flink CDC and running a first pipeline

The README points to the Quickstart Guide on the documentation site for detailed setup, and the release artifacts are published under the project's releases. The Dockerfile in the repository shows the intended shape of a deployment: it starts FROM flink, unpacks a flink-cdc binary tarball into /opt/flink-cdc, sets FLINK_CDC_HOME to that directory, and copies connector jars into /opt/flink/usrlib. The build argument FLINK_CDC_VERSION defaults to 3.7-SNAPSHOT in that file, so anyone building from source should expect to override it with a released version.

Once the distribution is unpacked and FLINK_CDC_HOME is set, a pipeline is a YAML file submitted through the CLI. The README's example reads a MySQL database and writes to Doris:

yaml
source:
  type: mysql
  hostname: localhost
  port: 3306
  username: root
  password: 123456
  tables: app_db.\.*

sink:
  type: doris
  fenodes: 127.0.0.1:8030
  username: root
  password: ""

The transform and route blocks go in the same file. This one renames a column and filters rows, then sends orders to a separate sink table:

yaml
transform:
  - source-table: app_db.orders
    projection: id, order_id, UPPER(product_name) as product_name
    filter: id > 10 AND order_id > 100

route:
  - source-table: app_db.orders
    sink-table: ods_db.ods_orders
  - source-table: app_db.\.*
    sink-table: ods_db.others

The pipeline block closes the file with the job name, parallelism and the schema evolution policy:

yaml
pipeline:
  name: Sync MySQL Database to Doris
  parallelism: 2
  schema.change.behavior: evolve

The README states that this file is submitted via the flink-cdc.sh CLI. What you should see is a Flink job running under the given name at the requested parallelism, with rows from app_db.orders landing in ods_db.ods_orders and everything else matching the pattern landing in ods_db.others.

If you prefer SQL, the second route is to place the SQL connector jar in FLINK_HOME/lib/ and register the source as a table in the Flink SQL Client. The README gives this DDL, with the connector identified as mysql-cdc and credentials typed directly into the WITH clause:

sql
CREATE TABLE mysql_binlog (
  id INT NOT NULL,
  name STRING,
  description STRING,
  weight DECIMAL(10,3),
  PRIMARY KEY(id) NOT ENFORCED
) WITH (
  'connector' = 'mysql-cdc',
  'hostname' = 'localhost',
  'port' = '3306',
  'database-name' = 'inventory',
  'table-name' = 'products'
);

After that statement the table behaves like any other Flink table, so a SELECT with UPPER(name) runs against live change data. The third route, DataStream, is a Maven dependency on an artifact such as flink-connector-mysql-cdc, versioned by a property the README writes as ${flink-cdc.version}.

Where Flink CDC is the wrong tool

The first limitation is operational weight. Every pipeline is a Flink job, so adopting Flink CDC means adopting Flink: a cluster, its resource configuration, its checkpointing setup and its upgrade cadence. The repository carries flink-cdc-flink1-compat and flink-cdc-flink2-compat modules, which is a plain signal that Flink's own major versions diverge enough to need compatibility shims. Teams that do not already run Flink are taking on that whole surface to move rows between two systems.

The second is that the API layers do not line up. The pipeline connectors and the SQL connectors are different lists with different sources and sinks. A source available in the pipeline API may have no SQL connector, and vice versa. The README lists pipeline connectors in one section and SQL connectors in another, and the connector overview page is the only place the full matrix lives. If your source-sink pair spans both lists, the YAML API is not an option for that pair.

The third is documentation depth in the README itself. It shows one source, one sink, one transform, one route and one pipeline block. It does not document failure behaviour, restart semantics, how schema.change.behavior values other than evolve are handled, or what happens when a routed table does not exist at the sink. Those answers may exist on the documentation site, but the README does not carry them, and a pipeline that silently drops a table is worse than one that fails loudly.

Finally, the configuration style puts credentials in plain text in YAML or SQL. The README's examples use root with a literal password and an empty Doris password. That is fine for a quickstart and unacceptable for production, and the README does not discuss secret management.

Flink CDC compared with Debezium and Kafka Connect

The comparison people reach for is Debezium, usually running under Kafka Connect. The architectural difference is what sits in the middle. A Debezium deployment turns database changes into Kafka topics, and everything downstream consumes those topics; the topic log is the integration point and the durability boundary. Flink CDC runs the capture inside a Flink job and writes to the sink from that job, so the Flink pipeline is the integration point.

That changes what you have to operate. The Debezium route requires Kafka and Connect, plus whatever consumes the topics. The Flink CDC route requires Flink, but no Kafka unless Kafka is your sink, which it can be since a Kafka pipeline connector is listed. If you already have Kafka, Debezium slots into infrastructure you own. If you already have Flink, Flink CDC avoids adding a broker tier just to move data.

A second difference is the schema and topology handling baked into the pipeline model. Routing many source tables onto one sink table, including sharded tables, and applying a schema.change.behavior policy are first-class concepts in the YAML API. In a Debezium deployment those are conventions you build around the topic structure. That is not a defect in either tool; it is where each one puts the complexity. Flink CDC puts it in the pipeline definition, Debezium puts it in the topic and consumer design.

The SQL API also has no direct Debezium equivalent. Declaring a CDC source as a Flink table and joining it against other Flink tables is a capability that comes from Flink SQL, not from the capture layer. If your work is mostly enrichment and joins rather than transport, that is the strongest argument for this project over a pure capture agent.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-21. Releases follow a steady cadence: release-3.6.0 on 2026-03-31, release-3.5.0 on 2025-09-26 and release-3.4.0 on 2025-05-19. Roughly two releases a year, with the most recent one about six months before the last commit. Anyone planning an upgrade should treat the release page as the source of truth rather than the master branch, since the repository's own Dockerfile defaults to 3.7-SNAPSHOT, which is a development version and not something to deploy.

The upgrade cost is dominated by the Flink compatibility modules. Two separate compat directories exist for Flink 1 and Flink 2, so a Flink major upgrade is a real migration rather than a version bump in a properties file. Connector jars are versioned alongside the core, and the SQL connectors are distributed separately from the pipeline connectors, so an upgrade touches several artifact coordinates at once. Budget for a staging cluster that mirrors your source and sink versions rather than testing in place.

Flink CDC is licensed under Apache-2.0, the same licence as Apache Flink itself. The practical implication is that redistribution and modification are permitted under the licence terms, and the repository carries LICENSE and NOTICE files at the top level. This is not legal advice, and the connector ecosystem is worth checking separately: each connector artifact on Maven Central carries its own licence metadata, and a sink connector may pull in a client library with different terms than the core project. Verify the licence of every connector in your pipeline, not just the core.

Editorial conclusion

Adopt Flink CDC when you already operate a Flink cluster and need full-database sync, sharded-table merging or schema evolution into a lakehouse sink such as Paimon or Iceberg; skip it when a single-table capture agent with no Flink dependency would be enough. Before committing, verify that the pipeline connector for your target sink is listed in the connector overview, that your source database privileges allow the capture mode you intend, and that the version you download from the release page matches the Flink version you run.

Frequently asked questions

What is Apache Flink CDC?

It is a distributed data integration tool for real-time and batch data, built on top of Apache Flink, that prioritizes end-to-end data integration and adds full database synchronization, sharding table synchronization, schema evolution and data transformation. It exposes three API layers: a declarative YAML pipeline API, Flink SQL connectors, and a DataStream API.

How does Flink CDC differ from Kafka Connect?

Flink CDC runs capture inside a Flink job and writes directly to the sink, so the Flink pipeline is the integration point rather than a Kafka topic log. The README lists Kafka among the pipeline connectors, so Kafka can still be the destination, but it is not a required middle tier.

What alternatives to Flink CDC exist?

Debezium is the common alternative, typically deployed under Kafka Connect, where database changes become Kafka topics that downstream systems consume. The difference is where the complexity lives: Flink CDC puts routing, transforms and schema evolution in the pipeline definition, while a Debezium deployment puts them in the topic and consumer design.

Official sources

  1. apache/flink-cdc on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apache-flink-cdc.svg)](https://hysenlabs.com/projects/apache-flink-cdc)