Library / SDK
zhisheng17/flink-learning avatar
zhisheng17/flink-learning

flink-learning: a Java sample repository for reading Flink by example

flink learning blog. http://www.54tianzhisheng.cn/ 含 Flink 入门、概念、原理、实战、性能调优、源码解析等内容。涉及 Flink Connector、Metrics、Library、DataStream API、Table API & SQL 等内容的学习案例,还有 Flink 落地应用的大型项目案例(PVUV、日志存储、百亿数据实时去重、监控告警)分享。欢迎大家支持我的专栏《大数据实时计算引擎 Flink 实战与性能优化》

15,101 stars3,924 forksJavaApache-2.0

At a glance

What is it?
zhisheng17/flink-learning is a multi-module Java project that pairs Apache Flink sample code with a long-running blog series. It is a reading and reference repository, not a library you add to a build.
Who is it for?
Adopt flink-learning if you want runnable Java samples for connectors, CDC, SQL or monitoring and you are willing to align your Flink version with the branch you check out. Do not adopt it if you need a supported library, a stable API or current release coverage: the README's own upgrade notes stop at Flink 1.14.2 and the repository is a collection of examples rather than a maintained artifact.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 147 days ago.
What is it written in?
Mainly Java, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What flink-learning is, and the reader it is written for

This repository is a teaching corpus. Its modules sit next to a blog series at 54tianzhisheng.cn that walks from Flink introduction through configuration, Data Sources and Sinks, transformations, windows, time semantics, parallelism and slots, JobManager high availability, and source-code reading of the local and standalone session startup paths. The code exists to support those articles.

The audience is a Java engineer who has decided to learn Flink and wants working examples to read, modify and run locally. The README addresses that reader directly: it explains how to build, links the blog posts, and points at a paid column plus a knowledge community for deeper material. If you are evaluating Flink for a production platform and need a supported dependency, this is the wrong shape of project. Nothing here is published as an artifact you would declare in a pom.xml.

Module layout: connectors, CDC, SQL, monitoring, project cases

The top level is a Maven aggregator. Alongside pom.xml sit flink-learning-basic, flink-learning-common, flink-learning-core, flink-learning-connectors, flink-learning-cdc, flink-learning-sql, flink-learning-monitor, flink-learning-extends, flink-learning-datalake, flink-learning-k8s, flink-learning-configuration-center and flink-learning-project. The topic list attached to the repository names the systems the connector modules touch: Kafka, Elasticsearch, HBase, MySQL, Redis, RocketMQ, RabbitMQ, ClickHouse, InfluxDB, OpenTSDB, Loki and Spark.

That breadth is the point and also the maintenance problem. Each connector module is a separate integration surface with its own client library, and a Flink version bump can invalidate several of them at once. The README's change log records exactly that pattern: 1.9.0 in September 2019, 1.10 in February 2020, 1.13.2 in August 2021 with the note that APIs changed substantially and part of the code was removed from master and moved to a feature branch, then 1.14.2 in December 2021. The last push to the repository was on 2026-05-06, but the documented version history ends at 1.14.2, so the version you find in a given module is something to check rather than assume.

Building the whole tree with Maven

The README gives one build path. It suggests adding the Aliyun central mirror to your Maven settings.xml, which matters if you are building from a network where Maven Central is slow, then running a single command from the repository root. The mirror block is quoted in the README as follows.

xml
<mirror>
  <id>alimaven</id>
  <mirrorOf>central</mirrorOf>
  <name>aliyun maven</name>
  <url>https://maven.aliyun.com/repository/central</url>
</mirror>

With that in place, the build command is one line, and the README says a successful build prints a result image rather than describing the expected output in text.

bash
mvn clean package -Dmaven.test.skip=true

Expect a multi-module reactor build. Because the aggregator covers every module in the tree, a single connector module that no longer compiles against its client library will fail the whole run. The README does not document a way to build one module in isolation, and it does not document rollback either; the branch pointers in the change log are the only fallback it names.

A first concrete use: pick one module and read its entry point

The README does not walk through running a specific job, so the practical path is to build the tree with the command above and then open the module that matches what you want to learn. The examples module is the natural starting point because the blog list includes a WordCount walkthrough and a piece on how a Flink project runs. The README does not give a submission command of its own; it points at the blog post on running a Flink project for that step.

What you get after a successful build is a jar per module under that module's target directory. From there you submit it to a Flink cluster the way the blog post describes, or run the main class from your IDE against the Flink dependencies the module already declares. Because the README documents no per-module build flag, the root build is the only path it actually supports.

One caveat before you start: the README's own advice, repeated in the 2020 change note, is to check that your Flink version matches the code's and to look for dependency conflicts if a job fails on your cluster. That note was written about the 1.10 code, but the reasoning has not stopped applying.

The books, papers and conference decks are part of the repository

Three top-level directories are not code: books, paper and the Flink-Forward-Asia-2019-PPT, Flink-Forward-Asia-2020-PPT, Flink-Forward-Asia-2021-PPT and Flink-Forward-2020 folders. The README lists four Flink books it once distributed (Introduction to Apache Flink, Learning Apache Flink, Stream Processing with Apache Flink by Flink PMC members, and Streaming System) and states plainly that the downloads were removed for copyright reasons, with a pointer to older branches. The paper directory holds a paper.md collecting stream-processing engine papers.

Treat this as a reading list with an expiry note, not as a maintained archive. If you came for the PDFs, the README tells you where they are not, and chasing old branches for copyrighted material is not a supported path. The conference decks are a different case: they are the author's own slides and remain in the tree.

Where this repository stops being useful

The failure mode is version drift, and it is structural rather than accidental. A learning repository tracks Flink releases by hand, and the change log shows the author doing that work at intervals of roughly a year before stopping after 1.14.2. Flink's DataStream and Table APIs have moved since; a module that compiled against 1.14.2 may not compile against a current release, and the README offers no compatibility matrix to tell you which module is on which version.

The second limit is that examples are not a runtime. There is no scheduler, no state backend configuration, no checkpoint tuning and no operational tooling here. If your question is how to size a TaskManager or how to recover from a checkpoint failure in production, the repository's answer is a blog post, not code you can run. That is a legitimate split, but it means the repository cannot be the last stop for anyone past the learning stage.

How it differs from Flink's own training material

The obvious alternative is the Apache Flink project's own training exercises and documentation, which track the current release and are maintained by the community that ships the engine. The difference in approach is the direction of the dependency: official material is versioned with Flink, so a sample is expected to work against the release it ships with. flink-learning is versioned by a single author's writing schedule, so a sample works if you happen to be on the Flink version that module was written against, and the README's branch pointers (feature/flink-1.10.0, feature/flink-1.8.0) exist precisely because that assumption breaks.

What flink-learning adds is breadth of a specific kind: a connector for nearly every system in the topic list, plus the large project cases the description names (PV/UV counting, log storage, real-time deduplication at scale, monitoring and alerting). Official training rarely covers that many third-party sinks. If you want a worked example of writing Kafka data into HBase, Redis, InfluxDB or RocketMQ, this repository has one and the official material probably does not.

Licence and the cost of keeping up

The repository is Apache-2.0, which permits commercial use and modification provided the licence and notices are preserved. That covers the code. It does not cover the books the README once linked, and the README says so itself. It also does not cover the paid column or the knowledge community; those are separate commercial offerings and the Apache-2.0 grant says nothing about them.

The upgrade cost is the interesting part. If you fork this repository to use as a starting point, you inherit the version-matching problem: each module pins its own Flink and client-library versions, and moving the whole tree to a newer Flink release means touching every module whose API changed, which is what the 2021 note describes happening at 1.13.2. Budget for that before treating the code as a base. If you only need one connector pattern, copy that module's approach and rewrite it against your own Flink version rather than adopting the module as-is.

Editorial conclusion

Adopt flink-learning if you want runnable Java samples for connectors, CDC, SQL or monitoring and you are willing to align your Flink version with the branch you check out. Do not adopt it if you need a supported library, a stable API or current release coverage: the README's own upgrade notes stop at Flink 1.14.2 and the repository is a collection of examples rather than a maintained artifact. Before you invest time, check the branch that matches your cluster's Flink version, confirm the module you need still exists on master, and run mvn clean package -Dmaven.test.skip=true to see which modules actually compile in your environment.

Frequently asked questions

Is Flink a replacement for Kafka?

No, and this repository treats them as separate roles. Flink is the stream processing engine; Kafka appears in the topic list and the connector modules as a source and sink that Flink jobs read from and write to.

What is Flink used for?

The repository's description frames it as a real-time computation engine, and the blog series covers DataStream API, Table API and SQL, connectors, metrics, windows, time semantics and monitoring. The project cases listed include PV/UV counting, log storage, real-time deduplication and alerting.

How do I build flink-learning?

The README suggests adding the Aliyun central mirror to your Maven settings.xml and then running mvn clean package -Dmaven.test.skip=true from the repository root. Because the root is a multi-module aggregator, a single failing module can fail the whole build.

Which Flink version does flink-learning target?

The README's change log records upgrades to 1.9.0, 1.10, 1.13.2 and finally 1.14.2 in December 2021. Older code was moved to branches such as feature/flink-1.10.0 and feature/flink-1.8.0, and the README advises checking that your Flink version matches the code's.

Can I download the Flink books linked in the flink-learning README?

No. The README states that the book downloads were removed for copyright reasons and points readers to older branches, so the current tree does not distribute them.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. README
  5. zhisheng17/flink-learning on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zhisheng17-flink-learning.svg)](https://hysenlabs.com/projects/zhisheng17-flink-learning)