Apache Paimon: a lake format for streaming updates with Flink and Spark
Apache Paimon is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations.
At a glance
- What is it?
- Apache Paimon is an Apache-licensed lake format that pairs a table format with an LSM structure so Flink and Spark can read and write the same tables in streaming and batch. This article covers what it solves, how to build and use it, and where it stops being the right tool.
- Who is it for?
- Adopt Apache Paimon if your pipelines already run on Flink or Spark and you need continuous updates to land in an open table format rather than being rewritten as nightly batches. Do not adopt it if your team has no Flink or Spark runtime and no appetite for managing a table store alongside object storage.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Java, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Paimon solves: streaming updates in a lake table
Batch lake tables assume data arrives in files that are written once and read many times. Streaming pipelines break that assumption. Events keep arriving, keys keep getting updated, and the downstream reader wants the current value of a row rather than the history of every file that ever contained it. Apache Paimon is aimed at exactly that gap. The README describes it as a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations, and states that it combines a lake format with an LSM structure, bringing realtime streaming updates into the lake architecture.
The intended user is a data engineer who already runs Flink or Spark and wants one table that both engines can read, with updates visible as they are committed rather than after a compaction job finishes. The project's former name was Flink Table Store, and the README says it was developed from the Flink community, so the centre of gravity sits with Flink users first. Spark support exists in the same repository, which matters if your organisation has both engines and does not want two separate storage designs.
One thing the README is explicit about: Paimon tracks issues in GitHub and prefers contributions as pull requests. That is a governance statement as much as a workflow one. There is no external tracker to watch, so the repository is the place where design discussion and bug reports land.
How the LSM structure and the lake format fit together
The mechanism the README names is the combination of a lake format with an LSM structure. An LSM tree buffers incoming writes in memory, flushes them as sorted files, and merges those files in the background. Applied to a lake table, that means a stream of updates does not have to be rewritten into one large file before it becomes queryable; the commits accumulate and compaction merges them later. The lake format layer is what makes those files addressable as a table for readers such as Flink and Spark, and the README credits Apache Iceberg as a source of design concepts in the architecture.
The repository layout supports that reading. There are separate top-level modules for the engines and for storage concerns: paimon-flink and paimon-spark for the engine integrations, paimon-core for the table implementation, paimon-common for shared code, paimon-filesystems for storage backends, paimon-format for file formats, and paimon-api for the interfaces. Beyond that, the layout shows the project reaching outward rather than staying a two-engine format. There are modules for Hive (paimon-hive), Iceberg (paimon-iceberg), Python (paimon-python), plus paimon-arrow, paimon-vector, paimon-vortex, paimon-lance, paimon-lumina, paimon-mosaic, paimon-full-text and paimon-service. What each of those does is not described in the README, and the README does not document their maturity, so treat the module list as a map of where work is happening, not as a feature guarantee.
A practical consequence of the LSM design is that writes and reads are decoupled in time. A writer commits small changes; a reader sees a consistent snapshot. The cost is background work: compaction has to run, and someone has to own its scheduling and its resource footprint. The README does not cover compaction tuning, so that knowledge lives in the documentation site rather than in the repository front page.
Building Apache Paimon from source
The README gives build instructions rather than a binary install. JDK 8 or 11 is required for building the project, and Maven version 3.6.3 or later is required. There is no documented package manager command, so the source build is the path the repository describes. The first command is the standard Maven install with tests skipped:
mvn clean install -DskipTestsRun it from the repository root. It compiles every module in the reactor, which is a large tree, so expect the build to take a while and to need a populated local Maven repository. Dropping -DskipTests runs the test suites instead, which is the right choice if you are validating a change rather than just producing artifacts.
The README also names a formatting command, and it is worth running before you open a pull request because the project applies it to both Java and Scala:
mvn spotless:applyIf you work in an IDE, the README gives one setup detail that is easy to miss. Mark paimon-common/target/generated-sources/antlr4 as Sources Root, otherwise the generated parser sources will not resolve and the project will look broken when it is not.
For a first real use, the README points to https://paimon.apache.org for background and documentation. That is where the engine-specific setup lives: the repository front page does not give a Flink or Spark session example, a catalog configuration, or a connector artifact coordinate. Follow the documentation site for the version of Flink or Spark you run, and check the releases page for the matching Paimon release rather than assuming the master branch matches your engine.
Where Paimon is the wrong choice
Paimon assumes you have a streaming or batch engine to talk to it. If your stack is a single-node analytics database and your data fits on one machine, adding a lake format plus an engine plus object storage is three moving parts where you needed none. The README frames the project around Flink and Spark, and nothing in it suggests a standalone query path.
The LSM side carries its own cost. Compaction is continuous background work, and it competes for the same storage and compute as your queries. Teams that expect a write-once table will find that the table keeps working after the write, which is either the point or an operational surprise depending on who is on call. The README does not discuss compaction scheduling, so budget time to read the documentation before you size a cluster.
Version alignment is the other trap. Paimon ships engine connectors that are built against particular Flink and Spark releases. The README does not publish a compatibility matrix, and the master branch is not a supported artifact for production. If you cannot state which Paimon release matches your engine version, you are not ready to deploy. The repository also carries an AGENTS.md file at the top level alongside the usual project files; the README does not describe it, so do not assume it documents user-facing behaviour.
Apache Paimon compared with Apache Iceberg
The README states that Paimon's architecture refers to some design concepts of Iceberg, which makes Iceberg the natural comparison and also the reason the two are not simply interchangeable. Iceberg is a table format built around immutable data files and metadata snapshots. Writes produce new files; readers resolve a snapshot. Updates to existing rows are handled by rewriting the affected files or by merge-on-read behaviour at query time, and the format itself does not carry an LSM tree.
Paimon adds that LSM structure on top of a lake format. The difference in approach shows up in the write path: Paimon can absorb a stream of small updates and merge them later through compaction, which is what makes the realtime lakehouse framing in the README possible. Iceberg's strength is breadth of engine and vendor support, and it is the safer default when your priority is a format many systems already read. Paimon is the more specific bet: you take on compaction and a smaller integration surface in exchange for streaming updates landing in the table as they happen. The repository even contains a paimon-iceberg module, so the two are not positioned as mutually exclusive in the codebase, though the README does not explain what that module does.
Maintenance, releases and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-09-23. The most recent release listed is Apache Paimon 2.0.0, published on 2026-08-07. Those two facts together describe a project with recent release activity and ongoing commits, and the README does not publish a support policy, a release cadence, or a list of which versions still receive fixes. Pin a release tag rather than tracking master, and read the release notes for the version you pick, because the README does not describe upgrade steps or backward-compatibility guarantees between releases.
Upgrade cost is dominated by the engine connectors. A Paimon upgrade that moves the connector also moves the Flink or Spark version it was built against, so an upgrade is rarely a single dependency bump. Budget for testing both the writer and the reader path, and for the compaction behaviour changing under you.
The code in the repository is licensed under the Apache Software License 2, per the README. That is a permissive licence, and the repository ships LICENSE, NOTICE and copyright.txt files at the top level, which is standard Apache practice. Redistributing the artifacts means carrying those notices with them. This is a description of what the repository states, not legal advice; check the NOTICE file and your own obligations before you ship a bundled build.
Editorial conclusion
Adopt Apache Paimon if your pipelines already run on Flink or Spark and you need continuous updates to land in an open table format rather than being rewritten as nightly batches. Do not adopt it if your team has no Flink or Spark runtime and no appetite for managing a table store alongside object storage. Before committing, check the docs for the connector version matching your engine release, confirm the JDK 8 or 11 and Maven 3.6.3 or later build requirements against your CI image, and read the licence and NOTICE files if you plan to redistribute the artifacts.
Frequently asked questions
What is Apache Paimon?
It is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations. The README states that it combines a lake format with an LSM structure to bring realtime streaming updates into the lake architecture.
How do I build Apache Paimon from source?
JDK 8 or 11 and Maven 3.6.3 or later are required. The README gives mvn clean install -DskipTests to build and mvn spotless:apply to format Java and Scala code.
Is Apache Paimon the same thing as Flink Table Store?
Yes. The README states that Paimon's former name was Flink Table Store and that it was developed from the Flink community.
Which engines can read and write Apache Paimon tables?
The README names Flink and Spark for both streaming and batch operations, and the repository also contains paimon-hive, paimon-iceberg and paimon-python modules. The README does not document the maturity or scope of those additional modules.
What licence is Apache Paimon released under?
The code in the repository is licensed under the Apache Software License 2, according to the README, and the repository includes LICENSE, NOTICE and copyright.txt files at the top level.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-paimon)