# Apache Fluss: streaming storage between Kafka and the lakehouse

> Apache Fluss is an incubating streaming storage layer for real-time analytics, with Flink and Spark connectors and a columnar Arrow-based format. Here is what the repository documents, and where it stays silent.

**apache/fluss** — Apache Fluss is a streaming storage built for real-time analytics.

- Repository: https://github.com/apache/fluss
- Website: https://fluss.apache.org/
- Stars: 2,182 · Forks: 634
- Language: Java
- License: Apache-2.0
- Published: 2026-08-04 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-fluss

## The gap Apache Fluss targets: streaming data that the lakehouse cannot query fast

Streaming platforms and lakehouses solve different halves of the same problem. A log-based broker moves events quickly but is not a table store; a lakehouse stores tables cheaply but serves them with latency measured in minutes, not milliseconds. Teams that need both end up maintaining a pipeline between the two, and that pipeline is where freshness is lost.

Apache Fluss positions itself in that gap. The README describes it as "a streaming storage built for real-time analytics & AI which can serve as the real-time data layer for Lakehouse architectures," and says it "bridges the gap between data streaming and data Lakehouse." The intended reader is an engineer who already runs Apache Flink or Apache Spark and wants streaming tables to be queryable directly rather than only after a write into object storage.

The name is German for river, pronounced /flus/, and the README uses the river metaphor explicitly: data "continuously converging, distributing and flowing into lakes." That is marketing, but it does describe the architecture: Fluss is meant to sit upstream of the lake, not replace it.

## How Fluss stores and serves data: tables, Arrow columns and changelogs

The repository layout is the clearest evidence of how the system is put together. There is a server module (fluss-server), a client (fluss-client), shared code (fluss-common), an RPC layer (fluss-rpc), a metrics module, and separate connector modules for Flink (fluss-flink), Spark (fluss-spark) and Kafka (fluss-kafka). There is also fluss-lake, fluss-filesystems, fluss-gateway and a Rust implementation under fluss-rust. That split tells you Fluss is a server process that clients and compute engines talk to over its own protocol, with the lake and filesystem modules handling the connection outward.

Two design choices in the README matter more than the feature list. The first is columnar streaming: "Based on Apache Arrow it allows database primitives on data streams and techniques like column pruning and predicate pushdown." A columnar wire format is what makes predicate pushdown possible, so an engine reading a Fluss table can skip data it does not need instead of deserializing every row. The second is compute-storage separation, where "stream processors focus on pure computation while Fluss manages state and storage." The README names the stateful operations this enables: deduplication, partial updates, delta joins and aggregation merge engines.

Fluss also generates changelogs. The README calls this "an append-only history of state and decision evolution," aimed at auditing and reproducibility. If your pipeline needs to answer what a value was at a given point, that is the mechanism, and it is a different proposition from a broker that only retains the latest record per key.

## Building Apache Fluss from source and running a first Flink job

The README documents one install path: build from source. It lists the prerequisites as a Unix-like environment (Linux, macOS, Cygwin or WSL), Git, Maven 3.8.6 or newer, and Java 11. Clone the repository and build with the Maven Wrapper, which the README says "ensures the correct Maven version is used":

```bash
git clone https://github.com/apache/fluss.git
cd fluss
./mvnw clean package -DskipTests
```

The README states the result: "Apache Fluss is now installed in build-target." That directory is where the packaged artifacts land, and it is what you would deploy or point a local cluster at.

For a first real use, the README links a Flink quickstart rather than reproducing the steps. The repository also carries a docker/ directory and a helm/ directory, so container and Kubernetes deployment paths exist in the tree, but the README itself does not document either one. The honest summary is that the source build is the only install procedure the README gives, and the Flink quickstart at fluss.apache.org/docs/quickstart/flink/ is where the first-job walkthrough lives. Read it there rather than guessing at connector configuration from the module names.

If you only want to try the project, note that -DskipTests skips the test suite. That is the README's own command, and it is fine for producing artifacts, but it means the build you get has not validated itself on your machine.

## Where Apache Fluss is the wrong choice

The project is incubating. The release names carry the -incubating suffix (v0.9.1-incubating is the most recent, dated 2026-05-04, following v0.9.0-incubating and v0.8.0-incubating), which is the Apache Software Foundation's marker for a project that has not yet graduated. That status is not a defect, but it means interfaces, configuration keys and deployment artifacts can change between minor versions in ways a graduated project would avoid. If you need a storage layer you can install once and leave alone for two years, this is not that yet.

There is a second, narrower limitation in the README itself. It lists Apache Flink and Apache Spark as integrations and then says StarRocks is "coming soon." Anyone whose real-time analytics stack is StarRocks-centric should treat Fluss as unavailable for now, regardless of what the feature list promises.

The README is also silent on several things an operator needs. It does not document rollback, upgrade procedures between the 0.8 and 0.9 lines, capacity planning, or failure recovery. The repository has modules that suggest answers exist (fluss-metrics for observability, fluss-test-coverage for test infrastructure), but the README does not connect them to an operational procedure. Treat the documentation site as the place to check before you assume any of this is handled.

## Apache Fluss compared with Apache Kafka plus a table format

The obvious alternative is the combination most teams already run: Apache Kafka as the log, plus a table format such as Apache Iceberg or Apache Hudi on object storage, with Flink or Spark doing the write. That stack is mature and widely deployed. Its weakness is exactly the gap Fluss targets: the table is only as fresh as the ingestion job's commit interval, and reading recent data means reading through the same pipeline that writes it.

Fluss changes the arrangement by making the storage layer itself the thing engines query, with columnar reads and pushdown, instead of inserting a log plus a separate table store. The README frames this as using "tables as a single abstraction to unify real-time and historical data across engines." In the Kafka-plus-lakehouse stack, real-time and historical data usually live in two systems with two query paths.

That difference cuts both ways. Fluss asks you to adopt a new storage system and its own client protocol, and the fluss-kafka module suggests Kafka interoperability is a concern the project takes seriously rather than something you get for free. If your workload is plain event transport with no table semantics, Kafka alone is simpler and has a far longer operational track record. Fluss earns its place when low-latency table reads are the requirement, not when a log is.

## Licence, maintenance and what upgrades cost you

Apache Fluss is licensed under the Apache License 2.0, and the repository carries both LICENSE and NOTICE files plus binary variants (LICENSE-bin, NOTICE-bin) and a copyright.txt. The README points to the LICENSE file in the repository for the full text. Apache-2.0 is a permissive licence with an explicit patent grant, which is normally the least complicated option for commercial use, but the repository also ships bundled dependencies and the NOTICE files exist for a reason. If you redistribute Fluss, read those files rather than assuming the top-level licence covers everything in the distribution. This is a description of what the repository contains, not legal advice.

The last push to the default branch was on 2026-05-04, the same date as the v0.9.1-incubating release. The repository is not archived. The release cadence visible in the published releases is roughly one minor release every four to five months across 0.8, 0.9.0 and 0.9.1, with a patch release following the 0.9.0 line. That cadence is the practical upgrade cost: pin a version, read the release notes before moving, and expect the -incubating suffix to mean breaking changes are possible between minors. The README does not document a rolling upgrade procedure, so verify that path against the docs before you plan a production upgrade.

## Conclusion

Adopt Apache Fluss if you already run Apache Flink or Apache Spark and want a storage layer that keeps streaming tables queryable without a separate ingestion pipeline; the repository ships connectors for both. Do not adopt it if you need a mature, widely deployed system with a long operational record, or if you cannot run Java 11 and Maven 3.8.6 or newer. Before committing, verify three things against the current docs: the exact deployment artifacts you intend to run, the state of the StarRocks connector the README still marks as coming soon, and whether the changelog and merge-engine features you need are documented for the version you plan to pin.

## FAQ

### What does "Fluss" mean in English?

The README states that Fluss is German for river, pronounced /flus/. The project uses that name as a metaphor for streaming data flowing into lakes.

### What is Apache Fluss?

It is a streaming storage built for real-time analytics and AI, which the README describes as a real-time data layer for lakehouse architectures. It integrates with Apache Flink and Apache Spark, with StarRocks listed as coming soon.

### How do I install Apache Fluss?

The README documents building from source: clone the repository and run ./mvnw clean package -DskipTests with Java 11 and Maven 3.8.6 or newer, which installs Fluss into build-target. The repository also contains docker/ and helm/ directories, but the README does not document those deployment paths.

### What is the Fluss app?

That is a different subject. Apache Fluss is a Java streaming storage project for real-time analytics, not a mobile application, and the README describes no app.

### What is a fluss device?

The README describes no device. Apache Fluss is a Java streaming storage project distributed as source, and the only install path it documents is a Maven build into build-target.

### What is the article of "Fluss" in German?

The README does not discuss German grammar. It only states that Fluss is German for river and gives the pronunciation /flus/.

## Sources

- [Official documentation](https://fluss.apache.org/)
- [Official README](https://github.com/apache/fluss#readme)
- [Project repository](https://github.com/apache/fluss)
- [Release notes](https://github.com/apache/fluss/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-fluss
