LakeSoul: A Rust-Core Lakehouse That Bundles Compaction, RBAC and Vector Search
LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.
At a glance
- What is it?
- LakeSoul is an LF AI & Data sandbox project that pairs a Rust metadata and IO core with Spark, Flink, Presto, Ray, Daft and PyArrow bindings. It is a full platform, not only a table format, and that scope is the main thing to weigh before adopting it.
- Who is it for?
- Adopt LakeSoul when you want upserts, streaming CDC and a single metadata core across Spark, Flink and Python engines, and you are willing to run PostgreSQL as part of the stack. Do not adopt it if you need a table format only, or if your compute engines fall outside the support matrix.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
The gap LakeSoul targets between a table format and a working lakehouse
Apache Iceberg is described in the README as the de facto open table format, and LakeSoul positions itself past that line. The pitch is that a table format alone leaves you assembling a catalog, a compaction service and an authorization layer yourself. LakeSoul ships those pieces as part of the project: automated disaggregated multi-level compaction, fine-grained RBAC that includes an S3 proxy authorization layer, OLAP queries, vector retrieval, and multimodal processing through Ray and Daft.
The intended user is a data platform team running both batch and streaming pipelines that need row-level upserts and concurrent updates, not just append-only writes. The README states LakeSoul implements incremental upserts for both row and column and allows concurrent updates, and that it uses an LSM-Tree-like structure for hash-partitioned tables with a primary key. That combination (primary key, upsert, streaming CDC) is the workload the project is built around. Teams that only append Parquet files and read them with a query engine are not the audience.
How the Rust core, PostgreSQL metadata and LSM-style upserts fit together
The architecture is split between a native core and engine bindings. The Cargo workspace lists the crates that make up that core: lakesoul-metadata, lakesoul-io, lakesoul-datafusion, lakesoul-flight, lakesoul-s3-proxy, lakesoul-console, lakesoul-vector, lakesoul-ivm and postgres-lakesoul, among others. The README states metadata management and file format IO are implemented entirely in Rust, with bindings for Java, Python and C++. The claim that follows is consistency: the same ACID guarantees and upsert semantics across engines, rather than per-engine reimplementations.
ACID control and metadata scaling come from PostgreSQL. The README states LakeSoul scales metadata management and achieves ACID control by using PostgreSQL, and that Postgres RBAC and row-level security policies implement permission isolation for metadata. Physical data isolation is handled by the S3 proxy authorization layer. Storage options named in the README are HDFS and S3.
On the file side there are three physical format modes: vortex-compact, which the README calls the default, parquet, and vortex. The vortex format is described as suitable for storing multimodal data and vector embeddings, which is where the vector search and AI training story connects to the table layer. Compute support in the README matrix covers Spark 3.5, Flink 1.20, Presto 0.296 with velox, Ray 2.55, Daft 0.7+, DuckDB, PyArrow 16+ and Pandas 2.0+, with read and write support varying by engine: Presto is read-only and DuckDB is listed as standalone read.
Installing LakeSoul and running a first Python read
The README does not inline install commands. It points to the Quick Start page at lakesoul-io.github.io for setting up a local test environment, and to the python/examples directory in the repository for Python data processing and model training examples. The repository also carries a justfile with test recipes that show the shape of a local setup, including a metadata cleanup target that connects to PostgreSQL on localhost:5432.
That recipe names the connection parameters the test environment expects, which is the clearest hint in the repository about how the pieces connect:
PGPASSWORD=lakesoul_test psql -h localhost -p 5432 -U lakesoul_test -f script/meta_cleanup.sql lakesoul_testThe command runs psql against a local PostgreSQL instance with user and database both named lakesoul_test, applying script/meta_cleanup.sql. If you are standing up a test environment from the Quick Start, this is the metadata store you are talking to.
For a first real read, the support matrix lists PyArrow 16+ and Pandas 2.0+ as standalone readers, so the smallest useful experiment is reading a LakeSoul table into an Arrow table or a DataFrame from Python. The repository keeps those examples under python/examples rather than in the README, so follow the example files there for the exact call sequence rather than guessing at API names. The Java and Spark path is the other common entry point: the justfile runs Maven test suites against the lakesoul-spark module, and the Quick Start is the documented route for a local Spark setup.
If you plan to build from source, the workspace uses Rust edition 2024 and pins DataFusion 55 and Arrow 59 as workspace dependencies, so the toolchain expectations are visible in Cargo.toml and rust-toolchain.toml before you start.
Where LakeSoul is the wrong tool
The scope is the risk. LakeSoul is a platform with a metadata service, a compaction service and an authorization layer, and PostgreSQL sits in the critical path for ACID control. A team that wants a library that writes Parquet plus a manifest and nothing else will find this heavier than necessary. Iceberg-style adoption is a format decision; LakeSoul adoption is closer to operating a service.
The engine matrix is another boundary. Presto 0.296 with velox is read-only, and DuckDB is listed as standalone read with no write column. If your query engine is not on the matrix, the README gives no path. Spark 3.5 and Flink 1.20 are the named versions, so running newer or older majors is unverified territory rather than a documented supported configuration.
The documentation is also uneven in the repository itself. The README links out to the doc site for Quick Start, Flink SQL usage, Flink CDC whole-database sync and multi-stream merge, but the top-level README does not document rollback or disaster recovery procedures. The justfile shows how the maintainers test, not how an operator recovers a table. Before committing, check the doc site for the operational chapters, because the repository root alone will not answer them.
LakeSoul compared with a plain Iceberg plus service stack
The honest comparison is with Iceberg as the table format plus separately chosen components for compaction and access control. LakeSoul's README frames the difference directly: instead of assembling and maintaining separate catalogs, compaction services and auth layers, the project bundles them. The trade is control for integration. With Iceberg you pick a catalog and a compaction job and an auth layer that fit your environment; with LakeSoul those choices are made for you, and the metadata layer is PostgreSQL with Postgres RBAC and row-level security.
A second difference is the engine consistency argument. LakeSoul's Rust core is shared across Java, Python and C++ bindings, which the README contrasts with per-language reimplementations and behavioral divergence between bindings. That matters most when the same table is written by Flink streaming and read by Spark or a Python training job in the same pipeline, which is exactly the Flink CDC plus PyTorch pattern the tutorials describe.
A third is the file layer. The default vortex-compact mode and the optional vortex format are not standard Parquet, and the README ties vortex to multimodal data and vector embeddings. If your downstream consumers expect Parquet files they can read without LakeSoul, the default mode is a decision point, not a detail.
Maintenance, releases and what the Apache-2.0 licence covers
The repository is not archived, and the last push was on 2026-09-20, so the project is being worked on. Recent releases are v4.0.0 on 2026-09-04 and py-v2.0.0 on 2026-09-08, with py-v1.0.2 before them on 2025-09-26. The version in the Cargo workspace is 4.0.0, matching the Java-side release. LakeSoul is an LF AI & Data sandbox project, donated by DMetaSoul in May 2023, and the README carries an OpenSSF Best Practices badge.
Upgrade cost is where the binding story cuts both ways. A single Rust core means one place to fix format or metadata bugs, but it also means the Java, Python and C++ bindings move with that core, and the release history shows separate Python versioning from the main line, so Python users track a different version stream. The workspace pins DataFusion 55 and Arrow 59, which are fast-moving upstreams; expect the build to follow them.
The licence is Apache-2.0, stated in the README, the Cargo workspace and the SPDX headers in the source files. That is a permissive licence with an explicit patent grant, and it does not impose copyleft obligations on your own code. This is a description of the licence text as it appears in the repository, not legal advice; if your organization has rules about bundled services or the LF AI & Data sandbox stage, check them against the actual LICENSE file.
Editorial conclusion
Adopt LakeSoul when you want upserts, streaming CDC and a single metadata core across Spark, Flink and Python engines, and you are willing to run PostgreSQL as part of the stack. Do not adopt it if you need a table format only, or if your compute engines fall outside the support matrix. Verify first that PostgreSQL and S3 or HDFS are available, that your engine versions match the matrix, and that you can operate the compaction and metadata services it expects.
Frequently asked questions
What is LakeSoul?
LakeSoul is an end-to-end, realtime cloud-native lakehouse framework from the lakesoul-io organization, with a Rust metadata and IO core and bindings for Java, Python and C++. The README describes it as going beyond a table format to include compaction, RBAC and multimodal data processing.
Which compute engines can read and write LakeSoul tables?
The README support matrix lists Spark 3.5, Flink 1.20, Ray 2.55 and Daft 0.7+ with both read and write, Presto 0.296 as read-only, and DuckDB, PyArrow 16+ and Pandas 2.0+ as Python-side readers.
Does LakeSoul need PostgreSQL?
The README states LakeSoul scales metadata management and achieves ACID control by using PostgreSQL, and that Postgres RBAC and row-level security policies implement metadata permission isolation. The justfile test recipe connects to PostgreSQL on localhost:5432 with user and database lakesoul_test.
What licence is LakeSoul released under?
LakeSoul is released under Apache-2.0, as stated in the README, the Cargo workspace metadata and the SPDX headers in the source files.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lakesoul-io-lakesoul)
Community notes