Lakekeeper: An Iceberg REST Catalog in Rust with Built-In Access Control
Apache Iceberg REST Catalog in Rust — access control, credential vending and audit for every engine and AI agent. Apache 2.0.
At a glance
- What is it?
- Lakekeeper is an Apache-licensed implementation of the Apache Iceberg REST catalog specification, written in Rust as a single stateless binary. It enforces access control and credential vending once at the catalog layer, eliminating the need to duplicate permission rules in each compute engine.
- Who is it for?
- Lakekeeper is the right choice for data engineering teams that want a self-hosted Iceberg REST catalog with uniform access control across Spark, Trino, PyIceberg, and StarRocks, without deploying a JVM service. Its stateless design means horizontal scaling is straightforward once Postgres is provisioned.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Access Control Problem Lakekeeper Solves
Apache Iceberg defines a table format, not a catalog with access control. When multiple compute engines, say Spark and Trino, both read and write to the same Iceberg tables, each engine enforces its own permissions separately. Rules must be duplicated across engines, and an engine that lacks a fine-grained authorization plugin has no coverage at all.
Lakekeeper shifts that enforcement to the catalog layer. Define access control once in Lakekeeper, and every engine that connects to it goes through the same policy check before any data is read. Every request is checked and recorded. The README frames this as the core value proposition: define access control once, in the catalog, and enforce it across every compute engine.
The implementation is built on the apache/iceberg-rust crate, written in Rust, and ships as a single all-in-one binary with no JVM or Python runtime dependency. The stateless design means multiple replicas can sit behind a load balancer, with all state in Postgres.
Key Features: Storage Access, OpenID, and Change Events
Lakekeeper's feature set addresses four recurring problems in lakehouse deployments.
Storage access management (credential vending and remote signing) lets compute engines access data in object storage without holding long-lived credentials. Lakekeeper vends temporary credentials scoped to the specific tables a query needs. All major cloud storage providers are supported: S3 on AWS (with and without role assumption and session tags), S3-compatible stores like MinIO and Ceph, Cloudflare R2, Alibaba Cloud OSS, Azure ADLS Gen2, Microsoft OneLake (including Fabric and private-link endpoints), and Google Cloud Storage.
OpenID integration is a one-variable setup. Setting `LAKEKEEPER__OPENID_PROVIDER_URI` in the environment points Lakekeeper at an OIDC discovery endpoint. From that point, tokens issued by your identity provider are valid for catalog access without writing custom authentication code. Kubernetes service account tokens are also accepted simultaneously.
Change events, emitted as CloudEvents, allow downstream systems to react to every table mutation. The event backend is pluggable: NATS and Kafka are both supported and documented.
Change approval goes further by letting an external system veto a proposed table change before it is committed. This is the `ContractVerification` trait, which can be used to block changes that would violate a data contract or break a quality SLO.
Getting Started with the Minimal Docker Compose Example
The quickstart uses a pre-built Docker Compose file that wires Lakekeeper together with a query engine. The README provides three commands:
git clone https://github.com/lakekeeper/lakekeeper.git
cd lakekeeper/examples/minimal
docker compose upAfter the stack starts, two interfaces are available. The Jupyter notebook environment at `http://localhost:8888` contains example notebooks for Spark, PyIceberg, and other engines. The Lakekeeper UI at `http://localhost:8181` shows the catalog state, warehouses, and access policies through the open-source console.
The `examples/` directory includes several more specific scenarios: `access-control-simple/`, `access-control-advanced/`, `agentic-medallion/`, `agentic-memory/`, and integration examples for Kafka Connect and Fluss. These go beyond the minimal case and demonstrate the credential vending and event flows in context.
For production deployments, the documentation at docs.lakekeeper.io covers the Helm chart (available on Artifact Hub), high-availability configuration, and storage profile details. The repository's Cargo workspace is split into crates for the catalog core, storage backends, authorization, event sinks, and secrets stores, so specific components can be replaced by implementing the corresponding Rust trait.
Authorization: OpenFGA by Default, Replaceable by Design
The default authorization system uses OpenFGA, an open-source fine-grained authorization engine developed by Auth0. Lakekeeper ships an `authz-openfga` crate that integrates with it. Policies are expressed as relationship tuples: warehouse, namespace, and table access can be granted independently.
For teams that already operate a different authorization system, say Open Policy Agent or a custom RBAC service, Lakekeeper exposes the `Authorizer` trait. Implementing a handful of methods in Rust connects Lakekeeper's catalog operations to any external policy engine.
This is a real engineering decision with a real cost. Replacing the default OpenFGA integration means writing and maintaining Rust code. The benefit is that authorization stays inside the existing corporate governance stack rather than adding another service to operate.
Fine-grained access at the table and namespace level, combined with credential vending, means that a Spark job connecting to Lakekeeper never sees credentials for tables it is not permitted to read. The job cannot elevate its own privileges by inspecting the storage tokens, because those tokens are issued per-query and scoped to specific paths.
What Lakekeeper Does Not Cover
Lakekeeper is an Iceberg REST catalog server. It is not a query engine, a data transformation tool, or a notebook environment. It does not replace Spark, Trino, Flink, or PyIceberg; it is the catalog layer those engines connect to.
The only documented database backend is Postgres 15 or higher. MySQL, SQLite, and other databases are not supported. Running Lakekeeper in a production environment without Postgres is not a documented configuration.
The Kubernetes Operator mentioned in the README is listed as currently in development. Teams adopting Lakekeeper on Kubernetes in the near term will use the Helm chart directly rather than a higher-level operator abstraction.
The secret stores are Postgres and kv2 (HashiCorp Vault user-password auth). Other Vault authentication methods and other secret management systems require implementing the `SecretsStore` trait in Rust.
The `rust-toolchain.toml` in the repository pins to Rust 1.94 or higher. Deploying from source requires a compatible Rust toolchain.
Lakekeeper vs. Project Nessie and Unity Catalog
Two widely-deployed alternatives occupy similar space: Project Nessie and Unity Catalog.
Project Nessie is a Git-for-data catalog from Dremio, written in Java. It also implements the Iceberg REST catalog spec and supports multi-table transactions with branch-and-merge semantics similar to Git. The difference in approach is that Nessie's branching model is its central feature, while Lakekeeper's central feature is access control and credential vending. Nessie requires a JVM; Lakekeeper is a single Rust binary.
Unity Catalog is the open-source catalog from Databricks. It originates from the Databricks Data Intelligence Platform and is tightly integrated with that ecosystem. Running it as a standalone service outside Databricks is possible but the project is primarily designed for Databricks-centric workflows. Unity Catalog also implements the Iceberg REST spec, but its authorization model and storage integration are optimized for the Databricks runtime.
Lakekeeper is designed for platform-independent deployments. If your data platform already centers on Databricks or Dremio, using their native catalog avoids the overhead of a third service. If you need to serve Iceberg data consistently to engines across cloud vendors without a platform dependency, Lakekeeper's Rust implementation and pluggable trait system give you more control.
Release Cadence and Licensing
The release cadence is high. Recent releases include v0.13.6 on 2026-09-22, v0.13.5 on 2026-09-16, and v0.13.4 on 2026-09-10: roughly one release per week. The last push to the repository was on 2026-09-27. The project keeps a CHANGELOG and uses release-please for automated release management, which the `release-please/` directory in the repository confirms.
The minor version history (0.13.x) indicates the project is in active development and has not yet declared a stable 1.0 API. Teams integrating Lakekeeper in automation pipelines should pin to a specific version rather than following latest, since the weekly cadence means behavior can change quickly.
The license is Apache-2.0. The code can be used, modified, and incorporated into commercial products without licensing fees or copyleft obligations. The dependency on OpenFGA carries its own Apache-2.0 license. Postgres 15, required for the catalog backend, is PostgreSQL-licensed (similar to MIT). The storage integrations for cloud providers rely on their respective SDKs, which carry separate terms.
Editorial conclusion
Lakekeeper is the right choice for data engineering teams that want a self-hosted Iceberg REST catalog with uniform access control across Spark, Trino, PyIceberg, and StarRocks, without deploying a JVM service. Its stateless design means horizontal scaling is straightforward once Postgres is provisioned. Teams that rely on Databricks or a managed lakehouse platform should evaluate Unity Catalog or their platform's native catalog first, since Lakekeeper is designed for platform-independent deployments. Before going to production, verify that your Postgres version is 15 or higher and that your identity provider's OIDC discovery endpoint is reachable by the catalog at `LAKEKEEPER__OPENID_PROVIDER_URI`.
Frequently asked questions
What is Lakekeeper?
Lakekeeper is an Apache-licensed implementation of the Apache Iceberg REST catalog specification, written in Rust. It provides fine-grained access control, credential vending for cloud storage, and multi-tenant support, all in a single stateless binary that uses Postgres as its only backend.
How does Lakekeeper compare to Unity Catalog for Iceberg?
Both implement the Iceberg REST catalog spec, but they come from different ecosystems. Unity Catalog originates from Databricks and is primarily designed for the Databricks platform. Lakekeeper is a platform-independent Rust binary designed to serve any compute engine, including Spark, Trino, PyIceberg, and StarRocks, without a JVM or Databricks dependency.
Is Lakekeeper the same thing as Apache Iceberg?
No. Apache Iceberg is a table format specification that defines how data files and metadata are organized in object storage. Lakekeeper is a catalog server that implements the Iceberg REST catalog API, meaning it tracks which tables exist, where their metadata files live, and who is allowed to access them. Iceberg tables can exist without Lakekeeper, but a catalog like Lakekeeper is needed for multi-engine access control and credential vending.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lakekeeper-lakekeeper)