DuckLake: A SQL-Based Open Lakehouse Format Built on DuckDB and Parquet
DuckLake is an integrated data lake and catalog format
At a glance
- What is it?
- DuckLake is a DuckDB extension that implements an open Lakehouse format using SQL and Parquet. It stores table metadata in a catalog database (a DuckDB file or PostgreSQL) and table data in Parquet files, and exposes time travel, schema evolution, change data feed, and partitioning through standard SQL statements attached to a running DuckDB session.
- Who is it for?
- DuckLake suits teams that already use DuckDB for analytics and want Lakehouse features (versioned tables, Parquet-on-cloud storage, schema evolution) without adopting a separate catalog service or switching query engines. It is not a production-ready replacement for Apache Iceberg in a multi-engine environment where Spark and Trino also need to read the same tables.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DuckLake Is and the Gap It Fills
A data lakehouse combines the low-cost object storage of a data lake with the transactional guarantees and query performance of a data warehouse. Existing open Lakehouse formats such as Apache Iceberg and Delta Lake require a separate catalog service and a query engine that understands their metadata format. For teams already using DuckDB, this adds infrastructure.
DuckLake solves this by storing all metadata in a catalog database that DuckDB already knows how to query. The catalog can be a local DuckDB file or a PostgreSQL database. Data is stored in Parquet files in any location DuckDB can read, including local disk, S3, or other object stores accessible through DuckDB extensions. The result is a Lakehouse that runs entirely within a DuckDB session, using SQL for everything from table creation to schema changes to time travel queries.
The extension is published by the DuckDB organization and the last push was on 2026-09-25. There are no GitHub releases yet; the main branch is the current development state.
Installation and Initial Setup
DuckLake installs through DuckDB's extension mechanism using a single SQL statement:
INSTALL ducklake;To install the latest development version from the nightly channel:
FORCE INSTALL ducklake FROM core_nightly;After installation, a DuckLake database is attached through the ATTACH statement. The following example stores metadata in a file called metadata.ducklake and data in Parquet files under the file_path/ directory:
ATTACH 'ducklake:metadata.ducklake' AS my_ducklake (DATA_PATH 'file_path/');
USE my_ducklake;Once attached, the database behaves like any other DuckDB database. Tables are created, queried, updated, and dropped using standard SQL. The catalog database stores all version history, schema changes, and transaction logs. The Parquet files store the actual row data.
The DuckLake documentation website at ducklake.select provides guidance on choosing a catalog database, whether DuckDB file or PostgreSQL.
Core Operations: Insert, Update, Time Travel, and Schema Evolution
After attaching a DuckLake database, standard DML and DDL work as expected. Creating a table and inserting rows:
CREATE TABLE my_ducklake.my_table(id INTEGER, val VARCHAR);
INSERT INTO my_ducklake.my_table VALUES (1, 'Hello'), (2, 'World');Updates work through the standard UPDATE syntax:
UPDATE my_ducklake.my_table SET val='DuckLake' WHERE id=2;Time travel uses the AT clause with a version number:
FROM my_ducklake.my_table AT (VERSION => 2);Schema evolution works through standard ALTER TABLE:
ALTER TABLE my_ducklake.my_table ADD COLUMN new_column VARCHAR;The change data feed function shows which rows changed in a given version range:
FROM my_ducklake.table_changes('my_table', 2, 2);This returns rows with a snapshot_id, rowid, change_type (insert, update, or delete), and the column values, useful for incremental processing pipelines that need to know exactly what changed between two versions.
Catalog Database Options and PostgreSQL Support
DuckLake separates metadata from data. The metadata catalog is a database that stores table schemas, version history, snapshot records, and file lists. The default catalog is a DuckDB database file, which works well for single-machine or local development use cases.
For shared or production environments, DuckLake supports PostgreSQL as the catalog database. The test suite in the Makefile includes a specific test configuration for this: `./build/release/test/unittest --test-config test/configs/postgres.json`. This allows multiple clients to share the same catalog while each querying the Parquet data files independently.
The DuckLake documentation site lists guidance on choosing between a local DuckDB catalog and PostgreSQL. The Makefile also includes an S3-backed test configuration, indicating that Parquet data can be stored on S3-compatible object storage.
The examples/ directory includes a minio-demo-server/ example that demonstrates DuckLake with MinIO as a local S3-compatible storage backend.
How DuckLake Compares to Apache Iceberg
Apache Iceberg is the most widely discussed alternative. Both Iceberg and DuckLake store table data as Parquet files and track versions in a metadata layer. The fundamental difference is the catalog design and the query engine model.
Iceberg was designed for multi-engine environments. Its catalog can be implemented over Hive Metastore, AWS Glue, Nessie, or a REST catalog, allowing Apache Spark, Trino, Flink, and DuckDB to all query the same tables. This generality requires running an external catalog service.
DuckLake puts the catalog inside a SQL database that DuckDB manages directly. There is no separate catalog server to operate. This reduces operational overhead for DuckDB-centric workloads but means that a Spark job cannot natively read a DuckLake table without going through DuckDB as an intermediary.
Teams that work exclusively in DuckDB and want Lakehouse semantics without adding infrastructure will find DuckLake simpler to operate. Teams that need the same tables to be readable by multiple query engines should evaluate Iceberg instead.
Building from Source and Running Tests
DuckLake is a C++ DuckDB extension. Building from source requires fetching the DuckDB submodule and using the extension build toolchain:
git submodule init
git submodule update
make pull
makeThe bundled DuckDB shell runs the built extension:
./build/release/duckdbThe test suite uses DuckDB's own test runner:
./build/release/test/unittestIndividual test files or patterns can be targeted:
./build/release/test/unittest test/sql/transaction/create_conflict.test
./build/release/test/unittest "test/sql/partitioning/*"The Makefile also provides a target that runs DuckDB's own core test suite using DuckLake as the storage backend, which tests compatibility between DuckLake's implementation and DuckDB's standard query behavior.
The Makefile includes references to ninja for faster parallel builds: `make GEN=ninja release`.
Maintenance Status and License
The repository is not archived. The last push was on 2026-09-25, and active development appears ongoing. The DuckLake website at ducklake.select provides documentation, and the extension is installable directly from DuckDB's extension registry with INSTALL ducklake.
There are no GitHub releases at this stage. The core_nightly channel provides access to the most recent unreleased build. Production deployments should check the DuckLake website for its documented stability level before relying on the format for long-term data storage.
The extension is licensed under MIT. The repository includes a CONTRIBUTING.md and an AGENTS.md, indicating that the project has set up AI coding assistant context for contributors. The CMakeLists.txt and Makefile coordinate the C++ build with DuckDB's extension toolchain.
Editorial conclusion
DuckLake suits teams that already use DuckDB for analytics and want Lakehouse features (versioned tables, Parquet-on-cloud storage, schema evolution) without adopting a separate catalog service or switching query engines. It is not a production-ready replacement for Apache Iceberg in a multi-engine environment where Spark and Trino also need to read the same tables. Before deploying in production, check the DuckLake website for its current stability status: the repository has no releases yet and the nightly build channel is labeled core_nightly.
Frequently asked questions
Why use DuckLake instead of plain DuckDB or Parquet files?
DuckLake adds transactional versioning, time travel, and schema evolution on top of DuckDB and Parquet. A plain DuckDB database file is a single file that cannot be shared across multiple writers easily. DuckLake separates metadata from data, allowing the catalog to live in PostgreSQL while data stays in Parquet files on object storage.
Is DuckLake a Lakehouse format?
Yes. The README describes DuckLake as an open Lakehouse format built on SQL and Parquet. It stores metadata in a catalog database and data in Parquet files, and supports the standard Lakehouse capabilities of time travel, schema evolution, and change data feed.
What are the differences between DuckLake and DuckDB?
DuckDB is the query engine. DuckLake is an extension for DuckDB that adds a Lakehouse format. Without DuckLake, DuckDB reads Parquet files directly but has no versioning or catalog layer. DuckLake adds transactional table management, time travel, and an explicit catalog that tracks all changes.
How do I install DuckLake?
Run INSTALL ducklake; inside a DuckDB session. To install the latest development build, run FORCE INSTALL ducklake FROM core_nightly; instead.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/duckdb-ducklake)