Library / SDK
delta-io/delta-rs avatar
delta-io/delta-rs

delta-rs: a native Rust implementation of Delta Lake with Python bindings

A native Rust library for Delta Lake, with bindings into Python

3,328 stars673 forksRustApache-2.0

At a glance

What is it?
delta-rs reads and writes Delta Lake tables without a JVM or a Spark cluster. It ships as the deltalake crate on crates.io and the deltalake package on PyPI. The Python surface is small, the Rust surface is lower level, and the feature table is where the real decisions live.
Who is it for?
Adopt delta-rs when you need to read or write Delta tables from a Python or Rust process that has no Spark session, and check the feature table for the specific operations your pipeline depends on before committing. Do not adopt it as a drop-in replacement for Spark's full Delta connector, and do not assume every table feature is implemented.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem delta-rs solves: Delta tables without a JVM

Delta Lake is a storage format that runs on top of existing data lakes. The README describes it as compatible with processing engines like Apache Spark and providing ACID transaction guarantees, schema enforcement, and scalable data handling. The usual way to touch a Delta table is through Spark, which means a JVM, a cluster or at least a local Spark install, and a dependency graph that is awkward inside a small Python service or a Rust binary.

delta-rs exists to remove that requirement. The project describes itself as a native Rust library for Delta Lake with bindings into Python, and the README states the aim is to provide native low-level APIs aimed at developers and integrators, plus a high-level operations API that lets you query, inspect, and operate your Delta Lake with ease. That is a different audience from Spark's: people embedding table reads and writes in an application, a scheduled job, or a data tool rather than running a general-purpose analytics engine.

The repository layout reflects the split. The Cargo workspace lists members as crates/* and python, so the Rust crates and the Python binding are built from one tree. The Python package and the Rust crate are released separately, which is why the recent releases are tagged python-v1.6.5, python-v1.6.4 and python-v1.6.3 rather than a single project-wide version.

How delta-rs is put together: workspace crates, Arrow, and a kernel dependency

The workspace Cargo.toml pins a Rust toolchain requirement of 1.94.1 and edition 2024, so building from source needs a recent compiler. The workspace declares a dependency on delta_kernel, a package named buoyant_kernel versioned 0.28.1 with a range constraint of <0.28.100, enabled with the arrow-59 and internal-api features. A second dependency, delta_kernel_default_engine (package buoyant_kernel_engine, version 0.28.0), is pulled with the rustls feature and default-features disabled. The commented-out lines in the same file show path and git variants of the same dependencies, which is how the maintainers point local builds at a checkout of the kernel repository.

Arrow is the data layer. The workspace pins arrow, arrow-arith, arrow-array, arrow-buffer, arrow-cast, arrow-ipc, arrow-json and arrow-ord at version 59, with arrow-array carrying the chrono-tz feature. The Python side is therefore not a separate implementation: a DataFrame written by the Python API becomes Arrow data, and the same table can be reopened from Rust. The README shows exactly this, opening from Rust a table that was written from Python and printing the active file URIs.

The practical consequence is that delta-rs is a library, not a service. There is no server process, no port to open, and no daemon to supervise. Table state lives in the Delta log on your storage backend, and every process that opens the table reads that log directly. Concurrency is therefore governed by the storage layer's atomicity guarantees, not by anything delta-rs runs in the background.

Installing delta-rs and writing a first Delta table

The README gives two installation commands, one per language. For Python, the package is deltalake on PyPI. For Rust, the crate is deltalake on crates.io.

bash
pip install deltalake
bash
cargo add deltalake

The quick start imports DeltaTable and write_deltalake from deltalake, builds a small pandas DataFrame with an id column and a value column, and writes it to a local directory. Reading it back and converting to pandas should return a frame equal to the one written, which is what the README's assertion checks.

python
from deltalake import DeltaTable, write_deltalake
import pandas as pd

df = pd.DataFrame({"id": [1, 2], "value": ["foo", "boo"]})
write_deltalake("./data/delta", df)

dt = DeltaTable("./data/delta")
df2 = dt.to_pandas()

assert df.equals(df2)

The same directory can be opened from Rust with open_table, which takes a Url. The README's example converts an absolute directory path with Url::from_directory_path and then calls get_file_uris to list the active files in the table. Because open_table is async, the example runs inside a tokio runtime.

rust
use deltalake::{open_table, DeltaTableError};
use url::Url;

#[tokio::main]
async fn main() -> Result<(), DeltaTableError> {
    let delta_path = Url::from_directory_path("/abs/data/delta").unwrap();
    let table = open_table(delta_path).await?;

    let files: Vec<_> = table.get_file_uris()?.collect();
    println!("{files:?}");

    Ok(())
}

For object storage, the repository ships a docker-compose.yml with localstack on ports 4566 and 8080 (SERVICES=s3,dynamodb), fake-gcs on port 4443, and azurite on port 10000. Those are the backends the project's own tests run against, so they are the fastest way to check that a credential configuration works before pointing it at real storage. The README also points to a Delta Lake image on DockerHub for trying things out without a local install.

Where delta-rs stops: the feature table is the real contract

The README's feature section does not enumerate capabilities. It links to a feature table in the documentation, and the link targets for writer-rs and onelake-rs point at open GitHub issues rather than at documentation. That is the honest signal in this repository: the project maintains a table of what is done, semi-done and open, and it does not claim parity with the Spark implementation.

The wrong way to adopt delta-rs is to assume that because a table opens, every operation on it will work. Delta Lake's protocol has grown table features over time, and a library that implements a subset will read some tables and reject others. The feature table is the document that tells you which. If your pipeline depends on a specific write mode or a specific table feature, that table is the first thing to read, and it is more informative than any benchmark.

The second limitation is environmental. Building the Rust crates from source requires Rust 1.94.1 per the workspace manifest, and the dependency range on the kernel crate is pinned below 0.28.100, so a kernel release outside that window will not resolve without a change to Cargo.toml. The Makefile's check target runs cargo fmt, then cargo clippy with the azure, datafusion, s3, gcs, glue and hdfs features, then the Python Makefile. The full test feature set is integration_test, azure, datafusion, s3, gcs, glue and hdfs, and the coverage target explicitly skips read_table_version_hdfs, test_read_tables_hdfs and test_read_tables_lakefs. Those skips are a fair indication of which paths get less continuous exercise.

delta-rs compared with Spark and with query engines like DuckDB and Polars

The comparison people reach for is delta-rs versus Spark. The difference is architectural rather than a matter of degree. Spark is a distributed execution engine that happens to speak Delta; delta-rs is a Delta implementation that you embed. Spark gives you a SQL planner, a scheduler and a cluster. delta-rs gives you table operations and hands the data to Arrow, which means the compute has to come from somewhere else, whether that is pandas, DataFusion, Polars, Daft or your own code. If your job is a large distributed transformation expressed in SQL, Spark is the tool; delta-rs is not trying to be it.

The second comparison is with engines that have their own Delta readers, such as DuckDB and Polars. The README lists both under integrations, along with DataFusion, ballista, Dask, Ray, AWS SDK for Pandas and datahub. Those integrations are a different arrangement: the engine reads the table through its own code path or through delta-rs, and the choice matters for which table features are honored. delta-rs's own value is that the same crate backs both the Rust API and the Python API, so a table written from a Python notebook and a table written from a Rust service go through one implementation rather than two.

A third option is Delta Lake on the JVM through Spark's libraries invoked from Python via PySpark. That keeps full feature coverage and costs you a JVM in the process, which is precisely the cost delta-rs was built to avoid. The trade is coverage for footprint, and the feature table is where you measure whether the trade is acceptable for your tables.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-09-23. The most recent Python releases are python-v1.6.5 and python-v1.6.4, both dated 2026-09-21, with python-v1.6.3 on 2026-08-23. Those are separate tags from the Rust crate versions, so upgrading the Python package and upgrading the Rust crate are two independent decisions, and a bug fixed in one may not be present in the other at the same moment.

The licence is Apache-2.0, declared in the workspace Cargo.toml and shown by the PyPI badge in the README. Apache-2.0 permits commercial use and modification and includes a patent grant; it also requires that you preserve the licence and attribution notices when redistributing. That is a summary of the licence text, not legal advice, and if you are redistributing delta-rs inside a product you should read LICENSE.txt in the repository and get your own counsel.

Upgrade cost is dominated by the dependency pins rather than by API churn. The kernel dependency is constrained to a version range, Arrow is pinned at version 59 across eight crates, and the minimum Rust version is 1.94.1. A workspace that already uses a different Arrow major version will have to reconcile that, because Arrow types appear in delta-rs's public API. The repository's own Makefile is the reference for what a supported build looks like.

Running the project's own test infrastructure locally

The Makefile is built to reproduce CI behavior on a workstation. The default goal is help, which prints the documented targets. The setup-dat target downloads the Delta Acceptance Tests archive at version 0.0.3 from the delta-incubator/dat releases, extracts it, and moves the output to dat/v0.0.3. The coverage target depends on setup-dat and runs cargo llvm-cov across the workspace with the full feature set, writing an lcov report and rendering it with genhtml.

bash
make setup-dat
make check
make coverage

Anyone evaluating delta-rs for a storage backend should run the integration suite against the docker-compose services first. The localstack service expects AWS_ACCESS_KEY_ID=deltalake and AWS_SECRET_ACCESS_KEY=weloverust, which are the credentials the project's tests use. If a backend configuration does not work there, it will not work against real object storage, and debugging it locally is considerably cheaper than debugging it in a pipeline.

Editorial conclusion

Adopt delta-rs when you need to read or write Delta tables from a Python or Rust process that has no Spark session, and check the feature table for the specific operations your pipeline depends on before committing. Do not adopt it as a drop-in replacement for Spark's full Delta connector, and do not assume every table feature is implemented. Verify first: whether the table you intend to write requires a protocol feature the feature table still lists as open, and whether your storage backend is covered by the integration tests in the Makefile's DEFAULT_FEATURES list.

Frequently asked questions

What is delta-rs?

It is a native Rust implementation of Delta Lake with bindings into Python, published as the deltalake crate on crates.io and the deltalake package on PyPI. It provides low-level APIs for developers and integrators plus a higher-level operations API for querying and inspecting tables.

How do you use delta-rs in Python?

Install it with pip install deltalake, then import DeltaTable and write_deltalake. The README's quick start writes a pandas DataFrame to a local directory with write_deltalake, reopens it with DeltaTable, and converts it back with to_pandas.

How does delta-rs compare with Spark for Delta Lake?

Spark is a distributed execution engine that reads and writes Delta tables; delta-rs is a Delta implementation you embed in a Python or Rust process, with no JVM and no cluster. Compute in the delta-rs case comes from whatever consumes the Arrow data, such as pandas, DataFusion or Polars.

How does delta-rs compare with Polars?

Polars is a query engine and appears in the README's integrations list, while delta-rs is the table implementation. The two are complementary rather than competing: delta-rs handles reading and writing the Delta table, and Polars can consume the data.

How does delta-rs compare with Iceberg?

The README describes Delta Lake as an open-source storage format that runs on top of existing data lakes and is compatible with engines like Apache Spark. It does not discuss Iceberg, so no comparison between the two formats can be drawn from the project's own documentation.

Official sources

  1. delta-io/delta-rs on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/delta-io-delta-rs.svg)](https://hysenlabs.com/projects/delta-io-delta-rs)