Library / SDK
delta-io/delta avatar
delta-io/delta

Delta Lake: ACID Storage for Lakehouse Architectures

Project brief: An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs.

9,032 stars2,188 forksScalaApache-2.0

At a glance

What is it?
Delta Lake is an open-source storage framework built on top of cloud object storage that brings ACID transactions, schema enforcement, and time travel to data lakes. It integrates with Apache Spark as its primary engine and exposes connectors for Flink, Trino, PrestoDB, and Hive.
Who is it for?
Delta Lake is a sound choice for organisations already using Apache Spark who need ACID guarantees on cloud object storage. The transaction log protocol is open and formally specified, which means other compute engines can read Delta tables without going through Spark.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Scala, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Problem Delta Lake Solves in Object Storage

Plain object storage (S3, GCS, Azure Data Lake Storage) provides no atomicity for multi-file writes. A job that writes a Parquet dataset to S3 by appending new files leaves a window during which readers may see a partial state. Deletes and updates require rewriting files, and two concurrent writers can corrupt each other's output without a locking mechanism.

Delta Lake addresses these problems by placing a transaction log alongside the data files in the same storage directory. Every write, delete, or update is recorded as a versioned JSON entry in that log before any data files change. Readers consult the log to determine which files are part of the current table state. This design means that the storage system itself only needs to provide atomic single-file writes, which all major cloud stores support.

How the Transaction Log Works

The Delta transaction log lives at `<table path>/_delta_log/`. Each commit adds a numbered JSON file describing the set of files added and removed in that transaction. Periodically, the log is checkpointed into a Parquet file to speed up log replay on large tables.

Delta Lake guarantees serializability for concurrent reads and writes. The concurrency documentation describes optimistic concurrency control: two writers each record their intent in separate log entries, and the system detects and rejects conflicting operations. For storage systems that support atomic rename, this is enforced at the file system level. The documentation notes that the underlying storage must provide atomic file visibility, mutual exclusion on writes, and consistent directory listings.

Backward compatibility is guaranteed: newer versions of Delta Lake can always read tables written by older versions. Forward compatibility is not guaranteed. When new protocol features are introduced, the `Protocol` action in the transaction log records updated minimum reader and writer version numbers, and an older client may refuse to read a table produced by a newer client.

Installing Delta Lake for Spark and Running a First Query

The Python package for Spark integration is `delta-spark`:

bash
pip install delta-spark

The setup.py shows that `delta-spark` requires pyspark between 4.0.1 and 4.2.0 and Python 3.10 or later. The version is taken from `version.sbt`, which tracks the Scala build version. The latest release documented in the repository is v4.4.0, released 2026-08-20.

The `delta-rs` repository in the delta.io organisation provides a Rust implementation with Python and Ruby bindings for use outside of Spark. The README links to it but it is maintained as a separate project.

Documentation for getting started with Scala, Java, and Python is at docs.delta.io. The README does not inline Spark usage examples; it points to the Quick Start Guide for that.

Connectors: Flink, Trino, PrestoDB, and Hive

Delta Lake ships connectors for multiple compute engines beyond Spark. The Apache Flink connector at `connectors/flink/` is marked as preview in the README. It allows Flink jobs to write to Delta tables but is not described as production-ready.

PrestoDB reads Delta tables through its own connector, documented at `prestodb.io/docs/current/connector/deltalake.html`. Trino reads from and writes to Delta tables through its connector at `trino.io/docs/current/connector/delta-lake.html`. Apache Hive reads Delta tables through the Hive integration documented at docs.delta.io.

Delta Standalone is a JVM library that allows JVM-based projects, including Flink, Hive, Beam, and PrestoDB, to read and write Delta tables without going through Spark. It implements the Delta transaction log protocol to provide transactional guarantees at the metadata level. The Delta Rust API at `docs.rs/deltalake` provides low-level access to Delta tables for Rust, Python, and Ruby code.

The full list of integrations is maintained at delta.io/integrations.

API Stability and Upgrade Cost

Delta Lake distinguishes between two types of public API. The Java, Scala, and Python API classes and methods documented in the API docs are treated as stable public APIs. Internal classes are subject to change without notice. Spark-based APIs accessed through `DataFrameReader` and `DataFrameWriter` are stable within a major version.

The Apache Spark compatibility matrix is the main upgrade constraint. Each Delta Lake major version is tied to specific Spark versions, and the matrix is documented at docs.delta.io/latest/releases.html. When a Spark upgrade is planned, the Delta Lake version must be checked against the new Spark version before upgrading either.

Release 2.0.0 introduced breaking changes. Version 4.4.0 was released on 2026-08-20. The changelog for breaking changes is in the CONTRIBUTING.md and the online release notes.

Storage Requirements and Concurrency Limitations

Delta Lake's ACID guarantees depend on the storage layer meeting three conditions. First, files must become visible atomically: a reader sees either the complete file or nothing. Second, only one writer can create a file at the final destination. Third, once a file appears in a directory, all subsequent listings must return it. The documentation on storage configuration at docs.delta.io covers which cloud stores meet these conditions and what configuration is needed for those that do not by default.

Some storage backends require additional configuration to guarantee mutual exclusion. The `storage-s3-dynamodb/` directory in the repository contains a DynamoDB-based log store for S3, which provides the mutual exclusion that plain S3 does not guarantee on its own.

Delta Lake does not currently provide row-level security or column masking at the storage layer. Those controls are applied by the compute engine, not by Delta Lake itself.

Apache Iceberg as an Alternative Format

Apache Iceberg is the other widely adopted open table format for cloud data lakes. Both Delta Lake and Iceberg provide ACID transactions, time travel, and schema evolution on top of object storage. The key architectural difference is the transaction log format: Delta Lake uses JSON commit files in `_delta_log/`, while Iceberg uses a metadata tree of JSON and Avro manifest files.

The repository contains an `iceberg/` directory, which suggests ongoing work on Iceberg compatibility or migration tooling. The delta.io organisation also maintains `delta-sharing`, a separate protocol for sharing Delta tables across organisations without moving data.

Teams choosing between the two formats should consult the current state of engine support in their compute stack. Both formats are widely supported, but specific engines may have better read or write performance with one format over the other depending on the version.

Editorial conclusion

Delta Lake is a sound choice for organisations already using Apache Spark who need ACID guarantees on cloud object storage. The transaction log protocol is open and formally specified, which means other compute engines can read Delta tables without going through Spark. Teams that want to adopt the Lakehouse pattern with strong transactional semantics and are comfortable running Spark or PrestoDB or Trino will find Delta Lake a production-ready foundation. The wrong scenario is a small pipeline that does not use Spark and needs a simpler format: Delta Lake's value is tied to the ecosystem of compatible engines. Before upgrading, check the Delta Lake and Apache Spark compatibility matrix on docs.delta.io, as major version boundaries have historically introduced breaking changes.

Frequently asked questions

Does Delta Lake guarantee ACID transactions on S3?

Yes, with additional configuration. Plain S3 does not guarantee mutual exclusion on writes, so Delta Lake on S3 requires either a log store that uses DynamoDB for locking or an S3 configuration that meets the requirements. The repository includes a DynamoDB-based log store in `storage-s3-dynamodb/`.

What is the difference between delta-spark and delta-rs?

delta-spark is the Python package for using Delta Lake with Apache Spark. delta-rs is a separate Rust implementation with Python and Ruby bindings that provides access to Delta tables without requiring Spark. Both are in the delta.io GitHub organisation but are maintained as separate projects.

Does Delta Lake support reading from engines other than Spark?

Yes. Trino, PrestoDB, and Apache Hive all have connectors that read Delta tables. Delta Standalone is a JVM library for projects that need to read or write Delta tables directly. The Rust API provides access for non-JVM systems.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/delta-io-delta.svg)](https://hysenlabs.com/projects/delta-io-delta)