Model or dataset
datachain-ai/datachain avatar
datachain-ai/datachain

DataChain: typed, versioned datasets over S3, GCS and Azure

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

2,819 stars161 forksPythonApache-2.0

At a glance

What is it?
DataChain indexes files in object storage into versioned, Pydantic-typed datasets stored in a local SQLite database. It is a good fit for teams whose pipelines keep recomputing the same embeddings and metadata.
Who is it for?
Adopt DataChain if your pipelines repeatedly read the same object-storage files to recompute embeddings, image metadata or other derived columns, and you want that work registered as a named dataset instead of a scratch script. Skip it if you need a shared multi-writer warehouse or a system with a published stability guarantee, since pyproject.toml still declares Development Status 2 - Pre-Alpha.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem DataChain targets: repeated computation over the same storage files

A pipeline that reads images or documents from S3, GCS or Azure usually recomputes everything on every run. The files have not changed, but nothing in the pipeline knows that, so the same embedding model runs over the same bytes again. DataChain's answer is to keep a typed index of what was already computed. The README describes it as "a Python library that turns files in S3, GCS, and Azure into versioned, typed datasets, queryable at warehouse speed." The audience is engineers building ML and multimodal pipelines, and, in the optional agent workflow, coding agents that need a stable description of a dataset rather than a directory listing. The README states that bytes never leave your storage: DataChain stores metadata and file pointers, not copies of the files.

How DataChain works: read_storage, map, save and the local SQLite Dataset DB

A pipeline is a chain of operations that ends in a saved dataset. dc.read_storage() enumerates files under a URI and yields file objects. A .map() call applies a Python function to each file and writes the return value into a new column. The return type of that function, a Pydantic model, becomes the dataset schema, so each field is a queryable column. .save() registers the result under a name and version, for example [email protected].

According to the README, every .save() registers the dataset in the Dataset DB, which holds schemas, versions, lineage and processing state and is kept locally in a SQLite database at .datachain/db. Pipelines then reference datasets by name rather than by path. The README states that when the code or the input data changes, the next run bumps the dataset version, and that re-runs only process new or changed files. On top of that store, the README describes vector search over the same rows, so embeddings do not need a separate vector database. The compute layer is parallel Python with async I/O, checkpoint recovery and incremental updates, distributed on Studio; the README claims sub-second filter, join and group_by over millions of typed records locally and hundreds of millions on Studio. Those scale numbers come from the project's own documentation and are not independently verified here.

Installing DataChain and building a first dataset

DataChain is a Python package on PyPI and requires Python 3.10 or later, as declared in pyproject.toml. Install it with pip:

bash
pip install datachain

The README's create_dataset.py example reads JPEG files from a public bucket and maps each one to a Pydantic model. Note the anon=True flag, which the example uses because the bucket is public, and delta=True, which makes later runs skip unchanged files:

python
from PIL import Image
import io
from pydantic import BaseModel
import datachain as dc


class ImageInfo(BaseModel):
    width: int
    height: int


def get_info(file: dc.File) -> ImageInfo:
    img = Image.open(io.BytesIO(file.read()))
    return ImageInfo(width=img.width, height=img.height)


ds = (
    dc.read_storage(
        "s3://dc-readme/oxford-pets-micro/images/**/*.jpg",
        anon=True,
        update=True,
        delta=True,  # re-runs skip unchanged files
    )
    .settings(prefetch=64)
    .map(info=get_info)
    .save("pets_images")
)
ds.show(5)

After the run, ds.show(5) prints the file path, file size and the nested info fields as dotted columns (file path, file size, info width, info height). The dataset is registered as [email protected]. Run the same script again and, with delta=True, the unchanged files are skipped rather than reprocessed.

There is also an optional agent path. datachain skill install --target claude installs a skill for Claude Code; the README lists cursor, codex, copilot and pi as other targets. The README's agent example copies a reference image with datachain cp --anon s3://dc-readme/fiona.jpg . and then prompts the agent to find similar dogs filtered by breed, mask availability and width. The agent output shown in the README is a ranked table with breed names and distance values. Treat that table as the project's illustration, not as a benchmark.

Where DataChain gets in the way: local state, pre-alpha status and storage assumptions

The Dataset DB is a local SQLite file at .datachain/db. That is convenient on one machine and awkward on a team: two engineers running the same pipeline in different working directories get different databases, and the README does not document a shared or remote Dataset DB for the open source package. The README mentions Studio for distributed execution and MCP access to the same datasets, but it does not describe how dataset state is reconciled between a laptop and Studio. If your team expects a single shared catalog, that gap matters.

The packaging metadata is blunt about maturity. pyproject.toml carries the classifier "Development Status :: 2 - Pre-Alpha", and the version is still 0.59.x. The release cadence is fast (0.59.6 on 2026-08-17, 0.59.7 on 2026-08-24, 0.59.8 on 2026-09-10), which is reassuring for maintenance but also means APIs can move. The last push to the repository was on 2026-09-10.

There is also a dependency footprint to weigh. DataChain pulls in pandas, pyarrow, numpy, fsspec with s3fs, gcsfs and adlfs, SQLAlchemy and litellm, the latter pinned to >=1.83.0,<1.92 with a comment in pyproject.toml explaining that 1.92.0 adds a Rust extension without wheels for some platforms. If you only need to list files and compute one column, a short boto3 script plus a Parquet file is less machinery. DataChain earns its place when the derived columns are expensive and you want the index, the schema and the lineage kept together.

DataChain compared with DVC, which shares its lineage

DataChain is not the first tool from this lineage. Its pyproject.toml depends on dvc-data and dvc-objects, and the author field points at the DVC project's maintainer, so the storage-hashing machinery underneath is shared. The difference is the unit of work. DVC versions files and directories as artifacts and tracks them through a Git repository, which suits models, checkpoints and dataset snapshots that you want to reproduce exactly. DataChain versions the result of a computation: a schema, a set of typed columns and the file pointers they came from. You query that result with filter, join and group_by instead of checking out a directory and writing your own loader. If your problem is "which version of this model file was used", DVC is the closer fit. If your problem is "which images already have embeddings and which breed metadata", DataChain's dataset model matches it better.

Maintenance, licensing and what an upgrade costs

The repository is not archived and the last push was on 2026-09-10, one week before this writing, so the project is being worked on. The release history shows three patch releases in August and September 2026, which suggests fixes land quickly but also that pinning versions is wise. Because the package is still at 0.59.x with a pre-alpha classifier, an upgrade can change the chain API or the on-disk Dataset DB layout; the README does not document a migration path for .datachain/db, so a schema change in a future release could mean rebuilding datasets from storage. That rebuild is a re-run of the pipeline, not a data loss, since the files stay in your bucket.

The licence is Apache-2.0, declared in pyproject.toml with a LICENSE file in the repository root. That is a permissive licence with an explicit patent grant, and it does not oblige you to publish modifications. The README does not describe any separate commercial licence for the open source package; Studio is mentioned as a hosted option but its terms are not covered by the repository files. This is a description of the licence text, not legal advice.

Editorial conclusion

Adopt DataChain if your pipelines repeatedly read the same object-storage files to recompute embeddings, image metadata or other derived columns, and you want that work registered as a named dataset instead of a scratch script. Skip it if you need a shared multi-writer warehouse or a system with a published stability guarantee, since pyproject.toml still declares Development Status 2 - Pre-Alpha. Before committing, run the README's create_dataset.py example against one of your own buckets with delta=True and confirm that a second run reports no changed files.

Frequently asked questions

What is DataChain?

It is a Python library that indexes files in S3, GCS, Azure or a local filesystem into typed, versioned datasets. Each dataset is registered in a local SQLite Dataset DB with its schema, version and lineage, and pipelines reference datasets by name instead of by path.

How do I install DataChain?

The README gives pip install datachain, and pyproject.toml requires Python 3.10 or later. An optional agent skill is added with datachain skill install --target claude, with cursor, codex, copilot and pi listed as other targets.

Where does DataChain store its dataset metadata?

The README states that the Dataset DB, which holds schemas, versions, lineage and processing state, is kept locally in a SQLite database at .datachain/db. The file bytes themselves stay in your own storage.

Does DataChain work with S3, GCS and Azure?

Yes. The README lists S3, GCS, Azure and local filesystems as supported, and pyproject.toml depends on fsspec with s3fs, gcsfs and adlfs. The README's example passes anon=True when reading from a public bucket.

Official sources

  1. datachain-ai/datachain on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datachain-ai-datachain.svg)](https://hysenlabs.com/projects/datachain-ai-datachain)