Daft: Distributed DataFrame Engine for Multimodal AI Data Pipelines
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
At a glance
- What is it?
- Daft is a Python data engine built on a Rust core that processes images, audio, video, and structured data in a single DataFrame API. It runs locally or distributes across Ray or Kubernetes clusters without changing application code, and it supports built-in AI operations including LLM prompts and embedding generation.
- Who is it for?
- Daft suits data engineers and ML practitioners who process structured data alongside images, audio, or video and want to scale that processing to a distributed cluster without rewriting their code. The Apache 2.0 licence permits commercial use.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Daft was built to solve in AI data workloads
Data pipelines for machine learning increasingly need to handle more than rows and columns. Processing a dataset that mixes product descriptions, product images, audio recordings, and structured metadata typically requires gluing together several libraries, each with its own datatype, memory model, and execution environment. Pandas works well for structured data but represents images and other binary types as Python objects with no native processing. Ray Data adds distributed processing and supports multimodal types, but the README's comparison table notes it has no query optimizer. Daft was built to close that gap: a single DataFrame API, backed by Arrow memory, with a query optimizer and first-class support for images, audio, video, and embeddings alongside structured columns. The stated goal is to remove the need to context-switch between tools when the input data contains more than one modality.
The feature set listed in the README includes running LLM prompts, generating embeddings, and classifying data at scale using OpenAI, Hugging Face Transformers, or custom models. All of those operations are expressed as DataFrame calls, which means they pass through the same query optimizer and distributed execution plan as a SQL filter or a join.
How Daft's Rust core and Python API fit together
Daft is built using Maturin, which compiles a Rust extension and exposes it through a Python package. The build system is declared in pyproject.toml with maturin as the build backend. The src/ directory contains the Rust source, while the daft/ directory holds the Python layer. The Cargo.toml at the root references a large set of internal crates covering catalog integration, CSV parsing, compression, Delta Lake, Iceberg, Avro, text processing, and more; each crate lives under src/ as a separate Rust package. This structure means the Python API is a thin wrapper over a multi-crate Rust engine rather than a Python library with optional native extensions.
The consequence for users is that installing daft installs a compiled binary. There is no pure-Python fallback path. On supported platforms the prebuilt wheel from PyPI installs without a Rust toolchain. Building from source requires maturin (version 1.5.0 to below 2.0.0) and the Rust toolchain pinned in rust-toolchain.toml. For Apple Silicon (M1 and later) the Makefile notes specific environment variables for grpcio compilation, so teams using that hardware should consult the Makefile before building from source.
Installing Daft and enabling optional integrations
The README documents a single pip command as the primary install path:
pip install daftPython 3.10 or higher is required. The base install brings in pyarrow (16.0.0 to below 26.0.0), tqdm, and packaging. Optional extras are declared in pyproject.toml and cover a wide range of integrations: audio processing (soundfile, librosa), cloud storage (aws, azure, google), streaming (kafka), data catalog (iceberg, deltalake, unity, gravitino), vector search (lance, turbopuffer), AI (openai, transformers, huggingface), and distributed execution (ray). A combined extra named all installs every one of them. For advanced installations such as building from source or adding extras beyond the base package, the README points to the installation guide at docs.daft.ai.
Native multimodal types and the comparison to Pandas and Polars
The README includes a comparison table that evaluates Daft against Pandas, Polars, Modin, Ray Data, PySpark, and Dask across six dimensions: query optimizer, multimodal support, distributed execution, Arrow backing, vectorized execution engine, and out-of-core processing. Daft claims yes across all six. Pandas has no query optimizer and no distributed execution, and while it has optional Arrow backing since version 2.0, multimodal types remain Python objects. Polars has a query optimizer and vectorized execution but no distributed execution and represents multimodal data as Python objects.
The distinction between a native column type and a Python object matters because a Python-object column bypasses the vectorized execution engine. Each item is processed in Python rather than in the Rust core, which removes the performance advantages of the columnar layout. Daft's claim is that images, audio, video, and embeddings are processed as typed Arrow arrays through the same Rust engine as structured data. That claim is central to the project's value proposition for mixed-modality pipelines, though the README does not describe the internal encoding of those types in detail.
Telemetry and dependency version pinning
Daft collects non-identifiable usage data through Scarf. The README describes the collected data as metadata-only, with no session IDs or user identifiers, and states it is not sold. To disable telemetry, set the environment variable DO_NOT_TRACK to true before running any Daft code. The project's telemetry documentation is at docs.daft.ai/en/stable/telemetry/.
The pyproject.toml pins optional dependencies tightly. The audio extra pins soundfile to >=0.13.0,<0.14.0 and librosa to >=0.11.0,<0.12.0. The huggingface extra pins datasets to <4.9.0. The deltalake extra requires >=1.6.0,<1.7.0. The google extra pins pillow to ==12.2.0. Teams that also use these libraries for other purposes will need to verify that these version windows are compatible with their other requirements before installing. This is a structural constraint of shipping a large number of integrations from one package: the version windows that are tested and known to work narrow over time as upstream libraries change. Checking for conflicts before committing to an install is worth doing explicitly, especially on teams that already use a specific version of PyArrow or Pillow for image work. The project is under Apache 2.0, which permits commercial use and redistribution.
When to choose Polars or PySpark instead of Daft
Polars is the most relevant comparison for teams evaluating Daft as a DataFrame engine. Both are built on Rust, both use Apache Arrow for memory representation, and both include a query optimizer and a vectorized execution engine. According to the README's comparison table, Polars does not support distributed execution and represents multimodal data as Python objects. Polars is a strong fit for single-machine workloads that do not include images, audio, or video: its API is mature, single-node performance is well-documented, and its dependency footprint is smaller than Daft's.
PySpark covers the distributed case and has a query optimizer, but the README table notes that PySpark has no native multimodal support and its vectorized execution relies on Pandas UDFs. Modin offers a Pandas-compatible API with distributed execution but has no query optimizer and does not use Arrow as its storage format. Daft's claim is that it handles both scale and modality natively in one API. For a team that already runs Spark infrastructure and does not process image or video data, migrating to Daft involves a real adoption cost with limited gain. The evaluation should start with a concrete inventory of data types in the pipeline and whether single-machine processing is already a bottleneck.
Editorial conclusion
Daft suits data engineers and ML practitioners who process structured data alongside images, audio, or video and want to scale that processing to a distributed cluster without rewriting their code. The Apache 2.0 licence permits commercial use. Python 3.10 or higher is required. Before adopting Daft, test that its optional dependency version pins do not conflict with the rest of your environment, and confirm whether your team needs distributed execution: if single-machine scale is sufficient and multimodal data types are not in scope, Polars offers a lighter adoption path. The last push was on 2026-09-26, and the latest release is v0.7.25 from 2026-09-11.
Frequently asked questions
What is Daft and what can it do?
Daft is a Python data engine backed by Rust that processes structured data, images, audio, and video in a single DataFrame API. It supports built-in AI operations such as LLM prompts, embedding generation, and classification, and it scales from a single machine to Ray or Kubernetes clusters.
How do I install Daft and what Python version is required?
Install Daft with pip install daft. Python 3.10 or higher is required. Optional extras declared in pyproject.toml (such as ray, aws, and huggingface) add support for distributed execution and cloud storage.
Does Daft support SQL queries?
The pyproject.toml lists sql as a named optional extra, suggesting SQL support is available through that install path. Full documentation for the SQL interface is at docs.daft.ai.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/eventual-inc-daft)