Apache DataFusion: An Extensible Rust Query Engine Built on Apache Arrow
Apache DataFusion SQL Query Engine. Out of the box," DataFusion offers SQL and DataFrame APIs, excellent [performance], built-in support for CSV, Parquet, JSON, and Avro, extensive customization, and a great community.
At a glance
- What is it?
- Apache DataFusion is a Rust-based SQL query engine that uses Apache Arrow as its in-memory columnar format, providing SQL and DataFrame APIs with vectorized, multi-threaded execution. It is not an end-user database; it is a library for teams building custom query engines, analytic platforms, and data pipelines who want a working execution engine as a starting point.
- Who is it for?
- DataFusion suits teams building domain-specific query engines, new database platforms, query language interpreters, or data pipelines in Rust who want a production-quality execution foundation rather than writing a query planner and vectorized executor from scratch. It is not the right choice for teams who need a ready-to-run database with a server, or for Python-first data science workflows where DuckDB or Polars offer simpler installation.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Apache DataFusion Is and Who It Is For
DataFusion is not a database you deploy and query. It is a Rust library that provides the internals a database needs: a SQL parser, a query planner, an optimizer, and a vectorized multi-threaded execution engine. Teams that need to build a custom query engine, an analytic platform tailored to a specific workload, or a new query language start from DataFusion rather than writing these components from scratch.
The README lists domain-specific query engines, new database platforms, data pipelines, and query language implementations as the primary use cases. It explicitly invites developers to start from a fully working engine and then customise the features specific to their needs. The project documents a known-users list at datafusion.apache.org/user-guide/introduction.html#known-users covering deployed systems built on DataFusion.
Python and Java interfaces exist as separate sub-projects. DataFusion Python at github.com/apache/datafusion-python and DataFusion Java at github.com/apache/datafusion-java expose the SQL and DataFrame APIs to non-Rust developers. DataFusion Ballista scales DataFusion across a cluster, and DataFusion Comet is an Apache Spark accelerator built on DataFusion.
Architecture: Arrow In-Memory Format, Columnar Execution, and Partitioned Sources
DataFusion uses Apache Arrow as its in-memory data format. Arrow is a columnar memory layout designed for analytical workloads: each column is stored contiguously in memory, which makes vectorized operations efficient because the CPU can process multiple values in a single instruction pass.
The execution engine is columnar, streaming, multi-threaded, and vectorized. Columnar means execution works on column batches rather than row-by-row. Streaming means results are produced incrementally without materialising the full dataset. Multi-threaded means the engine partitions work across CPU cores. Vectorized means batch operations use SIMD instructions where available.
Partitioned data sources allow the engine to split input files or datasets across threads. DataFusion has built-in support for CSV, Parquet, JSON, and Avro file formats. The architecture separates the query planner from the execution layer: almost all points of the pipeline are extensible, including additional data sources, query languages, functions, and custom operators. This extensibility is the core selling point for teams building systems on top of DataFusion.
Adding DataFusion to a Rust Project
DataFusion is published to crates.io as the datafusion crate under the Apache-2.0 license. Add it to a Rust project's Cargo.toml in the [dependencies] section. The package offers several feature flags that control which optional capabilities are compiled in.
Default features include nested_expressions for functions that work with arrays and nested types, compression for reading files compressed with xz2, bzip2, flate2, and zstd, crypto_expressions for cryptographic functions like md5 and sha256, datetime_expressions for date and time functions like to_timestamp, and encoding_expressions for encode and decode functions.
Feature flags let teams include only what they need, which matters for compile times and binary size in constrained environments. The full feature list is in the Cargo.toml at the crate level, and the features documentation is at datafusion.apache.org.
For users who want to run SQL queries from the command line without writing Rust code, DataFusion ships a CLI tool. Installation instructions for the CLI are at datafusion.apache.org/user-guide/cli/installation.html.
SQL and DataFrame APIs
DataFusion exposes two query interfaces. The SQL API accepts standard SQL strings and executes them against registered data sources. The DataFrame API provides a programmatic chained interface for constructing query plans in Rust code, similar in concept to Spark DataFrames or Pandas DataFrames.
Both APIs are documented with Rust examples in the repository under datafusion-examples/ and in the user guide at datafusion.apache.org/user-guide/example-usage.html and datafusion.apache.org/user-guide/dataframe.html. The Python DataFrame API is documented separately at datafusion.apache.org/python/.
DataFusion supports full query planning, including joins, aggregations, window functions, and subqueries. Custom scalar and aggregate functions can be registered with the context and called from SQL or the DataFrame API. Custom table providers allow DataFusion to query data sources beyond the built-in file formats, which is how teams integrate proprietary storage systems.
Limitations and Where DataFusion Is Not the Right Tool
DataFusion is a library, not a deployable service. It has no built-in server, no authentication layer, no persistence layer, and no replication. Building a queryable database service requires adding those components. Teams that need a ready-to-run analytical database for ad-hoc queries should look at DuckDB, which is a self-contained analytical database with SQL support that runs in-process or as a command-line tool. DuckDB has broader language bindings than DataFusion and does not require writing Rust code to use it.
For Python data science workflows, Polars is another columnar DataFrame library that exposes a Python API and handles common analysis tasks without requiring knowledge of Rust or query engine internals. DataFusion's advantage over both is that it is designed to be embedded and customised at the library level, which DuckDB and Polars are not.
DataFusion's extensibility also means that the more you customise it, the more maintenance responsibility you take on. Custom operators, data sources, and functions must be kept compatible with DataFusion's evolving internals. The repository has no GitHub releases listed, meaning version management happens through the Cargo.toml version in the repository rather than GitHub release assets.
The Broader DataFusion Ecosystem
Several sub-projects extend DataFusion for specific deployment scenarios. DataFusion Ballista distributes query execution across a cluster of nodes, making DataFusion viable for datasets that exceed single-machine memory. DataFusion Comet integrates DataFusion as an execution accelerator for Apache Spark, allowing Spark workloads to run on DataFusion's vectorized engine while retaining the Spark API surface. DataFusion Python wraps the Rust library with a Python API, and DataFusion Java does the same for JVM applications.
The project roadmap is maintained through GitHub issues. The current discussion at the time of the last push is the DataFusion 2026 Q3-Q4 Roadmap Discussion at github.com/apache/datafusion/issues/22882. The contributor guide, communication channels, and architecture documentation are all at datafusion.apache.org/contributor-guide.
Maintenance and License
The last push to the repository was on 2026-09-25. The repository has no GitHub release tags; version management is handled through the Cargo workspace. The project is an Apache Software Foundation project licensed under the Apache-2.0 license, which permits commercial use, modification, and distribution. The Apache-2.0 license requires preserving copyright notices and includes a patent grant, which is relevant for commercial deployments. The LICENSE.txt and NOTICE.txt files in the repository contain the full terms.
Editorial conclusion
DataFusion suits teams building domain-specific query engines, new database platforms, query language interpreters, or data pipelines in Rust who want a production-quality execution foundation rather than writing a query planner and vectorized executor from scratch. It is not the right choice for teams who need a ready-to-run database with a server, or for Python-first data science workflows where DuckDB or Polars offer simpler installation. The right first step is to check the datafusion crate on crates.io, review the Cargo.toml feature flags for the capabilities your use case requires, and read the architecture documentation at datafusion.apache.org/contributor-guide/architecture.html.
Frequently asked questions
What is Apache DataFusion?
Apache DataFusion is an extensible SQL query engine written in Rust that uses Apache Arrow as its in-memory format. It provides SQL and DataFrame APIs, a full query planner and optimizer, and a vectorized multi-threaded execution engine. It is a library for building query engines and analytic platforms, not a standalone database server.
What is DataFusion used for?
DataFusion is used for building custom query engines, new database platforms, domain-specific analytic systems, data pipelines, and query language implementations in Rust. It provides a working execution engine so teams do not need to write a query planner, optimizer, and vectorized executor from scratch.
How does DataFusion compare to DuckDB on benchmark performance?
DataFusion is a Rust library for building custom query engines, not a benchmarkable database product in itself. Its vectorized, Arrow-native execution engine is designed for high throughput. The README links to benchmark.clickhouse.com for external performance comparisons. DuckDB is a ready-to-run analytical database with its own optimizer; the two serve different audiences.
How does DataFusion compare to Apache Spark?
DataFusion runs as a single-node in-process library; Apache Spark is a distributed cluster computing framework. DataFusion Ballista extends DataFusion for distributed execution, and DataFusion Comet is a Spark execution accelerator. For single-machine analytics, DataFusion avoids the JVM and cluster management overhead that Spark requires.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/apache-datafusion)