# apache/arrow-rs: The Rust Implementation of Arrow and Parquet

> arrow-rs is the official Rust implementation of the Apache Arrow columnar format and the Apache Parquet file format. It splits into separate crates for arrays, IPC, Flight, CSV, JSON and Parquet, which matters when you only need part of the stack.

**apache/arrow-rs** — Official Rust implementation of Apache Arrow

- Repository: https://github.com/apache/arrow-rs
- Website: https://arrow.apache.org/
- Stars: 3,621 · Forks: 1,338
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-arrow-rs

## What arrow-rs is for, and who ends up depending on it

Apache Arrow defines an in-memory columnar layout; Apache Parquet defines an on-disk columnar format. arrow-rs implements both in Rust. The README describes it as the "Native Rust implementation of Apache Arrow and Apache Parquet", and the repository is a Cargo workspace rather than a single library. That distinction drives most adoption decisions. A service that only decodes Parquet files does not need the Flight protocol code, and a service that only moves Arrow record batches over gRPC does not need the Parquet writer. The audience is Rust engineers building query engines, data loaders, analytics services, or interchange layers that must speak to Python, Java or C++ processes using the same formats. The homepage points at arrow.apache.org, which is the umbrella project site, not a Rust-only portal.

## Crate layout: arrow, arrow-flight, parquet and parquet_derive

The README's table names four crates with published API docs: arrow for core functionality (memory layout, arrays, low level computations), arrow-flight for the Arrow-Flight IPC protocol, parquet for the Parquet columnar file format, and parquet_derive for deriving RecordWriter and RecordReader for arbitrary, simple structs. The workspace members list in Cargo.toml is much longer and shows how the core is factored: arrow-array, arrow-buffer, arrow-schema, arrow-data, arrow-cast, arrow-csv, arrow-json, arrow-ipc, arrow-arith, arrow-ord, arrow-select, arrow-string, arrow-row and arrow-cmp, plus parquet-variant, parquet-variant-compute, parquet-variant-json and parquet-geospatial. In practice you depend on the umbrella crates and let Cargo pull the pieces. The README also records that object_store used to live here and has moved to a separate repository, arrow-rs-object-store. If you are following an older tutorial that adds object_store as a workspace member of arrow-rs, that layout no longer applies.

## Installing arrow-rs and getting to a first Parquet read

The README does not include a cargo add line, a Cargo.toml dependency snippet or a worked example, so there is no install command to copy from it. What the README does give is the set of published crates and where each one lives: the table lists arrow on crates.io with docs at docs.rs/arrow/latest, arrow-flight on crates.io with docs at docs.rs/arrow-flight/latest, parquet on crates.io with docs at docs.rs/parquet/latest, and parquet_derive on crates.io with docs at docs.rs/parquet-derive/latest. Those crates.io pages are where the install instructions and version numbers live, and the README points to arrow.apache.org/rust for the development version of the API documentation. The README states that arrow-rs and parquet are built and tested with stable Rust, so a current stable toolchain is the expected baseline. For struct-to-Parquet mapping, parquet_derive provides the derive macros, and its own README is the place to check which field types are supported. The README itself does not walk through a first program, so the docs.rs pages for arrow and parquet are the reference for the reader and writer APIs.

## Versioning, MSRV and the monthly release train

arrow-rs follows Semantic Versioning and releases approximately monthly. The README is explicit about the cadence: new major versions with potentially breaking API changes at most once a quarter, incremental minor versions in the intervening months, and all releases cut from main. The planned schedule in the README lists 59.2.0 as a minor release with no breaking API changes, 60.0.0 as a major release with potentially breaking changes, then 60.1.0 and 60.2.0 as minors, and 61.0.0 as the next major. That is a real cost for downstream projects. A dependency that moves a major version every quarter forces you to track the changelog rather than pin and forget. The MSRV policy softens this: the minimum supported Rust version is a rolling value that can only move in major releases, must be at least six months old when selected, and stays fixed across the minor releases between majors. If a Rust hotfix lands for the current MSRV, the README says the MSRV is updated to the specific minor version containing the applicable hotfixes.

## Where arrow-rs does not fit

Platform support is the clearest boundary. The README states that only little-endian platforms are officially supported and tested in CI, that big-endian platforms are not tested and may not work correctly, and that fixes for them are welcome on a best-effort basis with no compatibility guarantee. If you ship to a big-endian target, this is not a supported configuration. The panic versus Result guidance is the second boundary, and the README's excerpt ends mid-sentence at "In general, use panics for bad states", so read the full section in the repository before deciding how much defensive error handling to wrap around a call. A third consideration is scope. This is a format and compute implementation, not an execution engine. It gives you arrays, buffers, casting, sorting, selection and file readers; it does not give you a query planner, a scheduler or a storage layer. Teams expecting a DataFrame runtime with lazy evaluation will find the API lower level than they want.

## arrow-rs compared with Polars

Polars is the obvious alternative for Rust users, and the difference is one of layer. Polars is a DataFrame library with its own expression API and query planning; arrow-rs is the columnar substrate underneath work like that, exposing arrays, buffers and record batches directly. If your job is to load a Parquet file, filter rows and aggregate, Polars gives you the operations as a language. With arrow-rs you assemble the same pipeline from compute kernels and reader APIs, which is more code but leaves you in control of memory layout and of exactly when a copy happens. The two are not mutually exclusive: a DataFrame engine can sit on top of an Arrow implementation. Choose arrow-rs when you are building the engine, the interchange layer or the file reader, and choose a DataFrame library when you are writing an application that consumes data.

## Licence and the cost of tracking main

arrow-rs is licensed under Apache-2.0, and the repository carries LICENSE.txt and NOTICE.txt alongside an ASF-style file header in every source file. For most consumers that is a permissive licence with a patent grant, and the NOTICE file is the part to preserve if you redistribute. This is not legal advice; check your own obligations. The upgrade cost is the more practical concern. Releases come from main on a monthly schedule, majors land at most quarterly, and the README states that how breaking API changes are handled is described in CONTRIBUTING.md under the breaking-changes section. A team that pins a version and upgrades once a year will accumulate several majors of drift, and the changelog is the record of what moved. The mitigation is to read the release issue linked in the planned schedule table before the upgrade, since each version in that table links to its own GitHub issue.

## Conclusion

Adopt arrow-rs when you need Arrow arrays, Parquet reading and writing, or Arrow Flight inside a Rust service, and you can absorb a major version roughly once a quarter. Do not adopt it if you need big-endian support, since the README states only little-endian platforms are officially supported and tested in CI. Before you commit, check the MSRV stated in the release notes against your toolchain, and read the contributing guide's breaking-changes section to see how API changes are staged.

## FAQ

### What is arrow-rs?

It is the official Rust implementation of Apache Arrow, an in-memory columnar format, and Apache Parquet, a columnar file format. The repository is a Cargo workspace whose published crates include arrow, arrow-flight, parquet and parquet_derive.

### How do you install arrow-rs in a Rust project?

The README does not give an install command; it lists the published crates, including arrow, arrow-flight, parquet and parquet_derive, with links to their crates.io and docs.rs pages, which is where the dependency instructions live.

### How does arrow-rs relate to Parquet?

The same repository provides both: the arrow crates cover the in-memory columnar format and the parquet crate covers the Parquet columnar file format. The README states that arrow and parquet are released on the same schedule with the same versions.

### How does arrow-rs compare with Polars?

arrow-rs exposes arrays, buffers, record batches and file readers at a lower level, while Polars is a DataFrame library with its own expression API. A DataFrame engine can be built on top of an Arrow implementation rather than replacing it.

## Sources

- [apache/arrow-rs on GitHub](https://github.com/apache/arrow-rs)
- [License: Apache-2.0](https://github.com/apache/arrow-rs/blob/main/LICENSE)
- [Project website](https://arrow.apache.org/)
- [README](https://github.com/apache/arrow-rs/blob/main/README.md)
- [Releases](https://github.com/apache/arrow-rs/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-arrow-rs
