Framework
vortex-data/vortex avatar
vortex-data/vortex

vortex-data/vortex: a columnar file format built around pluggable encodings

An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.

3,221 stars227 forksRustApache-2.0

At a glance

What is it?
Vortex is a Rust columnar format and toolkit, Apache-2.0 licensed under LF AI & Data, that separates logical schema from physical layout and ships a pluggable encoding system. The file format is declared stable from 0.36.0; the library APIs are not.
Who is it for?
Adopt Vortex if you are building a data system on object storage in Rust or Python and want the encoding layer to be yours, not the format's. Do not adopt it if you need a frozen library API today, or if your stack is already committed to Parquet and Iceberg with no room for a second format.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Vortex actually replaces

The README frames Vortex as a columnar file format and toolkit for data systems backed by object storage. The target is not the query engine but the layer underneath it: how bytes are laid out on disk or in an object store, how much of a file a reader has to fetch to answer a point lookup, and how much of the encoding logic a system author is allowed to change. That last part is the real subject. Vortex describes a pluggable encoding system, a pluggable type system, pluggable compression strategy and pluggable layout strategy, with an architecture the README says is modeled after Apache DataFusion's extensible approach. If you have ever wanted a compression scheme that Parquet does not have, this is the pitch.

The audience follows from that. This is a library for people writing storage engines, query engines, DataFrame runtimes or format converters. It is not an end-user analytics tool, and there is no server to deploy. The README lists integrations with Arrow, DataFusion, DuckDB, Spark, Pandas and Polars, with Apache Iceberg marked as coming soon, which tells you the intended shape: Vortex sits behind an engine you already use rather than in front of it.

Logical schema, physical layout, and the encoding stack

The design splits into a logical layer that defines data types and schema, and a physical layer that handles encoding and storage. Encodings come in two kinds. Built-in encodings are compatible with Apache Arrow's memory format, which is what makes the zero-copy conversion claim possible; extension encodings are the optimized schemes. The README names RLE and dictionary as examples. Cascading compression is supported, meaning nested encoding schemes, and the repository layout backs this up: the workspace carries a long list of encoding crates, among them fastlanes, runend, sequence, alp, datetime-parts, fsst, pco, sparse, zigzag, zstd, bytebool, parquet-variant and onpair.

That list is the most informative thing in the repository. A format that ships a dedicated crate for FastLanes bit-packing and another for ALP floating-point compression is making a bet that different data wants different physical representations, and that the format should not have to be revised to add one. The cost is visible too: each encoding is code that has to be maintained, tested and kept readable across versions. The README also mentions lazy-loaded summary statistics for optimization, which is the mechanism a reader uses to skip data without decoding it.

On performance, the README claims 100x faster random access reads versus modern Apache Parquet, 10-20x faster scans, 5x faster writes and similar compression ratios. Those are the project's own numbers, published without the benchmark methodology in the README itself; the repository points to bench.vortex.dev and to a bench-orchestrator directory for reproducing comparisons. Treat the ratios as a hypothesis to test on your data, not as a settled result.

Installing Vortex and reading a file

The README gives three entry points. For Rust, all features are exported through the main vortex crate:

bash
cargo add vortex

For Python, the package is vortex-data:

bash
uv add vortex-data

There is also a command line tool, vx, for browsing the structure of Vortex files. The README recommends the pre-built binary path, with a source build and a no-install route through uvx as alternatives:

bash
cargo binstall vortex-tui
cargo install vortex-tui --locked
uvx --from vortex-data vx --help

Once installed, the documented usage is a single subcommand against a file path. Running it should print the file's structure rather than its rows, which is the point of the tool: it is an inspector for layout and encodings, not a query shell.

bash
vx browse <file>

The README does not document a writing example, a schema-inspection flag, or an output format for vx browse, so plan on reading the docs site for anything beyond this. For development against the repository itself, the README lists prerequisites on macOS (flatbuffers and protobuf for .fbs and .proto files, duckdb for benchmarks), a Rust toolchain via rustup, submodule initialization with git submodule update --init --recursive, and uv sync --all-packages. It also notes that rust-toolchain.toml pins the toolchain used for development and CI, and that building Vortex as a dependency only requires a toolchain satisfying the Rust version compatibility policy.

The stability boundary is narrower than it sounds

The README is unusually direct about this, and it is the first thing to internalize. Library APIs may change from version to version. What is stable is the file format: from release 0.36.0 onward, all future releases should be able to read files written by any earlier version at or above 0.36.0. So the compatibility promise covers bytes on disk, not the functions you call to produce them. A dependency bump can break your build while leaving your data readable, which is the opposite of the trade-off most storage libraries make.

There are two other constraints worth checking before you commit. The Rust version compatibility policy states that Vortex supports the four most recent stable minor releases. Writing the latest stable release as 1.N, that means 1.N through 1.N-3 all build Vortex. The MSRV is declared once as rust-version in the root Cargo.toml and inherited by every crate in the workspace, and the README says that value, not the policy document, is the source of truth. An MSRV older than 1.N-3 is acceptable; newer does not meet the policy. CI enforces it in a Rust (MSRV) job. If you are pinned to an old toolchain, check the declared value rather than assuming.

Second, the README states that Apache Iceberg support is coming soon, which means it is not there yet. If your catalog and table semantics run through Iceberg, Vortex is not a drop-in replacement for your current format today. The README also does not document a rollback path for format changes or a downgrade procedure, so a team writing Vortex files in production has no documented way back to an older writer.

How it compares to Parquet and to Arrow IPC

The honest comparison is with Apache Parquet, because that is what the README benchmarks against and what most object-storage data systems already use. The difference is not primarily compression ratio, which the README describes as similar. The difference is where extensibility lives. Parquet's encoding set is defined by the format specification; adding a new one means changing the spec and every reader. Vortex puts encodings in separate crates and treats the physical layout as a pluggable strategy, so a new scheme is a new crate rather than a format revision. The cost of that freedom is ecosystem reach: Parquet readers exist everywhere, and Vortex readers exist where someone has integrated the library.

Against Arrow IPC the split is cleaner. Arrow IPC is an interchange format for in-memory arrays; the README describes Vortex's built-in encodings as compatible with Arrow's memory format and advertises zero-copy conversion to and from Arrow arrays. Vortex is the on-disk counterpart, adding compression, layout and statistics that IPC does not attempt. If your problem is moving arrays between processes, Arrow IPC is the right tool and Vortex adds nothing. If your problem is storing those arrays cheaply on object storage and reading them back selectively, that is the gap Vortex targets.

Maintenance cost, release cadence and licensing

The last push to the default branch develop was on 2026-09-20, and the most recent releases are 0.86.1 and 0.86.0 on 2026-09-11, following 0.85.0 on 2026-08-21. The repository is not archived. The cadence implied by those dates is fast, and combined with the README's statement that library APIs may change between versions, that cadence is a real cost: upgrading means reading release notes and expecting breakage in your own code. Budget for it.

Licensing is Apache-2.0, and the project is an Incubation Stage project at LF AI & Data, part of the Linux Foundation, formerly at spiraldb. The repository carries a LICENSE file, a LICENSES/ directory and a REUSE.toml, which suggests REUSE-style licence metadata is maintained per file. The practical implication for adopters is that Apache-2.0 is permissive and carries a patent grant, but the repository also contains Java, CUDA and DuckDB integration crates, and those dependencies have their own licences that you should review for your distribution. That is a review task, not legal advice.

The upgrade surface is larger than a single crate. The workspace lists dozens of members, from vortex-array and vortex-file through vortex-datafusion, vortex-duckdb, vortex-cuda and a Java binding, plus the encoding crates. If you depend on the top-level vortex crate you inherit that graph; if you depend on an integration crate you inherit its engine's version constraints as well.

Where Vortex is the wrong choice

If your team needs a frozen API surface and cannot absorb breaking changes on a monthly-ish cadence, the README's own stability note disqualifies Vortex for now. The file format is stable; your code against it is not.

If your data platform is built on Iceberg tables, the README lists Iceberg as coming soon, so the integration you need does not exist yet. Building on Vortex today means either waiting or maintaining the bridge yourself.

If your workload is small files and simple scans, the encoding machinery is overhead you will not recover. The claimed gains are in random access and scan throughput on object storage, and the README does not claim a win on compression ratio, so a workload that is already I/O-bound on network transfer has little to gain.

If you are not in Rust or Python, check the integration list first. Arrow, DataFusion, DuckDB, Spark, Pandas and Polars are named; anything outside that set means going through the FFI layer or the Java binding, and the README does not document those paths in detail.

Editorial conclusion

Adopt Vortex if you are building a data system on object storage in Rust or Python and want the encoding layer to be yours, not the format's. Do not adopt it if you need a frozen library API today, or if your stack is already committed to Parquet and Iceberg with no room for a second format. Before committing, verify the declared MSRV in the root Cargo.toml against your toolchain, and confirm that the readers you depend on can open files written by the version you plan to write with.

Frequently asked questions

Is the Vortex file format stable across versions?

The README states that from release 0.36.0, all future releases should be able to read files written by any earlier version at or above 0.36.0. The library APIs are explicitly excluded from that promise and may change from version to version.

How do I install Vortex for Rust or Python?

The README gives cargo add vortex for the Rust crate and uv add vortex-data for the Python package. All Rust features are exported through the main vortex crate.

What is the vx command line tool in Vortex?

vx is the command line UI for browsing the structure of Vortex files, installed from the vortex-tui crate or run without installing via uvx --from vortex-data vx --help. The documented usage is vx browse <file>.

Which Rust version does Vortex require?

The README's compatibility policy says Vortex supports the four most recent stable minor releases, so with the latest stable written as 1.N, versions 1.N through 1.N-3 all build Vortex. The MSRV is declared as rust-version in the root Cargo.toml, and the README says that value is the source of truth.

Does Vortex support Apache Iceberg?

The README lists Apache Iceberg under integrations with the note that it is coming soon, so it is not available yet.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. vortex-data/vortex on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vortex-data-vortex.svg)](https://hysenlabs.com/projects/vortex-data-vortex)