# Vaex: out-of-core DataFrames for billion-row tables in Python

> Vaex memory-maps HDF5 and Arrow files and evaluates expressions lazily, so statistics and histograms run without loading the table into RAM. It is a strong fit for single-machine exploration of very large tabular files, and a poor fit for anything that needs a full materialized copy.

**vaexio/vaex** — Out-of-Core hybrid Apache Arrow/NumPy DataFrame for Python, ML, visualization and exploration of big tabular data at a billion rows per second 🚀

- Repository: https://github.com/vaexio/vaex
- Website: https://vaex.io
- Stars: 8,508 · Forks: 598
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/vaexio-vaex

## The problem Vaex targets: tables larger than memory

Pandas loads a table into RAM. When a file is 100 GB and the machine has 16 GB, that model stops working, and the usual answer is a distributed cluster. Vaex takes the opposite route: the README describes it as a library for lazy out-of-core DataFrames, and its stated design is memory mapping, a zero memory copy policy and lazy computation. The intended user is a data scientist or analyst on a single machine who needs summary statistics and visual exploration of a table that does not fit in memory. The README frames the payoff as delaying the point at which a cluster becomes necessary. That framing matters: Vaex is not trying to replace a warehouse or a distributed engine. It is trying to make one laptop or one workstation useful for longer on data that is already on disk.

## Memory mapping, lazy expressions and the filter that does not copy

The mechanism has three parts. First, data lives in a file format that supports memory mapping. The README names HDF5 and Apache Arrow, and says opening such files is instant because the file is mapped rather than read. Second, expressions are lazy: the README says feature engineering does not waste memory or time because data is transformed when needed. A column expression is a description, not an array, until something asks for values. Third, filtering does not duplicate the table. According to the README, filtering and evaluating expressions keep the data untouched on disk and stream it only when required, so a filtered view is a selection over the mapped file rather than a new in-memory frame. Aggregations such as mean, sum, count and standard deviation are computed on an N-dimensional grid, and the README states this reaches more than a billion rows per second. The README also describes parallelized groupby operations, especially on categorical columns, and joins that do not materialize the right table, which it says saves gigabytes of memory. Remote streaming from S3 is documented as working with memory mapping. The honest reading of all this: Vaex is fast where the operation can be expressed as a pass over mapped columns, and the cost model changes completely the moment an operation needs a materialized result.

## Installing Vaex and running a first aggregation

The README gives two installation routes, pip and conda. Both install the vaex meta-package, which pulls in the core, visualization, HDF5 and other components.

```bash
pip install vaex
```

The conda route is the alternative the README lists, and is the one to prefer if you want the dependency resolution handled by conda-forge.

```bash
conda install -c conda-forge vaex
```

For a source install, the repository's setup.py is explicit that it exists only for local development or installing from source, and it names the archive URL directly.

```bash
pip install https://github.com/vaexio/vaex/archive/refs/heads/master.zip
```

That setup.py builds a meta-package named vaex-meta and marks it Private :: Do Not Upload to pypi server, so this path is for working copies, not for redistribution. Once installed, the README's own description of the workflow is: open a file, then ask for statistics. The README does not print a full code sample, so the shape of a first session is an import, an open call on an HDF5 or Arrow file, and a statistic such as a mean or a histogram. What you should expect to see is that the open call returns quickly even for a large file, because the file is mapped and not read, and that the statistic triggers the actual pass over the columns. If your source data is CSV, the README points to its documentation on converting data efficiently from CSV files, pandas DataFrames and other sources. That conversion step is not optional in practice: the memory-mapping advantage applies to HDF5 and Arrow, not to a CSV you read row by row.

## Where Vaex is the wrong tool

The clearest limitation is the join. The README presents the no-copy join as a feature, and for memory it is. But it also means the result is not a conventional materialized table, so code that expects to index into a fully joined frame, or to mutate it in place, will not behave the way the same code behaves on pandas. The second limitation is format. Everything the README promises about instant opening rests on HDF5 and Arrow. A CSV pipeline gets none of it until you convert, and the conversion is a real step with its own cost. Third, the project's release history is thin at the top: the most recent release listed is vaexpaper_v1 from 2018-03-29, and the repository's last push was on 2026-04-01. A reader deciding today should treat that gap between release tagging and commit activity as something to verify rather than assume. Fourth, the README itself flags that Remote DataFrames have documentation coming soon, so that capability is not something you can plan around from the README alone. Finally, if your problem is genuinely distributed, Vaex is not the answer by design; the README's own pitch is that it delays the need for a cluster, not that it removes it.

## Vaex vs Polars: different bets on where data lives

The comparison people search for is Vaex vs Polars, and the difference is architectural rather than a matter of speed. Polars is a DataFrame library built around Arrow-backed in-memory execution with a query engine; you load data into a frame and the engine optimizes the plan. Vaex keeps the data on disk, memory-maps it, and streams columns when a computation needs them, which is why the README emphasizes opening files instantly and not copying on filter. Put differently: Polars asks how fast can we execute a query over data we hold, and Vaex asks how much data can we query without holding it. If your table fits in RAM and you want a modern query engine, the in-memory model is simpler and avoids the format conversion step. If your table does not fit and you cannot get a cluster, the memory-mapped model is the one that lets you keep working. The two are not mutually exclusive in a pipeline, but choosing between them for the primary exploration layer comes down to whether your data is already sitting in HDF5 or Arrow files on disk.

## Maintenance, packaging and licence costs

The repository is not archived, and its last push was on 2026-04-01. That is the only maintenance signal available here, and it is a commit date, not a release date; the newest release in the list is the 2018 paper-tagged version. Anyone planning a long-lived dependency should look at the CHANGELOG.md and RELEASE.md files in the repository, which exist but are not reproduced in the README. On upgrade cost, Vaex is split into many packages: setup.py enumerates vaex-core, vaex-viz, vaex-hdf5, vaex-server, vaex-astro, vaex-jupyter, vaex-ml, vaex-graphql, vaex-contrib and vaex. The meta-package installs each with the [all] extra, so an upgrade moves the whole set together rather than one library. The Makefile shows the project's own uninstall path is a pip uninstall driven by a grep over installed vaex packages, which tells you the maintainers expect the multi-package layout to be managed as a unit. The licence is MIT, which is permissive and places few obligations on redistribution; the repository also carries a licenses/ directory and a SECURITY.md, and the setup.py comment about not uploading the meta-package to PyPI is a packaging constraint rather than a licence one. This is a description of what the files say, not legal advice.

## Conclusion

Adopt Vaex when your data already sits in HDF5 or Arrow on disk and you want statistics, groupby aggregations and histograms without a cluster. Do not adopt it if your workflow depends on pandas idioms such as a materialized join result or heavy row-wise mutation; the README notes that joins do not copy the right table, which changes what you get back. Before committing, verify that your columns convert cleanly to HDF5 or Arrow, because the memory-mapping advantage only applies to those formats, and check the last push date against your own maintenance expectations.

## FAQ

### How do I install Vaex?

The README gives two commands: pip install vaex, or conda install -c conda-forge vaex. For a source checkout, setup.py documents pip installing the master archive zip, but it labels that meta-package as private and not for upload to PyPI.

### What file formats does Vaex open fastest?

The README names HDF5 and Apache Arrow, and says opening huge files is instant because of memory mapping. CSV, pandas DataFrames and other sources are supported through a conversion step documented separately, and that conversion is what gets you into the memory-mapped formats.

### Does Vaex copy my data when I filter or join?

According to the README, filtering and evaluating expressions keep the data untouched on disk and stream it only when needed. It also states that joins do not materialize the right table, which saves memory but means the joined result is not a conventional copied table.

## Sources

- [License: MIT](https://github.com/vaexio/vaex/blob/master/LICENSE)
- [Project website](https://vaex.io)
- [README](https://github.com/vaexio/vaex/blob/master/README.md)
- [Releases](https://github.com/vaexio/vaex/releases)
- [vaexio/vaex on GitHub](https://github.com/vaexio/vaex)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vaexio-vaex
