# Velox: a C++ execution engine library for building your own query engine

> Velox is not a database and not a SQL engine. It is the vectorized execution layer you embed when you already have a query plan and need columnar operators, functions and memory management underneath it.

**facebookincubator/velox** — A composable and fully extensible C++ execution engine library for data management systems.

- Repository: https://github.com/facebookincubator/velox
- Website: https://velox-lib.io/
- Stars: 4,217 · Forks: 1,616
- Language: C++
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookincubator-velox

## What Velox actually is, and the gap it fills

Velox is a C++ library, not a service. The README describes it as a composable execution engine that supplies reusable components for data management systems, and it names batch, interactive, stream processing and AI/ML as the workloads it targets. The important constraint appears in the same paragraph: Velox does not ship a SQL parser, a dataframe layer, or a query optimizer, and the documentation says it is usually not meant to be used directly by end users. It takes a fully optimized query plan as input and performs the computation described by that plan.

That places the library at a specific layer. Everything above execution (parsing, binding, planning, cost-based optimization) is your problem. Everything below it (columnar memory layout, vectorized expression evaluation, relational operators, connectors, serializers, memory arenas, task and driver scheduling, spilling, caching) is what Velox offers. If you are building an engine and have already solved the planning side, this is the part you would otherwise spend years writing. If you have not solved the planning side, Velox does not shorten that work at all.

The governance section is worth reading before you commit to the dependency. The README states Velox was created by Meta and is developed in partnership with IBM/Ahana, Intel, Voltron Data, Microsoft, ByteDance and other companies, with maintainers listed on the project site and technical governance described in a separate document. A library that sits inside other companies' engines has to keep an extension surface stable, and the README lists eight extension points: custom types, simple and vectorized functions, aggregate functions, window functions, operators, file formats, storage adapters and network serializers.

## The component stack: vectors, expression eval, operators, I/O

The Type system supports scalar, complex and nested types including structs, maps and arrays. On top of it sits Vector, described in the README as an Arrow-compatible columnar memory layout with Flat, Dictionary, Constant and Sequence/RLE encodings, plus a lazy materialization pattern and support for out-of-order writes. Those encodings are the reason the rest of the stack can be vectorized: a dictionary-encoded column avoids materializing values until an operator actually needs them.

Expression Eval runs expressions over that Vector/Arrow data in a fully vectorized fashion. Functions are grouped as scalar, aggregate and window implementations that the README says follow Presto and Spark semantics. That semantic choice matters more than it sounds. If your engine exposes Presto or Spark SQL, function behaviour lines up with what your users expect. If you expose a different dialect, you are adopting someone else's null handling, coercion and string semantics, and reconciling the differences becomes your work.

Operators cover scans, writes, projections, filtering, grouping, ordering, shuffle/exchange, hash, merge and nested loop joins, and unnest. I/O is a connector interface for sources and sinks, with support for ORC/DWRF, Parquet and Nimble file formats and storage adapters for S3, HDFS, GCS, ABFS and local files. Network serializers are pluggable, with PrestoPage and Spark's UnsafeRow given as examples. Resource management covers memory arenas, buffer management, tasks, drivers, thread pools for CPU and thread execution, spilling and caching.

Read that list as a checklist against your own engine. The pieces you already have are duplicated effort; the pieces you lack are the reason to adopt. The README also points to velox/examples for integration patterns against the different component APIs, which is the honest starting point before you write anything.

## Installing Velox and building a first target

Velox is distributed as source. The README gives the clone as the first step, and everything after that is dependency installation and a CMake build.

```bash
git clone https://github.com/facebookincubator/velox.git
cd velox
```

The README then says the first step is to install dependencies, with details in CMake/resolve_dependency_modules/README.md, and notes that the project ships scripts to set up dependencies for a given platform. Two environment variables control where that work lands. DEPENDENCY_DIR sets where packages are downloaded and built, defaulting to deps-download in the current working directory. INSTALL_PREFIX sets where packages are installed, defaulting to deps-install on macOS and to the system location such as /usr/local on Linux. The README explicitly discourages using /usr/local on macOS because certain Homebrew versions use it, and suggests exporting INSTALL_PREFIX into your shell profile so later builds find the packages.

```bash
export INSTALL_PREFIX=/Users/$USER/velox/deps-install
```

BUILD_THREADS controls dependency build parallelism and overrides the default number of parallel compile and link processes. The README also notes that DEPENDENCY_DIR and INSTALL_PREFIX can be shared with Velox clients such as Prestissimo by pointing them at a common directory, which is the practical reason to set them deliberately rather than accept the defaults.

The top-level Makefile wraps the CMake configuration. Its variables include VELOX_BUILD_MINIMAL, VELOX_BUILD_TESTING, TREAT_WARNINGS_AS_ERRORS and ENABLE_WALL, and it passes them through as -DVELOX_BUILD_MINIMAL, -DVELOX_BUILD_TESTING, -DTREAT_WARNINGS_AS_ERRORS and -DENABLE_ALL_WARNINGS. The Makefile comments state that VELOX_BUILD_TESTING defaults to ON and VELOX_BUILD_MINIMAL defaults to OFF, and that setting minimal to ON restricts the build to a minimal set of components, possibly overriding other options.

```bash
make release
```

Expect the first build to be long and to fail on missing system packages before it succeeds. The compiler matrix is not optional. Minimum versions are gcc 11 and clang 15 on Linux and clang 15 on macOS. Recommended combinations are CentOS 9 or RHEL 9 with gcc 12, Ubuntu 22.04 with gcc 11, and macOS with clang 16. Alternatives listed include CentOS 9 or RHEL 9 with gcc 11, Ubuntu 20.04 with gcc 11, and Ubuntu 24.04 with clang 15. If your toolchain is older than the minimum, stop before you start.

There is also a container path. The docker-compose.yml file states it is used only for running services, with docker compose run adapters-cpp given as the invocation, and that building images requires docker bake instead, with target names defined in docker-bake.hcl. The compose file warns that using the compose target names for builds skips layer caching and tags images incorrectly for multi-arch, and says to do that for local use only. One service, ubuntu-cpp, sets VELOX_DEPENDENCY_SOURCE to BUNDLED to build dependencies from source, and the base block sets NUM_THREADS with a default of 8 and mounts the repository at /velox.

## Where Velox is the wrong tool

The clearest failure mode is adoption for the wrong layer. If you want to query Parquet files with SQL today, Velox gives you nothing to type. There is no parser, no optimizer and no end-user interface, and the README states this directly rather than leaving it to be discovered. Teams that pick Velox expecting a query engine will spend their first months building the planner they assumed was included.

The second constraint is the build. This is a large C++ library with a dependency graph you compile yourself, and the README's dependency documentation is a separate file for a reason. A pinned compiler matrix, a long first build, and environment variables that must be set consistently across every consumer of the library are the entry cost. If your team does not already build C++ at this scale, that cost is real and recurring.

Semantics are a third boundary. Functions follow Presto and Spark semantics, and the serializers named in the README are PrestoPage and Spark's UnsafeRow. If your system speaks neither dialect, you are either translating at the edges or accepting behaviour that differs from your own specification. Neither is fatal, but both are work that the component list does not do for you.

Finally, consider the extension surface as a maintenance commitment rather than a feature. Custom types, functions, operators, file formats, storage adapters and serializers are all things you can add, which also means all of them are things you now own. Any engine-specific specialization you write is code your team maintains against a library that evolves.

## Velox compared with embedding DuckDB or DataFusion

The natural alternative for many teams is an embedded engine that already includes the layers Velox omits. DuckDB and Apache DataFusion both ship a SQL front end and a planner alongside their execution machinery, so a team can go from files to results without writing a parser or a cost model. Velox deliberately does not compete there. The README's own framing is that Velox is used by developers integrating and optimizing their compute engines, which describes a different buyer.

The difference shows up in what you inherit. With an embedded engine you inherit a dialect, a catalog story and an optimization pipeline, and your control over plan shape is whatever the engine exposes. With Velox you inherit vectors, expression evaluation, operators, connectors, serializers and resource management, and you keep full control of the plan because you produced it. That is the trade: less to build above the plan, complete authority over the plan itself.

There is a middle position worth naming. Prestissimo is mentioned in the README as a Velox client that can share the same DEPENDENCY_DIR and INSTALL_PREFIX, which is a reminder that Velox already sits under at least one production query engine rather than being a greenfield bet. If your architecture resembles that one, the comparison is not Velox against an embedded database but Velox against writing the execution layer yourself.

## Licence, releases and the cost of staying current

Velox is licensed under Apache-2.0, with the licence text at LICENSE and a NOTICE.txt at the repository root. Apache-2.0 permits commercial use and modification and includes a patent grant, but it also carries attribution and notice obligations, and the NOTICE file exists for that purpose. Whether those obligations affect your distribution model is a question for your own counsel, not something a component list settles.

The repository is not archived. No release information was available for this article, and the README's Getting Started section describes building from source rather than consuming a versioned package, so treat the source tree as the artifact. The pyproject.toml defines a PyVelox project with scikit-build-core and setuptools_scm, requires Python 3.9 or newer, and depends on pyarrow, which indicates Python bindings and extensions exist alongside the C++ library. Nothing in the repository describes a published wheel or a stable ABI, so plan for source builds and pin whatever commit you build against.

Upgrade cost is dominated by the dependency graph and the extension points. Every custom type, function, operator, file format, storage adapter and serializer you add is code that must keep compiling as the library moves. The recommended compiler combinations are the ones the project tests, and drifting outside them shifts build breakage onto your team. The project's own communication channels are the Slack workspace, GitHub Issues and Discussions, with access to Slack described as requiring a comment on a specific discussion, so support is community-shaped rather than a vendor contract.

## Conclusion

Adopt Velox if you are writing a C++ engine and already own a parser, planner and optimizer, because the library starts at the optimized plan. Do not adopt it if you want a SQL endpoint or a dataframe API, since the README states it provides neither. Before committing, verify the compiler matrix for your platform, confirm your dependency build works with the DEPENDENCY_DIR and INSTALL_PREFIX layout, and check which connectors and file formats your workload needs against the I/O component list.

## FAQ

### What is Velox used for?

Velox is a C++ execution engine library used by developers to build data management systems for batch, interactive, stream processing and AI/ML workloads. It takes an already optimized query plan and performs the computation, providing vectors, expression evaluation, functions, operators, I/O connectors and resource management.

### How to install Velox?

Clone the repository with git clone https://github.com/facebookincubator/velox.git, then install dependencies as described in CMake/resolve_dependency_modules/README.md. The README notes setup scripts that use DEPENDENCY_DIR and INSTALL_PREFIX, and the top-level Makefile drives the CMake build.

### How to use Velox?

The README states Velox takes a fully optimized query plan as input and performs the described computation, and that it is usually not meant to be used directly by end users because it has no SQL parser, dataframe layer or query optimizer. Integration examples against the component APIs are in velox/examples.

### What is Velox?

Velox is a composable execution engine distributed as an open source C++ library, created by Meta and developed in partnership with IBM/Ahana, Intel, Voltron Data, Microsoft, ByteDance and other companies. It provides reusable components for building data management systems rather than a finished query engine.

## Sources

- [Official documentation](https://velox-lib.io/)
- [Official README](https://github.com/facebookincubator/velox#readme)
- [Project repository](https://github.com/facebookincubator/velox)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookincubator-velox
