Nimble: Meta's Columnar File Format Built for Wide Machine Learning Tables
New and extensible file format for storage of large columnar datasets.
At a glance
- What is it?
- Nimble is a C++ columnar file format from Meta designed to handle tables with thousands of columns, as found in feature engineering and model training workloads. It is a declared replacement for Apache Parquet and ORC, but it ships without versioning guarantees and is still under active development.
- Who is it for?
- Nimble suits teams at Meta's scale that need a columnar format optimized for feature stores and training tables with thousands of columns, where Parquet's metadata overhead becomes a real cost. Anyone outside that use case should weigh the lack of versioning guarantees carefully: the README states stability will be provided in a future stable release, not today.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem Nimble Solves: Metadata Overhead on Wide Tables
Parquet and ORC were designed when a few hundred columns was considered wide. Feature engineering tables and ML training datasets at large scale routinely have thousands of columns, sometimes tens of thousands. Parquet stores row-group metadata in a way that scales poorly with column count, and accessing a small slice of a wide table forces the reader to parse a disproportionately large metadata section.
Nimble addresses this by redesigning metadata organization specifically for the wide-table case. It uses FlatBuffers rather than Thrift or Protobuf to serialize metadata, which allows random access into large metadata sections without deserializing the entire structure. It also uses block encoding instead of stream encoding, which gives predictable memory usage during reads: the reader knows how much memory a block requires before it begins decoding, rather than discovering it mid-stream.
Architecture: Decoupled Encoding and Extensibility
Nimble's central architectural choice is to separate stream encoding from the underlying physical layout. In Parquet, the encoding scheme is tightly bound to the column type and the storage format. Nimble allows encodings to be extended by library users and applied recursively, which the README calls cascading or composite encoding.
A column can pass through one encoding to reduce its range, then a second encoding to exploit the reduced range. Because encodings are pluggable, teams can introduce custom encodings for their own data distributions without forking the format itself. The README explicitly discourages re-implementing the format specification, calling fragmentation from competing implementations a problem it observed in similar past projects. The design intent is a single unified library with language bindings.
The README also describes planned SIMD and GPU-friendly encoding paths for highly parallel hardware. That work is listed as not yet implemented, along with metadata structures to help decoders schedule kernel work without accessing data streams directly. These are stated design goals, not shipped features.
Building Nimble from Source
Nimble ships a self-sufficient CMake build system. The basic build sequence is:
git clone [email protected]:facebookincubator/nimble.git
cd nimble
makeThis clones the repository and builds a release configuration. The build system will either locate or compile its main dependencies: gtest, glog, Folly, Abseil, and Velox. Velox is included as a Git submodule and is a significant dependency. On Ubuntu 22.04 the following system packages must be present before the build will succeed:
sudo apt install -y \
git \
cmake \
flatbuffers-compiler \
protobuf-compiler \
libflatbuffers-dev \
libgflags-dev \
libunwind-dev \
libgoogle-glog-dev \
libdouble-conversion-dev \
libevent-dev \
liblz4-dev \
liblzo2-dev \
libelf-dev \
libdwarf-dev \
libsnappy-dev \
libssl-dev \
bison \
flex \
libfl-dev \
pkg-config \
clang \
clang-formatThe README states that builds have been tested with clang 15 and 16. To force a dependency to compile locally rather than using the system version, prepend the source variable to the make invocation:
folly_SOURCE=BUNDLED makeThis pattern applies to other dependencies in the build system.
The Velox Dependency: Coupling and Maintenance
Nimble's codebase is tightly coupled with Velox, Meta's open-source C++ execution engine for data processing. Velox is included as a Git submodule pointing to a specific commit, and the README tracks how far behind the submodule is from Velox's main branch with a badge. Velox itself has significant system dependencies and a long compilation time, so the first build on a fresh machine will take time.
This coupling is acknowledged as temporary. The README states that decoupling is planned but has not happened. For teams that already use Velox, this is a natural fit. For teams without a Velox dependency, adopting Nimble means adopting Velox as well, which is a large C++ codebase with its own system requirements.
Advancing the Velox submodule when changes in Nimble depend on newer Velox code requires a specific workflow:
git -C velox checkout main
git -C velox pull
git add veloxAfter updating, tests must pass and the change goes through a pull request. A pre-commit hook automatically updates the Velox comparison link in the README.
Stability Status and Versioning Guarantees
The README is direct on this point: Nimble does not provide stability or versioning guarantees. File formats written today may not be readable by future versions of the library. The README states that stability guarantees will be provided with a future stable release, but gives no timeline.
This matters for any production use case that involves persisting data for later reads. A columnar format is typically chosen because the data will outlive the writing application. Using Nimble in that role today means accepting the possibility of a migration step when the stable release arrives. Teams that need a stable, versioned format for long-lived data should treat Nimble as a technology preview rather than a production dependency.
The last push to the repository was on 2026-09-21, which shows the project is under ongoing development. There are no GitHub releases published; the build system is the only entry point. The README notes that many of the described features are still under design or active development, and it asks users to proceed at their own risk.
Comparison with Apache Parquet
Parquet is the dominant open columnar format for analytical workloads and has stable support across the entire Hadoop and Arrow ecosystem, including Spark, Trino, DuckDB, and virtually every cloud data warehouse. Its row-group and column-chunk structure is well understood, and multiple independent reader implementations exist in Python (pyarrow), Java, C++, Go, and Rust. The format has shipped stable versioning and compatibility guarantees for years.
Nimble deliberately takes the opposite position on fragmentation: it discourages independent format implementations and instead asks developers to use the C++ library and write language bindings. This means Nimble's ecosystem today is a single implementation, without the broad language support Parquet has accumulated. For a team evaluating whether to build a new pipeline around Nimble, the absence of Python, Java, or Rust readers is a real constraint that the README does not address.
ORC, another alternative, is similarly mature within the Hive ecosystem. Neither Parquet nor ORC was designed with the wide-table use case as a primary concern. Nimble's stated advantage is precisely that design focus: lighter metadata organization for thousands of columns, FlatBuffers for faster metadata access, and pluggable cascading encodings. Those advantages are specifically relevant to feature stores and training pipelines that already operate at Meta's scale, not general-purpose analytical workloads where Parquet's ecosystem breadth matters more than metadata overhead.
Editorial conclusion
Nimble suits teams at Meta's scale that need a columnar format optimized for feature stores and training tables with thousands of columns, where Parquet's metadata overhead becomes a real cost. Anyone outside that use case should weigh the lack of versioning guarantees carefully: the README states stability will be provided in a future stable release, not today. Before adopting Nimble, verify that your build environment can compile against clang 15 or 16 and that you can accept the tight coupling with the Velox submodule, which the README acknowledges will be addressed in future work.
Frequently asked questions
Does Nimble provide any versioning or stability guarantees?
The README states explicitly that Nimble does not yet provide stability or versioning guarantees, and that they will be provided with a future stable release. Using it today means accepting potential breaking changes.
What compilers are supported for building Nimble?
The README states that Nimble builds have been tested with clang 15 and clang 16. Other compilers are not mentioned.
Can Nimble files be read by tools that support Apache Parquet?
Nimble is a separate format from Parquet and is not compatible with Parquet readers. It is positioned as a replacement, not a superset, and requires the Nimble C++ library.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookincubator-nimble)