Library / SDK
facebook/openzl avatar
facebook/openzl

OpenZL: compression that gets fast when you tell it what your data looks like

A novel take on lossless data compression

3,264 stars173 forksC++NOASSERTION

At a glance

What is it?
Meta's OpenZL inverts the usual compressor trade-off. Instead of one generic algorithm guessing at your format, you describe the data and OpenZL builds a codec for it, which pays off on large structured datasets and costs you a description to maintain.
Who is it for?
OpenZL is worth the extra work when your data has a shape you can describe and the volume is large enough that throughput matters, which is why Meta uses it on AI workloads in production. A generic compressor such as zstd is the better answer for blobs with no exploitable structure, and for a small pipeline the description you would have to write and maintain may cost more than the ratio it buys back.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The idea is a description instead of a guess

The framing in the README is direct: OpenZL delivers high compression ratios while preserving high speed, a combination the README says is out of reach for generic compressors. The mechanism for getting there is that OpenZL takes a description of your data and builds from it a compressor specialised to that format.

That is the whole design in one sentence, and it explains both the promise and the cost. Generic compressors work by discovering redundancy statistically at runtime, which means they pay a cost on every byte and cannot exploit something you already know. A description moves that knowledge ahead of time. Once the engine knows a record is four little-endian integers followed by a variable-length string, it can skip straight to field three instead of scanning, and it can encode the integers with a width chosen for their actual range rather than a general-purpose bit packer.

The scope is narrower than a general compression library. The README positions it for engineers dealing with large quantities of specialised datasets, naming AI workloads as the example, and requiring high speed for their processing pipelines. It consists of a core library plus tools to generate specialised compressors, all compatible with a single universal decompressor, which is the property that keeps storage-side code simple even when the compressor was built for one dataset.

Building the library and the zli tool with make

The compiler requirements are ordinary C11 and C++17, with cmake 3.20.2 or newer when you use the cmake path. The default build is short:

sh
make

The `Makefile` shows what that produces. Its first recipe is the default one and it builds a target called `zli`, the command line tool, and the common repository-wide definitions come from `build-scripts/make/zldefs.make`, with macros for generating targets in `build-scripts/make/multiconf.make`. So the default target is the tool, not a bare library.

Parallelism is automatic: the build is multi-threaded by default, detecting the local core count, and you can override it with the standard make flag. Build type is selected with a variable:

sh
make lib BUILD_TYPE=DEV

`BUILD_TYPE=DEV` gives a debug build with asserts and ASAN and UBSAN enabled, while `BUILD_TYPE=OPT`, the default, is optimised with asserts disabled. The Makefile has an `ADAPTIVE` mode as well, where each target sets its own flags, which is a nice piece of engineering: production binaries and the benchmark get `-g0 -O3` and `-DNDEBUG`, while `gtests` keeps debug symbols and gains `-DZL_ENABLE_ASSERT`. The full list is available through `make help` and `make show-config`.

Dependencies are vendored per platform, including on macOS

The Makefile's dependency block is more informative than it looks, because it shows what OpenZL builds on: zstd, lz4, xgboost, and googletest for the test targets. Those live under `deps/` and the Makefile branches on platform to pick the right artifact name, with `.dll` paths for Windows, `.dylib` paths for Darwin and `.so` otherwise. The same block also defines the static variants, `libopenzl.a`, `libopenzl.so`, `libgtest.a`, `libzstd.a`, `liblz4.a` and `libxgboost.a`, so you can link either way.

The xgboost dependency is the interesting one, since it is not a compression library at all. Its presence says the codec selection step uses learned prediction rather than hand-tuned heuristics alone, which fits the project's pitch of building an optimised compressor from a description of the data.

Vendoring also means the `.gitmodules` file matters. Submodules are checked into the repository rather than resolved from a package manager, so a shallow clone without submodule initialisation will not build, and the pinned versions of zstd and lz4 are whatever the repository was tested against rather than whatever your distribution ships. That is a deliberate choice for a project whose whole value is bit-exact reproducibility, and it is also a supply-chain surface to look at before you vendor it internally.

The cmake path with tests, benchmarks and named build modes

cmake is supported alongside make, and the documented sequence builds everything and runs the tests:

sh
mkdir build
cd build
cmake -DCMAKE_BUILD_TYPE=Release -DOPENZL_BUILD_TESTS=ON ..
make -j
make -j test

The `OPENZL_BUILD_TESTS=ON` flag pulls in the testing dependencies and builds the unit and integration tests, with `OPENZL_BUILD_BENCHMARKS=ON` doing the same for benchmarking. There is a `benchmark/` directory at the top level for that.

What is unusual is `OPENZL_BUILD_MODE`, a set of predefined modes rather than just the cmake default. `none` is the default and defers to `CMAKE_BUILD_TYPE`. The rest are `dev`, `dev-nosan`, `opt`, `opt-asan`, `dbgo` and `dbgo-asan`, where the last two are the useful middle ground: optimised builds that keep asserts enabled. The README attaches a caution to switching modes, because a stale cache silently keeps the old configuration; `cmake --fresh -DOPENZL_BUILD_MODE=dev-nosan ..` is the documented way out.

The variable list is long but well separated by scope. Compiler and flag variables apply to OpenZL and its dependencies, while the `OPENZL_C_COMPILE_OPTIONS`, `OPENZL_CXX_COMPILE_DEFINITIONS` family applies to OpenZL only. Two that are easy to miss: `OPENZL_SANITIZE_ADDRESS=ON` sanitises OpenZL without sanitising its dependencies, which avoids a long tail of reports from vendored code, and `OPENZL_COMMON_FLAGS` passes extra flags to every target.

Windows needs clang-cl, and there is a script to tell you why

The Windows story has a clear recommendation and a clear reason. OpenZL uses modern C11 features that MSVC does not fully support, so the README recommends `clang-cl` for full compatibility, with MinGW-w64 as the GNU-toolchain alternative. MSVC itself is listed as limited support, with a specific symptom: C2099 errors.

The recommended Windows configure line is:

cmd
cmake -S . -B build -DCMAKE_C_COMPILER=clang-cl -DCMAKE_CXX_COMPILER=clang-cl
cmake --build build --config Release

Rather than making you guess, the repository ships a detection script for both shells:

cmd
./build-scripts/cmake/detect_windows_compiler.ps1
./build-scripts/cmake/detect_windows_compiler.bat

That is the pattern to notice about the build documentation generally. Wherever there is a portability decision, there is a script or a matrix next to it. For editor support the README asks for the `cmake-tools` and `clangd` extensions and, importantly, a generated `compile_commands.json`, which it gives a fallback command for when the CMake Tools extension will not cooperate. It also lists when to regenerate that file: after cloning, after adding or removing source files, and after modifying `CMakeLists.txt`.

SDDL2 is where the real leverage sits

The v0.2.0 release notes explain what SDDL is, since the acronym is never expanded in the README and it turns out to be the heart of the project. SDDL has been rebuilt as a real compiler rather than the thin runtime of the original demo: a parser feeds a semantic analyzer, the analyzer hands a typed AST to an optimizer, and the optimizer drives a code generator that emits VM bytecode.

The payoff is called instant-parse. When a record's layout can be fully determined from parameters and constants alone, the engine jumps directly to any field without scanning the preceding bytes, which the notes connect to zero-copy access and multi-GB/s throughput. That is the concrete mechanism behind the speed claim, and it only works because the layout was described rather than guessed.

The language grew alongside the toolchain: `when` blocks for conditional layouts, parameterised and anonymous records, member access on record fields, and bitwise and logical operators. The new semantic analysis phase catches undefined references, type mismatches and other errors before the codec runs.

The other v0.2.0 changes are worth naming too, because they mark where the project is going: a native LZ codec became the default for serial inputs, very large inputs are chunked automatically, and the graph visualizer was improved. A visualizer for codec graphs tells you something about the intended audience, which is people tuning a format description rather than people calling a compression API once.

On maturity, the README is unusually candid in both directions. It says the project is under active development and that the API, the compressed format, and the set of codecs and graphs are all subject to change. It also commits to two stability properties: payloads compressed with any release-tagged version stay decompressible by new releases for several years, and new releases can produce frames compatible with at least the previous release. Commits on the `dev` branch, which is the default branch here, carry no guarantees at all. The repo is not archived, the last push was on 2026-09-23, and v0.2.0 shipped on 2026-05-07 after v0.1.0 in October 2025. There is also a v0.0.23 tag that exists purely to reproduce the numbers in the whitepaper and blog post.

Editorial conclusion

OpenZL is worth the extra work when your data has a shape you can describe and the volume is large enough that throughput matters, which is why Meta uses it on AI workloads in production. A generic compressor such as zstd is the better answer for blobs with no exploitable structure, and for a small pipeline the description you would have to write and maintain may cost more than the ratio it buys back. Start with the quickstart guide and the `examples/` directory rather than the codec internals, build with `make` to get the `zli` tool and the library in one step, and pin a release tag rather than tracking `dev`, since the API and the frame format are both explicitly still moving.

Frequently asked questions

What is OpenZL and how is it different from a generic compressor?

OpenZL is a compression framework where you describe the format of your data and it builds a specialised compressor from that description, rather than having one general algorithm infer the structure at runtime. The result is a codec tuned to a specific format that still decompresses with a single universal decompressor.

How do you build OpenZL from source?

Run `make` in the repository root, which builds the `zli` tool and the core library using vendored dependencies under `deps/`. A cmake path is also documented, where `OPENZL_BUILD_TESTS=ON` adds the unit and integration tests. You need a compiler with C11 and C++17 support.

Does OpenZL build on Windows?

Yes, with a caveat. Because OpenZL uses modern C11 features, the README recommends `clang-cl`, with MinGW-w64 as an alternative, and lists MSVC as limited support because of C2099 errors. Detection scripts are included under `build-scripts/cmake/` for both PowerShell and Command Prompt.

Is the OpenZL format and API stable across releases?

The README describes the project as under active development with the API, compressed format, and codecs subject to change. It does commit to forward compatibility: data compressed by any release-tagged version stays decompressible by later releases for several years. Commits on the `dev` branch carry no guarantees.

What is SDDL in OpenZL?

SDDL is the schema description language used to describe your data layout, and as of v0.2.0 it is a real compiler: parser, semantic analyzer, typed AST, optimizer, and a code generator emitting VM bytecode. Its instant-parse mode lets the engine jump to a field without scanning preceding bytes when the layout is known from constants.

Official sources

  1. facebook/openzl on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebook-openzl.svg)](https://hysenlabs.com/projects/facebook-openzl)