# google/brotli: the stream format, the bindings, and when to pick it over gzip

> Brotli is a generic-purpose lossless compressor with a formal RFC, MIT licensing, and first-party bindings in C, Python, Java, Go, C# and JavaScript. Its stream format carries no checksum and no original length, which shapes how you should deploy it.

**google/brotli** — Brotli compression format

- Repository: https://github.com/google/brotli
- Stars: 14,894 · Forks: 1,362
- Language: TypeScript
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-brotli

## What brotli compresses, and who ends up using it

Brotli is a general-purpose lossless compressor. The README describes it as combining a modern variant of LZ77, Huffman coding and second order context modeling, and says it is similar in speed to deflate while producing denser output. That sentence is the whole pitch, and it is worth reading literally: the target is the same slot deflate occupies, not the slot a specialist codec occupies.

The people who end up using it are usually not choosing a codec for its own sake. They are shipping a web response, a static asset bundle, a package archive, or an on-disk blob, and they want fewer bytes on the wire or on disk without paying a large CPU penalty. The repository reflects that: alongside the C implementation there are csharp/, go/, java/, js/ and python/ directories, plus a research/ directory for work that has not landed in the format. If you are writing a service in any of those languages, there is a first-party binding rather than a wrapper you have to trust.

The specification is RFC 7932, which matters more than it sounds. A format with an RFC can be reimplemented independently, and the README points at exactly that: an independent decoder by Mark Adler based entirely on the format specification, and a JavaScript decoder port. You are not locked to Google's code to read your own data.

## The stream format has no checksum and no length field

This is the design decision that should drive your deployment, and it is stated plainly in the README: brotli is a stream format, it does not contain meta-information such as checksums or uncompressed data length, and it is possible to modify raw ranges of the compressed stream without the decoder noticing.

That is not a bug report. It is a boundary. Brotli compresses bytes and gives you bytes back. It does not tell you whether those bytes are the ones that were compressed, and it does not tell you how large the result will be before you finish decoding. If your pipeline currently relies on a container format to detect corruption, brotli will not replace that container. You keep the checksum at the layer above: TLS, a package manager's digest, a filesystem with checksums, or your own hash alongside the payload.

The missing length field has a second consequence that bites in practice. You cannot pre-allocate an output buffer from the header. Decoders must grow their output incrementally, and any API that wants a fixed-size destination buffer has to be told the size by some other channel. This is why brotli shows up inside formats that already carry framing rather than as a standalone archive format.

There is a third consequence, less obvious. Because the format is a stream and not a container, there is no place to record the encoder settings, the original filename, or a dictionary identifier. Reconstructing those means recording them yourself.

## Installing brotli and compressing a file from the shell

The README gives distribution packages as the normal route. On Debian-based systems the package is called brotli, and on macOS Homebrew provides the same name.

```bash
apt install brotli
```

```bash
brew install brotli
```

That installs both the library and a command-line tool. The README does not document the tool's flags, so the first real use is to compress a file and check the size against the original, then reverse the operation and confirm the bytes come back unchanged. Because the format itself carries no checksum, verify the round trip with a hash rather than by eyeballing the output. If the result does not reproduce the original byte for byte, stop there.

## Python, and why the binding installs differently from the CLI

The Python module is published separately. The README gives the release install and the tip-of-the-tree install, and the two are not equivalent: one tracks a published version, the other builds from the repository.

```bash
pip install brotli
```

```bash
pip install --upgrade git+https://github.com/google/brotli
```

The second command needs a working compiler and the build dependencies declared in pyproject.toml, which lists setuptools and pkgconfig. That pkgconfig requirement is the detail people miss on slim containers and on Windows: the extension is built against the C library, so a bare Python image without build tooling will fail at install time rather than at import time.

There is a second, easily confused package. The README's related projects section mentions a JavaScript port of the brotli decoder installable via npm install brotli. That is a decoder port, not the C library, and it is a different artifact from the Python module despite the shared name. The repository's own version string comes from c/common/version.h, read by setup.py, so a locally built module reports the version of the C sources it was compiled against.

## Building from source with CMake or vcpkg

If your distribution's package is too old, the README documents two build routes. The CMake route is the one most projects embed.

```bash
mkdir out && cd out
cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_INSTALL_PREFIX=./installed ..
cmake --build . --config Release --target install
```

Note the install prefix: everything lands under ./installed relative to the build directory, not in /usr/local. That is convenient for vendoring and inconvenient if you forget it and then wonder why the linker cannot find the library.

The vcpkg route bootstraps the package manager and then installs the brotli port.

```bash
git clone https://github.com/Microsoft/vcpkg.git
cd vcpkg
./bootstrap-vcpkg.sh
./vcpkg integrate install
./vcpkg install brotli
```

The README notes that the vcpkg port is maintained by Microsoft team members and community contributors, and that if the version is out of date you should file an issue or pull request on the vcpkg repository rather than on brotli. That is a useful division of responsibility to know before you open a bug in the wrong place. Bazel is also listed, with a pointer to the Bazel site and no further instructions in the README.

## Where brotli is the wrong tool

The README's own warning is the first limitation: no checksum, no length, and raw ranges can be modified undetectably. If you need a self-verifying archive, brotli alone is not it. Wrap it, or use a format that already frames and verifies its contents.

The second limitation is the codec's own terms of comparison. The README says brotli is similar in speed to deflate with denser compression. Similar, not faster. If your bottleneck is CPU rather than bandwidth, brotli does not buy you throughput; it buys you bytes at roughly the same cost. Higher quality settings push further in the direction of size at the expense of time, and the README does not publish a table of settings, so any quality-level choice is yours to measure rather than one the project hands you.

The third is that brotli is a general-purpose compressor. It is not a specialist. The README's benchmark links point at general compression benchmarks, not at domain-specific ones. For data that is already compressed, for very small payloads where framing overhead dominates, or for workloads where a different codec's decompression speed is the deciding factor, the general-purpose framing is a mismatch. The README does not claim otherwise, and it is worth resisting the urge to treat "denser than deflate" as a universal win.

The fourth is operational: because the format is a stream, there is no version negotiation inside the data. Sender and receiver must agree out of band on the fact that brotli is in use, which is exactly what content-encoding negotiation does at the HTTP layer and what a file extension does on disk.

## How brotli differs from zstd and from gzip

The honest comparison is with zstd, not gzip. Gzip is deflate plus a container, and the README's claim is that brotli beats it on density at comparable speed. That is a like-for-like comparison of codecs, and it is the one brotli was designed to win.

Zstd is a different design centre. Where brotli's README emphasises a stream format deliberately stripped of meta-information, zstd's ecosystem is built around framed formats that carry content size, checksums and dictionary identifiers. If your requirement is "the compressed artifact must be self-describing and self-verifying", the two projects are answering different questions, and the answer is not a density measurement. Brotli's README does not position the project against zstd at all; the benchmark links it offers are general compression benchmarks, and readers who want that comparison should go to those benchmarks rather than to this README.

The comparison with gzip is simpler and more actionable. Deflate is embedded in decades of infrastructure, so brotli usually arrives as an additional encoding rather than a replacement. You serve brotli where the client advertises support and fall back to deflate or identity otherwise. That fallback path is not optional, and it is the part of a brotli rollout that actually takes engineering time.

## Maintenance, licensing and what upgrading costs you

The repository is not archived, and the last push was on 2026-09-18, three days before this writing. The most recent release listed is v1.2.0 from 2025-10-27, preceded by v1.2.0rc2 on 2025-10-21. A release candidate followed by a final release within a week is the shape of a project that tags rather than one that ships continuously.

The licence is MIT, stated in the README and present as a LICENSE file at the repository root. MIT is permissive: it lets you embed the library in proprietary software provided you keep the notice. The README also points at SECURITY.md for vulnerability reporting and CONTRIBUTING.md for changes, which is where you should look before opening an issue.

The upgrade cost is dominated by one thing: the wire format is fixed by RFC 7932, so upgrading the library does not invalidate data you already compressed. Old streams stay readable. What can change is the API surface of the bindings. The Python packaging reads its version from c/common/version.h, so a binding upgrade is tied to a C source upgrade, and the tip-of-the-tree install command builds whatever is on master at that moment. Pinning to a released version rather than tracking master is the cheaper habit. The repository also carries an sbom.cdx.json, which is useful if your supply chain tooling wants a software bill of materials without you generating one.

## Conclusion

Adopt brotli when you control both ends of a transfer and want denser output than deflate at comparable speed, and when you can carry integrity checking at another layer. Do not adopt it where a stream must be self-describing or self-verifying, and do not adopt it for tiny payloads or for data that is already compressed. Before shipping, verify three things against your own corpus: the chosen quality level and window size, whether your runtime exposes the window parameter at all, and whether anything downstream assumes a checksum that brotli does not provide.

## FAQ

### What is brotli used for?

It is a general-purpose lossless compression format, described in the README as combining a modern variant of LZ77, Huffman coding and second order context modeling. It is used wherever you want denser output than deflate at comparable speed, such as compressing files, assets or transferred data.

### Is brotli better than gzip?

The README states that brotli is similar in speed to deflate but offers more dense compression. So the gain is size, not throughput, and the comparison is against the codec rather than against the gzip container that wraps deflate.

### How do I install brotli?

The README gives distribution packages as the usual route: apt install brotli on Debian-based systems and brew install brotli on macOS. It also documents building from source with CMake, installing through vcpkg, and pip install brotli for the Python module.

### How do I install brotli on Windows?

The README does not give a Windows-specific package. The documented routes that work there are the Python module via pip install brotli, the vcpkg port, or a build from source with CMake, and the README notes that the Dart framework it links to ships prebuilt binaries for Win, Linux and Mac.

### How do I install brotli on Ubuntu?

Ubuntu is Debian-based, and the README gives apt install brotli as the example for Debian-based distributions. That installs the library and the command-line tool together.

### Is there a brotli package for Python?

Yes. The README documents pip install brotli for the latest release and pip install --upgrade git+https://github.com/google/brotli for the tip-of-the-tree version. Note that the npm package named brotli mentioned in the related projects section is a separate JavaScript decoder port, not the same artifact.

## Sources

- [google/brotli on GitHub](https://github.com/google/brotli)
- [Issues](https://github.com/google/brotli/issues)
- [License: MIT](https://github.com/google/brotli/blob/master/LICENSE)
- [README](https://github.com/google/brotli/blob/master/README.md)
- [Releases](https://github.com/google/brotli/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-brotli
