# elfshaker: 400 GiB of object files in 100 MiB, if you compile them the right way

> elfshaker packs directories of pre-link object files into compressed archives with on-demand extraction, and the README attributes a 400 GiB to 100 MiB result to compiling with -ffunction-sections so unchanged sections stay byte-identical. It is Apache-2.0 from Arm engineers, has one release from November 2021, and tests CI on a single Ubuntu.

**elfshaker/elfshaker** — elfshaker stores binary objects efficiently

- Repository: https://github.com/elfshaker/elfshaker
- Stars: 2,340 · Forks: 50
- Language: Rust
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/elfshaker-elfshaker

## A version control system for build output

elfshaker describes itself as a low-footprint, high-performance version control system fine-tuned for binaries, and it is a CLI tool written in Rust. That framing is unusual and accurate. It is not a git replacement and not a remote store. It takes snapshots of a directory, compresses them into pack files, and hands individual files back on demand.

The headline number is the pitch: 400 GiB down to 100 MiB with 1s access time, when applied to clang builds. The linked manyclangs project holds prebuilt LLVM toolchains, and the README says elfshaker allows few-second access to any commit of clang from it. Bisection over LLVM is accelerated by a factor of 60x according to the README, with each pack holding about 1,800 builds at roughly 100 MiB, and extracting a single build taking 2-4s on modern hardware.

The audience is narrow and worth stating plainly: anyone with a large archive of completed builds they need to reach into. The canonical case is bisection, where you need to check out revision N of a compiler, build something with it, and see whether the result changed. If you have to rebuild from source every time, bisection is bounded by compilation. If you can extract the artifact, it is bounded by an unpack.

The original authors are Peter Waller and Vesko Karaganev, both at Arm Limited, and the copyright line in the manifest reads 2021 Arm Limited or its affiliates and Contributors. That tells you the origin, and it also tells you the funding question, which the project does not address.

## Where the compression actually comes from

The README has a section headed Applicability, phrased as how on earth do you get such a phenomenal result, and the answer is that the numbers only hold for a specific kind of input.

It works particularly well for the presented use case because storing pre-link object files has three properties: there are many files, most of them do not change very often so there are a lot of duplicate files, and when they do change the deltas of the binaries are not huge.

The first two are what a deduplicating pack format exploits. The third is the one you have to engineer for, and the README explains how: compile object code with the `-ffunction-sections` and `-fdata-sections` flags. The effect is that inserting a function into a translation unit does not cause all of the addresses to change across the whole object file.

That is the crux. Without those flags, an insertion shifts everything after it, so the bytes of an object file change wholesale even though most of its content is identical. With them, the object file is divided into sections that are addressed independently, so a change to one function leaves the others byte-identical and therefore compressible against each other.

Which means the headline ratio is not a property of elfshaker alone. It is a property of elfshaker plus a compiler flag choice. A team that stores object files compiled without those flags will see a fraction of this, and the README is honest that the applicability section is what qualifies the claim.

## Linked executables are the bad case, and the README says why

The same section explains the failure mode, and it is the most useful paragraph in the README for anyone deciding whether this applies to them.

If you looked at the binary delta on the linked executable from such a change, it will be large, because all of the absolute addresses after the insertion will change, and references to those addresses will change. These address changes are not handled well by compression algorithms, resulting in a poor compression ratio.

The magnitude is given. Compressing many revisions of clang executables together gives a ratio of something like 20 percent, which the README calls pretty good. elfshaker achieves something closer to 0.01 percent, or 10,000x, amortized across many builds.

So the honest reading of that headline is: about 100x better than compressing the linked binaries, which is itself already better than nothing. The 10,000x is relative to the uncompressed size of the object files, not relative to a naive compressor on the same files.

The design conclusion follows directly. elfshaker is a tool for storing pre-link object files built with section-level addressability, and it is a poor choice for storing linked executables, shared libraries or any artifact whose layout depends on absolute addresses. That is not a bug in the tool; compression of address-dense binaries is a genuinely hard problem and no pack format solves it.

## store, pack, extract

The quickstart is three commands, and the second one is where the compression happens.

`elfshaker store <snapshot>` captures the state of the current working directory into a named snapshot. `elfshaker pack <pack name>` captures all loose snapshots into a single pack file, and the README is explicit that this is what gets you the compression win. `elfshaker extract <snapshot>` restores the state of a previous snapshot into the current working directory.

That staging is the whole design. A snapshot on its own is a loose directory of files, useful but not compressed against its neighbours. Packing is what lets similar files across many snapshots deduplicate and compress together, which is why the README insists on the second step.

The README also tells you to consult the installation and usage guides first and to make sure you know what you are doing, which is an unusual thing to say about a three-command tool. That caution is about storage semantics rather than difficulty: snapshots are local, packs are the durable artefact, and a workflow that stores and extracts without ever packing gets none of the benefit.

There is a workflow document with more detail, and a talk from the 2021 LLVM Developers' Meeting recorded alongside the quickstart.

## One release in 2021, and a manifest that says 1.0.0-rc1

The release list contains a single entry: v0.9.0, published on 2021-11-18.

The manifest in the repository root says something different.

```toml
[package]
name = "elfshaker"
version = "1.0.0-rc1"
edition = "2018"
```

So the source tree has been at a 1.0 release candidate for a long time and that release candidate has never been tagged or published. The last push to the repository was on 2026-07-08, so commits are still landing, and the project is not archived.

That combination needs interpreting rather than guessing. It is consistent with a tool that works, has a stable format, and is maintained opportunistically by people whose day job is not this repository. It is also consistent with a project whose user base is small enough that nobody has felt the need for a release.

Either way, the practical advice is the same. If you adopt it, build from source or pin to v0.9.0 and test, and do not expect `cargo install elfshaker` to give you the current tree. The manifest's `edition = "2018"` is another signal: that is two editions behind current practice, so this code has not been touched with a modern toolchain.

The project does state a stability position. The file format and directory structure is stable, and the intention is that pack files created with the current version remain compatible with future versions. It pairs that with an unusually direct request to kick the tyres and do your own validation, and says there may be things the maintainers have not needed yet.

## The dependency list is the file format

There is no separate specification document for the pack format, so the dependency list is the closest thing to a description of it.

```toml
[dependencies]
zstd = { version = "0.13.3", features = ["zstdmt"] }
crossbeam-utils = "0.8"
walkdir = "2.5.0"
clap = { version = "4", features = ["color", "env", "error-context", "help", "usage"] }
sha1 = "0.10"
hashbrown = "0.15.3"
rmp-serde = "0.15.5"
fs2 = "0.4.3"
winapi = "0.3.9"
```

Compression is zstd with the `zstdmt` feature, which enables the multi-threaded implementation. That is consistent with a tool whose selling point is throughput over very large archives, and it means the compressor uses every core you give it.

Integrity and indexing come from `sha1` for content hashing, `hashbrown` for the hash table mapping digests to stored entries, and `rmp-serde` for MessagePack serialisation, which is a compact binary format well suited to a pack index.

`walkdir` enumerates the tree being stored and `fs2` provides file locking, which is what stops two processes writing the same pack. `winapi` at 0.3.9 is the Windows path, and `libc` covers the Unix equivalents.

Two things stand out. `ureq` at 2.12.1 is an HTTP client, so something in this tool talks to a server, and the README does not describe what. And `rand` at 0.9.1 alongside `chrono` suggests the metadata layer carries timestamps and possibly identifiers.

Around the Rust source, the repository carries Nix packaging in `flake.nix`, `flake.lock` and `elfshaker.nix`, a `rustfmt.toml`, a `contrib/` directory, and `test-scripts/`.

## CI on one distribution, and where it does not apply

The compatibility section is short and worth reading closely.

The platforms used for CI tests are listed as Ubuntu 20.04 LTS, singular. Then a distinction: the aim is to support all popular Linux platforms, macOS and Windows in production.

So continuous integration covers one distribution on one release, and everything else is a stated intention rather than a tested claim. The officially supported architectures are AArch64 and x86-64, which is the narrower and more trustworthy of the two statements.

There is a further asymmetry in the wording. Production support is claimed as an aim, and architecture support is claimed as official. Neither is the same as a compatibility matrix, and the README does not provide one. For a tool that writes a persistent archive you may want to keep for years, that gap matters more than it would for a stateless CLI.

Contact is through Gitter at the elfshaker community, and there is a `SECURITY.md`. One more detail is telling about how far the project got: the README contains a commented-out TODO pointing at an API reference at elfshaker.github.io/docs, so a generated documentation site was planned and not built.

## Where elfshaker is the wrong tool, and a compilation cache

Three cases where it does not apply, all of them from the README's own reasoning.

Linked executables, shared libraries, and any artifact with absolute addresses. The README quantifies this as a compression ratio around 20 percent rather than 0.01 percent, and explains why: address changes defeat compression.

Object files compiled without `-ffunction-sections` and `-fdata-sections`. You get duplicates that are not byte-identical, so the deduplication that produces the headline number does not happen.

And a workflow that never packs. Loose snapshots are stored without compression against each other, which is the README's own framing of why step two of the quickstart exists.

The alternative approach to the same underlying problem is a compilation cache, of the kind ccache or sccache provide. The difference in approach is that a compilation cache discards artifacts once they have been reused, keeping only a hash-indexed subset to answer the next identical invocation, whereas elfshaker keeps everything and lets you extract any revision you have. A cache is smaller and automatic; elfshaker is explicit, complete, and sized by how much history you keep.

Choose elfshaker when you need to reach a specific old build on demand, which is bisection and benchmark archaeology. Choose a compilation cache when you merely want to stop recompiling the same thing twice, which is everyday incremental work. Both apply at once in many projects, and they do not conflict, because one answers different questions from the other.

## Conclusion

Adopt elfshaker when you need to extract a specific historical build on demand, such as during compiler bisection, and only if your object files are compiled with -ffunction-sections and -fdata-sections, since the compression result depends on it. Do not use it for linked executables or shared libraries, where absolute addresses defeat compression. Verify first that the version you build matches the pack files you already hold, because the only tagged release is v0.9.0 from 2021-11-18 while the manifest sits at 1.0.0-rc1, and validate the format yourself as the README asks.

## FAQ

### What is elfshaker and what problem does it solve?

elfshaker is a Rust CLI described as a version control system fine-tuned for binaries. It stores snapshots of a directory into highly compressed pack files and provides fast on-demand access, which the README says allows few-second access to any commit of clang from the manyclangs archive.

### How do I store and restore a snapshot with elfshaker?

Run elfshaker store with a snapshot name to capture the working directory, elfshaker pack with a pack name to gather all loose snapshots into one compressed file, and elfshaker extract with the snapshot name to restore it. The README stresses that the pack step is where the compression benefit comes from.

### Why do I need -ffunction-sections and -fdata-sections?

The README credits those compiler flags for the compression ratio. They stop an inserted function from shifting every address in the object file, which is what leaves unchanged sections byte-identical and therefore compressible against each other.

### Does elfshaker compress linked binaries well?

No. The README says address changes after an insertion are handled poorly by compression algorithms, giving a ratio of something like 20 percent when many clang executables are compressed together, versus something closer to 0.01 percent for pre-link object files.

### Which platforms does elfshaker support?

CI tests run on Ubuntu 20.04 LTS. The project aims to support popular Linux platforms, macOS and Windows in production, and officially supports AArch64 and x86-64 architectures.

### When was elfshaker last released?

The only listed release is v0.9.0 from 2021-11-18, while the manifest declares version 1.0.0-rc1, so that release candidate has never been published. The repository is not archived and the last push was on 2026-07-08.

## Sources

- [elfshaker/elfshaker on GitHub](https://github.com/elfshaker/elfshaker)
- [Issues](https://github.com/elfshaker/elfshaker/issues)
- [License: Apache-2.0](https://github.com/elfshaker/elfshaker/blob/main/LICENSE)
- [README](https://github.com/elfshaker/elfshaker/blob/main/README.md)
- [Releases](https://github.com/elfshaker/elfshaker/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/elfshaker-elfshaker
