Open-source project
google/bloaty avatar
google/bloaty

google/bloaty: a size profiler for binaries, and how to read its first report

Bloaty: a size profiler for binaries

5,552 stars378 forksC++Apache-2.0

At a glance

What is it?
Bloaty attributes the bytes of an ELF, Mach-O, PE/COFF or WebAssembly binary to compileunits, symbols, sections and segments. It is a diagnostic tool for people who already have a binary that is too big and want to know which code is responsible.
Who is it for?
Adopt Bloaty if you ship a compiled artifact and need to explain its size to someone, especially if you want a size diff in CI. Do not adopt it if your problem is runtime memory, allocation churn or binary startup time; Bloaty reads files, it does not observe a running process.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The question Bloaty answers, and who is asking it

A binary that has grown past its budget is a problem with no obvious owner. The linker reports a total, the build system reports a delta, and neither says which source file is responsible. Bloaty exists to close that gap. It produces a size profile of a binary, ranked by whatever dimension you choose, so the answer to "what is making this big" is a list of names with byte counts attached.

The audience is narrow and specific. People who ship a compiled artifact and have to defend its size: embedded and mobile engineers working against a flash or download budget, release engineers who watch an artifact grow between versions, and library authors who want to know how much of their code a consumer actually pulls in. The README's own example is Bloaty profiling itself, which is a fair illustration of the intended use: a C++ program with vendored third-party code, where the interesting question is how much of the total comes from protobuf, capstone and re2 rather than from src/.

It is not a general performance tool. Nothing in the documentation suggests it observes a running process, measures allocation, or says anything about speed. The unit of analysis is the file on disk.

How Bloaty attributes bytes: custom parsers, data sources, hierarchical profiles

The mechanism described in the README is a set of custom parsers for ELF, DWARF and Mach-O. Bloaty does not shell out to readelf or nm and post-process the output; it parses the formats itself, with the stated goal of accurately attributing every byte of the binary to the symbol or compileunit that produced it. That ambition is why the tool needs a disassembler: the README says it will disassemble the binary looking for references to anonymous data, which is how bytes with no symbol of their own get attached to something meaningful.

The output is organized around data sources. compileunit, symbol, section and segment are named in the README as examples, and the report shown there uses compileunits. A data source is the dimension along which bytes are bucketed, and the same binary can be profiled along several of them without recompiling anything.

Data sources compose. The README lists hierarchical profiles, meaning multiple data sources can be combined into a single report, so a top-level breakdown by segment can be expanded into compileunits underneath it. There are also custom data sources, described as regex rewrites of built-in ones, for munging or bucketing results into categories the binary does not natively express. Regex filtering works in the other direction, removing parts of the binary that do or do not match a pattern. The repository carries config.bloaty and custom_sources.bloaty at the top level, which is where that configuration lives.

One design consequence worth naming: attribution quality depends on what the binary still contains. Compileunit attribution needs DWARF. Symbol attribution needs a symbol table. Strip the binary and the same command produces a different, coarser report. The README addresses this with support for separate debug files, so the binary under test can be stripped while debug data stays available for analysis.

Building Bloaty with cmake and reading a first report

The README gives cmake as the build path and bundles libprotobuf, re2, capstone and pkg-config as Git submodules. It prefers system versions of those four if they are present, and protoc comes from the bundled libprotobuf. Everything else is a submodule. If the clone was not recursive, the README says to initialize the submodules first:

bash
git submodule update --init --recursive

Then configure and build. The README's example uses Ninja as the generator and installs into the default prefix:

bash
cmake -B build -G Ninja -S .
cmake --build build
cmake --build build --target install

After that, bloaty is on the path. The first useful command is the one from the README, profiling a binary by compileunit. Here it is run against bloaty itself, which is the example the README shows:

bash
bloaty bloaty -d compileunits

The output has two size columns, FILE SIZE and VM SIZE, and rows ranked by percentage with a TOTAL line at the bottom. In the README's run, TOTAL is 29.5Mi of file size and 6.69Mi of VM size, and the largest single row is third_party/protobuf/src/google/protobuf/descriptor.cc at 17.2 percent of file size. The two columns diverge sharply for some entries, which is the point of showing both: capstone's M68KDisassembler.c is 1.3 percent of file size but 15.9 percent of VM size. Note also the [163 Others] row at the top, which aggregates everything below the display threshold.

To see where a binary grew between two builds, use the size diff mode the README lists among the features. It is described as being aimed at CI tests, which is the natural place for it: compare the artifact from the current commit against a stored baseline and fail or comment when the delta is large.

Where Bloaty's attribution breaks down

The most common failure is not a bug. It is a stripped binary. Compileunit-level output depends on DWARF being present, and a release build that has been run through strip will not produce the report shown in the README. Bloaty's answer to this is the separate debug files feature, but that requires the build to have produced and retained the debug data in the first place. If your release pipeline discards it, you have to change the pipeline before Bloaty can help.

The disassembly step has a cost that the README does not quantify. It says Bloaty will disassemble the binary looking for references to anonymous data. On a large C++ binary with heavy template instantiation, that is a lot of code to walk, and the README gives no timing figures. Treat analysis time on a multi-hundred-megabyte binary as unknown until you measure it yourself.

Attribution is also inherently approximate for anonymous data. Bytes that no symbol, section or debug entry claims end up in an aggregate row, which is why the README's own output shows [163 Others] taking 34.8 percent of file size. That row is not noise, it is the honest remainder, but it means the top of the report is sometimes a bucket rather than a file. If your binary is mostly stripped of debug info, most of the report will be that bucket.

Finally, the format support is tiered. The README lists ELF and Mach-O without qualification, and marks PE/COFF and WebAssembly as experimental. If you are profiling a Windows or WebAssembly artifact, you are on the experimental path and should expect less complete attribution.

Bloaty against linkers and compiler size flags

The obvious alternative is not another profiler. It is the toolchain itself: the linker's map file and the compiler's size-reporting flags. A linker map already tells you which object files and which sections contributed how many bytes, and it comes out of the build you were going to run anyway. Many compilers can emit a per-function or per-section size breakdown too.

The difference in approach matters. A linker map is a build artifact that reflects the link step exactly, including anything the linker discarded, and it costs nothing extra to produce. Bloaty works on a finished binary after the fact, which means it can analyze a file you did not build, a file someone sent you, or a release artifact pulled from a package repository. It also gives you dimensions a map file does not: compileunit attribution through DWARF, and the disassembly-based attribution of anonymous data.

So the split is roughly this. If you control the build and just want to know which object files are large, the map file is cheaper and always available. If you need to explain a binary you did not build, or you need the size attributed to source-level units rather than object files, Bloaty is the tool that does that. The README's size diff mode also has no direct equivalent in a map file, since a map file describes one link, not a comparison between two.

Maintenance, licence and what an upgrade costs you

The last push to the repository was on 2026-09-10, so the project is not abandoned, but the release history is thin: v1.0 in August 2018 and v1.1 in May 2020, with nothing tagged since. The README's install instructions are the cmake path, not a package manager, and the description of a statically-linked C++ binary that is easy to copy around suggests the intended deployment is a binary you build and carry, not a dependency you resolve.

That shapes the upgrade cost. There is no version to bump in a manifest. If you follow the README, you build from source and you own the result, which means the cost of staying current is the cost of rebuilding, including the bundled submodules for libprotobuf, re2, capstone and pkg-config. Because Bloaty prefers system versions of those four when available, the same source tree can build against different dependency versions on different machines, and a build that works on one may behave differently on another. Pinning is your responsibility, not the project's.

The licence is Apache-2.0, per the LICENSE file at the repository root. That is a permissive licence, and the README adds that this is not an official Google product, which is worth reading carefully if you were assuming the Google name implies support. The README directs support to GitHub issues and PRs and asks that tests accompany contributions, pointing at tests/README.md. There is no stated compatibility guarantee between versions and no documented deprecation policy, so treat the output format as something to verify rather than assume if you parse it in CI.

Editorial conclusion

Adopt Bloaty if you ship a compiled artifact and need to explain its size to someone, especially if you want a size diff in CI. Do not adopt it if your problem is runtime memory, allocation churn or binary startup time; Bloaty reads files, it does not observe a running process. Before trusting a report, verify that the binary still carries the debug information or symbol table your chosen data source needs, because a stripped binary silently changes what Bloaty can attribute, and check whether your platform is in the experimental column: PE/COFF and WebAssembly are marked experimental in the README while ELF and Mach-O are not.

Frequently asked questions

What is google/bloaty?

It is a size profiler for binaries. The README says it shows a size profile of a binary so you can understand what is taking up space inside, using custom ELF, DWARF and Mach-O parsers to attribute bytes to the symbol or compileunit that produced them.

How do I use google/bloaty on a binary?

Build it with cmake, then run bloaty against the file and pick a data source with -d. The README's example is bloaty bloaty -d compileunits, which prints FILE SIZE and VM SIZE columns ranked by percentage with a TOTAL row.

Does google/bloaty run on Windows binaries?

The README lists PE/COFF as a supported file format but marks it experimental, alongside WebAssembly. ELF and Mach-O are listed without that qualification, so Windows and WebAssembly artifacts are on the less-tested path.

Does google/bloaty need debug information in the binary?

Compileunit-level attribution depends on the debug data being available, which is why the README lists separate debug files as a feature: the binary under test can be stripped while debug data remains available for analysis. Without it, the report falls back to coarser data sources.

What is the licence for google/bloaty?

The repository carries a LICENSE file and the project metadata identifies the licence as Apache-2.0. The README also states that this is not an official Google product.

Official sources

  1. google/bloaty on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/google-bloaty.svg)](https://hysenlabs.com/projects/google-bloaty)