# bap: the tags are four years old, the code is not

> Carnegie Mellon's Binary Analysis Platform lifts binaries into an architecture-neutral representation, builds a control flow graph, and runs your analyses as passes over it. The thing that will confuse you first is the release history: the newest tag is a development build from January 2023, while the repository itself was pushed to in May 2026.

**BinaryAnalysisPlatform/bap** — Binary Analysis Platform

- Repository: https://github.com/BinaryAnalysisPlatform/bap
- Stars: 2,261 · Forks: 286
- Language: OCaml
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/binaryanalysisplatform-bap

## One command from binary to control flow graph

The demonstration is one line, and it is the whole pitch.

```bash
bap /bin/echo -d
```

Running the utility on a binary disassembles it, lifts the instructions into an intermediate representation that does not depend on the architecture, builds a control flow graph, and applies your staged analysis as passes over that program. The dump option writes the result out in a format you choose. It is the default command, so you do not even have to name the operation.

The README makes the comparison directly: unlike the usual disassembly utility, this one builds the control flow graph for you. That sentence is the reason the project exists.

Here is why that matters. An analysis that works on x86 instructions is an x86 analysis. The moment you want the same analysis on MIPS or PowerPC or ARM, you are reimplementing the front end, and the front end is where the bugs live, because it is the code that has to be right about calling conventions, about which register holds what after a call, and about where the boundaries of a function are when the binary has been stripped. If a pass is written against the lifted representation instead, it is written once.

The architecture list in the overview covers x86 in both widths, ARM, MIPS and PowerPC, with new architectures added through plugins. So the shape of the project is a front end per architecture, a shared middle, and your analysis at the top. That is the same architecture as a compiler, and it is the reason a plugin can be architecture-independent.

Two interpreters ship alongside: a standard one, and a microexecution one that runs the program a step at a time rather than symbolically. The symbolic executor is the third, and it is where the solver interface lives.

## Primus Lisp, and why a DSL earns its keep here

The project has its own domain-specific language, and the README lists four jobs for it: implementing analyses, specifying verification conditions, modelling functions by writing stubs, and interfacing with the SMT solver.

Three of those are ordinary uses for a DSL in this domain. Writing analyses in the host language would mean recompiling the framework to add a pass. Specifying a verification condition is the same thing in a different hat. And a solver interface is what turns a symbolic executor into anything other than an interpreter with extra steps.

The fourth one is the interesting one, and it is the one that decides whether you can use this framework at all.

Binary analysis stalls on the C library. A binary calls a function, the analyser has to decide what that function does, and unless somebody has described it, the analyser knows nothing and everything downstream inherits that ignorance. Real programs spend a large fraction of their calls into functions the analyser has never seen, and a static analysis of a stripped binary is mostly an analysis of the unknown.

So writing stubs is the bulk of the work, and a language that makes a stub a few lines rather than a plugin is the difference between analysing a binary and analysing nothing. That is the highest-leverage thing in this project and it is one line of the overview.

The other reason the DSL exists is the solver boundary. If the language can express a constraint and hand it to an SMT solver, then a symbolic execution path can mix concrete execution with symbolic values inside one analysis. If it cannot, you are running two systems and shuttling constraints between them, and every shuttle is a place to lose an assumption.

It is also worth saying what the DSL does not do. It is not a general-purpose language for writing tools around the framework; the README describes embedding options separately for that, and they are a different set of choices.

## Sixty packages, and the three tools it overlaps

The root of this repository is a list of package manifests, one per component, and reading the names is the fastest way to understand the scope.

There is a package per architecture, so the lifters. There are frontends for the object formats and debug information. There are small analyses, each in its own package, which is a deliberate choice: a dataflow analysis or a constant tracker is a package so it can be depended on or dropped without taking the framework. And there is a core theory package and a knowledge package, which are the abstract semantics rather than any particular instruction set.

Then there are three names that tell you where this project sits in its own ecosystem: a package for Ghidra, one for a second disassembler, and a plugin package for IDA. The framework ships integrations with the three tools whose job it overlaps.

That is the most informative thing in the list. A binary analysis platform that also has plugins for the incumbent tools is not trying to replace them at the interface level; it is trying to be the thing they call. And the list includes a package called for the future, which in a research platform is usually the experimental bucket, which is where a framework accumulates the thing it has not yet designed properly.

The build treats all of these as peers, and the documented way to leave some out is to delete their manifest files before you create a switch:

```bash
rm bap-{extra,radare2,ghidra,primus-symbolic-executor,ida}.opam
```

That is the configuration surface. The package list is not generated from something else; it is the thing, and removing a file is how you say no to a component. For a framework with this many moving parts it is an unusually honest design, and it also means a local clone is a working configuration rather than a copy of upstream.

## The tags stopped in 2023, the channel did not

Read the release history and you would conclude this project is abandoned. That conclusion is wrong, and understanding why is the single most useful thing in the README.

There are three tags visible. A stable release from July 2022, an alpha of the same version from two days earlier, and a development build of the next version from January 2023. The repository itself was pushed to in May 2026, so there are roughly three and a half years of commits past the newest alpha.

The reason is in the installation section. The rolling releases are described as automatically updated every time a commit to the master branch happens, and they are published through a separate package repository with a testing channel. To use them you create a switch pointed at that channel:

```bash
opam switch create bap-testing --repos \
    default,bap=git+https://github.com/BinaryAnalysisPlatform/opam-repository#testing 4.14.1
```

So the tags are a historical record and the package channel is the current state. For this project, the release history is not a maintenance signal.

It has a practical consequence for the install instructions, which are four years old. The pre-built Debian packages are fetched from a path containing the 2022 version number, and the compiler version pinned in the quick setup is from 2022 as well. Anyone following the documented binary install today gets a 2022 build, and the README gives no indication that a newer one exists.

That is a real documentation gap rather than a project problem, and it is the kind that costs an afternoon: you install the packages, they work, and you have no reason to think you are three years behind. The check is cheap. Look at whether you are on the stable channel or the testing one before you evaluate anything else.

## A build system that is an OCaml program

The Makefile in this repository is a shim, and reading it tells you about the project's ergonomics.

Almost every target is one line that invokes a compiled setup program. Build, install, reinstall, uninstall, clean, distclean and test are all the same shape, passing a flag. The one piece of logic in the file is a conditional that decides whether setup runs quietly, and it recognises an unusual number of spellings of false:

```make
SETUP = ./setup.exe -quiet
```

The idiom underneath is a string comparison against a list that includes empty, zero, no, disable and false. Six ways to say no to verbose output, because a build that prints a thousand lines is unusable in a terminal and somebody kept typing the other spelling.

Two things in this file are worth copying regardless of language. There is a target that runs the formatter over everything and then asserts the working tree is unchanged, which means the style check fails when reformatting produced a diff. That is a much better gate than one that only reports. And the test target is not a directory in this repository: there is a separate test suite repository, cloned by its own target, and the check target runs it against a pinned revision. Tests versioned independently of the code they test, pinned by commit.

There is a second build system as well, an OMakefile alongside the Makefile, which is the residue of a project that has moved build tooling twice and kept both entry points. And the manual path needs a compiler at a pinned version plus a local switch:

```bash
git clone git@github.com:BinaryAnalysisPlatform/bap.git && cd bap
opam switch create . --deps-only
dune build && dune install
```

One sharp edge is documented with a fix. If the LLVM configuration executable is not on your path, you point the build at yours explicitly, and the example given is a version-11 executable, which tells you how old that integration is.

## What it was funded to do, and who used it

The origin story and the user list are the best available evidence of what this framework is actually good at, and both point the same way.

It was developed at Carnegie Mellon University's CyLab, and the sponsors listed are the United States Department of Defence, Siemens, Boeing, ForAllSecure and the Korean government. That is a research and national-security funding profile, not a commercial product profile, and it is consistent with the design choices: an intermediate representation, a solver interface, and a language for writing stubs.

The user list is more specific. The framework is described as a backbone for a number of projects, three of which are named. One is the winning entry from the Cyber Grand Challenge, which is the public competition that asks automated systems to find and exploit vulnerabilities in purpose-built binaries, and which is the hardest public test of this class of tool that exists. One is a set of tools from a national laboratory focused on assessing software. The third is a checker for a common weakness enumeration at a research institute.

So the framework has been stressed by people whose entire objective was to break binaries, which is the strongest available evidence that its abstractions survive adversarial use. And the three projects are all analysis tools rather than exploitation tools, which says the framework is used where the question is what a program does rather than how to make it do something else.

One more detail about adoption, because it decides who can use this. There are three ways in. Use it as a framework driven by the command line and extended with plugins. Embed it as a library in an OCaml application. Or embed it from any other language through C bindings. There is also minimal Python support, described as being there to make it easier to start learning.

The C bindings are the important one. A framework written in a language your team does not use is a research artefact, and a framework with foreign function bindings is infrastructure. The bindings are why a project that started as a Cyber Grand Challenge entry point is now a component inside tools at a national laboratory.

## Where bap is the wrong tool

Five cases, and the first two are about prerequisites rather than about capability.

If you do not have an OCaml toolchain and do not want one, the cost is real: a package manager, a pinned compiler, a local switch, and a build that is itself a program you must compile before it can build anything. The escape hatch is the C bindings, and they work, but they put a boundary between your analysis and the framework, which means you are now debugging two languages.

If you are on a current compiler and a current LLVM than the documentation assumes, expect friction. The quick setup pins a compiler version and the LLVM note names a version-11 executable as the example. These are the kinds of pin that a research framework sets once and does not revisit, because the person who set it had that toolchain.

If you want a graphical environment, use one of the tools it plugs into. This is a framework and a library; there is no interface here, and the integrations for the three commercial and open-source tools exist precisely because people wanted the framework's capabilities inside somebody else's window.

If you want to find a string in a binary, use a string extraction utility. Everything described here, including the lift and the control flow graph, is in service of an analysis you have to write.

And if your binary is enormous, this is in-process, single-machine analysis, so the memory ceiling is your own machine's. The README describes running many experiments in parallel from a graphical interface, which is parallelism over programs rather than over one large program, and nothing in it addresses the case of a single analysis that will not fit.

The honest summary is that this is infrastructure for people who already know binary analysis. If you are starting, the tutorial and the example toolkit are the way in, and the command in this README is the fastest way to see what the lift looks like.

## Conclusion

Adopt bap if you are writing analyses that must work across architectures, because the lift and the control flow graph are done for you and a pass written once runs on every backend. Get it from the rolling opam channel rather than from the tags, since the documented binary packages are from July 2022 and the newest tag is from January 2023. Budget for an OCaml toolchain and an LLVM configuration step, or embed it through the C bindings if your analysis lives in another language. And read the package list before you commit, because it includes integrations for the three tools this one overlaps with.

## FAQ

### What is the Binary Analysis Platform?

It is Carnegie Mellon University's framework for analysing binary programs, written in OCaml. It disassembles a binary, lifts the instructions into an intermediate representation independent of the architecture, builds a control flow graph, and applies user-defined analyses as passes over that program. It includes a standard interpreter, a microexecution interpreter and a symbolic executor.

### Which architectures does bap support?

x86, x86-64, ARM, MIPS and PowerPC, with additional architectures supported through plugins. Because analyses run over the lifted representation rather than over instructions, an analysis written once runs against every backend the framework supports.

### How do I install bap from source?

Use the OCaml package manager, which the README recommends. Initialise it with a pinned compiler version, install the package, then evaluate the environment activation command. If installing system dependencies through the operating system package manager fails, install the missing one manually and repeat the install.

### Why does bap have a domain-specific language called Primus Lisp?

It is used to implement analyses, specify verification conditions, model functions by writing stubs, and interface with the SMT solver. Writing stubs is the highest-leverage use, because a static analysis of a binary is limited by how much of the C library the analyser has been given models for.

### How do I get the current version of bap?

From the rolling package channel rather than from the GitHub tags. The newest visible tag is a development build from January 2023 and the newest stable is from July 2022, while the repository itself was pushed to in May 2026. The rolling channel is updated on every commit to the master branch and is selected by creating a switch pointed at the project's testing repository.

### Can I use bap from a language other than OCaml?

Yes, through C bindings, which the README describes as allowing the framework to be embedded in a user application written in any other language. There is also minimal Python support, described as being there to make it easier to start learning the framework.

## Sources

- [BinaryAnalysisPlatform/bap on GitHub](https://github.com/BinaryAnalysisPlatform/bap)
- [Issues](https://github.com/BinaryAnalysisPlatform/bap/issues)
- [License: MIT](https://github.com/BinaryAnalysisPlatform/bap/blob/master/LICENSE)
- [README](https://github.com/BinaryAnalysisPlatform/bap/blob/master/README.md)
- [Releases](https://github.com/BinaryAnalysisPlatform/bap/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/binaryanalysisplatform-bap
