# Boltz: open source biomolecular interaction and binding affinity prediction

> Boltz is a family of MIT-licensed models for biomolecular interaction prediction, with Boltz-2 adding binding affinity output. Here is how it installs, what the affinity fields mean, and where the documentation stops.

**jwohlwend/boltz** — Official repository for the Boltz biomolecular interaction models

- Repository: https://github.com/jwohlwend/boltz
- Stars: 4,232 · Forks: 905
- Language: Python
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/jwohlwend-boltz

## The gap Boltz fills: structure and affinity in one open model

Structure prediction and affinity prediction have historically been separate steps with separate tooling. Boltz-2 is described in the README as a biomolecular foundation model that jointly models complex structures and binding affinities, rather than treating affinity as a downstream calculation on a frozen structure. The README states that Boltz-2 is the first deep learning model to approach the accuracy of physics-based free-energy perturbation methods while running 1000x faster, a claim taken from the project's own description rather than an independent benchmark.

The intended audience is early-stage drug discovery work where in silico screening has to be cheap enough to run broadly. The README frames the speed difference as what makes accurate in silico screening practical at that stage. Boltz-1, the earlier model in the family, is described as the first fully open source model to approach AlphaFold3 accuracy. Both the code and the weights are released under the MIT license, which the README says covers academic and commercial use.

## How Boltz-2 works: YAML inputs, MSA generation, two affinity outputs

The unit of work is a YAML file describing the biomolecules to model and the properties to predict. The README states that input_path can point to a single YAML file or to a directory of YAML files for batched processing. The examples/ directory in the repository shows the range of supported cases: prot.yaml for a single protein, multimer.yaml for several chains, ligand.yaml and ligand.fasta for small molecules, pocket.yaml, cyclic_prot.yaml, prot_custom_msa.yaml, prot_no_msa.yaml, and affinity.yaml.

The --use_msa_server flag delegates multiple sequence alignment generation to a server rather than requiring a precomputed alignment. The README notes that when that server requires authentication, credentials can be supplied in one of two ways, with details in docs/prediction.md rather than in the README itself. Automatic MSA generation carries its own citation requirement: the README asks users to cite the ColabFold paper when they use it.

Affinity output is deliberately split into two fields trained on largely different datasets with different supervision. The README positions affinity_probability_binary as a binder-versus-decoy detector for hit discovery, with a value from 0 to 1 representing the predicted probability that the ligand is a binder. affinity_pred_value is positioned for ligand optimization, hit-to-lead and lead optimization, reporting log10(IC50) derived from an IC50 measured in μM. Using the wrong field for the wrong stage is the most likely way to misread a Boltz run.

## Installing Boltz and running a first prediction

The README recommends a fresh Python environment. The package requires Python 3.10 up to but not including 3.13, according to pyproject.toml. The recommended install is from PyPI with the CUDA extra:

```bash
pip install boltz[cuda] -U
```

For a build that tracks the repository directly, the README gives a clone and editable install:

```bash
git clone https://github.com/jwohlwend/boltz.git
cd boltz; pip install -e .[cuda]
```

On CPU-only or non-CUDA hardware, drop the [cuda] extra. The README warns that the CPU version is significantly slower than the GPU version, so this is a fallback rather than an equivalent path. The cuda extra pulls in the cuequivariance packages listed in pyproject.toml, which the README ties to NVIDIA cuEquivariance acceleration on recent NVIDIA GPUs.

A first run uses the predict subcommand and points at an input file:

```bash
boltz predict input_path --use_msa_server
```

The README states that the boltz command runs the latest version of the model by default. To see every available option, run boltz predict --help. Input formats are documented in docs/prediction.md, and the files under examples/ are the fastest way to see a working YAML before writing your own.

## Where Boltz-2 is the wrong tool, and where the docs go quiet

The clearest limitation is stated by the project itself. The README marks updated evaluation code for Boltz-2 as coming soon, and updated training code for Boltz-2 as coming soon as well. If your work depends on reproducing the reported affinity numbers, or on retraining the model, the repository does not currently give you that. The training instructions that exist are described as being for Boltz-1 for now.

The affinity fields are also easy to misuse. affinity_probability_binary and affinity_pred_value come from different datasets and different supervision, so treating them as interchangeable numbers is a mistake the README explicitly warns against. A pipeline that ranks compounds by affinity_pred_value at a hit-discovery stage is using the field outside its stated purpose.

Hardware is a real constraint. The CUDA path depends on cuequivariance packages, and the README offers no CPU performance figure, only the statement that CPU is significantly slower. There is also no documented rollback or version pinning procedure in the README for moving between model versions, even though the CLI defaults to the latest model. Teams that need reproducible runs across months will have to work that out from the release history themselves.

## Boltz-1, Boltz-2 and the AlphaFold3 comparison

The natural alternative is AlphaFold3, and the difference is not only accuracy. The README describes Boltz-1 as the first fully open source model to approach AlphaFold3 accuracy, and Boltz-2 as going beyond both AlphaFold3 and Boltz-1 by jointly modeling structures and binding affinities. The practical distinction is licensing and scope: Boltz ships code and weights under MIT, which the README says covers commercial use, and it produces affinity output in the same run as the structure.

A second point of comparison is within the family. Boltz-1 covers structure; Boltz-2 adds affinity. The README's evaluation section says scripts and structural predictions for Boltz-2, Boltz-1, Chai-1 and AlphaFold3 on the project's test benchmark dataset, plus affinity predictions on the FEP+ benchmark, CASP16 and the MF-PCBA test set, are still to come. Until those land, comparing Boltz-2 against any of those models means running your own evaluation rather than citing the project's.

A third path exists for non-NVIDIA hardware. The README states that Boltz runs on Tenstorrent hardware through a fork by Moritz Thüning at github.com/moritztng/tt-boltz. That is a community fork, not the main repository, so its maintenance and feature parity are separate questions from the ones covered here.

## Licence, maintenance and the cost of staying current

Boltz is released under the MIT license, and the README states that both the model and the code can be freely used for academic and commercial purposes. The repository's LICENSE file is at the top level. Two attribution obligations sit alongside that: the README asks users to cite the Boltz-2 and Boltz-1 papers when using the code or models, and to cite the ColabFold paper when using automatic MSA generation. Those are citation requests, not licence terms, but they are stated in the README and are easy to miss. This is not legal advice; check the licence text and your own obligations.

The dependency list in pyproject.toml pins several packages to exact versions, including hydra-core 1.3.2, pytorch-lightning 2.5.0, numpy below 2.0, and rdkit at or above 2024.3.2. Exact pins make upgrades predictable but also mean a Boltz upgrade can move several transitive dependencies at once. The Python range is capped below 3.13, so a project on a newer interpreter cannot install it without changing Python versions.

The last push to the repository was on 2026-05-29, and the most recent release listed is v2.2.1 from 2025-09-08, described as minor bug fixes. The two releases before it, v2.2.0 and v2.1.1, mention contact conditioning fixes, new potentials, and new CUDA kernels from NVIDIA. Older releases are not annotated with upgrade or migration notes in the available release history, so pinning a version and reading the release description before moving is the only documented signal.

## Conclusion

Adopt Boltz if you need an MIT-licensed model that outputs both complex structure and a binding affinity estimate from a single YAML input, and if you can run it on CUDA hardware, since the README states the CPU version is significantly slower. Do not adopt it if you need the Boltz-2 evaluation or training code, because the README marks both as coming soon, or if you expect a documented upgrade path between model versions, because the README does not document one. Verify first that your input YAML matches one of the files under examples/, then confirm which affinity field fits your stage: affinity_probability_binary for binder-versus-decoy triage, affinity_pred_value for comparing binders during optimization.

## FAQ

### What is Boltz used for?

Boltz predicts biomolecular interactions. Boltz-2 jointly models complex structures and binding affinities, and the README positions the affinity output for hit discovery and for ligand optimization stages such as hit-to-lead and lead optimization.

### How do you install Boltz?

The README recommends a fresh Python environment and installing from PyPI with pip install boltz[cuda] -U, or cloning the repository and running pip install -e .[cuda]. On CPU-only or non-CUDA hardware you remove the [cuda] extra, though the README notes the CPU version is significantly slower.

### How do you use Boltz-2?

You run boltz predict input_path --use_msa_server, where input_path is a YAML file or a directory of YAML files describing the biomolecules and the properties to predict. The README says the boltz command runs the latest version of the model by default, and boltz predict --help lists all options.

### How do you install Boltz-2?

Boltz-2 ships in the same boltz package, currently at version 2.2.1 in pyproject.toml, so installation is the same command as for the package as a whole: pip install boltz[cuda] -U, or an editable install from the cloned repository.

## Sources

- [Issues](https://github.com/jwohlwend/boltz/issues)
- [jwohlwend/boltz on GitHub](https://github.com/jwohlwend/boltz)
- [License: MIT](https://github.com/jwohlwend/boltz/blob/main/LICENSE)
- [README](https://github.com/jwohlwend/boltz/blob/main/README.md)
- [Releases](https://github.com/jwohlwend/boltz/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jwohlwend-boltz
