# AlphaFold 2: Running the Open Source Inference Pipeline Locally

> DeepMind's AlphaFold repository ships the inference code, model parameters and database scripts behind the CASP14 structure predictor. It is a Linux and NVIDIA GPU undertaking measured in terabytes, not a pip install.

**google-deepmind/alphafold** — Open source code for AlphaFold 2.

- Repository: https://github.com/google-deepmind/alphafold
- Stars: 14,865 · Forks: 2,951
- Language: Python
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-deepmind-alphafold

## What the AlphaFold repository actually gives you

This is the inference pipeline of AlphaFold v2, released under Apache-2.0, plus the model parameters and the scripts that fetch the sequence databases the model searches against. It predicts a protein's three-dimensional structure from its amino acid sequence. The intended audience is a research group or a computational biology engineer with GPU hardware, not a developer looking for a library to import into a web service.

The README is explicit about the boundary: you need a machine running Linux, and other operating systems are not supported. The repository also carries an implementation of AlphaFold-Multimer for complexes, which the README describes as a work in progress that is not expected to be as stable as the monomer system. Two supporting artifacts ship alongside the code: a technical note for v2.3.0 and a CASP15 baseline set of predictions with documentation of manual interventions. If you publish findings from this code or its parameters, the README asks you to cite the AlphaFold paper and, where relevant, the AlphaFold-Multimer paper.

## How the pipeline runs: Docker, databases, templates

The architecture is a container wrapped by a Python driver. `docker/run_docker.py` builds the command that starts the AlphaFold container, mounts the downloaded genetic databases and your FASTA input, and writes results to an output directory. Inside the container the pipeline searches the sequence databases, retrieves structural templates, and runs the neural network on the GPU.

The database list is the part that catches people out. AlphaFold queries BFD, MGnify, PDB70, PDB in mmCIF form, UniRef90, and UniRef30 (formerly UniClust30). Two more, PDB seqres and UniProt, are needed only for AlphaFold-Multimer. The README states the full download is 556 GB and that full installation requires up to 3 TB of disk space, with SSD storage recommended. The template search is bounded by a date you supply on the command line, so you control which structures the model is allowed to see. That single flag is the difference between a fair retrospective benchmark and a leaky one.

## Installing AlphaFold locally and running a first prediction

The README's path assumes Docker with GPU support. Install Docker, then the NVIDIA Container Toolkit for GPU access, and set Docker up to run as a non-root user. Clone the repository and enter it.

```bash
git clone https://github.com/deepmind/alphafold.git
cd ./alphafold
```

The databases come from a script that needs `aria2c` and `rsync`. On Debian-based distributions `aria2` is available from the package manager. The README recommends running the download in the background because of the size, and warns that the download directory must not be a subdirectory of the repository, otherwise the Docker build copies the databases into its build context.

```bash
sudo apt install aria2
scripts/download_all_data.sh <DOWNLOAD_DIR> > download.log 2> download_all.log &
```

Before building anything, confirm the container runtime can see your GPU. The README says the output should show a list of your GPUs; if it does not, the NVIDIA Container Toolkit setup is the first thing to recheck.

```bash
docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi
```

Build the image, then install the driver's Python dependencies. The README notes you may want a virtual environment to avoid conflicts with the system Python.

```bash
docker build -f docker/Dockerfile -t alphafold .
pip3 install -r docker/requirements.txt
```

Finally, run a prediction against a FASTA file. `--fasta_paths` points at your sequence, `--max_template_date` caps template search, `--data_dir` is the database directory and `--output_dir` is an absolute path that must exist and be writable. The default output directory is `/tmp/alphafold`. When the run finishes, the output directory holds the predicted structures.

```bash
python3 docker/run_docker.py \
  --fasta_paths=your_protein.fasta \
  --max_template_date=2022-01-01 \
  --data_dir=$DOWNLOAD_DIR \
  --output_dir=/home/user/absolute_path_to_the_output_dir
```

## The storage and GPU bill is the real adoption cost

Nothing here is hard to type. The difficulty is resource commitment. Up to 3 TB of SSD for databases and a 556 GB download sit in front of the first prediction, and the README recommends SSD specifically. A modern NVIDIA GPU is required, and the README notes that GPUs with more memory can predict larger protein structures, which is a polite way of saying that the size of the protein you care about may be capped by the card you own.

The repository does acknowledge a reduced-database path, pointing to the genetic databases documentation for running with fewer databases. That is a real option for smaller work, but the README does not quantify what accuracy you give up, so you would be trading a known cost for an unmeasured one. There is also a known Docker build failure documented in the README, a GPG error against the NVIDIA CUDA repository where signatures cannot be verified, with a workaround linked in a GitHub issue. Expect the build to need that fix on some hosts.

## AlphaFold-Multimer is the part to be careful with

Complex prediction is where the repository asks for patience. The README states plainly that the AlphaFold-Multimer implementation represents a work in progress and is not expected to be as stable as the monomer system. If your project depends on protein-protein interaction prediction, that sentence should shape your planning: treat multimer output as something to inspect rather than to trust by default, and budget time for runs that do not behave like the monomer pipeline.

There is a second constraint worth naming. Multimer needs PDB seqres and UniProt databases on top of the standard set, so the storage figure and download time grow if you intend to use it. The README documents an upgrade path under the heading for updating an existing installation, which matters because the container image and the database set move together. An older clone with newer databases, or the reverse, is not a configuration the README describes.

## Where AlphaFold is the wrong tool

If you want one structure this afternoon, this repository is the wrong instrument. The install alone is a multi-hour download before the first prediction, and the output needs interpretation. The README itself points to community-supported versions of AlphaFold as a slightly simplified alternative, which is a candid admission that the full local pipeline is not the only way in.

Windows and macOS users are out entirely; the README restricts the software to Linux. Teams without an NVIDIA GPU are out as well, since the GPU check is a required step rather than an optional acceleration. And anyone who needs a supported, versioned API with a stability guarantee will find that the most recent tagged release in the repository is v2.3.2 from 2023-04-05, while the main branch has continued to receive commits, the last push being 2026-04-22. The tags and the branch are not moving at the same pace.

## AlphaFold versus a hosted structure service

The practical alternative for many users is not another local package but a hosted prediction service, including the AlphaFold Server and the AlphaFold Protein Structure Database, both of which appear in the repository's own framing of community-supported routes. The difference in approach is where the compute and the databases live. A hosted service keeps the 3 TB, the GPU and the container maintenance on someone else's infrastructure, and you submit a sequence and receive a structure. The local repository keeps the sequence on your hardware, lets you pin `--max_template_date` for reproducible retrospective evaluation, and gives you the model parameters directly rather than a result page.

That trade is the whole decision. Hosted services are the right default for occasional queries and for teams without GPUs. The local pipeline is for groups that need to run many sequences, control the template cutoff, or keep unpublished sequences inside their own network.

## Licence, maintenance and upgrade cost

The code is Apache-2.0, and `pyproject.toml` declares the same licence for the package. That is a permissive licence, but it does not settle the terms attached to the model parameters or to the genetic databases you download, which come from separate providers with their own conditions. The README's citation request is not a licence term, but it is a stated expectation for publications that use the code or the parameters.

On upgrades, the repository is a moving target in a specific way. The most recent tagged release listed is v2.3.2 from 2023-04-05, yet the main branch received a push on 2026-04-22. Anyone tracking the branch rather than a tag is tracking unreleased code. The README documents a procedure for updating an existing installation, and the dependency pins in `requirements.txt` are exact versions, including `jax==0.4.26`, `tensorflow-cpu==2.16.1` and `numpy==1.24.3`. Exact pins make a rebuild reproducible and make a partial upgrade fragile. Budget for a full rebuild, databases included, rather than an in-place dependency bump.

## Conclusion

Adopt this repository if you have a Linux host with a modern NVIDIA GPU, roughly 3 TB of SSD for genetic databases, and a reason to keep sequences off third-party servers; the Docker path documented in the README is the intended route. Do not adopt it for a single structure you need today, for Windows or macOS, or if you expect a stable multimer pipeline, since the README labels AlphaFold-Multimer a work in progress that is not expected to be as stable as the monomer system. Before committing storage, verify three things: that `docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi` lists your GPUs, that your download directory sits outside the repository clone, and that `--max_template_date` matches the template cutoff your experiment actually wants.

## FAQ

### How do I install AlphaFold locally?

The README requires Linux, Docker with the NVIDIA Container Toolkit, and up to 3 TB of disk for the genetic databases. You clone the repository, run `scripts/download_all_data.sh` into a directory outside the clone, build the image with `docker build -f docker/Dockerfile -t alphafold .`, install `docker/requirements.txt`, and run `docker/run_docker.py` with your FASTA file.

### How do I use AlphaFold to predict a protein structure?

Put your sequence in a FASTA file and pass it to `run_docker.py` with `--fasta_paths`, along with `--data_dir` pointing at the downloaded databases and `--output_dir` pointing at a writable absolute path. The predicted structures appear in the output directory when the run finishes.

### How does AlphaFold work?

The repository implements the inference pipeline of AlphaFold v2. It searches multiple genetic sequence databases, retrieves structural templates up to the date given by `--max_template_date`, and runs the neural network on an NVIDIA GPU inside a Docker container.

### Can I use AlphaFold for free?

The source code is released under Apache-2.0, so the software itself is freely available to download and run. Running it locally still requires your own Linux machine, NVIDIA GPU and up to 3 TB of storage, and the downloaded genetic databases come from separate providers with their own terms.

### How do I use AlphaFold-Multimer?

The repository includes an AlphaFold-Multimer implementation, but the README describes it as a work in progress that is not expected to be as stable as the monomer system. It also requires the PDB seqres and UniProt databases in addition to the standard set.

## Sources

- [google-deepmind/alphafold on GitHub](https://github.com/google-deepmind/alphafold)
- [Issues](https://github.com/google-deepmind/alphafold/issues)
- [License: Apache-2.0](https://github.com/google-deepmind/alphafold/blob/main/LICENSE)
- [README](https://github.com/google-deepmind/alphafold/blob/main/README.md)
- [Releases](https://github.com/google-deepmind/alphafold/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-deepmind-alphafold
