AlphaFold 3: running the inference pipeline locally
AlphaFold 3 inference pipeline.
At a glance
- What is it?
- AlphaFold 3 ships the inference code for biomolecular structure prediction, but the model parameters come with their own terms of use. Here is what the repository actually gives you, how to run a first prediction, and where it stops being the right tool.
- Who is it for?
- Adopt AlphaFold 3 if you need to run the inference pipeline on your own hardware under the weights terms of use, and you already have the databases and a GPU. Do not adopt it if you want a hosted, no-setup path: the repository points to alphafoldserver.com for non-commercial use and to the Gemini Enterprise Agent Platform on Google Cloud for commercial use.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What AlphaFold 3 solves, and who it is written for
AlphaFold 3 predicts the structure of biomolecular interactions, not just single protein chains. The README describes the package as "an implementation of the inference pipeline of AlphaFold 3", and the examples directory shows the range the input format is built to carry: protein complexes, DNA, RNA, modified residues, glycosylation, and small-molecule ligands such as the KRAS G12C complex with sotorasib and the streptavidin-biotin example given as SMILES. If your work is limited to single-chain protein folding, that breadth is overhead you pay for in setup and runtime.
The audience is narrower than the paper's readership. The repository expects you to supply your own model parameters, your own databases, and a machine with a GPU. It is built for groups who need to run predictions in-house, either because the data cannot leave their environment or because they want to script many predictions rather than submit them one at a time. Anyone who just wants a structure should read the hosted options first, because the local path has real prerequisites.
The pipeline is split into a data stage and an inference stage
The README names two flags that decide what actually runs. `--run_data_pipeline` defaults to `true` and covers genetic and template search. The README states plainly that this part is CPU-only, time consuming, and could be run on a machine without a GPU. `--run_inference` also defaults to `true` and requires a GPU. That split is the most useful architectural fact in the repository, because it means the expensive database search and the GPU-bound model evaluation can be scheduled separately, on different machines, as long as you pass the intermediate results along.
The data search itself is not implemented inside this repository. The top-level `fetch_databases.sh` script exists to pull the reference databases, and the acknowledgements list HMMER Suite, DSSP, and libcifpp as separate dependencies. So the pipeline is a coordinator: it prepares inputs, calls out to sequence and structure search tooling, then feeds the result into the JAX model. The Python dependencies in pyproject.toml confirm the model side: `jax==0.10.2`, `dm-haiku==0.0.17`, `tokamax==0.0.12`, and `rdkit==2025.9.4` for the chemistry handling. The build backend is `scikit_build_core`, and there is a CMakeLists.txt at the top level, so parts of the package are compiled rather than pure Python.
Installing AlphaFold 3 and running a first prediction
The README does not put installation steps in the top-level file. It points to `docs/installation.md` and says to see that document. What the README does give is the container invocation, and that is the path most readers will follow, because the Docker directory is present at the top level and the command below assumes an image already built and tagged `alphafold3`.
You need three host directories mounted: one for input JSON, one for output, and one holding the model parameters, plus a fourth for the databases. The parameters come from a single URL, https://storage.googleapis.com/alphafold3/af3.bin.zst, and the README states you may only use AlphaFold 3 model parameters if received directly from Google.
docker run -it \
--volume $HOME/af_input:/root/af_input \
--volume $HOME/af_output:/root/af_output \
--volume <MODEL_PARAMETERS_DIR>:/root/models \
--volume <DATABASES_DIR>:/root/public_databases \
--gpus all \
alphafold3 \
python run_alphafold.py \
--json_path=/root/af_input/fold_input.json \
--model_dir=/root/models \
--output_dir=/root/af_outputReplace the two angle-bracketed placeholders with real paths before running. The `--gpus all` flag is what makes the inference stage possible; without it, only the data pipeline can run.
The input format is JSON. The README gives a two-chain test case named `2PV7`, and the structure of that file is the template you will copy for your own jobs.
{
"name": "2PV7",
"sequences": [
{
"protein": {
"id": ["A", "B"],
"sequence": "GMRESYANENQFGFKTINSDIHKIVIVGGYGKLGGLFARYLRASGYPISILDREDWAVAESILANADVVIVSVPINLTLETIERLKPYLTENMLLADLTSVKREPLAKMLEVHTGAVLGLHPMFGADIASMAKQVVVRCDGRFPERYEWLLEQIQIWGAKIYQTNATEHDHNMTYIQALRHFSTFANGLHLSKQPINLANLLALSSPIYRLELAMIGRLFAQDAELYADIIMDKSENLAVIETLKQTYDEALTFFENNDRQGFIDAFHKVRDWFGDYSEQFLKESRQLLQQANDLKQG"
}
}
],
"modelSeeds": [1],
"dialect": "alphafold3",
"version": 1
}Note the `dialect` and `version` keys. They are part of the schema, and `docs/input.md` is where the full field set is documented. The `id` array is how a single sequence is declared as a homodimer, which is why the test case has two chains from one sequence string. Save the file as `fold_input.json` in the mounted input directory, run the container command, and the output lands in the mounted output directory. For the full flag list, the README says to run `python run_alphafold.py --help`.
The weights are not covered by the Apache-2.0 code licence
This is the part that catches people. The repository ships under Apache-2.0, and the LICENSE file is at the top level. But the model parameters are governed by a separate document, WEIGHTS_TERMS_OF_USE.md, and the README is explicit: you may only use AlphaFold 3 model parameters if received directly from Google. There is also an OUTPUT_TERMS_OF_USE.md and a WEIGHTS_PROHIBITED_USE_POLICY.md in the same directory. Three separate policy files sit next to the code licence.
The practical consequence is that a permissive code licence does not make the whole stack permissive. You can read, modify and redistribute the inference code. You cannot treat the weights the same way. The README also states that any publication disclosing findings from the source code, the parameters or the outputs must cite the Nature paper, and gives the BibTeX entry to use. If you are planning to publish, that citation requirement is not optional.
None of this is legal advice, and the terms of use documents are the authority. What can be said from the repository layout is that the licence situation is deliberately split, and anyone evaluating adoption should read WEIGHTS_TERMS_OF_USE.md before downloading af3.bin.zst, not after.
Where AlphaFold 3 is the wrong tool
The setup cost is the first limit. You need a GPU, the databases fetched by `fetch_databases.sh`, and the model parameters from Google. The README describes the data pipeline as time consuming, and it is CPU-only, so a laptop is not a realistic host for a full run. If you want a structure today and have none of that infrastructure, the local repository is the wrong entry point. The README itself points to alphafoldserver.com for non-commercial use, and notes that the server has a more limited set of ligands and covalent modifications. That trade is worth naming: the local pipeline covers more chemistry, the server covers less setup.
The second limit is scope. The examples show small-molecule ligands, modified RNA, methylated DNA, and glycosylation, but the repository is an inference pipeline. It does not train, fine-tune or evaluate models. If your question is about model behaviour under distribution shift, or about retraining on your own structures, this codebase does not address it.
The third is the known issues file. The README maintains a docs/known_issues.md and asks users to check it and the issue tracker before filing a new issue. That file is the honest place to look for failure modes, and it is more current than any summary. If a run fails in a way you cannot explain, the README's instruction is to check there first.
Alternatives and how they differ
The README names two alternatives directly, and the difference is deployment model rather than method. The first is alphafoldserver.com, described as available for non-commercial use with a more limited set of ligands and covalent modifications. You give up the local control, the scripting, and part of the chemistry coverage, and in exchange you skip the GPU, the database download and the parameter handling.
The second is the Gemini Enterprise Agent Platform on Google Cloud, which the README lists for commercial use. That is the path for organisations whose licence situation rules out the research terms attached to the downloadable weights.
Both alternatives run the same underlying model, according to the README's framing. The choice is about where the compute sits, who holds the weights, and which terms apply. There is no third option in the repository, and the README does not position the local pipeline as a replacement for either hosted service.
Maintenance, upgrades and what a version bump costs
The repository is not archived, and the last push was on 2026-09-21. Recent releases run v3.0.2 on 2026-04-20, v3.0.3 on 2026-06-09, and v3.0.4 on 2026-07-28, with pyproject.toml carrying `fallback_version = "3.0.4"` to match. That is a steady cadence of point releases rather than a frozen snapshot.
The upgrade cost is not trivial, because the dependency pins are tight. `jax==0.10.2`, `dm-haiku==0.0.17`, `rdkit==2025.9.4` and `tokamax==0.0.12` are exact versions, not ranges. The project declares supported environments in `[tool.uv]`: Linux x86_64, Linux aarch64, and macOS arm64. A macOS arm64 environment pulls `jax-mps==0.10.9` instead of the CUDA build, and the CUDA extra is conditional on `sys_platform != 'darwin'`. So an upgrade that moves JAX will move everything downstream of it, and you should expect to rebuild rather than just bump a number.
The wheel build is also constrained: `[tool.cibuildwheel]` builds `cp3*-manylinux_x86_64` against the `manylinux_2_28` image, and the Python requirement is `>=3.12`. If your cluster runs an older Python or an older glibc, that is a blocker before any model question comes up. There is an optional `gcsfs` extra, enabled with `uv sync --extra gcsfs`, for reading and writing inputs and outputs on GCS paths.
Editorial conclusion
Adopt AlphaFold 3 if you need to run the inference pipeline on your own hardware under the weights terms of use, and you already have the databases and a GPU. Do not adopt it if you want a hosted, no-setup path: the repository points to alphafoldserver.com for non-commercial use and to the Gemini Enterprise Agent Platform on Google Cloud for commercial use. Before you commit, verify three things: that you can download af3.bin.zst and accept the weights terms of use, that your input JSON passes the schema in docs/input.md, and that your dataset download completed via fetch_databases.sh. The pipeline is only as usable as those three checks.
Frequently asked questions
Is AlphaFold 3 free?
The code is Apache-2.0 licensed, but the model parameters are separate: the README states you may only use AlphaFold 3 model parameters if received directly from Google, subject to the weights terms of use. Non-commercial use is also offered at alphafoldserver.com, and commercial use through the Gemini Enterprise Agent Platform on Google Cloud.
Can I run AlphaFold 3 locally?
Yes. The README gives a docker run command that mounts input, output, model parameter and database directories, with a GPU passed through via --gpus all. The data pipeline stage is CPU-only, while the inference stage requires a GPU.
How do I install AlphaFold 3?
The README does not repeat the steps at the top level; it points to docs/installation.md. The repository also includes fetch_databases.sh for the reference databases and a docker/ directory for the container build.
How do I use AlphaFold 3?
You write an input JSON file in the alphafold3 dialect, then run python run_alphafold.py with --json_path, --model_dir and --output_dir. The README's test case is a two-chain file named 2PV7, and run_alphafold.py --help lists the remaining flags.
Is AlphaFold 3 open source?
The inference pipeline code is released under Apache-2.0, and the repository is not archived. The model parameters are not covered by that licence; they come with their own terms of use and prohibited use policy documents in the repository root.
What does AlphaFold 3 do?
It predicts biomolecular structures and interactions. The README describes the package as an implementation of the AlphaFold 3 inference pipeline, and the examples cover protein complexes, DNA, RNA, modified residues and small-molecule ligands.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/google-deepmind-alphafold3)