Open-source project
lucidrains/alphafold3-pytorch avatar
lucidrains/alphafold3-pytorch

alphafold3-pytorch: A Community PyTorch Implementation of AlphaFold 3

Implementation of Alphafold 3 from Google Deepmind in Pytorch

1,695 stars232 forksPythonMIT

At a glance

What is it?
alphafold3-pytorch is a community reimplementation of Google DeepMind's AlphaFold 3 protein structure prediction system in PyTorch, maintained by Phil Wang and contributors. It is not the official DeepMind code, and using it for research requires downloading and processing the full PDB dataset independently.
Who is it for?
alphafold3-pytorch is the right tool for researchers who need to study, modify, or extend AlphaFold 3's architecture in PyTorch without waiting for an official open-source release of DeepMind's training code. It is not suitable for production structure prediction without the full PDB dataset and substantial GPU infrastructure, and the repository is a reimplementation from the published paper rather than a port of DeepMind's internal code.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 51 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What This Repository Implements and For Whom

alphafold3-pytorch is a community implementation of the AlphaFold 3 architecture as described in the Nature paper published by Google DeepMind. The repository is maintained by Phil Wang (lucidrains) and a set of named contributors who have each added specific modules, dataset utilities, or bug fixes documented in the README's Appreciation section.

The target audience is AI and bioinformatics researchers who want to study, reproduce, or extend AlphaFold 3's architecture in PyTorch. AlphaFold 3 predicts the structure of proteins and their interactions with other molecules including DNA, RNA, and small molecules. The original paper, linked in the README, describes an architecture based on a diffusion model for atom coordinates rather than the end-to-end differentiable pipeline used in AlphaFold 2.

This is not an official DeepMind or Google release. Google DeepMind has released AlphaFold 3 as a server for prediction tasks, but this repository is a separate effort by the open-source community to reconstruct the training code from the paper. Anyone using it for published research should cite the original paper using the BibTeX entry included in the README.

The project is licensed under MIT and its current version in pyproject.toml is 0.8.3. The last push was on 2026-08-10.

Architecture: How Atom-Level Diffusion Differs from Sequence-Level Prediction

AlphaFold 2 worked at the residue level, predicting torsion angles and positions for amino acid residues in a protein chain. AlphaFold 3, as described in the paper and reflected in this codebase, works at the atom level and uses a diffusion process to predict three-dimensional atom coordinates directly.

The codebase reflects this design. The core input types include `atom_inputs` (per-atom features), `atompair_inputs` (pairwise features between atom pairs), `molecule_atom_indices`, `molecule_atom_lens`, and `Alphafold3Input`, which is a higher-level class that accepts protein sequences, atom positions, and molecule definitions. The `Alphafold3Input` class is one of the main usability additions in this community implementation: it handles the translation from sequence-level inputs to atom-level tensors.

Contributors from the README have added specific algorithmic pieces: Relative Positional Encoding and Smooth LDDT Loss by Joseph Kim; Weighted Rigid Align, Express Coordinates In Frame, Compute Alignment Error, and Centre Random Augmentation by Felipe Engelberger; confidence measures, clash penalty ranking, and sample ranking logic by xluo233; and the model selection score and unresolved RASA computation also by xluo233. The logic for translating `Alphafold3Input` to `BioMolecule` and saving to mmCIF was contributed by Dhuvi.

For users who want Triton kernel optimization, the README points to a separate project called MegaFold, which the README describes as an optimized version leveraging Triton kernels. A fork with full Lightning and Hydra support is maintained by Alex Morehead at a separate repository.

Installing alphafold3-pytorch and Running a First Pass

The simplest installation path is pip:

bash
pip install alphafold3-pytorch

The README also provides a Dockerfile for containerized use. The Dockerfile starts from a PyTorch base image with CUDA 12.1 and cuDNN 8:

bash
git clone https://github.com/lucidrains/alphafold3-pytorch . --branch main
python -m pip install .

The dependency list in pyproject.toml is long. Direct dependencies include torch (2.1 or later), biopython (1.83 or later), rdkit (2023.9.6 or later), einops, einx, gemmi, fair-esm, lightning, pydantic, polars, scipy (pinned to 1.13.1), and gradio for the frontend. On a clean environment, the full install can take several minutes.

The README's usage example shows how to instantiate the model with explicit dimension parameters and run it against mock tensor inputs. The higher-level interface uses `Alphafold3Input`:

python
from alphafold3_pytorch import Alphafold3, Alphafold3Input

train_alphafold3_input = Alphafold3Input(
    proteins = [contrived_protein],
    atom_pos = mock_atompos
)

This example uses mock data. Running the model against real protein structures requires PDB data in mmCIF format.

PDB Dataset Preparation: The Dominant Operational Cost

Training the model requires the Protein Data Bank (PDB) in mmCIF format. The README describes a multi-step dataset curation pipeline with significant infrastructure requirements.

First, download all first-assembly and asymmetric unit complexes from the RCSB PDB. The README recommends using AWS snapshots for reproducibility, citing the snapshot tag `20240101` as an example. The mmCIF files go into `data/pdb_data/unfiltered_assembly_mmcifs/` and `data/pdb_data/unfiltered_asym_mmcifs/`.

Second, run the filter scripts (`filter_pdb_train_mmcifs.py`, `filter_pdb_val_mmcifs.py`, `filter_pdb_test_mmcifs.py`) to remove structures that do not meet the training criteria. Third, run the clustering scripts (`cluster_pdb_train_mmcifs.py`, etc.) to prevent data leakage between train and test sets.

The README does not document the total disk space required for the full PDB download. The PDB is one of the largest public biological databases. The README also does not document the expected wall-clock time for the filtering and clustering steps, though it credits Milot Mirdita for optimizing the clustering script.

This preparation step makes the repository more appropriate for research groups with dedicated compute resources than for individual developers running a laptop.

Differences from Google DeepMind's Official AlphaFold 3 Server

DeepMind's AlphaFold 3 server accepts protein sequences and returns predicted structures through a web interface. It uses DeepMind's internally trained weights. This repository is a PyTorch reimplementation of the architecture from the paper, not a wrapper around DeepMind's server or a port of DeepMind's training code.

The practical consequence is that to run structure predictions with this codebase, you need to train or obtain weights yourself. The paper describes the training procedure, but the README does not document a checkpoint download path for pretrained weights. The README notes in several places that specific hyperparameters were corrected based on inconsistencies found between the paper and early code (Wei Lu is credited with catching several erroneous hyperparameters).

AlphaFold 2, in contrast, has an official open-source implementation from DeepMind that includes pretrained weights and has been widely validated for publication-quality structure prediction. AlphaFold 2 predicts only protein structures, while AlphaFold 3 extends to protein-DNA, protein-RNA, and protein-small molecule complexes. This codebase targets the AlphaFold 3 architecture, but AlphaFold 2's official implementation remains the safer choice for researchers who need validated, reproduced prediction quality.

Maintenance, Community, and Licence

The repository receives ongoing contributions. The README's Appreciation section lists over a dozen named contributors and describes specific modules or bug fixes each added. An optimized fork (MegaFold) and a Lightning plus Hydra fork are actively maintained by separate contributors.

The last push was on 2026-08-10. The pyproject.toml version is 0.8.3, with the most recent GitHub release tagged at 0.7.12 in September 2025. The gap between the pyproject.toml version (0.8.x) and the latest GitHub release (0.7.12) suggests the repository continues to accept patches on main without cutting formal releases on every merge.

The licence is MIT, which permits use in commercial and academic projects without restrictions on redistribution. The `.env.sample` documents environment variables for the training setup: `TYPECHECK=True`, `DEBUG=False`, `DEEPSPEED_CHECKPOINTING=False`, `USE_NIM=False`.

Researchers who want to engage with the community can find a Discord server linked in the README, where discussion about the implementation is hosted.

Editorial conclusion

alphafold3-pytorch is the right tool for researchers who need to study, modify, or extend AlphaFold 3's architecture in PyTorch without waiting for an official open-source release of DeepMind's training code. It is not suitable for production structure prediction without the full PDB dataset and substantial GPU infrastructure, and the repository is a reimplementation from the published paper rather than a port of DeepMind's internal code. The current version in pyproject.toml is 0.8.3, and the last push was on 2026-08-10. Before running training, verify that the PDB dataset preparation scripts (`filter_pdb_*_mmcifs.py` and `cluster_pdb_*_mmcifs.py`) complete successfully on your storage and compute setup.

Frequently asked questions

Can you explain how AlphaFold 3 works and what it can do?

AlphaFold 3 predicts three-dimensional structures of proteins and their interactions with DNA, RNA, and small molecules. It uses a diffusion model to generate atom coordinates directly, rather than predicting residue-level torsion angles as AlphaFold 2 did. This repository implements that architecture in PyTorch from the published Nature paper.

How exactly does AlphaFold work?

In alphafold3-pytorch, the model takes per-atom and pairwise atom features as inputs along with molecule-level data. A diffusion process generates the three-dimensional coordinates of each atom in the complex. Higher-level inputs such as protein sequences are converted to atom-level tensors by the Alphafold3Input class before being passed to the model.

Does alphafold3-pytorch include pretrained weights for inference?

The README does not document a pretrained weight download path. The repository is a reimplementation of the architecture from the published paper, intended for research and training rather than as a drop-in inference tool with ready-to-use weights.

Official sources

  1. Issues
  2. License: MIT
  3. lucidrains/alphafold3-pytorch on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lucidrains-alphafold3-pytorch.svg)](https://hysenlabs.com/projects/lucidrains-alphafold3-pytorch)