ColabFold: AlphaFold2 protein structure prediction without the database grind
Making Protein folding accessible to all!
At a glance
- What is it?
- ColabFold wraps AlphaFold2 and MMseqs2 so that a single sequence can produce a structure in a browser notebook, or on your own GPU through colabfold_batch. Here is what it does, how to install it, and where it stops being the right tool.
- Who is it for?
- ColabFold fits labs that want AlphaFold2-quality structures for monomers and complexes without maintaining terabytes of sequence databases, and it fits them best through the AlphaFold2_mmseqs2 notebook or a local colabfold_batch run. It is the wrong choice if you need guaranteed GPU time for long sequences, if your queries are not serial from a single IP, or if you need a documented rollback path, because the README does not describe one.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The database problem ColabFold removes
AlphaFold2 as originally released expects a large genetic sequence database on local disk, and the search step against that database is what makes a naive install expensive in both storage and time. ColabFold's answer is to replace that search with MMseqs2 and, in the notebook case, to run the search on a shared server. The project describes itself as "Making protein folding accessible to all via Google Colab", and the notebook table in the README is the clearest statement of intent: monomers, complexes, MMseqs2 and templates are all marked Yes for the AlphaFold2_mmseqs2 notebook.
The audience is therefore not the group that already runs AlphaFold2 in production. It is the structural biologist, the small lab, or the student who has a sequence and wants a structure today. The batch notebook and the pip-installable colabfold_batch exist for people who have moved past one-off predictions but still do not want to own the full database stack.
How the MMseqs2 pipeline and the notebooks fit together
There are two delivery paths and they share the same search idea. In the browser path you open a notebook, and the MSA is fetched from the ColabFold MMseqs2 MSA server rather than computed locally. The README is explicit about the etiquette of that shared resource: you may query it from a local computer "if you queries are serial from a single IP", and it asks that you not use multiple computers to query the server. That is a real constraint, not a formality. The server is a shared service and the project treats parallel querying from one user as abuse.
In the local path, the same MMseqs2 binary does the search against databases you download yourself. The README points to colabfold.mmseqs.com for those databases, and the wiki carries an MSA Server Database History page, which matters because a structure predicted against one database version is not automatically reproducible against another. The notebooks also differ in modelling approach: the README notes that AlphaFold2_advanced uses residue index jump while AlphaFold2_mmseqs2 and the batch notebook use the AlphaFold2-multimer model. Those are two different ways to get a complex, and swapping one for the other changes what the model is actually doing.
On top of the AlphaFold2 line, the README lists an AlphaFold3 (OpenFold3) notebook and a set of beta notebooks including RoseTTAFold2, Boltz, BioEmu and OmegaFold. The beta label is the project's own; treat those as moving targets.
Installing ColabFold and running a first batch prediction
The README offers two routes. The one-step installer script is LocalColabFold, which the README says supports Linux, macOS and Windows via WSL2. The alternative is a direct conda and pip install. The conda line pins python=3.13 and mmseqs2=18.8cc5c, and the pip line installs the alphafold and openmm extras together with a CUDA 12 build of JAX and OpenMM:
conda create -n colabfold -c conda-forge -c bioconda python=3.13 mmseqs2=18.8cc5c
conda activate colabfold
pip install colabfold[alphafold,openmm] jax[cuda12] openmm[cuda12]Note that the pyproject file states a wider Python floor, >=3.10, than the conda example uses. If you build from the repository rather than following the README verbatim, that is the constraint the packaging metadata actually enforces.
The Dockerfile is the third route and encodes the same idea for a container. It downloads the MMseqs2 GPU build for amd64 or arm64, installs the package with the alphafold and openmm extras plus a CUDA 12 JAX, and declares a cache volume:
ARG CUDA=cuda12
VOLUME cache
ENV MPLBACKEND=Agg
ENV MPLCONFIGDIR=/cache
ENV XDG_CACHE_HOME=/cache
RUN pip install --no-cache-dir \
".[alphafold,openmm]" \
"jax[${CUDA}]<0.12" \
"openmm[${CUDA}]"The environment variables matter if you run the image read-only or with a small writable layer: matplotlib and the cache directories are all redirected into /cache.
For a first real run, the batch notebook or colabfold_batch takes a FASTA file and produces predicted structures. The README's FAQ gives the practical ceiling: on GPUs with roughly 16 GB the maximum length is about 2000 residues, and the free Colab GPU is described only as "fingers-crossed". If you are working with a complex, decide first whether you want the multimer model or the residue index jump route, because that choice is made before the run, not after.
Where ColabFold breaks down
The length limit is the first wall. Roughly 2000 residues on a 16 GB GPU is a hardware ceiling, and the README ties it to whatever GPU Colab happens to hand you, which means the same input can succeed or fail depending on the day. For a large multi-domain protein or a big assembly, this is not a tuning problem you can solve with flags.
The second wall is the MSA server. Serial queries from a single IP are allowed; parallel queries from several machines are not. Any workflow that fans out across a cluster while pointing at the public server is outside what the project permits, and the honest fix is to download the databases and run MMseqs2 locally, which puts the storage and search cost back on you.
The third issue is reproducibility. Because the MSA server databases change over time, and because the wiki tracks that history separately, a prediction is tied to the database state at the moment it ran. The README does not document a rollback mechanism for pinning a run to a specific database version, so if you need bit-identical reruns across months, plan for local databases and record the version yourself.
Finally, the pLDDT-in-bfactor convention is a trap for downstream tools. The README warns that the bfactor column holds pLDDT values where higher is better, while Phenix.phaser expects a real bfactor where lower is better. Feeding a ColabFold model into molecular replacement without converting that column will mislead the software.
ColabFold compared with running AlphaFold2 or OpenFold directly
The comparison that matters is against DeepMind's own AlphaFold2 notebook, which the README lists in the same table. That notebook uses jackhmmer for the MSA and does not use MMseqs2 or templates. ColabFold's AlphaFold2_mmseqs2 notebook uses MMseqs2 and templates. So the difference is not the folding network, which is AlphaFold2 in both cases, but the search front end and what it costs to run. ColabFold's bet is that MMseqs2 plus a hosted MSA service is a better trade for most users than a full jackhmmer search against local databases.
Against OpenFold, the distinction is similar in spirit but different in kind: ColabFold is a pipeline and packaging effort around existing models, and the README even ships an AlphaFold3 (OpenFold3) notebook alongside the AlphaFold2 ones. It is not trying to be a reimplementation of the network. If your requirement is to modify the model internals, you are looking at the wrong layer of the stack.
Compared with the AlphaFold Server, the difference is control rather than accuracy. ColabFold runs the model on hardware you choose, which is what makes the local install and the batch path meaningful, at the cost of managing databases, GPUs and versions yourself.
Licence and the cost of staying current
The repository is MIT licensed, but the pyproject file states the licence as "MIT, but separate licenses for the trained weights". That distinction is the one to carry into any deployment decision: the code you clone and the weights you download are governed differently, and the README does not restate the weight terms. Read the weight licence at the source before you build a service on top of it. Nothing here is legal advice.
The maintenance picture is concrete. The last push to the default branch was on 2026-09-23, and the most recent release is v1.6.3 from 2026-09-14, preceded by v1.6.2 in July 2026 and v1.6.1 in March 2026. That is a steady release cadence, and the README points to a wiki change log per version.
The upgrade cost is real, though. The conda example pins mmseqs2=18.8cc5c and python=3.13, the packaging metadata allows Python from 3.10, and the Dockerfile pins jax[${CUDA}]<0.12. A version bump in any of those can change your environment, and because the MSA server databases move independently of the code, upgrading ColabFold does not by itself reproduce an older result. Budget for re-validating a known structure after each upgrade rather than assuming the output is stable.
Editorial conclusion
ColabFold fits labs that want AlphaFold2-quality structures for monomers and complexes without maintaining terabytes of sequence databases, and it fits them best through the AlphaFold2_mmseqs2 notebook or a local colabfold_batch run. It is the wrong choice if you need guaranteed GPU time for long sequences, if your queries are not serial from a single IP, or if you need a documented rollback path, because the README does not describe one. Before committing, check the v1.6.3 change log on the wiki, confirm your Python version satisfies the pyproject constraint of >=3.10, and read the separate licence that covers the trained weights rather than assuming the MIT licence on the code applies to everything you download.
Frequently asked questions
What are the key differences between AlphaFold and ColabFold?
ColabFold keeps the AlphaFold2 model but changes the search stage, using MMseqs2 instead of jackhmmer and offering templates in the AlphaFold2_mmseqs2 notebook. DeepMind's own notebook, listed in the same README table, uses jackhmmer and no templates.
Is ColabFold free to use?
The repository is MIT licensed, but the pyproject file notes separate licenses for the trained weights, so the code and the weights are not covered by the same terms. The free Google Colab GPU is described in the README only as "fingers-crossed", which is a statement about availability, not price.
How long does ColabFold take to run?
The README does not give a runtime figure. It gives a length limit instead: on GPUs with about 16 GB the maximum is roughly 2000 residues, and the limit depends on the free GPU Colab provides.
How do I install ColabFold?
The README points to LocalColabFold for a one-step installer covering Linux, macOS and Windows via WSL2, or you can create a conda environment with python=3.13 and mmseqs2=18.8cc5c and then pip install colabfold[alphafold,openmm] with jax[cuda12] and openmm[cuda12]. A Dockerfile is also provided in the repository.
Is ColabFold the same as AlphaFold?
No. ColabFold packages and runs AlphaFold2, and the README also lists an AlphaFold3 (OpenFold3) notebook, but the project itself is the pipeline around those models plus the MMseqs2 search and the MSA server. The folding network comes from elsewhere.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sokrypton-colabfold)