PaddleHelix: Baidu's Bio-Computing Toolkit for Molecular and Protein Models
Bio-Computing Platform Featuring Large-Scale Representation Learning and Multi-Task Deep Learning “螺旋桨”生物计算工具集
At a glance
- What is it?
- PaddleHelix bundles graph learning for molecules, drug-target and protein-protein interaction models, and the HelixFold family of structure predictors under one PaddlePaddle repository. Its breadth is real, but the install path and the licence terms both need reading before you commit.
- Who is it for?
- PaddleHelix fits teams already running PaddlePaddle who want pretrained compound and protein models, plus a HelixFold3 code path, inside one repository. Teams on PyTorch, or anyone who needs a permissive, clearly stated licence for a commercial product, should verify the LICENSE file and the HelixFold3 non-commercial wording before writing code against it.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What PaddleHelix actually covers
PaddleHelix is a collection of bio-computing models and pipelines built on PaddlePaddle, Baidu's deep learning framework. The repository topics list the intended scope: molecular representation learning, drug-drug interaction, drug-target interaction, protein-protein interaction, protein folding, docking, molecule design and RNA structure prediction. The apps/ directory is the map. It holds pretrained_compound for molecular property models such as GEM and GEM-2, protein_folding for the HelixFold family, molecular_docking for HelixDock, drug_target_interaction for BatchDTA, and protein_protein_interaction for a multimodal PPI model.
The audience is computational chemistry and structural biology groups that already accept PaddlePaddle as their framework. If your stack is PyTorch, the models here are not drop-in: the pretrained weights are Paddle checkpoints, and the training loops use Paddle APIs. The payoff for accepting that constraint is that a single repository carries both small-molecule representation models and a full AlphaFold-style structure prediction pipeline, with published papers behind most of them.
How the toolkit is layered
There are two distinct layers. The pahelix/ package is the reusable library: graph featurisation, data loaders and model layers that the apps import. setup.py builds a native extension into pahelix/toolkit/ through CMake, which is why the repository carries a CMakeLists.txt and a c/ directory at the top level. That extension is compiled, not pure Python, and the build script maps Windows platform specifiers to CMake generator arguments, so the project expects to be built on Linux, Windows or macOS from source rather than installed as a wheel.
The apps/ directory is the second layer: self-contained research pipelines, each with its own README, configuration and pretrained parameters. HelixFold and HelixFold-Single sit under apps/protein_folding, HelixFold3 under apps/protein_folding/helixfold3, HelixDock under apps/molecular_docking/helixdock. They share the library but not a common entry point. You do not run one binary; you enter the directory of the model you want and follow its instructions.
Installing PaddleHelix and running a first model
The repository points to installation_guide.md at the top level rather than putting steps in the README, so that file is the place to start. The setup.py build is the part most likely to bite: it invokes CMake and raises a RuntimeError with the message "CMake missing - probably upgrade to a newer version of Pip?" when the cmake binary is absent. Install CMake before anything else.
git clone https://github.com/PaddlePaddle/PaddleHelix.git
cd PaddleHelix
git checkout dev
pip install -r requirements.txt
python setup.py installThe checkout targets dev, the default branch. The install step compiles the native extension and places it under pahelix/toolkit/. If that directory is empty afterwards, the CMake build did not run and the graph featurisation calls will fail at import time.
The README gives no single canonical training command. Each model has its own instructions, so pick one app and read its README. HelixFold3, for example, is documented at apps/protein_folding/helixfold3, and the repository states that its initial release is available as open source for non-commercial academic research. The tutorials/ and research/ directories hold worked examples if you want to see the library used before committing to a full pipeline.
HelixFold3, HelixFold-Single and the licence boundary
The HelixFold line is the most visible part of the project. HelixFold is a reproduction of the AlphaFold 2 inference pipeline in PaddlePaddle, released in January 2022 with training and inference later fully opened. HelixFold-Single is described as MSA-free, relying only on primary sequences, and the README claims structures within seconds. HelixFold3 replicates AlphaFold3 capabilities, and the 2025.07.23 news entry announces HelixFold3.2 with improvements in protein-related tasks and structural quality.
Read the licensing carefully before planning a product around any of this. The repository's licence field is NOASSERTION, meaning GitHub could not classify the LICENSE file automatically, and the HelixFold3 announcement explicitly frames the open source release as being for non-commercial academic research. A paid API for academic and commercial use is mentioned in the 2024.11.08 entry, with a usage guide on the PaddleHelix site. Those two facts together suggest the open code and the commercial path are separate, but the README does not spell out the boundary. Check LICENSE and the HelixFold3 directory before assuming either way.
Where PaddleHelix is the wrong tool
The release cadence is the first warning. The most recent tagged release is v1.2.2 from 2023-08-01, with v1.1.0 in 2021 and v1.0 in 2021. Commits continue on dev, the last push was on 2026-03-31, but versioned releases are not how this project communicates change. If your dependency policy requires pinned, versioned artefacts, PaddleHelix does not offer them at the pace you may expect.
The second issue is framework lock-in. Every model here is a PaddlePaddle model. Porting a HelixFold checkpoint to PyTorch is not a configuration change; it is a reimplementation. Teams with existing PyTorch training infrastructure and no appetite for a second framework should treat PaddleHelix as a reference implementation to read, not a dependency to add.
The third is scope discipline. The repository is a research monorepo. There is no unified CLI, no shared configuration schema across apps, and no documented rollback or upgrade path between model versions. The README does not document migration between HelixFold3 and HelixFold3.2. Treat each app as a separate project with its own setup cost.
Alternatives and how they differ
OpenFold3 and Protenix appear in the search terms people use around this project, and both are plausible comparison points for the structure prediction side, though this repository does not benchmark against them. The practical difference is framework and packaging. PaddleHelix couples structure prediction to a broader bio-computing toolkit built on PaddlePaddle, and its HelixFold3 code is published for non-commercial academic research with a separate paid API for other uses. A team that only wants protein structure prediction and already runs PyTorch is likely to find a narrower, single-purpose predictor easier to integrate than a monorepo that also carries molecular property models and docking pipelines.
The reverse also holds. If you need small-molecule representation learning, drug-target affinity and structure prediction under one roof, with papers behind each model, PaddleHelix is unusually broad. That breadth is the reason to choose it, and the reason its install path is heavier than a single-model repository.
Maintenance and upgrade cost
The last push to the repository was on 2026-03-31, so development has not stopped. But the tagged releases tell a different story: v1.2.2 in August 2023, then nothing. New work arrives as news entries and directory changes, such as HelixFold3.2 under apps/protein_folding/helixfold3, not as version bumps you can pin against.
That has a concrete cost. Upgrading means tracking the dev branch and reading the app-level README for whichever model you use, because there is no changelog entry describing what changed between HelixFold3 and HelixFold3.2 beyond the qualitative claim of improved protein tasks and structural quality. If you vendor the code, record the commit hash you built from; the branch name alone will not identify a state.
On licensing, the NOASSERTION classification plus the non-commercial wording around HelixFold3 means the terms are not something to infer from the repository metadata. Read LICENSE and the HelixFold3 directory. This is a factual observation about what the repository states, not legal advice, and a commercial deployment warrants its own review.
Editorial conclusion
PaddleHelix fits teams already running PaddlePaddle who want pretrained compound and protein models, plus a HelixFold3 code path, inside one repository. Teams on PyTorch, or anyone who needs a permissive, clearly stated licence for a commercial product, should verify the LICENSE file and the HelixFold3 non-commercial wording before writing code against it. Start by reading installation_guide.md and building the pahelix toolkit extension, because the setup.py CMake build is the first thing that can fail.
Frequently asked questions
Is protein folding really solved?
The PaddleHelix README does not make that claim. It describes HelixFold as a reproduction of the AlphaFold 2 inference pipeline in PaddlePaddle, HelixFold-Single as an MSA-free pipeline that relies only on primary sequences, and HelixFold3 as replicating AlphaFold3 capabilities with accuracy comparable to AlphaFold3 on conventional ligands, nucleic acids and proteins.
What is the best software for protein structure prediction?
This repository does not rank prediction tools. What it offers is the HelixFold family under apps/protein_folding: HelixFold for AlphaFold 2 style inference, HelixFold-Single for MSA-free prediction from primary sequences, and HelixFold3 for biomolecular structures, with HelixFold3.2 announced in the 2025.07.23 news entry.
Has AI ever solved the protein folding problem?
The README does not answer that question directly. It presents HelixFold3 as achieving accuracy comparable to AlphaFold3 for conventional ligands, nucleic acids and proteins, and notes that the HelixFold3 code and model parameters were released under a non-commercial academic research framing.
How to predict protein structure?
With PaddleHelix you go to apps/protein_folding and pick a pipeline. HelixFold covers AlphaFold 2 style inference, HelixFold-Single predicts from primary sequences without an MSA, and HelixFold3 handles biomolecular structures. Each app directory carries its own README and parameters.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/paddlepaddle-paddlehelix)