microsoft/Graphormer: a molecular modeling backbone built on the Transformer
Graphormer is a general-purpose deep learning backbone for molecular modeling.
At a glance
- What is it?
- Graphormer packages a Transformer-based molecular modeling backbone with PyG, DGL, OGB and OCP interfaces, plus pre-trained PCQM4M and PCQM4Mv2 checkpoints. It is a research toolkit, not a plug-in prediction service.
- Who is it for?
- Adopt Graphormer if you are a researcher or engineer who can train and evaluate a molecular model yourself and wants a Transformer backbone with OGB and OCP data paths already wired in. Do not adopt it if you need a hosted prediction API or a maintained production service; the README points to Azure Quantum Elements for advanced pre-trained versions.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What microsoft/Graphormer actually is
Graphormer is a deep learning package for training custom models on molecule modeling tasks. The README frames the goal as accelerating research and application in AI for molecule science, naming material discovery and drug discovery as target areas. The repository is written in Python and released under the MIT licence.
The important word is training. This is not a service that takes a SMILES string and returns a property. It is a backbone plus training scripts, and the README directs readers to the online documentation at graphormer.readthedocs.io for getting started, training new models, and extending the code with new model types and tasks. The examples directory shows command line usage of common tasks, split into customized_dataset, oc20 and property_prediction.
The audience is narrow and technical: researchers who already have a dataset and a training budget, and who want the Transformer architecture rather than a message-passing graph neural network. If that description does not fit you, the rest of this article will mostly explain why the project is the wrong shape for your problem.
The Transformer-as-graph-encoder idea behind the architecture
The design question Graphormer answers is whether a Transformer, which has no built-in notion of graph structure, can be made to perform well on graph representation tasks. The NeurIPS 2021 paper listed in the citation block is titled Do Transformers Really Perform Badly for Graph Representation?, and that is the intellectual core of the project.
Rather than replacing attention with message passing, Graphormer keeps the Transformer backbone and injects structural information into it. The README highlights a fairseq backbone, which is the training framework the code is built around. Data enters through interfaces for PyG, DGL, OGB and OCP, so the graph datasets from those ecosystems can be fed into the same training loop.
The practical consequence is that the model, code and script used in the Open Catalyst Challenge are part of the repository, along with pre-trained models on PCQM4M and PCQM4Mv2. The README states that more pre-trained models are coming, which is a statement about intent rather than a current inventory. Treat the checkpoint list as something to verify rather than assume.
Installing Graphormer and running a first training job
The README gives exactly one installation instruction, under the heading Setup with Conda, and it is a single shell script. The script name is install.sh and it sits at the top level of the repository alongside the graphormer and distributional_graphormer directories.
bash install.shThat is the whole documented setup. There is no pip package name given in the README, no version pin, and no separate CUDA instructions. If the script fails, the README offers no troubleshooting path, so the online documentation is the next place to look.
For a first real use, the README points at the examples directory rather than walking through a command. The three example folders are customized_dataset, oc20 and property_prediction, and the README describes them as showing command line usage of common tasks. The repository layout is the only guide to what lives in them, since the README does not print expected output, loss values or accuracy numbers for any example. That absence is worth knowing before you start, because it means you will be judging success against your own expectations rather than the project's.
Where Graphormer stops and Azure Quantum Elements begins
The README carries a line that matters more than its size suggests: advanced pre-trained versions of Graphormer are available exclusively on Azure Quantum Elements. The word exclusively is doing real work there.
This means the open repository and the best-performing checkpoints are not the same artifact. What you can download and train here is the v2.0 code, the Open Catalyst Challenge model and scripts, and pre-trained models on PCQM4M and PCQM4Mv2. What you cannot get from this repository is whatever Microsoft considers its advanced pre-trained tier. For a team hoping to skip training and start from the strongest available weights, that boundary is the first thing to check.
It also shapes the maintenance story. The last push to the repository was on 2026-06-12, which is recent enough that the code is not abandoned, but the two most recent releases both date to 2024-04-03: pre-release and dig-v1.0, the Distributional Graphormer release. The README's own What's New section stops in March 2022. So the repository is receiving commits while its changelog and release notes have gone quiet, and the newest named release belongs to a different subproject than the v2.0 toolkit described at the top of the README.
Two toolkits in one repository: Graphormer v2.0 and Distributional Graphormer
The top-level layout contains both graphormer/ and distributional_graphormer/ as sibling directories, and the newest release is tagged dig-v1.0 for Distributional Graphormer. A reader arriving from a search for the NeurIPS paper may not realize the repository now hosts two related but distinct lines of work.
The v2.0 material is what the README's highlights describe: Open Catalyst Challenge code, PCQM4M pre-trained models, PyG, DGL, OGB and OCP interfaces, and a fairseq backbone. The Distributional Graphormer line is referenced only through the directory name and the release tag, since the README does not explain it in the body text. Anyone evaluating the project should decide which of the two they actually need before installing, because the documentation and examples are organized around the v2.0 training workflow.
This split is a mild usability problem rather than a defect. It does mean that a search result pointing at dig-v1.0 and a search result pointing at the v2.0 highlights can land on the same repository page and describe different software.
Limits, failure modes, and when to pick a different model
The most concrete limitation is documentation depth. The README's installation section is one line. It does not document supported Python versions, CUDA versions, GPU memory requirements, expected training time, or how to recover from a failed install. The readthedocs site is named as the primary documentation, which shifts the burden there, but nothing in the README states how current that site is relative to the code.
A second limit is the checkpoint gap described above. If your plan depends on starting from the strongest pre-trained weights, the README explicitly reserves those for Azure Quantum Elements, and the open checkpoints cover PCQM4M and PCQM4Mv2 along with the Open Catalyst work.
A third is the maintenance signal. Commits land, but the changelog has not been updated since March 2022 and the release list has been static since April 2024. For a research dependency that is tolerable. For a component in a pipeline you expect someone else to keep working, it is a risk you should price in.
On the alternative side, the natural comparison is a message-passing GNN built on PyG or DGL directly. The difference in approach is architectural: a GNN encodes structure through neighborhood aggregation, while Graphormer keeps the Transformer and injects structural information into the attention mechanism. If your graphs are small and your dataset is modest, a standard GNN is simpler to train and debug. Graphormer's case rests on the large-scale molecular benchmarks where the Transformer formulation was shown to compete, and if your problem does not look like PCQM4M or OC20, that case is weaker.
Licence and the cost of staying current
The repository is MIT licensed, which is permissive and places few restrictions on reuse in commercial or closed work. Note that the README also carries a trademarks section stating that use of Microsoft trademarks and logos in modified versions must not imply Microsoft sponsorship, and that contributions require a Contributor License Agreement. Those are project governance terms rather than licence terms, and they do not change the MIT grant on the code. This is a description of what the repository states, not legal advice; read LICENSE and the trademark section yourself before shipping a product.
The upgrade cost is the more practical question. Because there is no pip package documented in the README, the install path is install.sh against a cloned repository, which means your build is tied to a commit rather than a version number. The dependency chain runs through fairseq and the PyG, DGL, OGB and OCP interfaces, and the README does not pin versions for any of them. Upgrading PyTorch or CUDA therefore becomes a manual exercise with no documented compatibility matrix to consult. Budget for that before you build anything on top.
Editorial conclusion
Adopt Graphormer if you are a researcher or engineer who can train and evaluate a molecular model yourself and wants a Transformer backbone with OGB and OCP data paths already wired in. Do not adopt it if you need a hosted prediction API or a maintained production service; the README points to Azure Quantum Elements for advanced pre-trained versions. Before committing, verify the install.sh path against your CUDA and PyTorch versions, confirm which pre-trained checkpoints are actually downloadable, and check whether the dig-v1.0 Distributional Graphormer code fits your use case better than the v2.0 training scripts.
Frequently asked questions
What is microsoft/Graphormer used for?
It is a deep learning package for training custom models on molecule modeling tasks, aimed at material discovery and drug discovery research. The README describes it as a general-purpose backbone, with examples covering property prediction and the Open Catalyst dataset.
How does Graphormer differ from a graph attention network?
Graphormer keeps a Transformer backbone and injects structural information into it rather than relying on neighborhood aggregation. The NeurIPS 2021 paper cited in the repository is titled Do Transformers Really Perform Badly for Graph Representation?, which frames the comparison directly.
Are graph neural networks still used?
The repository does not discuss the broader state of graph neural networks. What it does show is that the PyG, DGL, OGB and OCP interfaces are supported, so Graphormer is designed to plug into the existing graph learning ecosystem rather than replace it.
Who invented the graph attention network?
The repository does not name the inventor of the graph attention network. Its citation block credits the Graphormer authors for the NeurIPS 2021 paper Do Transformers Really Perform Badly for Graph Representation? and the benchmarking technical report.
What is a graph neural network (GNN)?
The README does not define graph neural networks. It positions Graphormer against that class of model by keeping a Transformer backbone and adding structural information, with PyG and DGL interfaces for graph data.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-graphormer)