# MatterGen: Microsoft's diffusion model for inorganic crystal design, and how to run it

> MatterGen generates inorganic crystal structures across the periodic table and can be steered toward target properties such as band gap or magnetic density. It installs as a Python package with CUDA wheels, and its checkpoints are pulled from Hugging Face.

**microsoft/mattergen** — Official implementation of MatterGen -- a generative model for inorganic materials design across the periodic table that can be fine-tuned to steer the generation towards a wide range of property constraints.

- Repository: https://github.com/microsoft/mattergen
- Website: https://www.nature.com/articles/s41586-025-08628-5
- Stars: 1,835 · Forks: 352
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/microsoft-mattergen

## What problem MatterGen solves, and who it is for

Searching for a new inorganic solid with a specified property is normally a loop: propose a composition, guess a structure, relax it, compute the property, repeat. MatterGen inverts part of that loop. Instead of screening candidates, it samples crystal structures directly from a learned distribution, and the distribution can be shifted toward a property constraint. The README describes it as "a generative model for inorganic materials design across the periodic table that can be fine-tuned to steer the generation towards a wide range of property constraints."

The intended user is a researcher in solid-state chemistry, physics or materials informatics who is comfortable with Python, has a CUDA GPU, and already thinks in terms of CIF files and pymatgen Structure objects. The output is not a finished material. It is a set of candidate structures, in CIF and extxyz form, that you still have to judge. If you want a tool that predicts a property for a structure you already have, this is the wrong direction of travel: MatterGen proposes structures, it does not score yours.

## Diffusion over crystal structures, and the conditioning mechanism

The core is a denoising diffusion model operating on crystal representations rather than images. The README exposes this directly through the output options: when --record-trajectories is left at its default of True, the run writes generated_trajectories.zip containing one .extxyz file per generated structure holding "the full denoising trajectory for each individual structure." That trajectory is the generation process itself, frame by frame, which is unusually transparent for a generative model and useful when you want to see how a structure settled.

Conditioning is handled by fine-tuned checkpoints plus a guidance factor at sampling time. The README lists nine checkpoints: mattergen_base and mp_20_base unconditionally trained on Alex-MP-20 and MP-20 respectively, then chemical_system, space_group, dft_mag_density, dft_band_gap, ml_bulk_modulus, and two jointly conditioned models, dft_mag_density_hhi_score and chemical_system_energy_above_hull. The joint models matter because single-property conditioning is easy to satisfy in unphysical ways; asking for both a chemical system and an energy above the hull constrains the sample toward something more plausible.

The knob that controls how hard the model pushes toward the target is --diffusion_guidance_factor, which the README maps to the gamma parameter in classifier-free diffusion guidance. The README's own description of the trade-off is worth quoting in full: setting it to zero gives unconditional generation, and "increasing it further tends to produce samples which adhere more to the input property values, though at the expense of diversity and realism of samples." That is the central tension of the whole tool. A high guidance factor gets you numbers close to your target; it does not get you a material.

## Installing MatterGen and generating your first crystals

The README recommends uv and assumes Linux with a CUDA GPU. From a clone of the repository, the editable install path looks like this:

```bash
pip install uv
uv venv .venv --python 3.10
source .venv/bin/activate
uv pip install -e .
```

If you prefer the released package rather than the working tree, the README gives a PyPI route with matching pre-built PyTorch Geometric CUDA wheels:

```bash
uv pip install mattergen --find-links https://data.pyg.org/whl/torch-2.2.0+cu121.html
```

Datasets and checkpoints live in the repository through Git Large File Storage, so a plain clone can leave you with pointer files instead of weights. Check before anything else:

```bash
git lfs --version
```

If that prints nothing, install it with sudo apt install git-lfs followed by git lfs install, or pull an individual checkpoint with git lfs pull -I checkpoints/<model_name> --exclude="".

Unconditional sampling is one command. The README's example is:

```bash
export MODEL_NAME=mattergen_base
export RESULTS_PATH=results/
mattergen-generate $RESULTS_PATH --pretrained-name=$MODEL_NAME --batch_size=16 --num_batches 1
```

You should end up with generated_crystals_cif.zip (one .cif per structure), generated_crystals.extxyz (all structures as frames in one file), and, because trajectory recording defaults to on, generated_trajectories.zip. To condition on a property, switch the checkpoint and pass the target. The README's magnetic density example:

```bash
export MODEL_NAME=dft_mag_density
export RESULTS_PATH="results/$MODEL_NAME/"
mattergen-generate $RESULTS_PATH --pretrained-name=$MODEL_NAME --batch_size=16 --properties_to_condition_on="{'dft_mag_density': 0.15}" --diffusion_guidance_factor=2.0
```

The batch size is the practical lever: the README advises raising it to the largest your GPU can hold without an out-of-memory error. Passing --seed=42 resets the Python, NumPy and PyTorch random generators, though the README notes exact reproducibility is only expected with identical software and hardware because some CUDA operations are nondeterministic.

## Where MatterGen breaks down

The most consequential limitation is stated plainly in the README and is easy to miss: the released checkpoints "were re-trained using this repository, i.e., are not identical to the ones used in the paper. Hence, results may slightly deviate from those in the publication." If your plan is to reproduce a number from the Nature paper, the artifacts shipped here are not the artifacts that produced it. You are evaluating a re-trained model, and the README does not quantify the deviation.

The second limitation is the guidance factor. Because the README frames higher values as trading diversity and realism for adherence, a conditional run with an aggressive factor can return structures that match the requested property while being chemically odd. Nothing in the README describes a built-in physical validity filter, so that judgement is yours, and it is the expensive part.

Third, platform support. The install instructions assume Linux with CUDA, and the Apple Silicon note carries an explicit warning that the platform is experimental, requiring export PYTORCH_ENABLE_MPS_FALLBACK=1 before any training or generation run. The dependency set reinforces this: torch, torchvision and torchaudio are pinned to Linux-only markers in pyproject.toml. There is no documented CPU-only path, so a laptop without a supported GPU is not a realistic target.

Finally, the README is silent on several operational questions. It does not document a rollback procedure for a fine-tune that degrades the base model, and it does not describe how to resume an interrupted training run. If your workflow depends on either, budget time to read the training code rather than the documentation.

## MatterGen versus MatterSim

The comparison people reach for is MatterSim, and the two are complementary rather than competing. MatterSim appears in the dependency list as mattersim>=1.1, so MatterGen already depends on it. MatterSim is a machine-learned interatomic potential: you give it a structure and it gives you energies and forces, which is what you need for relaxation and property evaluation. MatterGen is a generator: you give it a constraint and it gives you structures that did not exist before.

The practical consequence is that they sit at opposite ends of the same pipeline. MatterGen proposes, MatterSim (or DFT) evaluates. Choosing between them is really choosing which end of the loop you are automating. If your bottleneck is that you have more candidate structures than you can afford to relax, MatterSim is the answer. If your bottleneck is that you cannot think of enough plausible candidates to begin with, MatterGen is the answer. The ml_bulk_modulus checkpoint is a useful illustration of the split: it is conditioned on a property predicted by a machine-learned model rather than by DFT, so the target you steer toward inherits whatever error that predictor carries.

## Licence, maintenance and the cost of upgrading

The project is MIT licensed, per both the LICENSE file and the license field in pyproject.toml, which is permissive and places few obligations on how you redistribute or modify it. Two caveats sit outside the code licence and are worth checking yourself rather than assuming: the model checkpoints and datasets are separate artifacts distributed via Git LFS and Hugging Face, and the repository ships a MODEL_CARD.md and a NOTICE file, which is where any additional terms would be recorded. This is not legal advice; if you plan to ship something commercial built on the weights, read those files.

The last push to the repository was on 2026-08-27, and the most recent tagged release is v1.0.3 from 2025-07-23. Those two facts point in different directions: the code is still being touched, but the packaged version you would pin has not moved in over a year. For an upgrade, that means pip install mattergen gives you 1.0.3 while the default branch may contain unreleased changes, so pinning the version and installing from PyPI is the more predictable path than tracking main.

The dependency list is where upgrade cost actually lives. numpy is pinned below 2.0, pytorch-lightning is pinned at 2.0.6, hydra-core at 1.3.1, and torch at 2.2.1 with Linux-only markers. Several of these pins are deliberate and annotated as such in pyproject.toml, for instance the numpy pin is commented as protecting against breaking changes in 2.0. Moving any one of them forward is not a version bump, it is a compatibility exercise across the whole stack.

## Conclusion

Adopt MatterGen if you are a computational materials researcher who already works with pymatgen and ASE structures and has a CUDA GPU available; the property-conditioned checkpoints let you test whether a target band gap or magnetic density is reachable before committing to a synthesis or DFT campaign. Do not adopt it if you need a CPU-only pipeline, if you expect the published paper's numbers to reproduce exactly (the README states the released checkpoints were re-trained and may deviate), or if you want a pretrained interatomic potential for a fixed composition rather than a generator. Before building anything on it, verify three things: that git lfs --version prints a version so the checkpoint files actually materialize after cloning, that --diffusion_guidance_factor values you choose still produce structures you consider chemically sensible rather than merely close to the target property, and that your evaluation path handles the CIF and extxyz outputs the same way your existing analysis scripts do.

## FAQ

### How do you install MatterGen?

The README recommends uv on Linux with a CUDA GPU: create a Python 3.10 virtual environment, activate it, then run uv pip install -e . from the repository. Alternatively, install the released package from PyPI with uv pip install mattergen --find-links https://data.pyg.org/whl/torch-2.2.0+cu121.html.

### How do you use MatterGen?

After installing, set MODEL_NAME and RESULTS_PATH and run mattergen-generate with --pretrained-name and --batch_size. For property-conditioned generation, pass --properties_to_condition_on as a dictionary and adjust --diffusion_guidance_factor, as in the README's magnetic density example targeting 0.15.

### What is MatterGen?

It is a generative model for inorganic materials design across the periodic table, described in the README as fine-tunable to steer generation toward a wide range of property constraints. It is the official implementation from Microsoft and is linked to a Nature paper.

### Is MatterGen open source?

Yes. The repository is public under the MIT licence, and both the LICENSE file and the license field in pyproject.toml record that. Note that the checkpoints and datasets are distributed separately through Git LFS and Hugging Face.

### What is the difference between MatterGen and MatterSim?

MatterSim is a machine-learned interatomic potential that evaluates structures, and it appears as a dependency of MatterGen. MatterGen generates new crystal structures from a learned distribution, optionally steered by a property constraint. In practice MatterGen proposes candidates and MatterSim or DFT evaluates them.

## Sources

- [License: MIT](https://github.com/microsoft/mattergen/blob/main/LICENSE)
- [microsoft/mattergen on GitHub](https://github.com/microsoft/mattergen)
- [Project website](https://www.nature.com/articles/s41586-025-08628-5)
- [README](https://github.com/microsoft/mattergen/blob/main/README.md)
- [Releases](https://github.com/microsoft/mattergen/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/microsoft-mattergen
