Open-source project
materialyzeai/maml avatar
materialyzeai/maml

maml: Python interfaces for machine learning on crystals and molecules

Python for Materials Machine Learning, Materials Descriptors, Machine Learning Force Fields, Deep Learning, etc.

466 stars98 forksJupyter NotebookBSD-3-Clause

At a glance

What is it?
maml wraps pymatgen, matminer, scikit-learn and TensorFlow into one package for materials feature generation and interatomic potentials. It is a glue layer with a real dependency cost, and the repository layout shows exactly where that cost lands.
Who is it for?
Adopt maml if you already work in the pymatgen and matminer ecosystem and want descriptors plus sklearn or keras model wrappers without writing the glue yourself. Do not adopt it if you need a self-contained potential trainer: maml.pes delegates to LAMMPS, GAP, MLIP and n2p2, so your environment has to carry those external codes.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 65 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap maml fills between materials data and ML libraries

A crystal structure is not a feature vector. Turning a CIF file into something scikit-learn can consume means computing composition averages, site-level statistics and local environment descriptors, and that work is separate from choosing a model. maml positions itself as the high-level interface for that middle layer. The README is explicit about the intent: the package "aims to provide useful high-level interfaces that make ML for materials science as easy as possible," and it states that the goal "is not to duplicate functionality already available in other packages." So maml is a coordination layer, not a new learning algorithm.

The audience follows from that. If you have structures in pymatgen objects and targets in a table, and you want to move between them without hand-rolling descriptor code, maml is aimed at you. If you are building a general-purpose deep learning pipeline on tensors that have nothing to do with atom positions, the package adds dependencies without adding capability.

What sits underneath: pymatgen, matminer, sklearn and keras

The dependency list in pyproject.toml is the architecture. maml depends on numpy, scipy, pymatgen, mp-api, scikit-learn>=1.6.1 and matgl>=3.0.3. Crystal and molecule manipulation comes from pymatgen, feature generation from matminer, and the learning algorithms from scikit-learn and, optionally, TensorFlow through the deep extra.

On top of that base, the README lists three descriptor families that go beyond the standard compositional and structural features: bispectrum coefficients, Behler-Parrinello symmetry functions, and Smooth Overlap of Atom Position (SOAP). Graph network features are listed as a fourth, covering composition, site and structure. These are the fine-grain local environment descriptors, and they are the part of the package that is not simply a re-export.

The applications layer is where the architecture gets interesting. maml.pes models potential energy surfaces and lists four potential types: Neural Network Potential (NNP), Gaussian approximation potential (GAP) with SOAP features, spectral neighbor analysis potential (SNAP), and Moment Tensor Potential (MTP). maml.rfxas builds random forest models that predict atomic local environments from X-ray absorption spectroscopy. maml.bowsr does structural relaxation with Bayesian optimization over a surrogate energy model. Each of these is a thin orchestration layer over external code, and that shapes the installation story.

Installing maml and running a first descriptor job

The base package installs from PyPI in one command. Nothing in the README suggests a source build is required for the core library.

bash
pip install maml

That gets you the feature generation and the sklearn and keras model interfaces. The Python floor is 3.11, set by requires-python in pyproject.toml, and the classifiers list 3.11 and 3.12. If you are on an older interpreter, pip will refuse before you get a confusing import error.

The potential energy surface code needs more than pip. The README states that running pes requires a LAMMPS installation, available from source or from conda:

bash
conda install -c conda-forge/label/cf202003 lammps

The SNAP potential ships with that LAMMPS installation. GAP and MTP are not bundled: the README says the GAP package is needed for GAP and the MLIP package for MTP. Fitting an NNP potential requires the n2p2 package. These are separate installations with their own build requirements, and maml does not vendor them.

For a full development environment the repository ships layered requirement files, and the README gives the order to install them:

bash
pip install -r requirements-ci.txt
pip install -r requirements-optional.txt
pip install -r requirements-dl.txt
pip install -r requirements.txt

For a first real use, the README points at the notebooks directory rather than inline examples. That is a deliberate choice: the package ships usage as Jupyter notebooks instead of docstring-sized snippets. The repository also has a nanoHUB tool and tutorial lecture linked from the README. Expect your first hour to be spent in a notebook, not in a REPL.

The external binary dependency is the real adoption cost

Everything in maml.pes is a wrapper around someone else's compiled code. LAMMPS for the runtime, GAP for GAP, MLIP for MTP, n2p2 for NNP. That is a reasonable design if your group already runs those codes, and a hard stop if it does not.

The failure mode is quiet. A missing LAMMPS binary or an unbuilt n2p2 does not surface as a Python import error at the top of your script. It surfaces when you reach the fitting or evaluation step, often after you have already spent time preparing training data. The README names the required packages but does not document what happens when one is absent, and it does not describe a fallback path. If your environment is a container you control, pin those binaries there. If you are on a shared cluster where you cannot install LAMMPS, maml.pes is the wrong tool for you, and the descriptor and sklearn parts of the package are the parts you can actually use.

The second boundary is scope. The README frames maml as not duplicating other packages, which means it will not paper over gaps in pymatgen or matminer. If a descriptor you need does not exist upstream, maml is unlikely to add it.

maml against matgl and the M3GNet line of work

The dependency on matgl>=3.0.3 is the clearest signal of how the ecosystem has moved. matgl is the graph deep learning library from the same research group, and maml now depends on it rather than carrying its own graph model implementations. If your interest is graph neural network potentials specifically, matgl is the direct route, and maml is the broader toolkit that includes it alongside classical descriptors and sklearn models.

The difference in approach is not just packaging. maml's value is breadth: bispectrum, Behler-Parrinello, SOAP and graph features under one interface, plus wrappers for four different potential formalisms. A library focused on one graph architecture gives you depth in that architecture and no descriptor zoo. Choosing maml means accepting the pymatgen and matminer dependency tree in exchange for that breadth. Choosing a narrower library means fewer moving parts but you write the descriptor plumbing yourself.

Maintenance, releases and the BSD-3-Clause terms

The last push to the default branch was on 2026-07-27. The most recent tagged release is v2025.4.1 from 2025-04-02, with v2024.6.13 and v2023.9.9 before it. The version string in pyproject.toml is 2025.4.3, ahead of the newest tag, which is normal for a repository that tags less often than it commits.

The release cadence is worth reading carefully. Three tags in roughly two years, with the newest more than a year behind the last push, means you should track the main branch or a commit hash rather than waiting for a release to pick up fixes. That is a real upgrade cost: there is no documented policy in the README for backporting fixes to tagged versions.

The licence is BSD-3-Clause, declared in pyproject.toml and in the LICENSE file at the repository root. That is a permissive licence, and the practical implication is that you can redistribute modified versions provided you keep the copyright notice and the disclaimer. The three-clause form also means you cannot use the names of the copyright holders to endorse derivative work. This is not legal advice; read the LICENSE file yourself if you are redistributing. Note that maml's dependencies carry their own licences, and LAMMPS in particular has its own terms, so a bundled distribution has more than one licence to satisfy.

Editorial conclusion

Adopt maml if you already work in the pymatgen and matminer ecosystem and want descriptors plus sklearn or keras model wrappers without writing the glue yourself. Do not adopt it if you need a self-contained potential trainer: maml.pes delegates to LAMMPS, GAP, MLIP and n2p2, so your environment has to carry those external codes. Before committing, verify that your Python version satisfies the requires-python >=3.11 floor in pyproject.toml, and check the maml.pes notebooks against your own LAMMPS build, since the README names the potentials but does not document a fallback when one of them is missing.

Frequently asked questions

What is maml in materials machine learning?

maml (MAterials Machine Learning) is a Python package that provides high-level interfaces for applying machine learning to materials science. According to the README, it converts crystals and molecules into features and supports sklearn and keras models, while relying on pymatgen and matminer rather than duplicating them.

How can AI and machine learning be used in materials science with maml?

The README lists descriptor generation for crystals and molecules, including bispectrum coefficients, Behler-Parrinello symmetry functions and SOAP, plus applications for potential energy surfaces, X-ray absorption spectroscopy and structural relaxation. Models come from scikit-learn or TensorFlow rather than being implemented in maml itself.

How do I install maml?

The README gives pip install maml for the base package from PyPI. Running the potential energy surface code additionally requires a LAMMPS installation, and the README shows a conda command for that.

Which Python versions does maml support?

pyproject.toml sets requires-python to >=3.11, and the classifiers list Python 3.11 and 3.12. Older interpreters will be rejected at install time.

What is the maml licence?

maml is BSD-3-Clause, declared in both pyproject.toml and the LICENSE file at the repository root. Dependencies such as LAMMPS carry their own licences.

Official sources

  1. Issues
  2. License: BSD-3-Clause
  3. materialyzeai/maml on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/materialyzeai-maml.svg)](https://hysenlabs.com/projects/materialyzeai-maml)