# DScribe: turning atomic structures into fixed-size descriptors

> DScribe is a Python package that converts atomic structures into numerical fingerprints for machine learning in materials science. It ships eight descriptors with C++ acceleration, and its main constraint is that all of them depend on the ASE Atoms interface.

**SINGROUP/dscribe** — DScribe is a python package for creating machine learning descriptors for atomistic systems.

- Repository: https://github.com/SINGROUP/dscribe
- Website: https://singroup.github.io/dscribe/
- Stars: 474 · Forks: 99
- Language: C++
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/singroup-dscribe

## What DScribe converts, and for whom

Machine learning models consume fixed-length vectors. Atomic structures do not arrive that way: they are lists of element types and coordinates, and the same molecule can be written in many equivalent ways. DScribe's job is the conversion step. The README describes it as a package "for transforming atomic structures into fixed-size numerical fingerprints", and those fingerprints are the descriptors used for machine learning, visualization and similarity analysis.

The intended user is a researcher or engineer working on atomistic systems who already has structures in memory and wants a feature matrix out of them. The dependency list confirms the audience: numpy, scipy, scikit-learn, joblib and ase. This is not a general chemistry toolkit. It is the feature-extraction layer that sits between a structure dataset and a model, and it assumes the rest of that stack is already in place.

## Eight descriptors, one create() call

The README's descriptor table lists Coulomb matrix, sine matrix, Ewald matrix, Atom-centered Symmetry Functions (ACSF), Smooth Overlap of Atomic Positions (SOAP), Many-body Tensor Representation (MBTR), Local Many-body Tensor Representation (LMBTR) and the Valle-Oganov descriptor. Each row also marks whether a spectrum variant exists and whether derivatives are implemented. All eight support both.

Every descriptor is constructed with its own parameters and then called through create(). The README example builds a CoulombMatrix with n_atoms_max=3 and permutation="sorted_l2", and a SOAP with species, r_cut, n_max, l_max and crossover. Those parameters are where the real decisions live: r_cut and n_max control how much of the neighbourhood is encoded and how large the output vector becomes. The library does not choose them for you.

Two output paths are visible in the example. create() returns a dense numpy array, and the README notes descriptors can also be produced as sparse arrays. For large systems with many atomic centres, that choice matters more than any parameter tuning.

The derivatives() method returns derivatives with respect to atomic positions. The README shows it called with return_descriptor=True, which returns both the derivative and the descriptor in one pass rather than computing the descriptor twice.

## Installing DScribe and computing a first SOAP descriptor

The README gives three installation routes. The shortest is pip, which is the one to try first because it avoids the C++ build entirely:

```bash
pip install dscribe
```

conda-forge is the alternative for environments already managed by conda:

```bash
conda install -c conda-forge dscribe
```

If you need to build from source, the README's sequence clones the repository, initialises the submodule and installs. The submodule step is not optional: setup.py points its include directories at dependencies/eigen, and the extension is compiled with cxx_std=11 and -O3. Skipping the submodule leaves the Eigen headers missing and the build fails.

```bash
git clone https://github.com/SINGROUP/dscribe.git
cd dscribe
git submodule update --init
pip install .
```

The README's quick example is the fastest way to confirm the install works. It builds three molecules with ase.build.molecule, sets up a SOAP descriptor and creates a descriptor for the water molecule centred on atom 0:

```python
from ase.build import molecule
from dscribe.descriptors import SOAP

samples = [molecule("H2O"), molecule("NO2"), molecule("CO2")]
soap_desc = SOAP(species=["C", "H", "O", "N"], r_cut=5, n_max=8, l_max=6, crossover=True)
soap = soap_desc.create(samples[0], centers=[0])
```

What you should see is a numpy array whose length is set by the descriptor parameters, not by the number of atoms in the molecule. That is the point of the library: H2O, NO2 and CO2 all produce vectors of the same shape, which is what a downstream model needs.

## Parallelism is per-process, and the ASE dependency is absolute

The README example shows n_jobs=3 passed to create() alongside a list of samples, and the comment says the work "can be parallelized across processes". That word is doing real work. This is joblib-style process parallelism, so the structures and the resulting arrays are serialised between processes. For a few hundred small molecules that overhead is invisible. For a dataset of large periodic cells, the cost of shipping each structure to a worker can rival the descriptor computation itself, and the fix is to chunk the sample list yourself rather than raise n_jobs.

The harder constraint is the input type. Every example in the README goes through ase.build.molecule, and the pyproject dependency list pins ase>=3.19.0. DScribe does not read a structure format on its own; it consumes ASE Atoms objects. If your structures are in a format ASE cannot parse, or if your pipeline deliberately avoids ASE, DScribe is the wrong layer and you would be writing an adapter before writing any science.

There is also a version floor. pyproject.toml sets requires-python to >=3.9 and lists classifiers through Python 3.13. Older interpreters are not supported, and the source build needs pybind11>=2.10.0 plus a working C++ compiler.

## Where DScribe is the wrong choice

The README does not document any rollback or migration path between descriptor versions, and it does not describe a way to serialise a fitted descriptor definition for reuse elsewhere. If your workflow requires reproducing a feature matrix months later from a stored configuration, you are relying on your own record of the constructor arguments, not on anything the package enforces.

The parameter surface is another boundary. SOAP alone takes species, r_cut, n_max, l_max and crossover, and the output dimensionality grows with n_max and l_max. Nothing in the README warns you about that growth. A reader who sets generous values on a large periodic cell will get a very wide feature matrix and a slow create() call, and the library will not tell them why.

Finally, DScribe is a descriptor library, not a model. It does not train, evaluate or predict. If what you actually need is a pretrained interatomic potential, this package is one input to that pipeline rather than a substitute for it.

## DScribe against hand-written symmetry functions

The obvious alternative is implementing symmetry functions yourself, typically ACSF variants, directly against your structure objects. That approach has one real advantage: no ASE dependency and no build step, so it fits a pipeline that already has its own structure representation.

The difference in approach is what you get for the cost. DScribe ships a compiled C++ extension built with pybind11, and its descriptor table covers eight representations with spectrum and derivative support marked per row. A hand-written implementation gives you exactly the functions you wrote and nothing else. It also gives you the bug surface: cutoff handling, periodic images and permutation invariance are all places where a small mistake produces a plausible-looking feature matrix that quietly degrades model accuracy. DScribe's test suite and coverage badge exist precisely because those details are easy to get wrong.

The trade is control against correctness. If your descriptor is genuinely non-standard, write it yourself. If it is one of the eight in the table, reimplementing it is hard to justify.

## Licence, releases and what maintenance looks like

DScribe is licensed under Apache-2.0, and pyproject.toml declares the licence as a file reference to LICENSE. Apache-2.0 is permissive and includes an explicit patent grant, which matters for a library that may end up inside a commercial materials pipeline. That is a description of the licence text, not legal advice; if the descriptor output feeds a patentable workflow, have counsel read the licence rather than this paragraph.

The version in pyproject.toml is 2.1.2. The repository is not archived, and the last push was on 2026-04-18. The README cites two papers, one on the original library and one titled "Updates to the DScribe library: New descriptors and derivatives", which is where the newer descriptor entries in the table came from.

Upgrade cost is concentrated in the compiled extension. A source install compiles C++ against Eigen and pybind11, so a Python or compiler upgrade can break the build even when the Python API is unchanged. The pip and conda routes avoid that, which is a reason to prefer them unless you are patching the extension itself.

## Conclusion

Adopt DScribe if your pipeline already speaks ASE Atoms and you need periodic-system descriptors such as SOAP or the Ewald matrix, including derivatives with respect to atomic positions. Do not adopt it if your structures live only in a format ASE cannot read, or if you need a descriptor the table does not list. Verify first that your Python is 3.9 or newer, that your ASE version is at least 3.19.0, and that the descriptor you pick supports the spectrum or derivative output your model needs.

## FAQ

### How do I install DScribe?

The README gives three routes: pip install dscribe, conda install -c conda-forge dscribe, or a source build that clones the repository, runs git submodule update --init and then pip install . The submodule step is required because the C++ extension includes Eigen headers from dependencies/eigen.

### What is DScribe SOAP and how do I use it?

SOAP is the Smooth Overlap of Atomic Positions descriptor, one of the eight in the README's table, and it supports both spectrum output and derivatives. You construct it with parameters such as species, r_cut, n_max, l_max and crossover, then call create() on an ASE Atoms object, optionally passing centers to restrict which atoms are described.

### Does DScribe work with structures that are not ASE Atoms objects?

No. Every example in the README builds structures with ase.build.molecule, and pyproject.toml lists ase>=3.19.0 as a dependency. DScribe consumes ASE Atoms objects and does not parse structure files on its own, so a different format needs an adapter first.

### Can DScribe compute derivatives with respect to atomic positions?

Yes. The README's descriptor table marks derivative support for all eight descriptors, and the quick example calls soap_desc.derivatives(samples, return_descriptor=True), which returns both the derivative and the descriptor.

### Which Python versions does DScribe support?

pyproject.toml sets requires-python to >=3.9 and lists classifiers for Python 3.9 through 3.13. The source build additionally needs pybind11>=2.10.0 and a C++ compiler.

## Sources

- [Issues](https://github.com/SINGROUP/dscribe/issues)
- [License: Apache-2.0](https://github.com/SINGROUP/dscribe/blob/master/LICENSE)
- [Project website](https://singroup.github.io/dscribe/)
- [README](https://github.com/SINGROUP/dscribe/blob/master/README.md)
- [SINGROUP/dscribe on GitHub](https://github.com/SINGROUP/dscribe)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/singroup-dscribe
