# SeisBench splits seismology machine learning into data, models and generate

> A GPL licensed Python toolbox that gives seismic datasets, pretrained models and data generation pipelines one API, with fourteen Colab notebooks as the real documentation. Two rough edges are worth knowing first: the notebooks can run ahead of a release, and the packaging metadata disagrees with itself on the Python floor.

**seisbench/seisbench** — SeisBench - A toolbox for machine learning in seismology

- Repository: https://github.com/seisbench/seisbench
- Stars: 423 · Forks: 118
- Language: Jupyter Notebook
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/seisbench-seisbench

## Three modules, and the gap between data and models is what generate fills

SeisBench is the Seismology Benchmark collection, a Python toolbox whose stated purpose is reducing the overhead of applying or developing machine learning to seismological tasks. The pitch is a unified API, both for reaching datasets and for training and applying models to seismic data.

That API splits into three core modules. `data` gives access to benchmark datasets and the functionality for loading them. `models` holds a collection of machine learning models for seismology, where you can create a model, load a pretrained one, or train on any dataset. `generate` holds the tools for building data generation pipelines.

The relationship between them is stated as a gap rather than a hierarchy: `generate` bridges the gap between `data` and `models`. That is the piece a researcher writing their own architecture usually ends up writing by hand, since a model in this field rarely consumes a dataset object directly.

The licence is GPL-3.0, the repository is not archived, and its last push was on 2026-09-24.

## Two install routes, and the README cuts off mid-sentence on the second

Installation is a single pip line, and the project says so plainly. The recommended route:

```bash
pip install seisbench
```

The alternative is a source install: clone the repository, switch to the repository root, and run:

```bash
pip install .
```

Both routes are prefaced by the same suggestion, which is to put SeisBench in a virtual environment, with conda named as an example. That advice matters more than usual here, because the dependency list is long and includes torch.

What the README does not finish is the sentence following `pip install .`, which begins by explaining that it will install SeisBench in and then stops. So the part of the source route that would tell you where the package lands, and any caveat attached to it, is not in the text.

The project distributes through both PyPI and a Zenodo DOI, and the documentation lives at seisbench.readthedocs.io. The README also links to the Python 3.10 release page, which is a hint about the intended interpreter rather than a stated requirement.

## The notebooks can run ahead of a release, and the project says so

The documented way in is a set of Colab notebooks reached through Open in Colab links, with cloning the repository and running the same examples locally offered as the alternative. The catch is stated as a note rather than buried:

```bash
pip install "seisbench[das] @ git+https://github.com/seisbench/seisbench"
```

An example notebook added very recently may rely on functionality that is not yet part of a numbered version, and replacing the installation line with that git command points the notebook at the development branch instead. The `[das]` extra is carried along, which matters because DAS support is an optional dependency group rather than part of the default install.

This is a small thing to document and a useful thing to have documented. It tells you the examples directory is a development surface and the numbered releases are a slower surface, and it tells you the project expects that gap to exist.

For anyone pinning a version in production, that is the sentence to read twice, because a notebook copied out of Colab will carry whatever install line it was written with.

## Event catalog building gets two notebooks, one per associator

The examples directory is the substance of this project. Fourteen notebooks sit under `examples/` alongside a README of their own, and they are numbered so that the order reads as a path.

The first three are the orientation set: dataset basics, the model API, and generator pipelines. One to one with the three modules, which is a sensible way to teach an API whose whole point is the seam between them.

The advanced set covers the work that usually dominates a seismology machine learning project. Applied picking deploys a model on streams. Training PhaseNet is a full training walkthrough. Creating a dataset covers building one rather than loading a benchmark. And event catalog construction gets two separate notebooks, one for GaMMA and one for PyOcto, which tells you the project treats associator choice as a real fork in the road rather than an interchangeable detail.

A miscellaneous group covers denoising, both applying DeepDenoiser and training a denoiser, training DKPN, and depth phases for earthquake depth. The depth phases notebook is the one that is clearly domain-specific rather than machine learning specific.

## DAS support lives behind an extra, and a release made it the headline

Distributed acoustic sensing is handled as a separate concern, with two tutorials for training and applying models on DAS data: applying DAS models, and training DeepSubDAS.

The dependency behind that is an optional group, and it is heavier than the name suggests:

```toml
das = ["xdas>=0.2.8", "pyarrow", "torchvision"]
```

So DAS pulls xdas, pyarrow and torchvision on top of the base install. That is why the development-branch install line in the previous section carries `[das]` even when you have no intention of touching fiber data, and why a plain `pip install seisbench` will not get you far if you do.

The release history suggests this was the main event rather than an afterthought. Version 0.12.0 is titled full DAS support, denoising and more picking models, and it is the most recent numbered release. Version 0.11.0 carried DAS support earlier with performance work and new datasets and models, and 0.10.0 introduced SkyNet, SeisDAE and a more powerful model API.

Two releases in a year, both with DAS in the title, is a project whose centre of gravity has moved toward fiber data.

## requires-python says 3.10 and the trove classifier says 3.9

The packaging metadata contradicts itself in two places, and both are in the file a user reads first.

The first is the Python floor. `requires-python` is set to `>=3.10`, while the trove classifier in the same block declares `Programming Language :: Python :: 3.9`. Only one of those is enforced at install time; the classifier is metadata that indexes and documentation read. The README's link to the Python 3.10 release page is consistent with the enforced field rather than with the classifier.

The second is numpy. The build-system requirements ask for `numpy>=2.0`, which is what has to be present to build, while the runtime dependency list asks for `numpy>=1.21.6`, which is what the installed package declares. Those two floors answer different questions, so the gap is defensible, but it means a build machine on numpy 1.x cannot build this source tree.

Neither mistake will break a working install. Both will cost somebody an afternoon, which is why they are worth reading before you file an issue.

## The version number comes from git tags, so a plain clone installs an odd version

SeisBench declares `dynamic = ["version"]` and configures setuptools_scm, so the installed version is derived from git metadata rather than written down. That is the right choice for a project whose releases are tags, and it has one consequence worth planning for: a clone with no tags reachable produces a version that is not one of the numbered ones.

That interacts with the development-branch install line from earlier. Installing from git is the documented way to get functionality that is not in a release, and it is also the way to end up on a version string that does not match anything in the release list.

The rest of the metadata is conventional and worth noting for provenance. Authors and maintainers are the same two people, from KIT and GFZ Potsdam, and the declared keywords are seismology, machine learning, signal processing and earthquake. Optional groups are `dev` with ruff and pre-commit, and `tests` with pytest, pytest-asyncio and pytest-benchmark.

Those testing extras are not decorative. Pytest is configured with a slow marker and a documented way to deselect it, which tells you the suite has benchmarks in it and that the project expects contributors to run a subset.

## A C extension is compiled during setup, so this is not a pure Python install

Alongside `pyproject.toml` and `requirements.txt` there is a third build file, `setup.py`, and its entire content is a compiled extension:

```python
Extension(
    "seisbench.ext.utils",
    sources=["seisbench/ext/utils.c"],
    include_dirs=[numpy.get_include()],
    extra_compile_args=["-O3", "-flto"],
)
```

So SeisBench ships C source that is compiled at install time, with numpy's headers on the include path and optimisation plus link-time optimisation requested. A wheel built ahead of time hides this. A `pip install .` from a clone does not, and it needs a working C toolchain and numpy headers present before the build starts.

The packaging configuration also declares a cibuildwheel section, which is the standard way to build wheels across platforms, and that section is cut off in the file as it can be read here. Combined with the truncated sentence after `pip install .`, the practical advice is unchanged: use the PyPI wheel unless you have a reason not to.

The rest of the repository follows the layout of a research library: the `seisbench/` package, `tests/`, `docs/`, `examples/`, `contrib/`, a `.readthedocs.yaml` for the documentation build, and a `.pre-commit-config.yaml`.

## Conclusion

Adopt SeisBench if your seismic work already spans dataset loading and model training, since the value is in the seam between those two and the fourteen notebooks demonstrate that seam end to end. Do not adopt it expecting a pure Python install, because a C extension is compiled during setup and a source install builds it with a compiler present. Two things to check first: which numbered release your notebooks were written against, since a recent notebook may need the development branch rather than the published package, and whether your interpreter satisfies requires-python, since that field says 3.10 while the trove classifier still says 3.9.

## FAQ

### What are the three core modules in SeisBench?

They are data for accessing and loading benchmark datasets, models for creating, loading and training seismology models, and generate for building data generation pipelines. The README describes generate as bridging the gap between data and models.

### How do I install SeisBench?

The recommended route is pip install seisbench. The alternative is to clone the repository and run pip install . at the root, and the project suggests using a virtual environment such as one made with conda for either route.

### What extra do I need for DAS data in SeisBench?

The das extra, which pulls xdas, pyarrow and torchvision. It is also included in the git install line the README offers for notebooks written ahead of a numbered release.

### Which Python versions does SeisBench support?

The packaging metadata is inconsistent: requires-python is set to >=3.10 while the trove classifier still declares Python 3.9. The enforced floor is the requires-python field.

## Sources

- [Issues](https://github.com/seisbench/seisbench/issues)
- [License: GPL-3.0](https://github.com/seisbench/seisbench/blob/main/LICENSE)
- [README](https://github.com/seisbench/seisbench/blob/main/README.md)
- [Releases](https://github.com/seisbench/seisbench/releases)
- [seisbench/seisbench on GitHub](https://github.com/seisbench/seisbench)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/seisbench-seisbench
