Library / SDK
stumpy-dev/stumpy avatar
stumpy-dev/stumpy

STUMPY: Matrix Profiles for Time Series Pattern and Anomaly Discovery in Python

STUMPY is a powerful and scalable Python library for modern time series analysis

4,158 stars371 forksPythonNOASSERTION

At a glance

What is it?
STUMPY computes the matrix profile, a nearest-neighbor summary of every subsequence in a time series, and exposes it through NumPy, Dask and GPU code paths. It is aimed at researchers and data scientists who need motif discovery, discord detection or segmentation without hand-rolled distance code.
Who is it for?
Adopt STUMPY if you are doing offline or streaming time series mining in Python and can accept a Numba-compiled dependency plus the need to choose a window size m yourself. Do not adopt it if you need a maintained, production-supported API with a stability guarantee; pyproject.toml still classifies the package as Development Status 3 - Alpha.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem STUMPY solves: nearest neighbors for every subsequence

Most time series questions reduce to a comparison between two windows of data. Does this week's sensor trace resemble last month's? Where does the shape break? Doing that by hand means sliding a window, computing pairwise distances, and deciding which neighbors count. STUMPY replaces that with one computation. The README describes the matrix profile as a way to, for every subsequence within your time series, automatically identify its corresponding nearest neighbor. The output is an array of distances plus an array of indices, and from those two arrays the library derives motifs (approximately repeated subsequences), discords (anomalies), shapelets, semantic segments, chains and snippets. The intended audience is stated plainly in the README: academic, data scientist, software developer, or time series enthusiast. In practice the library suits anyone who already has a NumPy array of measurements and wants pattern structure rather than a forecast. It does not fit problems where the question is what value comes next; there is no forecasting API in the README.

How the matrix profile computation is organized

The core entry point is stumpy.stump, which takes a one-dimensional time series and a window size m. The README's example uses m=50 and describes it as approximately how many data points might be found in a pattern. The result is a matrix profile array whose length is the number of subsequences, and each entry holds the z-normalized distance to that subsequence's nearest neighbor elsewhere in the series. Two design choices shape everything else. First, the heavy loops are compiled with Numba, which is why numba >= 0.61.2 appears in both requirements.txt and the project dependencies alongside numpy >= 1.24 and scipy >= 1.10. Second, the same algorithm is exposed at four scales: stump for a single series on one machine, stumped for a single series across a Dask distributed Client, gpu_stump for CUDA devices discovered through numba.cuda.list_devices, and mstump for multi-dimensional input where each row is a dimension and each column is a shared time index. The parallel variants are not separate implementations with different semantics; they are the same computation dispatched to workers or devices, which is why the README can present them as near-identical call signatures.

Installing STUMPY and running a first matrix profile

The README points at PyPI and conda-forge through its badges, and the repository ships pip.sh, conda.sh and environment.yml for its own build workflows. The declared requirement is Python 3.10 or newer. Install from PyPI:

bash
pip install stumpy

Then compute a profile on synthetic data. This mirrors the README's typical usage snippet, including the __main__ guard that the project uses consistently, which matters because Numba compilation interacts badly with some interactive and multiprocessing contexts:

python
import stumpy
import numpy as np

if __name__ == "__main__":
    your_time_series = np.random.rand(10000)
    window_size = 50  # Approximately, how many data points might be found in a pattern

    matrix_profile = stumpy.stump(your_time_series, m=window_size)

Expect the first call to be slow while Numba compiles the kernel, and later calls on the same shape to be much faster. To spread the same work across a Dask cluster, the README shows stumped taking a Client as its first argument:

python
import stumpy
import numpy as np
from dask.distributed import Client

if __name__ == "__main__":
    with Client() as dask_client:
        your_time_series = np.random.rand(10000)
        window_size = 50

        matrix_profile = stumpy.stumped(dask_client, your_time_series, m=window_size)

For multi-dimensional input, mstump returns two arrays rather than one, because each dimension gets its own profile and index set:

python
import stumpy
import numpy as np

if __name__ == "__main__":
    your_time_series = np.random.rand(3, 1000)
    window_size = 50

    matrix_profile, matrix_profile_indices = stumpy.mstump(your_time_series, m=window_size)

The README does not document expected runtimes for any of these, so treat the first successful run as the only reliable baseline for your own data.

Where STUMPY stops being the right tool

The most consequential limitation is the window size. Every downstream result, motif, discord, segment, inherits the choice of m, and the README offers only the loose guidance that it is approximately how many data points might be found in a pattern. Pick m too small and noise dominates the nearest-neighbor distances; too large and short recurring shapes disappear. The project acknowledges this by listing pan matrix profiles for selecting the best subsequence window size among its capabilities, which means the selection problem is real enough to have a dedicated API rather than a default. A second limitation is the dependency surface. Numba is a hard requirement, not an optional accelerator, and Numba pins to specific Python and NumPy ranges; the pyproject.toml requires Python >= 3.10 and numpy >= 1.24, so an environment locked to an older interpreter cannot install the current release at all. Third, the project classifies itself as Development Status 3 - Alpha despite a 1.x version number and a published JOSS paper, and the release history shows a gap between v1.13.0 in July 2024 and v1.14.0 in February 2026. If you need an API with a stability contract, that combination is a warning rather than a convenience. Finally, nothing in the README describes a distributed scheduler setup beyond passing a Dask Client, so cluster sizing, memory per worker and failure recovery are left to the Dask documentation, not to STUMPY's.

STUMPY against rolling-window distance code and general time series libraries

The obvious alternative is writing the sliding-window nearest-neighbor search yourself with NumPy and SciPy. That gives you full control over the distance function and avoids Numba entirely, and for a few thousand points it is entirely reasonable. The difference in approach is that hand-written code typically recomputes distances between overlapping windows from scratch, while STUMPY's implementation is built around the matrix profile formulation, which the README presents as the single computation from which motifs, discords, shapelets, segmentation, chains and snippets are all derived. You get one artifact that supports many questions instead of a bespoke distance matrix per question. A second comparison is with general-purpose time series libraries oriented toward forecasting. Those model the series and predict the next value; STUMPY does not forecast at all. It answers structural questions about what already happened, which is why anomaly detection and motif discovery appear in the project's topics while prediction does not. If your actual requirement is a forecast, STUMPY is the wrong layer, not a weaker one. The distributed and GPU variants are also a differentiator: the README shows the same call moving to Dask workers or to CUDA devices via gpu_stump with a device_id list, so scaling out does not mean rewriting the analysis.

Maintenance, releases and the licence situation

The repository is not archived, and the last push was on 2026-09-12, which is recent. Release cadence is uneven: v1.13.0 landed on 2024-07-09, v1.14.0 on 2026-02-03 and v1.14.1 on 2026-02-08, so the patch followed the minor within days while the minor itself followed a roughly eighteen-month gap. Upgrade cost is dominated by the Numba and NumPy floors rather than by STUMPY's own API. Moving to a new Python release means waiting for a Numba build that supports it, and the classifier list in pyproject.toml currently runs from 3.10 through 3.14. On licensing, the repository's LICENSE.txt is referenced by the badge as the licence source, pyproject.toml declares license = "BSD-3-Clause" with license-files = ["LICENSE.txt"], and the metadata shown for the repository reports NOASSERTION. Those two signals disagree, and the README's badge links to LICENSE.txt rather than stating a licence inline. Read LICENSE.txt in the repository before you rely on any particular term; that is a factual gap in the metadata, not a legal opinion, and I am not in a position to resolve it for you.

Editorial conclusion

Adopt STUMPY if you are doing offline or streaming time series mining in Python and can accept a Numba-compiled dependency plus the need to choose a window size m yourself. Do not adopt it if you need a maintained, production-supported API with a stability guarantee; pyproject.toml still classifies the package as Development Status 3 - Alpha. Before committing, verify that your Python version is 3.10 or newer, that numba >= 0.61.2 builds in your environment, and that the matrix profile for your own data at a candidate m actually separates the motifs you expect.

Frequently asked questions

What is STUMPY in Python?

STUMPY is a Python library that computes the matrix profile, which the README describes as identifying, for every subsequence in a time series, its corresponding nearest neighbor. It is used for motif discovery, anomaly detection, shapelets, segmentation, streaming data and related time series mining tasks.

How do you install STUMPY?

The README's badges point to PyPI and conda-forge, and the declared dependencies are numpy >= 1.24, scipy >= 1.10 and numba >= 0.61.2 with Python 3.10 or newer. Installing from PyPI with pip install stumpy pulls those in.

Does STUMPY support GPU or distributed execution?

Yes. The README shows stumped taking a Dask distributed Client for distributed execution, and gpu_stump taking a device_id list built from numba.cuda.list_devices for GPU execution. Multi-dimensional input is handled by mstump, with a distributed counterpart in mstumped.

How do you choose the window size m in STUMPY?

The README describes m as approximately how many data points might be found in a pattern, and its examples use m=50. The project also lists pan matrix profiles for selecting the best subsequence window size, which indicates the choice is not automatic.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. stumpy-dev/stumpy on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/stumpy-dev-stumpy.svg)](https://hysenlabs.com/projects/stumpy-dev-stumpy)