Open-source project
tslearn-team/tslearn avatar
tslearn-team/tslearn

tslearn: dynamic time warping and time series machine learning in Python

The machine learning toolkit for time series analysis in Python

3,182 stars388 forksPythonBSD-2-Clause

At a glance

What is it?
tslearn wraps scikit-learn's estimator API around time series distances such as DTW, plus clustering, classification and barycenter tools. It is a good fit when your sequences are short and unevenly sampled, and a poor fit when you need to forecast long horizons on large datasets.
Who is it for?
Adopt tslearn if you have a few hundred to a few thousand short, possibly variable-length sequences and you need DTW-based distances, elastic clustering or a scikit-learn compatible classifier without writing the alignment code yourself. Do not adopt it if your problem is long-horizon forecasting, streaming data, or millions of series: the DTW kernels are the bottleneck and the library does not offer forecasting models.
Can I use it commercially?
Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What tslearn solves, and for whom

Most Python machine learning assumes rows of fixed-width features. Time series break that assumption in two ways: series in the same dataset can have different lengths, and two series that look alike may be shifted in time relative to each other. A Euclidean distance between two lagged copies of the same signal is large even though the signals are the same shape. tslearn exists to make that class of problem work inside the scikit-learn idiom.

The audience is research and applied teams already comfortable with scikit-learn: the README states that models in tslearn follow the same API as scikit-learn and are compatible with utilities such as hyper-parameter tuning and pipelines. If you are doing time series classification, clustering or distance computation on datasets measured in hundreds or thousands of series, and you want elastic distances rather than raw point-to-point comparison, this is the target case. The library is written in Python and depends on numba for the heavy kernels, so it is not aimed at people who need a compiled C++ or Rust pipeline.

The 3D array contract and DTW-based distances

Everything in tslearn flows through one data structure. The README states that a dataset is a 3D numpy array with dimensions (n_ts, max_sz, d): number of time series, number of measurements per series, number of dimensions. Variable-length support is handled by padding to max_sz, which is why the scaling example in the README prints nan in the trailing slots of the shorter series. That padding is the mechanism behind variable-length support, and it is also the first thing to watch: models that do not mask padded values will treat them as real observations.

On top of that array, tslearn.metrics provides distance functions, with dynamic time warping as the headline one. DTW aligns two sequences by allowing non-linear warping of the time axis, so a shifted or stretched version of a shape can score as close. The standard constraint is a Sakoe-Chiba band, which limits how far the alignment may drift; an Itakura parallelogram is the alternative. That band is the main accuracy-versus-cost dial, and it is the parameter most users need to tune rather than accept the default.

Clustering builds on the same distances. tslearn.clustering offers TimeSeriesKMeans, which uses DTW as its metric, and KShape, which uses a shape-based distance on normalized series. Barycenters, in tslearn.barycenters, compute an average series under a given metric, which is what a DTW k-means needs as its centroid update. Classification, in tslearn.neighbors, includes KNeighborsTimeSeriesClassifier, which is the same idea as a standard k-NN but with an elastic distance in place of Euclidean.

Installing tslearn and running a first classification

The README lists three installation routes. The simplest is pip from PyPI. The second is conda-forge. The third installs from the Git archive of the main branch, which is useful if you need a change that has not been released yet. Python 3.10 or newer is required according to pyproject.toml.

bash
python -m pip install tslearn

Alternatively, if your environment is conda-managed:

bash
conda install -c conda-forge tslearn

With the package installed, the first step is getting data into the 3D array. The README's example builds three series of unequal length and labels them. Note the third series has five points while the first two have four, so the result is padded.

python
from tslearn.utils import to_time_series_dataset
my_first_time_series = [1, 3, 4, 2]
my_second_time_series = [1, 2, 4, 2]
my_third_time_series = [1, 2, 4, 2, 2]
X = to_time_series_dataset([my_first_time_series,
                            my_second_time_series,
                            my_third_time_series])
y = [0, 1, 1]

Before fitting anything that compares magnitudes rather than shapes, scale the series. The README uses TimeSeriesScalerMinMax, and its printed output shows the padded position of the short series coming out as nan, which is expected.

python
from tslearn.preprocessing import TimeSeriesScalerMinMax
X_scaled = TimeSeriesScalerMinMax().fit_transform(X)
print(X_scaled)

Then fit a classifier. This one is a 1-nearest-neighbour model over the DTW-style distance, and it reproduces the labels exactly on the three training series, which is what you would expect from k=1 on its own training set.

python
from tslearn.neighbors import KNeighborsTimeSeriesClassifier
knn = KNeighborsTimeSeriesClassifier(n_neighbors=1)
knn.fit(X_scaled, y)
print(knn.predict(X_scaled))

The output is [0 1 1]. Because the estimators follow the scikit-learn interface, the same object can go into a scikit-learn pipeline or a grid search without an adapter, which the README points to in its examples gallery.

Where tslearn is the wrong tool

The DTW computation is the cost centre. Computing a full distance matrix between n series is quadratic in n before you even account for the per-pair alignment cost, and the Sakoe-Chiba band only reduces the constant factor of the inner dynamic program. On a few thousand series this is manageable; on millions it is not, and no amount of numba compilation changes the asymptotics. If your dataset is large and your series are long, a Euclidean or shapelet-based approach, or a representation learning approach, will finish while DTW is still running.

There is also a scope boundary. tslearn covers classification, clustering, regression, distances, barycenters and preprocessing. It is not a forecasting library. The optional torch dependency and the tslearn.foundation module, described in pyproject.toml as being for re-use of pre-trained models, point at a different direction than classical statistical forecasting, but the README does not present a forecasting API, and anyone whose goal is predicting the next hundred points should be looking at a dedicated forecasting package instead.

A third limitation is the padding convention itself. Variable-length support via a padded 3D array means every downstream step has to be aware that some entries are filler. The README shows the nan values appearing after scaling but does not document how each estimator treats them. That is a real gap: verify masking behaviour on your own data before trusting a model trained on mixed-length series.

How tslearn differs from a general time series library

The natural comparison is with sktime, which also targets time series in Python and also builds on scikit-learn conventions. The difference in approach is what sits at the centre. tslearn is organised around elastic distances and the algorithms that consume them: DTW, soft-DTW, shape-based distances, barycenters, and the clustering methods whose centroid update depends on those distances. sktime is organised around a unified interface for many task types, including forecasting and time series regression, with a registry of estimators from multiple backends.

That means the two overlap on classification and clustering but diverge at the edges. If your work is aligning and clustering sequences, tslearn's distance-first design is the more direct path, and the barycenter module has no obvious counterpart in a general-purpose framework. If your work is forecasting, or if you need a single API across forecasting, classification and transformation, sktime's scope is broader and tslearn simply does not compete there. Choosing between them is mostly a question of whether the distance function is the centre of your problem or one component among several.

Maintenance, releases and the BSD-2-Clause licence

The repository is not archived, and the last push was on 2026-09-10, so the codebase is being touched recently. The release cadence is visible in the tags: v0.8.0 on 2026-02-19, v0.8.1 on 2026-03-13, and v0.9.0 on 2026-06-30. That is a steady but not rapid rhythm, roughly two to three releases a year, which is typical for a research-adjacent library.

Upgrade cost is driven by the dependency floor in pyproject.toml. tslearn requires scikit-learn 1.5 or newer, numpy 1.24.3 or newer, scipy 1.10.1 or newer, numba 0.61 or newer, joblib 1.2 or newer, and statsmodels 0.14 or newer. The scikit-learn floor is the one most likely to force an environment rebuild, because it is high enough that older pinned stacks will not satisfy it. Numba is the other risk: it pins to specific Python and numpy ranges, so a Python upgrade can leave you waiting on a numba release before tslearn will install. The optional extras in pyproject.toml are pytorch and foundation, both pulling torch, plus tests and docs groups; a plain pip install does not bring torch in.

The licence is BSD-2-Clause, declared in pyproject.toml with license-files pointing at LICENSE. That is a permissive licence: it allows use in closed-source products with the copyright notice retained, and it does not carry the patent grant that BSD-3-Clause or Apache-2.0 include. If your organisation treats a patent grant as a requirement, that difference matters and is worth raising with whoever handles licensing. This is a description of the licence text, not legal advice.

Editorial conclusion

Adopt tslearn if you have a few hundred to a few thousand short, possibly variable-length sequences and you need DTW-based distances, elastic clustering or a scikit-learn compatible classifier without writing the alignment code yourself. Do not adopt it if your problem is long-horizon forecasting, streaming data, or millions of series: the DTW kernels are the bottleneck and the library does not offer forecasting models. Before committing, verify three things on your own data: that TimeSeriesScalerMinMax plus to_time_series_dataset produce the 3D array your model expects, that the DTW variant you pick (the default sakoe_chiba, or itakura, or an explicit radius) changes accuracy in the direction you want, and that your installed scikit-learn is at least 1.5, since pyproject.toml pins that as the floor.

Frequently asked questions

How do I install tslearn?

The README gives three options: python -m pip install tslearn from PyPI, conda install -c conda-forge tslearn, or installing from the Git archive of the main branch with pip. Python 3.10 or newer is required.

What is tslearn used for?

It is a machine learning toolkit for time series in Python, covering classification, clustering, regression, distance metrics such as dynamic time warping, and barycenters. Its estimators follow the scikit-learn API so they work with scikit-learn pipelines and hyper-parameter tuning.

Does tslearn work with variable-length time series?

Yes. The README states that tslearn supports variable-length time series, and to_time_series_dataset accepts series of different lengths by padding them into a 3D array. The padded positions appear as nan after scaling in the README's own example.

Which scikit-learn version does tslearn require?

pyproject.toml lists scikit-learn>=1.5 as a dependency, along with numpy>=1.24.3, scipy>=1.10.1, numba>=0.61, joblib>=1.2 and statsmodels>=0.14.

Official sources

  1. License: BSD-2-Clause
  2. Project website
  3. README
  4. Releases
  5. tslearn-team/tslearn on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tslearn-team-tslearn.svg)](https://hysenlabs.com/projects/tslearn-team-tslearn)