# tsfresh: Automatic Time Series Feature Extraction in Python

> tsfresh turns any sampled series into hundreds of statistical features, then filters them with a multiple-testing procedure. It suits tabular machine learning on sensor and event data, and it is a poor fit when you need sub-second latency or learned representations.

**blue-yonder/tsfresh** — Automatic extraction of relevant features from time series:

- Repository: https://github.com/blue-yonder/tsfresh
- Website: http://tsfresh.readthedocs.io
- Stars: 9,435 · Forks: 1,286
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/blue-yonder-tsfresh

## The feature engineering problem tsfresh automates

Time series classification and regression usually start the same way: someone writes a rolling mean, a standard deviation, a peak count, then a few more by hand, and the model is trained on whatever that person thought of. tsfresh replaces that manual step with a fixed catalogue. The README states the package "automatically extracts 100s of features from time series" drawn from statistics, time series analysis, signal processing and nonlinear dynamics. The features range from simple ones (number of peaks, average value, maximal value) to more involved statistics such as the time reversal symmetry statistic. The intended audience is the data scientist who already has a working model pipeline and wants the input matrix generated rather than authored. The README frames this as freeing time from building features so it can go into other work. The second half of the package matters as much as the first: because most of those hundreds of features will be noise for any given task, tsfresh ships a filtering procedure that scores each feature against the target using hypothesis testing with a multiple test procedure, which the README says mathematically controls the percentage of irrelevant features that survive. That combination, generate broadly then prune statistically, is the whole design.

## How extract_features and select_features fit together

The mechanism is a two-stage pipeline over a long-format DataFrame. In the first stage, extract_features groups rows by an id column and computes the enabled calculators per group, producing one row per id and one column per feature. In the second stage, select_features takes that matrix plus the target vector and returns the subset of columns whose association with the target survives the multiple-testing correction. The naming is literal: the README describes the filtering procedure as evaluating "the explaining power and importance of each characteristic for the regression or classification tasks at hand." Two consequences follow. The extraction stage is embarrassingly parallel across ids, which is why the FRESH whitepaper cited in the README is about distributed and parallel extraction for industrial-scale data. The selection stage is not parallel in the same way, because the correction is computed across the whole feature set at once. The README also notes that systematic feature engineering lets you work with series of different lengths, since every series is projected into the same feature space. That is the property that makes the approach usable when some samples are short and some are long, a situation that breaks most fixed-window models.

## Installing tsfresh and running a first extraction

The package is a normal Python distribution, so it installs from PyPI. The README links a Binder badge that opens the notebooks directory, which is the fastest way to see the API without a local environment. The repository's Dockerfile shows the source install route it uses itself: the build stage copies the source tree into /source and runs pip3 install --prefix=/install . there, then copies /install into /usr/local in the base image. For a local install, the equivalent is a pip install of the package name. After installation, the two functions you need are importable from the top-level package, and the README's own description of the workflow is that features are extracted and then filtered before being used to construct statistical or machine learning models for regression or classification tasks. The README does not print the exact call signature in the excerpt available, so check the documentation at tsfresh.readthedocs.io for the current argument names before wiring the call into a pipeline. The README does not document rollback or partial-failure behaviour for the extraction step either, so if a run fails partway through a large corpus there is no described resume mechanism.

## Where tsfresh is the wrong tool

The extraction stage computes a large fixed catalogue before any selection happens. On a corpus with many ids and long series, that means the expensive calculators run across every group whether or not they will survive selection. The README's own framing acknowledges this: most extracted features will not be useful, and the filtering exists precisely because of that. The practical cost is that you pay for generation before you learn what to keep. There is no documented incremental mode that extracts only the features that previously survived selection, so iterating on a large dataset means either accepting the full cost each time or manually restricting the calculator set. The second limitation is structural. tsfresh projects each series into a fixed feature vector, which discards the ordering information beyond what the calculators encode. If your signal's meaning lives in a pattern that none of the statistical, spectral or nonlinear calculators capture, no amount of selection will recover it. A convolutional or recurrent model that learns its own representation is the better choice there, and the README does not claim otherwise. The third case is latency. Every prediction requires running the same extraction over the incoming window, so the per-window cost is the extraction cost. For streaming or embedded inference with tight budgets, that is the wrong architecture.

## tsfresh compared with featuretools

featuretools and tsfresh are both automated feature engineering libraries, and the comparison people search for is reasonable, but they operate on different data models. featuretools builds features across related tables using deep feature synthesis, deriving aggregations and transformations along entity relationships. Its unit of interest is a row in an entity table and the relational graph around it. tsfresh has no relational layer. Its unit is a single time series identified by an id, and its transformations are time series specific: peaks, autocorrelation, spectral energy, entropy-like statistics, and the time reversal symmetry statistic the README names. If your problem is "join these five tables and generate features across the joins," featuretools addresses that and tsfresh does not. If your problem is "I have one long sensor recording per unit and I need numbers describing each unit's signal," tsfresh addresses that directly and featuretools would require you to pre-aggregate into a table first. The selection stage is the other difference. tsfresh's filtering is tied to a supervised target and uses a multiple-testing correction; the README describes it as controlling the percentage of irrelevant features mathematically. featuretools leaves feature selection to whatever model or selector you apply afterward.

## Maintenance, licence and the cost of staying current

The repository is not archived, and the last push was on 2026-07-06. Releases are infrequent rather than continuous: v0.21.0 in February 2025, v0.21.1 in August 2025, and v0.21.2 in May 2026. That cadence matters for planning. Bug fixes and calculator additions arrive on a roughly annual rhythm, so a workaround you write today may sit in your codebase for a while. The build configuration shows a constraint worth knowing about before you pin versions: pyproject.toml caps setuptools below 70 because the project's PyScaffold version has known compatibility issues with setuptools 70 and above, and the comment states the upper bound can be removed once the project migrates to PyScaffold 4 or later. That cap propagates into any environment where tsfresh is built from source. The Dockerfile pins python:3.8, while the Makefile tests against Python 3.9 through 3.12, so the container image and the tested interpreter range do not agree. The licence is MIT, which permits commercial use and modification; the repository file is LICENSE.txt. That is a factual note about the licence text, not legal advice, and anyone embedding the library in a distributed product should read the file themselves. The README lists an AUTHORS.rst and a CHANGES.rst at the repository root, so release-level changes are tracked in-tree.

## Conclusion

Adopt tsfresh when your time series are already grouped into samples, your model is a tree ensemble or linear model, and you need an interpretable feature matrix rather than a learned representation. Skip it for streaming inference where the extraction cost per window matters, and skip it when a small hand-written set of features already describes the signal. Verify first that your long-format DataFrame carries an id column and a sort column, because the extraction contract depends on that shape, and check whether the default feature set includes the high-cost calculators you actually need before running it over a large corpus.

## FAQ

### How do I install tsfresh?

It is a Python package distributed on PyPI, so a pip install of the package name is the standard route. The repository's Dockerfile shows the source install it uses itself, running pip3 install --prefix=/install . inside the build stage. The README also links a Binder badge that opens the notebooks directory if you want to try the API without a local install.

### What does tsfresh do?

It extracts hundreds of features from time series automatically, covering basic characteristics such as number of peaks and average value plus more complex statistics like the time reversal symmetry statistic. It then filters those features against a regression or classification target using a multiple test procedure.

### How do I use tsfresh for feature extraction?

You extract features from a long-format DataFrame and then filter them, which the README describes as producing a set of features that can be used to construct statistical or machine learning models for regression or classification. The README excerpt does not print the exact call signature, so check the documentation at tsfresh.readthedocs.io for the current argument names.

### How does tsfresh differ from featuretools?

featuretools performs deep feature synthesis across related tables and their entity relationships, while tsfresh operates on individual time series identified by an id and applies time series specific calculators. tsfresh also includes a supervised filtering stage based on multiple hypothesis testing, which featuretools does not provide.

## Sources

- [blue-yonder/tsfresh on GitHub](https://github.com/blue-yonder/tsfresh)
- [License: MIT](https://github.com/blue-yonder/tsfresh/blob/main/LICENSE)
- [Project website](http://tsfresh.readthedocs.io)
- [README](https://github.com/blue-yonder/tsfresh/blob/main/README.md)
- [Releases](https://github.com/blue-yonder/tsfresh/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/blue-yonder-tsfresh
