PyPOTS: A Python Toolbox for Machine Learning on Partially-Observed Time Series
A Python toolkit/library for reality-centric machine/deep learning & data mining on partially-observed time series, with 50+ SOTA neural network models for scientific analysis tasks (imputation, classification, clustering, forecasting, anomaly detection, cleaning) on incomplete industrial irregularly-sampled multivariate TS with NaN missing values
At a glance
- What is it?
- PyPOTS gathers 50+ neural network models for imputation, classification, clustering, forecasting and anomaly detection on multivariate time series that contain missing values. It is a library for people whose data is already broken, not for people who can clean it first.
- Who is it for?
- Adopt PyPOTS if your input is genuinely multivariate, irregularly sampled and full of NaN values, and you want to compare several published models behind one API instead of reimplementing each paper. Do not adopt it if your series are complete and evenly spaced, or if you need a model that ships with a trained checkpoint out of the box; PyPOTS trains models, it does not hand you weights.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 16 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap PyPOTS is trying to fill
Most time-series libraries assume a rectangular array. Every timestamp present, every channel populated, and any gap handled by a pandas fill before the model ever sees the tensor. The PyPOTS README states the motivation plainly: sensors fail, communication drops, hardware malfunctions, so missing values are common in real-world series. The project calls this partially-observed time series, abbreviated POTS, and argues that the area lacked a dedicated toolkit.
That framing matters because it changes what the library optimizes for. A forecasting library that treats missingness as a preprocessing bug will push you toward interpolation, and interpolation invents data. PyPOTS instead treats the mask as part of the input. The intended audience is engineers and researchers working on industrial, clinical or scientific series where the gaps carry information and where deleting incomplete rows would remove most of the dataset.
The scope is deliberately broad. The README lists five task families: imputation, classification, clustering, forecasting and anomaly detection, on multivariate partially-observed series with NaN missing values. The pyproject.toml classifiers name healthcare and scientific research as intended audiences, which matches the kind of data where dropping NaNs is not an option.
How the model zoo is organized, and why the wrench icon matters
PyPOTS is not one architecture. It is a registry of published models grouped by task, and the README's algorithm table marks which model supports which task. That table is the real entry point to the project, and reading it carefully saves time.
The most important distinction in the table is the wrench symbol. Models whose names carry it, the README names Transformer, iTransformer and Informer as examples, were not proposed for POTS data in their original papers. They cannot accept time series with missing values directly, and the README is explicit that they cannot do imputation as published. To make them usable, the project applies the embedding strategy and training approach called ORT+MIT, the same combination used in the SAITS paper. In other words, some entries in the zoo are the original architecture plus a wrapper that the PyPOTS authors designed.
That is a design decision worth naming as a trade-off. It gives you a much larger menu of models than the POTS literature alone would provide, and it gives you one API across all of them. It also means a result from a wrench-marked model is not a reproduction of its paper. If you are comparing against published numbers, check whether the baseline was run with the same embedding and training treatment.
The repository layout reinforces the task-first organization: examples are split into examples/imputation, examples/classification, examples/clustering, examples/forecasting and examples/anomaly_detection, and the library code sits under pypots/. Version history shows steady additions rather than a single monolithic release. v1.3 added TKAN, v1.4 was a bug-fix release, and v1.5 is titled HELIX. The last push to the default branch was on 2026-09-10.
Installing PyPOTS and running a first imputation
The package is on PyPI and conda-forge, and the README's installation section points at the docs for the details. The Python version badge shows 3.8+, while pyproject.toml declares requires-python >= 3.9, so treat 3.9 as the floor. The install documentation carries a section about version limitations on dependencies, which is worth reading before you pin anything, because deep learning stacks tend to break in combination rather than individually.
A plain pip install is the shortest path, and it is the command the project documents for getting the package:
pip install pypotsAfter that, the README's usage section shows the shape of a typical run. Every model follows the same pattern: build the model with a configuration, fit it on the training split, then call the task-specific method. For imputation the call is impute, and the input is the partially-observed array itself, not a pre-filled version of it. The README's own example constructs a SAITS model and calls fit and impute on it, so copy the constructor arguments from the example matching your task in examples/imputation rather than inventing values.
For a first real use, start from the example directory for your task instead of from a blank file. Those scripts wire up a dataset, instantiate a model and call the right method, which removes the guesswork about what the input object must look like. Hyperparameter tuning is a separate concern: the README notes that NNI support existed from v0.2 until v2.0, and that in PyPOTS v2 this is reimplemented with Optuna. If you were relying on the older NNI path, plan for that migration.
Where PyPOTS is the wrong tool
The clearest failure mode is a mismatch between the library's premise and your data. PyPOTS exists because missingness is pervasive. If your series are complete and regularly sampled, you are paying for mask handling you do not need, and a conventional forecasting or classification library will give you a shorter path and a larger pool of tutorials.
The second limitation is that PyPOTS is a modeling library, not a deployment artifact. Nothing in the README describes pretrained checkpoints, a model registry or a serving layer. You construct a model, fit it on your data, and the weights live in your process. That is normal for research-oriented tooling, but it means the work of serializing, versioning and serving a fitted model is yours.
The wrench-marked models deserve a second warning here. They were not designed for missing input. PyPOTS makes them work through ORT+MIT, and the README's own wording is that without that treatment they cannot accept time series with missing values, let alone impute. If your evaluation protocol requires a faithful reproduction of a Transformer or Informer baseline, PyPOTS is not the tool for that comparison, because the thing you would be running is not the original model.
Finally, the README states that models will be continuously updated to handle tasks not currently supported. That is a promise about direction, not a guarantee about any specific cell in the algorithm table. Check the table for your task rather than assuming coverage.
PyPOTS versus a general forecasting library
The natural alternative is a mainstream forecasting library that assumes complete series and expects you to impute upstream. The difference is not the algorithm list, it is where the mask lives.
In the upstream-imputation approach, you run an imputer first, then hand a dense array to the forecaster. Errors from the imputation step propagate silently, and the forecaster has no way to know which values were observed and which were invented. In PyPOTS, the mask travels with the data into the model, so the network can weight observed positions differently from filled ones. That is the whole point of the ORT+MIT treatment for models that were not built for it.
The cost of that design is scope. A general forecasting library will have more mature tooling around backtesting, cross-validation and deployment, because that is what its users ask for. PyPOTS concentrates its effort on the missing-data path. If your gaps are small and easily interpolated, the general library is the better buy. If your gaps are the problem you are trying to solve, the general library is solving a different problem.
A second alternative is to implement a single POTS paper yourself. That is reasonable when you need exactly one architecture and full control over the training loop. PyPOTS becomes worth its weight when you want to compare several models under one interface, which is the situation the unified API is designed for.
Maintenance, releases and what the licence allows
The repository is not archived, and the last push to the default branch was on 2026-09-10, so the codebase is moving. Release cadence in the recent window is roughly monthly: v1.3 on 2026-03-26 added TKAN, v1.4 on 2026-04-27 was a bug-fix release, and v1.5 on 2026-05-05 is titled HELIX. That pattern suggests feature releases interleaved with fixes rather than long quiet periods.
The upgrade cost is concentrated in two places. The first is dependency pinning. The install documentation has a dedicated section on version limitations, which exists because neural network stacks constrain each other. The second is the tuning framework change: NNI was the hyperparameter-optimization backend from v0.2 until v2.0, and Optuna replaces it in v2. If your pipeline calls into the tuning path, that is a real migration, not a version bump.
On licensing, PyPOTS is BSD-3-Clause, stated both in the README badge and in pyproject.toml. That is a permissive licence, and it is the same family as much of the scientific Python ecosystem. Two things to check before you rely on it commercially. First, the licence covers PyPOTS itself, not the papers the models come from; if you are building on a specific architecture, the original publication and its code may carry separate terms. Second, the pyproject.toml declares the licence by file reference rather than by SPDX expression, so read the LICENSE file directly. This is a description of what the repository states, not legal advice.
Editorial conclusion
Adopt PyPOTS if your input is genuinely multivariate, irregularly sampled and full of NaN values, and you want to compare several published models behind one API instead of reimplementing each paper. Do not adopt it if your series are complete and evenly spaced, or if you need a model that ships with a trained checkpoint out of the box; PyPOTS trains models, it does not hand you weights. Before committing, verify the dependency pins in the install documentation for your Python version, and check whether the specific algorithm you want is marked as natively designed for POTS or as one of the models that require the ORT+MIT embedding and training strategy to accept missing input at all.
Frequently asked questions
What is time series imputation?
It is the task of filling in the values a time series is missing, and in PyPOTS it is one of five supported tasks alongside classification, clustering, forecasting and anomaly detection. PyPOTS models take the partially-observed series with NaN values as input and return an array with the missing positions filled.
How to detect anomalies in time series with PyPOTS?
Anomaly detection is one of the task families PyPOTS supports on multivariate partially-observed time series, and the repository ships a dedicated examples/anomaly_detection directory. The README's algorithm table marks which models are available for that task, so check the table before picking one.
How can I use Python to predict time series data with PyPOTS?
Forecasting is one of the supported tasks, and the examples/forecasting directory contains runnable scripts. The general pattern is to instantiate a model with a configuration, call fit on the training split, then call the task-specific method.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/wenjiedu-pypots)