TFB: benchmarking time series forecasting methods on 27 multivariate datasets
[PVLDB 2024 Best Paper Nomination] TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods
At a glance
- What is it?
- TFB is a Shell-driven benchmark harness from the decisionintelligence group that wraps deep learning and statistical forecasting baselines behind one configuration interface. It is built for researchers who need comparable numbers, not for production forecasting.
- Who is it for?
- Adopt TFB if you are comparing forecasting architectures under a shared protocol and need the baseline implementations in one tree; skip it if you need a deployed predictor, because the repository is a benchmark harness rather than a serving system.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 61 days ago.
- What is it written in?
- Mainly Shell, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The reproducibility problem TFB was built to attack
Forecasting papers are hard to compare. Each new architecture arrives with its own data loader, its own train/validation split, its own normalization, and its own choice of lookback window. Two papers can report the same metric on the same named dataset and still not be measuring the same thing. TFB, from the decisionintelligence group, takes the position that the comparison itself should be a piece of software. The repository collects baseline implementations under ts_benchmark/baselines, and the README lists the models that have been folded in, including SparseTSF, TimeKAN, xPatch, HDMixer, PatchMLP, Amplifier, DUET and PDF. Each entry carries a checkbox: a checked box means both the code and its results are in the OpenTS-Bench leaderboard, an unchecked box means only the code has landed. That distinction matters when you are deciding whether a number you see is reproducible from this repository alone.
The audience is narrow and specific. This is for graduate students and industrial researchers who are writing a forecasting paper and need a defensible table, or who are reviewing one and want to check a claim. It is not a forecasting library in the scikit-learn sense. There is no fit/predict API aimed at application developers, and the README does not present deployment or serving as a goal.
What sits inside ts_benchmark and how a run is assembled
The top-level layout is small: config, scripts, ts_benchmark, characteristics_extractor, docs, a Dockerfile, and two requirements files. The work happens in ts_benchmark, which holds the baseline implementations, and in scripts, which holds per-algorithm, per-dataset shell entry points. The README points at that scripts folder as the place where the hyperparameters ultimately selected for each algorithm on each dataset are recorded, and it notes that retested results may differ from those in the TFB paper. Read that as an admission that the paper table and the repository table are not guaranteed to agree, and that the scripts are the authoritative record of what was actually run.
A separate directory, characteristics_extractor, computes time series characteristics such as trend, seasonality, stationarity, shifting, transition and correlation. The README dates that release to 2025.04 and provides both Chinese and English documentation for it. This is the part of the project that answers a different question from accuracy: given a dataset, what is it like, and does that explain why one model family wins on it. The dependency list in requirements.txt shows the breadth of what has to be installed to cover the baseline set, from darts and statsmodels through torch and ray to lightgbm and reformer-pytorch. Ray in particular suggests the harness is designed to fan runs out across processes rather than execute them one at a time.
Installing TFB and running a first configuration
There is no published package. The README does not give a pip install line for TFB itself, so the path is to clone the repository and install requirements.txt into an environment. The Dockerfile shows the maintainers' own route: an ubuntu:20.04 base, python3 with python-is-python3 and python3-pip, a virtualenv at /env, and then pip install -r requirements-docker.txt against the copy placed at /home/TFB. Note that the Dockerfile installs requirements-docker.txt, not requirements.txt, so the two files are not interchangeable.
For a manual setup, the README's dependency file is the starting point:
python -m venv /env
source /env/bin/activate
pip install --upgrade pip
pip install -r requirements.txtAfter that, a run is driven from the scripts directory rather than from a Python entry point you write yourself. The README describes the scripts folder as holding the selected hyperparameters for each algorithm on each dataset, so the intended first use is to pick an existing script, read its configuration, and execute it. The documentation site at tfb-docs.readthedocs.io is described in the README as the detailed API documentation, and the config directory at the repository root holds the configuration files those scripts consume. Nothing in the README describes a single canonical command that runs every baseline; the scripts are per-experiment.
One practical feature is worth knowing before you plan an experiment. The README announces support for predicting only a subset of input variables, with both a Chinese and an English tutorial under docs/tutorials. If your target series is one column of a wide multivariate table, that is the documented path.
Where TFB stops being the right tool
The honest limitation is that TFB is a research artifact with a research artifact's maintenance shape. The last push to the repository was on 2026-07-31, and the README's news entries cluster around 2025, with the most recent dated item being the June 2025 announcement of the sibling projects TAB and TSFM-Bench. The baseline list itself contains an unchecked entry, TimeBridge, meaning its code is in the tree but its results are not in the leaderboard. If your comparison depends on TimeBridge numbers, this repository does not supply them.
Second, the paper and the repository diverge by design. The README states plainly that some algorithms were retested and that results may differ from the TFB paper, directing readers to the scripts folder for the final hyperparameters. That is more transparent than most benchmarks manage, but it also means a citation to the paper's table and a reproduction from this repository can produce different numbers without either being wrong.
Third, the dependency surface is heavy. torch, ray, darts, statsmodels, lightgbm and reformer-pytorch in one environment is a large install, and version drift across those packages is a realistic source of failure. If you only want to forecast your own series with one method, TFB imposes the cost of the whole benchmark for no benefit. Use the upstream implementation of that single method instead.
TFB against BasicTS and the OpenTS leaderboard
BasicTS is the closest thing to a direct alternative in this space, and the difference is in scope rather than in kind. BasicTS is built around a unified training and evaluation pipeline for time series tasks, and it is a reasonable place to start if you want one framework to train models on your own data. TFB's organizing principle is the benchmark table: the repository exists so that a fixed set of baselines can be compared on a fixed set of datasets under recorded hyperparameters, and the scripts folder is the evidence trail. If your goal is to train a model, BasicTS is the more natural fit. If your goal is to produce a comparison table that a reviewer can trace back to a script, TFB's structure is doing work that a general training framework does not.
The other comparison point is TFB's own leaderboard, OpenTS-Bench, hosted on the decisionintelligence GitHub Pages site. This is not an alternative implementation but a different consumption mode of the same results. The README links the leaderboard for multivariate forecasting, and the checked boxes in the baseline list indicate which models have results there. For a reader who wants to know which method currently leads on a dataset, the leaderboard is faster than running anything. The repository is for when you need to reproduce or extend the number.
Licence, upgrade cost and what a fork inherits
TFB is released under the MIT licence. That is permissive and short, and it permits reuse and modification with the licence text retained. One thing to check before you redistribute anything built on this tree: the baselines folder vendors implementations of published methods, and those upstream projects may carry their own terms. The repository-level LICENSE covers TFB's own code; it does not automatically relicense third-party code that has been copied in. This is a description of the file layout, not legal advice.
Upgrade cost is dominated by the dependency set rather than by TFB's own code. requirements.txt pins darts at 0.25.0 and reformer-pytorch at 1.4.4 while leaving torch, ray, numpy and pandas as lower bounds. A floating lower bound on torch means a fresh environment in a year may resolve to a version the baselines were never run against. The Dockerfile's use of requirements-docker.txt suggests the maintainers keep a separate, presumably more constrained, set for container builds, and that file is the one to prefer if you want a reproducible environment. Rebuilding the virtualenv is cheap; re-deriving which torch version a given baseline script was validated on is not, and the README does not document that mapping.
Editorial conclusion
Adopt TFB if you are comparing forecasting architectures under a shared protocol and need the baseline implementations in one tree; skip it if you need a deployed predictor, because the repository is a benchmark harness rather than a serving system. Before committing, verify that the scripts folder contains the hyperparameters you intend to reproduce, since the README states that retested results can differ from the paper and that the finally selected per-dataset hyperparameters live there.
Frequently asked questions
Does TFB ship a pip package I can install?
The README does not give a pip install command for TFB itself. The documented route is to clone the repository and install requirements.txt, or to build the provided Dockerfile, which installs requirements-docker.txt instead.
Why do TFB's numbers differ from the TFB paper?
The README states that some algorithms were retested and that the results may differ from those in the TFB paper. It directs readers to the scripts folder for the hyperparameters ultimately selected for each algorithm and dataset.
How many datasets does TFB cover?
A 2025.04 news entry in the README says TFB added PEMS03 and PEMS07, bringing the total to 27 multivariate datasets.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/decisionintelligence-tfb)