Library / SDK
thuml/OpenLTM avatar
thuml/OpenLTM

thuml/OpenLTM: a pipeline for pre-training and adapting large time-series models

Implementations, Pre-training Code and Datasets of Large Time-Series Models

560 stars63 forksJupyter NotebookMIT

At a glance

What is it?
OpenLTM collects implementations, pre-training code and datasets for large time-series models behind one script layout. It is a research codebase, not a forecasting library, and the README is thin on evaluation and maintenance.
Who is it for?
Adopt OpenLTM if you are reproducing or extending large time-series pre-training and you are comfortable reading model files and experiment code rather than an API reference. Do not adopt it if you need a supported forecasting library with a stable interface; the README itself points to Time-Series-Library for deep time series models and to Large-Time-Series-Model for out-of-the-box checkpoints.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What problem OpenLTM solves, and for whom

Pre-training a time-series foundation model involves more than a model definition. You need a data pipeline, a training loop, checkpoints, adaptation routines and a way to compare runs. OpenLTM packages those pieces together. The repository describes itself as "an open codebase aiming to provide a pipeline to develop and evaluate large time-series models", and the directory layout matches that claim: data_provider/, exp/, layers/, models/, scripts/ and a run.py entry point.

The intended reader is a researcher or engineer who already knows PyTorch and wants to reproduce or extend work such as Timer-XL, Timer, Moirai, Moment, TTMs, GPT4TS, Time-LLM or AutoTimes. The model checklist marks which of these are implemented and which are not. Sundial, LLMTime, Chronos, Time-MoE and a Google decoder-only model are listed unchecked, which tells you the checklist is a roadmap as much as an inventory.

This is not a forecasting product. There is no service, no API surface and no published benchmark table in the README. If you want to call a pre-trained model on your own series today, the README directs you elsewhere: to the Large-Time-Series-Model repository and to a HuggingFace collection of time-series foundation models.

How the pipeline is wired: data_provider, exp, models and scripts

The architecture is the familiar Time-Series-Library pattern, and the README confirms the lineage by recommending that library for deep time series models. Data loading lives in data_provider/. Model definitions live in models/, with timer_xl.py given as the reference file to copy. The experiment harness lives in exp/, and run.py is the entry point that ties a dataset, a model and a task together.

The extension path is explicit. The README says to add a model file to ./models, include the new model in Exp_Basic.model_dict inside ./exp/exp_basic.py, and create scripts under ./scripts. That model_dict is the registry the runner consults, so a model that is not registered there cannot be selected from a script. This is a plain, understandable design, and it is also why the project feels like a research scaffold: adding a model means editing a dictionary in framework code, not registering a plugin.

Scripts are grouped by purpose rather than by model. The README lists supervised training (one-for-one and rolling one-for-all forecasting), large-scale pre-training on UTSD and ERA5, and adaptation in full-shot and few-shot variants. Each script is a shell wrapper around run.py, so the arguments you care about are visible in the script text rather than hidden in a configuration system.

Installing OpenLTM and running a first training script

The README asks for Python 3.11 and a single requirements install. The pinned requirements.txt is short: einops, matplotlib, numpy, pandas, scikit-learn, torch==2.0.1 and transformers. Note the torch pin. If your environment already ships a newer PyTorch, you are choosing between matching the pin and adapting the code.

bash
pip install -r requirements.txt

Data goes into ./dataset. For pre-training you need UTSD, described as 1 billion time points in numpy format, or ERA5 for a domain-specific model. For supervised training or adaptation, the README points to the datasets from TSLib. Skip the pre-training download if you intend to start from the released checkpoint.

bash
bash ./scripts/supervised/forecast/moirai_ecl.sh

That command is the one-for-one forecasting example from the README. Expect the script to invoke run.py with a model name, a dataset and a task; the printed output is the training and evaluation log from the experiment harness.

bash
bash ./scripts/adaptation/few_shot/timer_xl_etth1.sh

The few-shot adaptation script follows the same pattern with a different argument set. Before running either, place the matching dataset under ./dataset, since the README treats that folder as the fixed location for downloaded data.

bash
bash ./scripts/pretrain/timer_xl_utsd.sh

This is the UTSD pre-training entry point. The README also mentions a checkpoint pre-trained on 260B time points, distributed as a PyTorch file that the project describes as easier to fine-tune than the HuggingFace equivalent, with load_pth_ckpt.ipynb showing how to load it.

Where OpenLTM stops being the right tool

The README is a usage guide, not documentation. There is no description of the evaluation protocol, no reported metrics, no explanation of how the one-for-one and rolling forecasting scripts differ in what they measure, and no guidance on which checkpoint suits which task. If your decision depends on accuracy numbers, this repository will not supply them.

Environment fragility is the second issue. The requirements pin torch==2.0.1 alongside numpy 1.26.0 and transformers without a version. A floating transformers version against a pinned older torch is a common source of import-time breakage, and the README does not describe a tested combination beyond the file itself. You should read requirements.txt as the supported set, not as a suggestion.

Scale is the third. Pre-training scripts assume UTSD or ERA5 on disk and a checkpoint released at 260B time points. That is not a laptop workload. The few-shot and supervised scripts are the realistic entry point for a single machine, and even those depend on datasets you must fetch from external links.

Finally, the repository is a collection. The README credits TTMs and other LLM4TS methods to an outside contributor and GPT4TS to another, which means quality and style vary between model folders. Treat each implementation as its own artifact.

OpenLTM compared with Time-Series-Library and Large-Time-Series-Model

The README makes the comparison for you. For deep time series models it recommends Time-Series-Library, and for out-of-the-box foundation models it recommends Large-Time-Series-Model and the HuggingFace collection. That is an unusually clear division of labour from the maintainers, and it should shape how you choose.

Time-Series-Library is the supervised deep-learning library. You pick an architecture, train it on a benchmark dataset, and compare against published baselines. OpenLTM sits one level up: its pre-training and adaptation scripts exist to produce and adapt foundation models, and the supervised scripts are there for reference points. If your goal is a well-performing forecaster on a known dataset with a known architecture, the sibling library is the shorter path.

Large-Time-Series-Model is the inference-oriented counterpart. It hosts checkpoints such as timer-base-84m for zero-shot use, which is what you want when you have no intention of running a pre-training job. OpenLTM is what you want when you need to reproduce the training, change the backbone, or fine-tune on your own data with the project's own scripts.

The practical difference is the unit of work. In the inference repositories the unit is a model call. In OpenLTM the unit is an experiment: a dataset folder, a registered model in exp/exp_basic.py, and a script under scripts/.

Maintenance, licence and the cost of upgrading

The repository is not archived and the last push was on 2026-03-22. The README's own news entries run from 2024.10 through 2025.05, covering the inclusion of large time-series models, the release of pre-training code and scripts, the GPT4TS and TTMs contributions, and the 260B Timer checkpoint. There are no retrieved releases, so versioning is by commit rather than by tag.

That has a direct cost. Without releases there is no changelog to read before pulling, and the pinned requirements mean an upgrade is a manual exercise: you move torch, numpy and pandas together and then re-run a script to see whether the experiment harness still imports. The model_dict registry in exp/exp_basic.py is the file most likely to conflict if you have added your own models, because upstream changes to that dictionary will collide with local edits.

Licensing is MIT, which is permissive and imposes few conditions on reuse. The README does not discuss the licences of the individual model implementations folded in from outside contributors, and several of the linked upstream projects are separate codebases with their own terms. If you plan to redistribute a modified version, check the licence of each model folder you keep rather than assuming the top-level MIT file covers everything. That is a factual gap in the repository, not a legal opinion.

Editorial conclusion

Adopt OpenLTM if you are reproducing or extending large time-series pre-training and you are comfortable reading model files and experiment code rather than an API reference. Do not adopt it if you need a supported forecasting library with a stable interface; the README itself points to Time-Series-Library for deep time series models and to Large-Time-Series-Model for out-of-the-box checkpoints. Before you invest time, verify the Python 3.11 and torch==2.0.1 combination against your hardware, confirm the dataset layout the scripts expect under ./dataset, and read exp/exp_basic.py to see which model names the runner will accept.

Frequently asked questions

What Python version and dependencies does OpenLTM need?

The README asks for Python 3.11 and a single pip install from requirements.txt. The pinned file lists einops 0.8.0, matplotlib 3.9.2, numpy 1.26.0, pandas 2.2.3, scikit-learn 1.5.2, torch 2.0.1 and transformers without a version.

Where does OpenLTM expect datasets to be placed?

Downloaded data goes in the ./dataset folder. The README points to UTSD and ERA5 for univariate pre-training, and to the datasets from Time-Series-Library for supervised training or model adaptation.

How do I add my own model to OpenLTM?

Add the model file to ./models, following the example of ./models/timer_xl.py, then include it in Exp_Basic.model_dict inside ./exp/exp_basic.py, and create the matching scripts under ./scripts.

Official sources

  1. Issues
  2. License: MIT
  3. README
  4. thuml/OpenLTM on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/thuml-openltm.svg)](https://hysenlabs.com/projects/thuml-openltm)