# A pinned transformers version, a context note with the number missing, and two code samples that stop mid-expression

> Time-MoE is the official code for the ICLR 2025 Spotlight paper on decoder-only time series foundation models, trained on a 300 billion point corpus. The published weights are 50M and 200M rather than the billion scale the headline claims, and the install instructions disagree with the dependency file.

**Time-MoE/Time-MoE** — [ICLR 2025 Spotlight] Official implementation of "Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts"

- Repository: https://github.com/Time-MoE/Time-MoE
- Website: https://arxiv.org/abs/2409.16040
- Stars: 1,004 · Forks: 115
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/time-moe-time-moe

## requirements.txt asks for transformers 4.57 while the README pins 4.40.1

The two instructions contradict each other. The prose says Time-MoE requires `transformers==4.40.1`, an exact pin, and the actual dependency file at the repository root says `transformers>=4.57.0`, a floor that is more than forty releases above it. The same file lists pyyaml, numpy, pandas, torch, scikit-learn, then `datasets>=2.18.0` and `accelerate>=0.28.0`. Which one you follow decides what you get, because the documented install is the file:

```bash
pip install -r requirements.txt
```

Following that gives you the floor and breaks the pin. Following the note gives you 4.40.1 and breaks the floor. Nothing in the repository reconciles the two, and since the model loads through `trust_remote_code=True` and depends on remote code written against a particular transformers release, the difference is not cosmetic.

## The context length note is missing the number it sets

The forecasting section opens with a sentence that lost its value. It reads that the `max_position_embeddings` for Time-MoE is set to during training, with nothing between set to and during training, and then the next clause states the maximum sequence length is 4096. The number is supplied by the following sentence instead of the slot where it belongs, so anyone scripting against that config key has to infer it rather than read it. The usable rule is the rest of the paragraph: context lengths up to 4096, autoregressive, arbitrary prediction horizons, and the sum of `context_length` and `prediction_length` should not exceed 4096. Anything longer is not supported out of the box, and the stated remedy is to fine-tune with the longer length you want.

## Both inference samples stop in the middle of an expression

The only worked example of making a forecast is two near-identical code blocks, and neither is a complete program. The first ends at `mean, std = seqs.me`, cut off inside the normalization step, so the tensor is never normalized and no forecast is produced. The second handles the already-normalized case and then trails off at a bare `prediction_length` under a `# forecast` comment, with no call that uses it. Both load `Maple728/TimeMoE-50M` with `device_map="cpu"` and a comment to switch to cuda, and both carry a commented variant that turns on `attn_implementation='flash_attention_2'`. The structure is right and the ends are missing, which tells you the shape of the API but leaves the reader to write the decoding loop.

## Billion scale in the headline, 50M and 200M in the download lines

The README opens by claiming the first work to scale time series foundation models up to 2.4 billion parameters trained from scratch. The weights it then links are `Maple728/TimeMoE-50M` and `Maple728/TimeMoE-200M`, base and large, both announced in October 2024 and both on Hugging Face. So the parameter counts a user can actually download are one and two orders of magnitude below the headline figure, and the inference example loads the smaller of the two. The corpus is on the same scale story from the other direction: Time-300B is described as over 300 billion time points spanning more than nine domains. Apache-2.0 covers the code, and there is no GitHub release for either checkpoint.

## Two unchecked TODO items define what the model cannot do

The scope of the model is stated as two open boxes at the top of the README, both unchecked: add covariate support, and enable fine-tuning of Time-MoE for forecasting with dynamic features while supporting time series classification. Read together they mean the shipped model handles univariate autoregressive forecasting and nothing else. There is no path to exogenous variables, and no classification objective. That matches the rest of the repository, which contains one training entry point, one distributed launcher and one evaluation script with no classification branch. The fine-tuning path that does exist takes your data as a list of observations per row and nothing else, so anything beyond a plain value sequence has to be encoded into the sequence itself or added to the model by hand.

## Benchmark files come from a Drive folder, and there is no package manifest

Evaluation is documented in two steps. Fetch the pre-processed datasets from a linked Google Drive folder, drop the contents under `./dataset`, then run the script, for ETTh1 as:

```bash
python run_eval.py -d dataset/ETT-small/ETTh1.csv -p 96
```

Nothing in the repository records a checksum or a revision for those files, so an evaluation run is only reproducible if you keep your own copy of the folder. The root listing has `.gitignore`, `LICENSE`, `README.md`, `figures/`, `main.py`, `requirements.txt`, `run_eval.py`, `scripts/`, `time_moe/` and `torch_dist_run.py`, with no setup.py and no pyproject.toml. That matters because the dataset example imports `from time_moe.datasets.time_moe_dataset import TimeMoEDataset`, an import that resolves only inside a clone, so the Time-300B loader is a source checkout feature rather than an installed one.

## Multi-node training is four exported variables and a launcher

Distributed fine-tuning hands off to `torch_dist_run.py`, and the inter-node setup is four exports the README asks you to fill in yourself: MASTER_ADDR, MASTER_PORT, WORLD_SIZE and RANK. Single node single or multiple GPU runs the same launcher without them, and plain CPU runs `main.py` directly. One argument is called out in bold for small datasets, `--stride 1`, added to the training command, and training from scratch takes `--from_scratch`. The launcher on a single GPU path is simply:

```bash
python torch_dist_run.py main.py -d <data_path>
```

Related to that, flash-attn is optional but recommended for speed and memory, pinned to `flash-attn==2.6.3`, with a second route that installs packaging and ninja first and sets `MAX_JOBS=64`, carrying a comment to replace 64 with the machine's CPU core count.

## No tags, and the news list stops five months before the last push

The repository has no GitHub releases at all, so there is no version to pin and nothing to diff against except commits. The news section lists four entries and the last one is February 2025, the ICLR Spotlight acceptance at Top 5.1 percent, with the two before it from October 2024 covering the Time-300B dataset and the two checkpoints, and the earliest from September 2024 for the arXiv preprint at 2409.16040. The branch itself was pushed on 2026-03-21, which is about six and a half months before now, leaving a stretch of commits with no announcement attached and no tag to name them by. The citation block at the bottom of the README is where this account stops, at a request to report mistakes, so the bibtex entry that would normally sit there is not in view.

## Conclusion

It fits research use of the two released checkpoints, where the 4096 step ceiling and the univariate autoregressive framing are acceptable, and it does not fit production work that needs covariates, classification heads, or a supported dependency set. Before building on it, install from requirements.txt rather than the pinned note and see which transformers version your code actually needs, expect to finish the inference example yourself since the published snippets are incomplete, and remember that the branch was last pushed on 2026-03-21 with no tagged releases at all.

## FAQ

### What is Time-MoE used for?

Univariate autoregressive time series forecasting with a decoder-only mixture of experts architecture, supporting arbitrary prediction horizons and context lengths up to 4096 steps. Covariates and time series classification are open TODO items rather than supported features.

### Which transformers version does Time-MoE need?

The README prose pins `transformers==4.40.1` while requirements.txt at the repository root asks for `transformers>=4.57.0`. The documented install command reads from requirements.txt, so the two instructions conflict.

### What model checkpoints are available for Time-MoE?

TimeMoE-50M and TimeMoE-200M, announced in October 2024 and hosted on Hugging Face under the Maple728 organisation, while the README headline refers to a 2.4 billion parameter model. There are no GitHub releases.

### How do you evaluate Time-MoE on a benchmark dataset?

Download the pre-processed datasets from the linked Google Drive folder, place them under ./dataset, then run `python run_eval.py -d dataset/ETT-small/ETTh1.csv -p 96`. The repository records no checksum for those files.

### Can Time-MoE sequences be longer than 4096 steps?

Not without changes. The README advises keeping the sum of context_length and prediction_length at or below 4096 and says to fine-tune the model with the longer length if you need more.

## Sources

- [Issues](https://github.com/Time-MoE/Time-MoE/issues)
- [License: Apache-2.0](https://github.com/Time-MoE/Time-MoE/blob/main/LICENSE)
- [Project website](https://arxiv.org/abs/2409.16040)
- [README](https://github.com/Time-MoE/Time-MoE/blob/main/README.md)
- [Time-MoE/Time-MoE on GitHub](https://github.com/Time-MoE/Time-MoE)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/time-moe-time-moe
