# mlforecast: scikit-learn style machine learning forecasting for millions of series

> Nixtla's mlforecast wraps lag and date feature engineering around any scikit-learn compatible regressor, then scales the same pipeline to Dask, Ray or Spark. It is a good fit when you already have a model you trust and want to forecast many series with it.

**Nixtla/mlforecast** — Scalable machine 🤖 learning for time series forecasting.

- Repository: https://github.com/Nixtla/mlforecast
- Website: https://nixtlaverse.nixtla.io/mlforecast
- Stars: 1,284 · Forks: 137
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/nixtla-mlforecast

## What mlforecast solves, and who it is for

The README states the motivation plainly: existing Python alternatives for machine learning models are "slow, inaccurate and don't scale well." mlforecast is the answer to that specific complaint. It is a framework for time series forecasting with machine learning models, and the README says it can scale to massive amounts of data using remote clusters.

The intended user is someone who already has a regressor they trust, typically LightGBM or XGBoost, and wants to apply it across many time series instead of one. The README says MLForecast includes efficient feature engineering to train any machine learning model with fit and predict methods, such as scikit-learn, to fit millions of time series. That framing matters: mlforecast is not a model. It is the plumbing between your data and a model you bring.

If you have a single series and no external variables, this is probably more machinery than you need. The README notes that if you are using only a single time series, you set the unique_id column to a constant value, which works but signals that the library is designed around the multi-series case.

## The mechanism: long-format data, lag transforms, one model over all series

The README describes the data setup as a pandas dataframe in long format, where each row is an observation for a specific series and timestamp. The example output shows four columns: unique_id, ds, y, and a static feature column. That is the entire contract. Everything else is derived.

You then instantiate an MLForecast object with the models and the features you want. The README says features can be lags, transformations on the lags, and date features, and that you can define transformations to apply to the target before fitting, which are restored when predicting. The import shown in the README is from mlforecast.lag_transforms import ExpandingMe, which is a truncated name in what the README shows, so treat the exact class list as something to confirm in the documentation rather than something to copy from here.

Each model is trained on all series, per the README. That is the key architectural decision. There is no per-series model fitting by default; you get one global model that sees series identity through the unique_id column and any static covariates you supply. Feature engineering is where the library claims its edge, calling its implementations the fastest for time series forecasting in Python. The repository lists coreforecast as a dependency, which is consistent with the feature computation living in a compiled layer rather than in pure pandas.

The API surface is deliberately small: fit and predict, matching scikit-learn. The README calls this out as a feature, and it means the learning curve is mostly about feature configuration, not about a new estimator interface.

## Installing mlforecast and running a first forecast

The README gives two install paths. From PyPI:

```bash
pip install mlforecast
```

Or from conda-forge:

```bash
conda install -c conda-forge mlforecast
```

The README points to an installation page for more detailed instructions. Note the Python requirement in pyproject.toml: requires-python is >=3.10, so 3.9 and below are out.

For a first run, the README uses a generator to produce synthetic data rather than asking you to bring your own. This is a reasonable way to confirm the install works before touching real data.

```python
from mlforecast.utils import generate_daily_series

series = generate_daily_series(
    n_series=20,
    max_length=100,
    n_static_features=1,
    static_as_categorical=False,
    with_trend=True
)
series.head()
```

What you should see is a dataframe with columns unique_id, ds, y and static_0, where unique_id values look like id_00 and ds values are daily timestamps starting at 2000-01-01. Then define your models as any scikit-learn compatible regressor:

```python
import lightgbm as lgb
from sklearn.linear_model import LinearRegression

models = [
    lgb.LGBMRegressor(random_state=0, verbosity=-1),
    LinearRegression(),
]
```

Finally, construct the MLForecast object with those models and your feature configuration. The README's construction snippet is truncated at the import line, so the full feature arguments are not visible in what the README shows; the documentation's quick start is the place to get the exact call. What the README does establish is the shape of the workflow: generate or load long-format data, declare models, declare features, then fit and predict.

## Where mlforecast stops being the right tool

The dependency list is the first real constraint. pyproject.toml pins pandas<3.0. If your environment is already on pandas 3.x, or you are planning a migration to it, mlforecast will hold you back until that pin is lifted. That is not a criticism of the choice, but it is a fact you should check against your own lockfile before adopting.

The second constraint is the global-model assumption. Because each model is trained on all series, a dataset where series have wildly different scales or dynamics will push the model toward a compromise unless you supply static covariates that let it separate them. The README lists support for exogenous variables and static covariates, which is the escape hatch, but it is on you to engineer those features. There is no built-in per-series model selection.

The third is the distributed story. Scaling to Dask, Ray or Spark is an optional extra, not the default. The optional dependency groups in pyproject.toml show dask, ray and spark each pulling in fugue plus lightgbm and xgboost. Ray is marked sys_platform != 'win32', so Windows users do not get the Ray path. If your plan depends on distributed training, that is a platform decision you make before writing code.

Finally, the project classifies itself as Development Status :: 4 - Beta in pyproject.toml. The version number is past 1.0, but the classifier is the project's own label, and it is worth taking at face value when you decide how much of your production pipeline to hang on it.

## How mlforecast differs from statsforecast and Prophet-style approaches

The README's own comparison is against other Python machine learning alternatives, not against classical statistical models. That distinction is the useful one. A tool like Prophet or a classical ARIMA-style library fits a model per series, usually a statistical one, and the model itself carries the temporal structure. mlforecast does the opposite: it converts the temporal structure into features and hands the result to a general-purpose regressor.

The practical difference shows up in three places. First, cross-series learning. A global LightGBM model can borrow strength across series, which helps when individual series are short. A per-series statistical fit cannot. Second, exogenous variables. The README lists support for exogenous variables and static covariates, which is the natural mode for a feature-based approach and often awkward in classical pipelines. Third, compute shape. Feature construction dominates, and the library's claim is that its implementations are the fastest in Python for this job. That is a claim from the README, not something verified here.

The cost of the feature-based approach is that you lose the interpretable decomposition a statistical model gives you. If a stakeholder needs a trend and seasonality component they can read off a chart, mlforecast is the wrong layer. If they need an accurate number and you have covariates, it is the right one.

Within the Nixtla family, the README's own links point to statsforecast in a tweet badge, so the two are siblings rather than competitors: statistical models in one, machine learning models in the other.

## Maintenance, versioning and licence

The repository is not archived, and the last push was on 2026-09-10. The release history shows v1.1.0 on 2026-07-10, v1.0.31 on 2026-03-10, and v1.0.3 on 2026-02-25. That pattern, a steady stream of patch releases between minor versions, suggests active maintenance, and the September push is recent enough to support that reading.

Upgrade cost is mostly dictated by the pandas<3.0 pin and by the optional extras. The core dependency set is small: cloudpickle, coreforecast>=0.0.15, fsspec, narwhals, optuna, pandas<3.0, scikit-learn, utilsforecast>=0.2.9. The presence of optuna in the core dependencies is notable; it means hyperparameter search tooling ships by default rather than as an extra. The distributed extras are heavier, since dask, ray and spark each bring their own cluster runtime plus lightgbm and xgboost.

Licensing is Apache-2.0, both in the repository LICENSE reference and in the pyproject.toml license field. Apache-2.0 is permissive and includes an explicit patent grant, which is usually what enterprises want. The repository also carries a THIRD_PARTY_LICENSES.md file, and the Makefile shows it is regenerated with pip-licenses and a filter script. If your legal review requires a dependency licence inventory, that file is the starting point, but it is a generated artifact and should be checked against the versions you actually install. Nothing here is legal advice.

The Beta classifier in pyproject.toml is the main signal to weigh against the release cadence. Frequent releases and a Beta label are not contradictory, but they do mean you should pin versions rather than float.

## Conclusion

Adopt mlforecast when you have many series, a regressor you already trust, and you want lag and date features built for you instead of hand-rolled. Skip it when you need a single univariate model with no external regressors, or when you cannot accept a pandas<3.0 pin. Before committing, verify that your target column and unique_id column follow the long-format convention the README shows, and run one fit and predict on a small slice of your own data.

## FAQ

### What is mlforecast?

mlforecast is a Python framework for time series forecasting using machine learning models, with the option to scale to large datasets using remote clusters. It provides feature engineering for lags, lag transformations and date features, then trains any scikit-learn compatible regressor on all series at once.

### Is there a Python library for forecasting like mlforecast?

Yes. mlforecast is a Python library for time series forecasting with machine learning models, and the README says it can scale to massive amounts of data using remote clusters. It requires Python 3.10 or newer according to pyproject.toml.

### Which AI is best for forecasting with mlforecast?

mlforecast does not ship a model of its own. The README says each model is trained on all series and that these can be any regressor following the scikit-learn API, with LightGBM and LinearRegression shown in its example. Which one performs better depends on your data, and the README points to an end-to-end walkthrough covering model training, evaluation and selection.

### What are the four types of forecasting models, and where does mlforecast fit?

The README does not enumerate forecasting model categories. It positions mlforecast as a framework for forecasting with machine learning models, in contrast to the Python machine learning alternatives it calls slow, inaccurate and hard to scale.

## Sources

- [License: Apache-2.0](https://github.com/Nixtla/mlforecast/blob/main/LICENSE)
- [Nixtla/mlforecast on GitHub](https://github.com/Nixtla/mlforecast)
- [Project website](https://nixtlaverse.nixtla.io/mlforecast)
- [README](https://github.com/Nixtla/mlforecast/blob/main/README.md)
- [Releases](https://github.com/Nixtla/mlforecast/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nixtla-mlforecast
