skforecast: scikit-learn estimators as time series forecasters
Python library for time series forecasting using scikit-learn compatible models, statistical methods, and foundation models
At a glance
- What is it?
- skforecast wraps any scikit-learn compatible regressor into a recursive, direct or multi-series forecaster with backtesting and prediction intervals. Here is how it installs, how the recursive mechanism works, and where it stops being the right tool.
- Who is it for?
- Adopt skforecast if you already have a scikit-learn workflow, want gradient boosting on lagged features, and need backtesting and prediction intervals in the same object. Do not adopt it if you want a single call that selects and fits a model for you, or if your series are so short that a lag window of 15 leaves almost no training rows.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap skforecast fills between scikit-learn and forecasting
scikit-learn regressors expect a two-dimensional feature matrix and a target vector. A time series is a single ordered column. The work of turning one into the other, choosing how many lags, shifting them without leaking future values, refitting at each step of a multi-step horizon, and scoring the result honestly, is the part that eats a week before you have a baseline. skforecast packages that work.
The README states the library works with any estimator compatible with the scikit-learn API, and names LightGBM, XGBoost, CatBoost and Keras as examples. The pyproject classifiers list Development Status 5 - Production/Stable and Python 3.10 through 3.14. The intended audience is developers and researchers who already know how to fit a regressor and want the forecasting scaffolding around it, not a new modelling language to learn.
How ForecasterRecursive turns lags into a supervised learning problem
The core mechanism is visible in the README quick example. You pass an estimator and a number of lags. The forecaster builds a training matrix where each row is the target at time t and the features are the observations at t-1 through t-15, then calls the estimator's normal fit. Prediction runs the model once to get t+1, appends that value to the lag window, and runs again for t+2. Errors compound across the horizon, which is the defining trade-off of the recursive approach.
The example uses `lags = 15` against a monthly demo series loaded by `load_demo_dataset()`. That is a concrete constraint worth internalising: 15 lags on a monthly series means 15 months of history are consumed before the first training row exists. The README does not state a minimum series length, so the practical floor depends on your lag choice and how many rows your estimator needs.
Beyond the recursive class, the repository topics list backtesting-forecasters, multi-series-forecasting, probabilistic-forecasting, exogenous-predictors and foundation-models. Those topic strings are the honest map of scope: the project is not one forecaster but a family, and the class you import determines whether you get one model per series or a single model with a series identifier as a feature.
Installing skforecast and running a first forecast
The README documents installation through PyPI and conda-forge, and the package name is the same in both. The Python version badges cover 3.10 through 3.14.
pip install skforecastThe conda-forge channel is the alternative the README links:
conda install -c conda-forge skforecastThen the quick example from the README, which needs LightGBM installed separately since it is not the library's own dependency:
import pandas as pd
from lightgbm import LGBMRegressor
from skforecast.recursive import ForecasterRecursive
from skforecast.datasets import load_demo_dataset
y = load_demo_dataset()
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 15
)
forecaster.fit(y=y)
predictions = forecaster.predict(steps=12)
print(predictions.head())The README shows the resulting object as a pandas Series indexed by month start, with the name `pred` and dtype float64. If you see that shape, the pipeline is wired correctly. If `load_demo_dataset()` fails, the download step is the first thing to check, because the demo series comes over the network rather than being bundled.
Backtesting, prediction intervals and what the README leaves out
The repository topics include backtesting-forecasters and prediction-intervals, and the About section claims validation methods for realistic performance evaluation. That is the strongest part of the pitch: a single fit and a single predict tells you almost nothing about whether a forecaster will hold up next quarter, and the library's answer is to score the model over successive historical windows.
What the README does not document is rollback, version pinning policy, or a migration path between major releases. The changelog file exists at the repository root, so release-by-release changes are traceable there, but the README itself gives no upgrade guidance. The release cadence visible in the tags is roughly every six to ten weeks, which means a pinned dependency is a real decision rather than a formality.
The recursive design has a second failure mode the documentation does not dwell on: because each prediction is fed back as a feature, a single bad step propagates. On series with regime changes or structural breaks, a direct multi-step approach that trains one model per horizon often behaves better. skforecast exposes both, and choosing between them is a modelling judgement the library cannot make for you.
Where skforecast is the wrong tool
If you want a library that inspects your series and picks a model, skforecast is not that. It is an adapter layer, and the estimator you hand it determines the quality ceiling. Passing a linear regression with 15 lags is a legitimate use and will produce a legitimate result, but nothing in the library will tell you that a seasonal naive baseline would have beaten it.
The library also assumes you can express your problem as a supervised regression on a pandas-indexed series. Irregular timestamps, hierarchical reconciliation across many aggregation levels, and series with heavy missing-value structure are not what the README describes. The topics list multi-series-forecasting but the README quick example is single-series, so anyone with thousands of short series should read the multi-series documentation before assuming the fit is as clean as the demo.
Finally, the README's own framing is machine learning first. If your series is short, or your horizon is long relative to the available history, a classical statistical model may be the more honest choice, and the ARIMA and SARIMAX topics suggest the library supports that path too, but the quick example steers you toward gradient boosting.
skforecast compared with Prophet, ARIMA and Sktime
The comparison people search for is skforecast against Prophet and against Sktime, and the difference is architectural rather than a matter of accuracy.
Prophet is a complete model. You hand it a dataframe with `ds` and `y` columns and it fits a decomposable trend plus seasonality curve with its own inference procedure. You get a forecast without choosing an estimator, and you give up the ability to swap in LightGBM or add your own lag features. skforecast does the opposite: it gives you no model at all and instead gives you the machinery to apply whatever regressor you bring.
ARIMA and SARIMAX, which appear in the repository topics, sit in a third position. They model the autocorrelation structure explicitly rather than treating lags as generic features, which makes them strong on short, stationary series and awkward when you want to mix in exogenous predictors of different types. skforecast's exogenous-predictors topic is the point of contrast: adding a promotion flag or a weather column is a column in a feature matrix, not a term in a likelihood.
Sktime is the closer neighbour, since it also builds a scikit-learn style forecasting API. The distinction the skforecast README draws is narrower and more explicit: it is about wrapping estimators, with backtesting and interval estimation attached to the forecaster object.
Licence, maintenance and the cost of upgrading
The licence is BSD-3-Clause, declared in pyproject.toml with the LICENSE and CITATION.cff files listed as licence files. That is a permissive licence, and it is compatible with commercial use without a copyleft obligation on your own code. It is not legal advice, and the CITATION.cff sitting alongside the licence file is a reminder that the project asks for academic citation as well as code reuse.
The repository is not archived, and the last push was on 2026-09-10. Releases in the repository are v0.22.0 on 2026-04-23, v0.23.0 on 2026-07-08 and v0.24.0 on 2026-08-24, while pyproject.toml declares version 0.25.0, so the working tree is ahead of the latest tagged release. For a pre-1.0 package, minor version numbers can carry breaking changes, and the changelog at the repository root is the file to read before bumping. Pinning to a tested version and moving deliberately is the lower-risk path.
Editorial conclusion
Adopt skforecast if you already have a scikit-learn workflow, want gradient boosting on lagged features, and need backtesting and prediction intervals in the same object. Do not adopt it if you want a single call that selects and fits a model for you, or if your series are so short that a lag window of 15 leaves almost no training rows. Before committing, check the changelog for the release that matches your pandas and scikit-learn versions, and confirm that the forecaster class you need (recursive, direct, or multi-series) is the one your horizon logic requires.
Frequently asked questions
How do I install skforecast?
The README documents PyPI and conda-forge as the two installation routes, with the package named skforecast in both. The Python version badges cover 3.10 through 3.14.
What is the best Python library for forecasting?
There is no single answer, and skforecast's own framing is narrow rather than universal: it is for forecasting with scikit-learn compatible models, statistical methods and foundation models. If your workflow is already scikit-learn based, that narrowness is the reason to pick it; if you want a self-contained model with no estimator choice, it is the reason not to.
Is time series forecasting hard?
The work skforecast removes is mechanical rather than conceptual: building lag features, refitting across a multi-step horizon, and scoring the result over historical windows. What it cannot remove is the modelling judgement, since the estimator you pass determines the quality ceiling.
What are the four types of forecasting models?
The repository does not categorise forecasting models into four types, so this question cannot be answered from what the project publishes. What the repository does show is a split by mechanism: recursive forecasters that feed predictions back as features, direct multi-step approaches, and multi-series forecasters.
Is Prophet better than ARIMA?
The repository does not compare Prophet with ARIMA. What it does show is that ARIMA and SARIMAX appear among its topics, while Prophet is not part of the project at all: skforecast wraps whatever scikit-learn compatible estimator you supply rather than shipping a single built-in model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/skforecast-skforecast)