# lazypredict: the example output reports the same number twice in every row

> A library that trains forty-odd models on your data and prints a leaderboard, extended in the current line with forecasting, foundation models and experiment tracking. Its example leaderboard reports ROC AUC identical to balanced accuracy on every single row, which no real evaluation produces.

**shankarpandala/lazypredict** — Lazy Predict help build a lot of basic models without much code and helps understand which models works better without any parameter tuning

- Repository: https://github.com/shankarpandala/lazypredict
- Stars: 3,350 · Forks: 363
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/shankarpandala-lazypredict

## ROC AUC equals balanced accuracy on every row of the example

The example in the documentation is a breast cancer classification run split fifty-fifty, and the table below it is meant to be what the library prints. Its two middle columns are identical in every row. For the linear support vector machine, balanced accuracy 0.987544 and ROC AUC 0.987544. For logistic regression, 0.98269 and 0.98269. For the nearest neighbour classifier, 0.95503 and 0.95503. All twenty-odd rows behave this way. That is not a property of classifiers, it is a property of arithmetic: balanced accuracy depends on the decision threshold, ROC AUC does not, so they coincide only by accident and never across a whole table. Whatever produced this table, it did not compute two independent metrics, and a reader deciding which model to use is reading a duplicated column.

## Three unrelated algorithms report byte-identical rows

The duplication goes further than the column. LinearSVC and SGDClassifier carry the same accuracy, the same balanced accuracy, the same area under the curve and the same F1, differing only in time taken, which reads as a plausible pair of converging linear models. PassiveAggressiveClassifier, LabelPropagation and LabelSpreading are a different matter: a passive-aggressive linear classifier, a semi-supervised label propagation and a transductive label spreading algorithm have nothing in common architecturally, and they report 0.975439 accuracy, 0.974448 balanced accuracy, 0.974448 area and 0.975464 F1 between them. Four identical columns across three different algorithm families points at a table assembled rather than measured. A timing column that does vary is consistent with a real run behind those numbers, so the most charitable reading is that the metric columns were filled in by something other than the evaluation.

## The manifest says 0.3.0 and the newest release is 0.3.1

The package metadata and the release list do not agree about the current version. The build file declares version 0.3.0, and the release list tops out at 0.3.1 from 2026-08-29, with 0.3.0 from 2026-03-15 and an alpha of 0.3.0 from the day before that. The default branch is not master either, it is dev, so what the build file describes is the development line rather than the last tag. The practical effect for a consumer is small but real: someone reading the repository sees a version number one release behind what they can install, and the alpha in the release list shows a prerelease that was published to the same channel as the finals. Build metadata lives in the pyproject file and the setup script is now three lines long, with a comment pointing at it, and a separate setup configuration file sits beside both.

## The classifiers promise Python 3.14 and the documentation stops at 3.13

The interpreter support claims are written down three times and one of them is wider than the others. The feature list says support for Python 3.9 through 3.13. The metadata requires 3.9 or newer with no upper bound. And the trove classifiers enumerate 3.9, 3.10, 3.11, 3.12, 3.13 and then 3.14 as well, six interpreter versions where the documentation names five. The missing upper bound in requires-python is the part that matters operationally, because it means the installer will not stop you on an interpreter the classifiers never claimed, and the gap between the classifier list and the documentation means the documentation is the narrower of the two claims rather than the authority.

## Six dependency groups exist that the installation section never names

The installation section documents four extras and an everything extra: boosting libraries, forecasting, forecasting with deep learning, forecasting with a foundation model, and all. The metadata defines twelve groups. Beyond the documented ones there are groups for experiment tracking, hyperparameter tuning with a Bayesian optimiser, plotting, model explanation with SHAP values, a separate interpretation library, an automated tuning framework, and Spark. None of those seven appears in the installation instructions, so a user who wants plots or explanations has to guess the group name, and the fact that the packages involved are named in the dependency list is the only clue. The metadata also adds a data-encoder dependency to the everything group that has no group of its own and is required by two of the documented categorical encoding options.

## The requirements file installs the optional extras unconditionally

Two files describe what a project needs, and they are not the same list. The metadata has six required dependencies: a command line library, scikit-learn, pandas, a progress bar, joblib and numpy. The requirements file at the root adds three boosting libraries, two forecasting libraries and a tuning library on top of that:

```
click
scikit-learn>=1.0
pandas>=1.3
tqdm>=4.0
joblib>=1.0
numpy>=1.21
xgboost>=1.5
lightgbm>=3.0
statsmodels>=0.13
pmdarima>=2.0
optuna>=3.0
```

So a developer who follows the requirements file installs the optional extras by default, while a user who installs the package gets none of them. That is a reasonable split for a repository requirements file and a common source of confusion when a bug reproduces locally and not on a user's machine. The experiment tracking package is absent from both, so its integration is optional in a way the requirements file does not hint at.

## One extra is narrower than the package that contains it

The forecasting installation lines carry annotations, and one of them constrains the interpreter:

```bash
pip install lazypredict[timeseries]          # statsmodels + pmdarima
pip install lazypredict[timeseries,deeplearning]  # + LSTM/GRU via PyTorch
pip install lazypredict[timeseries,foundation]    # + Google TimesFM (Python 3.10-3.11)
```

The foundation group, which pulls a pretrained time series model, is annotated as Python 3.10 to 3.11 only, while the package itself requires 3.9 or newer with no ceiling and its classifiers run to 3.14. Nothing in the metadata encodes that restriction: the dependency is listed with no environment marker, so the resolver on 3.12 or 3.13 will attempt the install rather than decline it. The constraint exists only in a comment on an installation line, which means it is discovered by reading the documentation and not by installing.

## Two release automation systems and four process documents at the root

The root directory tells two different stories about how the project is run. For releases there are two independent mechanisms: a release-please manifest and configuration, and a separate bumpversion configuration file, both of which will want to own the version string that the metadata also declares. For process there are four documents that are not code: a CI pipeline audit and plan, a library audit and improvement plan, an issue closure report and a roadmap, next to a change history, a contributing guide, an authors file, a citation file and a paper in both markdown and bibliography form. The build itself is the stock Makefile from a Python package template, down to the inline Python that prints the help screen and opens the coverage report in a browser. Its phony declaration covers seven targets while the file defines about ten, so lint, test, coverage and the release target are not declared phony and a file of the same name in the working directory would win.

## Conclusion

Use it for what it genuinely is good at, a first pass that tells you which model families are even in the running on your data before you spend time tuning. Do not quote its example table as evidence of anything, because the numbers in it cannot have come from an evaluation as described. Three things to check. The published version and the manifest disagree, so pin deliberately. The install surface is much wider than the installation section admits, with six dependency groups that exist in the metadata and not in the documentation, one of which narrows the supported Python range. And the requirements file installs the optional extras unconditionally, so a developer environment built from it is not the environment a user gets from the package index.

## FAQ

### what is lazy predict

A Python library that fits many standard models on your data in one call and prints a comparison table, so you can see which families are competitive before tuning anything. It covers classification, regression and time series forecasting, claims over forty built-in models including more than twenty forecasting models, and adds automatic seasonal period detection and built-in experiment tracking.

### How do I install lazypredict with the extra dependencies?

The package installs with pip or conda, and optional groups cover boosting libraries, forecasting, forecasting with LSTM and GRU through PyTorch, and forecasting with a pretrained foundation model. There is also an all group. The documentation names four extras plus that one, while the metadata defines twelve groups, so plotting, explanation, tuning and Spark support exist but are not documented in the installation section.

### Which Python versions does lazypredict support?

The feature list says Python 3.9 through 3.13 and the metadata requires 3.9 or newer with no upper bound, while the trove classifiers enumerate 3.9 through 3.14. One optional group, the pretrained foundation model, is annotated as usable only on Python 3.10 and 3.11.

### Can lazypredict use a GPU?

There is a flag for it, off by default, and the feature list names what it accelerates: XGBoost, LightGBM, CatBoost, the RAPIDS machine learning library, the LSTM and GRU models, and the foundation model. The GPU flag is a parameter on the classifier and regressor rather than a separate install, and the accelerating packages are optional rather than required.

## Sources

- [Issues](https://github.com/shankarpandala/lazypredict/issues)
- [License: MIT](https://github.com/shankarpandala/lazypredict/blob/dev/LICENSE)
- [README](https://github.com/shankarpandala/lazypredict/blob/dev/README.md)
- [Releases](https://github.com/shankarpandala/lazypredict/releases)
- [shankarpandala/lazypredict on GitHub](https://github.com/shankarpandala/lazypredict)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shankarpandala-lazypredict
