Model or dataset
sb-ai-lab/LightAutoML avatar
sb-ai-lab/LightAutoML

LightAutoML has two project names and gives macOS a different lightgbm than Linux

Fast and customizable framework for automatic ML model creation (AutoML)

1,479 stars74 forksPythonApache-2.0

At a glance

What is it?
LightAutoML is a Python AutoML framework that offers presets or hand-assembled pipelines over tabular, time series, image and text data. The repository documents itself more carefully than most: the manifest is honest about what it pins, the source tree carries two build systems and two CI systems, and the tutorial index has one entry describing the wrong notebook.
Who is it for?
LightAutoML earns its place if you want an interpretable, inspectable AutoML pipeline rather than a black box, and the WhiteBox path plus the ICE and PDP material show real attention to that. Two practical warnings.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The project has two names depending on which file you read

The Poetry manifest declares the project as `LightAutoML`, capitalised, with the version, the license and the documentation homepage alongside it. A second build file at the root declares it in lower case:

python
#!/usr/bin/env python

# we use poetry for our build, but this file seems to be required
# in order to get GitHub dependencies graph to work

import setuptools


if __name__ == "__main__":
    setuptools.setup(name="lightautoml")

The comment is candid about the reason: Poetry is the real build, and the setuptools file exists only so that GitHub's dependency graph tooling has something to read. The hedge in that comment is fair, since a graph that resolves the lower case name is not necessarily resolving the same artefact a user installs. The package index links at the top of the page use the lower case spelling.

Two build systems for one project is not unusual, but keeping the names in agreement would cost nothing, and the moment it matters is when a dependency scanner reports a name that differs from the one on PyPI.

Seven authors in the page, six in the manifest

Two author lists exist and they do not match.

The README credits seven people by name and links most of them to Kaggle profiles rather than to addresses. The Poetry manifest lists six, and the difference is not only count: one person named in the README has no entry in the manifest, and one name is spelled two ways across the two files, with a doubled letter in the manifest version.

The ordering also differs, and the manifest sorts differently from the page rather than following it. Alongside those differences, the manifest carries personal email addresses for each contributor while the page carries public profile links, so the two files publish different amounts of contact information about the same people.

None of this affects the code. It does mean that a reader reconstructing who worked on a release has two answers, and an automated tool reading author metadata will see a smaller set than a human reading the page.

The dependency floor is old and the versions split by interpreter

The manifest requires Python 3.8 or newer, and the classifiers run from 3.8 through 3.12, marked as production stable, OS independent and typed. Below that floor the manifest becomes conditional.

NumPy is unpinned below Python 3.10 and requires a specific minimum from 3.10 upward. SciPy is unpinned from 3.9 and is capped to an older release below that, which is the kind of upper bound that eventually conflicts with something else in the tree. The pattern is defensive rather than sloppy, but it does mean the resolved environment differs substantially depending on which interpreter you install under.

The sharpest split is not by interpreter but by platform. LightGBM is required at 2.3 or newer everywhere, except on macOS, where the manifest asks for 4.4 specifically and carries an inline comment linking to an upstream issue as the justification. A developer on a Mac and a developer on Linux therefore resolve different major versions of the same library from the same manifest, which is the kind of difference that surfaces as a model behaving differently on a laptop than in a container.

Three boosting libraries and PyTorch are required, not optional

The required dependency set is large for a tabular problem. LightGBM, CatBoost and XGBoost are all present, alongside Optuna for optimisation, a date and holiday library, a time series library capped at an older release, a graph library, and Jinja templating.

Two entries stand out for what they imply about the default install. PyTorch is a required dependency at a minimum version rather than an extra, so a user who only wants a binary classification on a small CSV still installs a deep learning stack. And a package named `autowoe` is required at a specific minimum, which is the weight of evidence encoding that the interpretable preset path depends on.

That second entry explains the shape of the tutorial list. One notebook is dedicated to building interpretable models with the WhiteBox preset, and another covers local and global interpretation of results using ICE and PDP approaches. Neither would be the default path without that dependency being pulled in for everyone, which is the trade this manifest makes: one dependency set that serves both the quick tabular preset and the interpretable pipeline.

Tags carry an unusual prefix and lag the branch by months

Three releases are listed, and each is tagged with a dot between the prefix and the number, while the release title inside drops the dot and the packaging manifest uses neither. Three spellings of one version is a small thing that will trip a script matching tags.

The dates are the more useful signal. The newest release is dated in late 2025, the one before it in early 2025, and the one before that at the end of 2024, so the intervals are irregular rather than slow or fast. The most recent push is dated at the end of September 2026 and the repository is not archived, which puts roughly ten months of commits outside the newest published release.

The packaging manifest still carries that older release number, so anyone building from a checkout gets a version string that does not describe the code. For a library whose selling point is reproducibility, that gap is the number to keep in mind before quoting a version in a paper.

One tutorial's description belongs to a different notebook

Eleven tutorial notebooks are indexed, from a basics introduction through custom pipelines, neural networks, relational data, a SQL data source, uplift modelling and time series. Each entry has a one line description, and one of them is wrong.

The notebook covering neural networks is described as an example of using the tabular preset with neural networks. The next entry, whose filename refers to relational data with a star schema, carries the identical sentence. The description was copied from the previous entry and never changed, so a reader deciding which notebook to open is told the relational one is about neural networks.

The Kaggle resource list above it has a smaller version of the same carelessness: one entry titled as custom machine learning pipeline elements inside existing ones appears twice, pointing at the same address both times.

Both are trivially fixable and both are the kind of error that survives in a project with this many resources, which is a fair proxy for how much of the documentation is maintained by hand.

Two CI systems, a linter set, and a demo that deletes its own report

The source tree carries configuration for two hosting services at once, a directory for GitHub and a directory for GitLab, alongside a tox file, a setup configuration, a pre-commit configuration and a script at the root that checks the documentation. The docs have their own readthedocs configuration and a Jupyter book config for building them.

Two operational notes in the tutorial section are worth carrying into any run. The first says the profiler increases work time and memory consumption, so it should be left off in production, and it is off by default. The second is the more awkward one: to look at the report after a demo finishes, you have to comment out the demo's last line, because that line is what deletes the report.

So the example scripts are written to clean up after themselves by default, and inspecting the output means editing the file first. For a tutorial that is the safer default, since nobody accumulates reports, but it means the quick tour you run on your own machine ends with nothing to look at unless you know to change one line.

Editorial conclusion

LightAutoML earns its place if you want an interpretable, inspectable AutoML pipeline rather than a black box, and the WhiteBox path plus the ICE and PDP material show real attention to that. Two practical warnings. The install is heavy for a tabular problem, since PyTorch sits in the required set next to three gradient boosting libraries, so read the dependency list before you commit to it in a constrained environment. And the release history is slow enough that a pin means something: the newest tag is many months behind the branch, and tags are written with an unusual prefix, so verify what you actually installed rather than trusting a version string.

Frequently asked questions

What is LightAutoML used for?

It builds machine learning models either from ready-made presets or by assembling a custom pipeline from reusable blocks, and it covers tabular, time series, image and text data. The documented entry point creates a preset with a task name and metric, fits it with column roles, and then predicts, returning out of fold predictions as well.

How do I install LightAutoML?

It is published on PyPI under the lower case name lightautoml. The build uses Poetry and requires Python 3.8 or newer, with classifiers covering 3.8 through 3.12, and the project is marked OS independent, typed and production stable.

Does LightAutoML need PyTorch installed?

Yes, PyTorch is a required dependency at a minimum version rather than an extra, alongside LightGBM, CatBoost, XGBoost and Optuna. A tabular-only run still installs the deep learning stack as part of the default dependency set.

How do I get an interpretable model out of LightAutoML?

There is a WhiteBox preset covered by a dedicated tutorial for building interpretable models, and another tutorial covering local and global interpretation with ICE and PDP. The default dependency set includes a package named autowoe, which supplies the weight of evidence encoding that this path uses.

Why did my LightAutoML demo not leave a report?

The demo's last line deletes the report, so the documented workaround is to comment that line out before running. A separate note says the profiler should stay off in production because it increases work time and memory consumption, and that it is disabled by default.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. sb-ai-lab/LightAutoML on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/sb-ai-lab-lightautoml.svg)](https://hysenlabs.com/projects/sb-ai-lab-lightautoml)