# AutoGluon: Automated Machine Learning in Three Lines of Python

> AutoGluon is an Apache-2.0 Python library that automates model selection and ensembling for tabular, time series and multimodal data. It is easy to start and hard to fully control, which is exactly the trade-off to weigh before adopting it.

**autogluon/autogluon** — Fast and Accurate ML in 3 Lines of Code

- Repository: https://github.com/autogluon/autogluon
- Website: https://auto.gluon.ai/
- Stars: 10,665 · Forks: 1,188
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/autogluon-autogluon

## What AutoGluon automates, and who it is actually for

The README states the problem plainly: from classic ML algorithms to foundation models, the options keep multiplying, and the question of which one to use is left to the practitioner. AutoGluon answers that question by searching over model families and combining the results, rather than asking you to pick one. Its stated goal is strong predictive performance with a few lines of code.

The audience is therefore not the researcher who already knows that gradient-boosted trees will win on a given table. It is the engineer or analyst who has a dataframe, a label column, and a deadline, and who would otherwise spend a week comparing libraries. The repository is a monorepo with separate top-level directories for tabular, timeseries, multimodal, core, features and common, which reflects the fact that AutoGluon is really three predictors sharing one set of training utilities rather than one monolithic tool.

That split matters when you evaluate it. TabularPredictor, TimeSeriesPredictor and MultiModalPredictor are documented as separate entry points with separate quickstarts and separate API pages. Adopting AutoGluon for tabular work does not commit you to its computer vision or forecasting code, but it does mean the install surface is wider than a single-purpose library's.

## How the predictor fits, trains and predicts

The mechanism visible in the README is a fit/predict pair. You construct a predictor with a label name, call fit on a training file or dataframe, and call predict on held-out data. The presets argument controls how much search and ensembling happens. The README's own example uses presets="best".

What happens inside fit, according to the project's description, is that AutoGluon finds the combination of models that works best for the use case. The repository layout supports this: autogluon.core holds the training and hyperparameter machinery, autogluon.features holds feature generation, and autogluon.tabular holds the predictor itself. The published packages depend on each other through exact version pins, and pyproject.toml describes these sibling dependencies as being supplied dynamically from a single source of truth in core/_setup_utils.py, so the caps land in both the workspace lock and the published wheels.

That is an unusually disciplined arrangement for a project of this size, and it has a practical consequence: mixing autogluon.tabular from one release with autogluon.core from another is not a supported configuration. The pins are exact, not ranges.

## Installing AutoGluon and running a first tabular fit

The README gives a single install command and states support for Python 3.10 through 3.13 on Linux, macOS and Windows. It also points to an installation guide for GPU support, Conda installs and optional dependencies, so the pip command below is the minimal path, not the only one.

```bash
pip install autogluon
```

After installation, the quickstart is three lines. The first constructs a TabularPredictor bound to a label column, the second fits it against a training CSV, and the third writes predictions for a test CSV.

```python
from autogluon.tabular import TabularPredictor
predictor = TabularPredictor(label="class").fit("train.csv", presets="best")
predictions = predictor.predict("test.csv")
```

The string passed as label must be a column present in train.csv. The README does not state what happens if it is absent, so check your column names before the first run rather than after. Expect fit to be the slow part: presets="best" is the setting the project uses to advertise accuracy, and it is not the setting to choose when you are still exploring the shape of your data. The documentation's tutorials, linked per predictor in the README table, are where the preset names and their trade-offs are spelled out.

## Where AutoGluon is the wrong tool

The honest limitation is the one implied by the design. An ensemble that searches over many model families costs time and memory in proportion to how many it keeps. The README does not publish a resource budget for presets="best", and it does not document a rollback or checkpointing story for a fit that is interrupted. If your training job has a hard wall-clock limit, you are choosing a tool whose headline setting is not designed around one.

Latency is the second constraint. A stacked ensemble is not a single model, so serving it is not the same as serving one gradient-boosted tree. If your inference path has a strict per-request budget, the ensemble that wins on accuracy may lose on the metric you actually care about.

Interpretability is the third. If a reviewer needs to see one model's coefficients or one tree's splits, AutoGluon's answer of combining models is the opposite of what you want. The README does not claim to produce a single interpretable model, and nothing in the project's description suggests it does.

Finally, scale. The repository is organised around pandas-style tabular and time series work. Nothing in the README describes a distributed training mode, so treat very large datasets as an open question to test rather than an assumed capability.

## AutoGluon against XGBoost and against AutoML in general

The comparison people actually search for is AutoGluon versus XGBoost, and the difference is one of scope. XGBoost is a single algorithm. You choose it, you tune it, you own the result. AutoGluon treats that algorithm as one candidate among many and adds a layer that decides which candidates to keep and how to weight them.

If you already know boosted trees are right for your problem, XGBoost gives you a smaller dependency, a shorter fit and a model you can explain. AutoGluon gives you a better chance of finding a configuration you would not have tried, at the cost of the search itself. Neither is strictly better; they answer different questions.

The second comparison is AutoGluon against the general idea of AutoML. AutoML is the category, not a competing library. AutoGluon is one implementation of it, and the README positions it as covering tabular, time series and multimodal data under one project. A narrower AutoML tool that only handles tables will have a smaller install and a smaller surface area. The reason to pick AutoGluon over such a tool is the breadth: one library, three predictors, one release cadence.

## Licence, releases and the cost of keeping up

AutoGluon is Apache-2.0, with a LICENSE and a NOTICE file at the repository root. Apache-2.0 permits commercial use and modification and includes an explicit patent grant, and the NOTICE file is the mechanism by which attribution obligations are carried. This is a description of the licence text, not legal advice; if your organisation has a policy on NOTICE handling, route it through whoever owns that policy.

Upgrade cost is where this project asks something of you. The release history shows v1.5.0 on 2025-12-19, then v1.6.0 on 2026-08-05 and v1.6.1 the following day. That is a long quiet stretch followed by a rapid pair of releases, which is a pattern worth knowing about when you plan a pin. The last push to the default branch was on 2026-09-10.

The pyproject.toml is the part that should shape your upgrade process. The workspace pins sibling packages with exact equality requirements, and the file documents an exclude-newer setting used to keep dependency resolution reproducible. In practice that means upgrading one AutoGluon package in isolation is not the intended path. Upgrade the set, and re-run your own evaluation rather than assuming a patch release is behaviour-neutral.

## Conclusion

Adopt AutoGluon when you have a labelled table, a deadline, and no strong prior about which model family will win; the three-line TabularPredictor path gets you a stacked ensemble without hand-tuning. Do not adopt it when training time, memory or inference latency is the binding constraint, or when you need a single interpretable model rather than a weighted ensemble. Before committing, verify two things yourself: that the installed version matches the release you read about (v1.6.1 was tagged on 2026-08-06), and that your machine can hold the ensemble that presets="best" produces on your data. If it cannot, drop to a lighter preset rather than discovering the limit mid-fit.

## FAQ

### What is AutoGluon used for?

It automates machine learning on tables, time series and multimodal data, aiming for strong predictive performance with a few lines of code. The README presents it as a way to avoid choosing between the growing number of model families yourself.

### How do I install AutoGluon?

The README gives pip install autogluon, and states support for Python 3.10 through 3.13 on Linux, macOS and Windows. It points to a separate installation guide for GPU support, Conda installs and optional dependencies.

### How do I use AutoGluon for tabular data?

Import TabularPredictor, construct it with the name of your label column, call fit on your training data with a preset such as "best", then call predict on your test data. The README shows exactly this three-line sequence.

### Is AutoGluon open source and free?

Yes. The repository is licensed under Apache-2.0 and ships a LICENSE and a NOTICE file at the root. Apache-2.0 permits commercial use and modification, subject to the licence's attribution terms.

### Is AutoGluon the best option?

That depends on your constraint. AutoGluon is built around searching over models and combining them, so it favours accuracy over fit time, memory and inference latency. If any of those is your binding limit, a single-algorithm library is the better fit.

### What is AutoGluon tabular?

It is the TabularPredictor component, one of the three predictors the README documents alongside TimeSeriesPredictor and MultiModalPredictor. It lives in the tabular directory of the monorepo and is imported from autogluon.tabular.

## Sources

- [autogluon/autogluon on GitHub](https://github.com/autogluon/autogluon)
- [License: Apache-2.0](https://github.com/autogluon/autogluon/blob/master/LICENSE)
- [Project website](https://auto.gluon.ai/)
- [README](https://github.com/autogluon/autogluon/blob/master/README.md)
- [Releases](https://github.com/autogluon/autogluon/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/autogluon-autogluon
