mljar-supervised: AutoML for tabular data that writes its own report
Python package for AutoML on Tabular Data with Feature Engineering, Hyper-Parameters Tuning, Explanations and Automatic Documentation
At a glance
- What is it?
- mljar-supervised trains and tunes tabular models across four modes and leaves a Markdown report behind. Here is how it installs, what the reports contain, and where the design gives out.
- Who is it for?
- Adopt mljar-supervised if you need a documented baseline on tabular data and want to inspect the pipeline rather than trust a black box. Skip it if you need unstructured data, streaming training, or a sub-second prediction service, since the package pulls in xgboost, lightgbm, catboost, shap and optuna-integration and trains many models per run.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 66 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem mljar-supervised solves, and who feels it
Most tabular machine learning work is not research. It is a data scientist preprocessing a CSV, trying LightGBM, then XGBoost, then CatBoost, tuning each, and writing up what happened. The README describes the package as one that "abstracts the common way to preprocess the data, construct the machine learning models, and perform hyper-parameters tuning to find the best model." That is the job it takes over.
The target user is a data scientist or analyst who has a DataFrame and a target column, and needs a defensible model plus a record of how it was built. The README lists the intended outcomes directly: explaining data through model reports, feature importance and SHAP explanations, trying many algorithms with hyperparameter tuning, creating Markdown reports, generating a web app for a trained model, and saving and re-running an analysis. That last item matters more than it sounds. An AutoML run that cannot be replayed is an anecdote, not a result.
The package is not aimed at people who want to skip model understanding. It is aimed at people who want the search automated but the evidence kept.
Four modes, one pipeline, and the greedy ensemble
The mode you pick determines how much search happens. Explain mode is for understanding data and produces decision tree visualizations, linear model coefficients, permutation importance and SHAP explanations. Perform builds pipelines meant for production. Compete trains heavily tuned models with ensembling and stacking for competitions. Optuna mode, available since version 0.10.0 according to the README, searches for highly tuned models when performance matters more than compute time.
The algorithm roster is fixed and named in the README: Baseline, Linear, Random Forest, Extra Trees, LightGBM, XGBoost, CatBoost, Neural Networks and Nearest Neighbors. Ensembling uses a greedy algorithm from the Caruana paper, and stacking builds a level 2 ensemble, available in Compete mode or by setting the stack_models parameter. Preprocessing covers missing value imputation, categorical conversion and target preprocessing. Feature work includes Golden Features, feature selection, and text and time transformations.
The transparency claim rests on output, not on internals. The README states the package "is no black box, as you can see exactly how the ML pipeline is constructed (with a detailed Markdown report for each ML model)." That is the honest way to read it: the mechanism is a search over a fixed set of learners, and the audit trail is a per-model report. If you want to change the search itself rather than its parameters, this is the wrong layer.
Installing mljar-supervised and running a first fit
Installation goes through pip or conda. The setup.py declares python_requires of 3.9 or newer and classifies 3.9 through 3.12, so check your interpreter before anything else. The package name on PyPI is mljar-supervised, and the README's installation section points at both PyPI and conda-forge badges.
pip install mljar-supervisedThe first real use is a supervised fit. The README's app example shows the shape of the API, with a results_path passed to the constructor and X and y passed to fit.
from supervised import AutoML
automl = AutoML(results_path="AutoML")
automl.fit(X, y)After the run, results_path holds the analysis. The README says each ML model gets a detailed Markdown report, so expect a directory of reports rather than a single file. If you want the explanations rather than a production model, set the mode to Explain in the constructor; the README lists Explain as the mode for understanding data, with decision tree visualization and SHAP output among its products. A first run is slow by design, since the package trains and tunes several algorithms before ensembling.
Turning a trained model into a prediction app
This is the part that separates mljar-supervised from most AutoML libraries. After fit, the README shows three calls: app() to generate the web app files, local_app() to start it locally, and publish_app() to publish it. Publishing creates an app URL on the first call and then reuses the last successfully published URL by default.
from supervised import AutoML
automl = AutoML(results_path="AutoML")
automl.fit(X, y)
automl.app()The generated app can include a single prediction dashboard, batch prediction from CSV files, downloadable predictions, feature importance plots, and feature context plots for single predictions. The app layer is built on Mercury, a separate project from the same organization. The README frames the whole path as CSV data to trained model to prediction web app.
Be clear about what this is and is not. It is a way to hand predictions to domain experts who do not work in Python. It is not a model serving framework with latency guarantees, versioned endpoints or autoscaling. The README does not document rollback for a published app, and publish_app() reusing the last URL means the app identity is stable while its contents change.
Where mljar-supervised is the wrong tool
The scope is tabular data, stated in the first sentence of the README. There is no image, audio or text-classification pipeline here beyond the text transformations mentioned for tabular columns. If your problem is a transformer fine-tune, this package will not help.
Compute cost is the second constraint. requirements.txt pins xgboost>=2.1.0, lightgbm>=4.3.0, catboost>=1.2.8, shap>=0.46.0 and optuna-integration>=3.6.0, among others. Every install carries all of those gradient boosting libraries plus their transitive dependencies. A run in Compete or Optuna mode trains and tunes many models, so the wall-clock cost scales with the search rather than with one fit. On a laptop with a modest CPU budget, a single dataset can occupy the machine for a long stretch. The README does not publish runtime estimates, so plan by measuring on your own data rather than by expectation.
The third case is latency. Nothing in the README describes a serving mode with a response time target. If you need a model behind a strict millisecond budget, train here, then export the winning model and serve it with the library that produced it.
Finally, the reports are the deliverable, and they are Markdown. Teams that need results piped into an experiment tracker will find that the README documents saving and re-running analyses through the package's own results_path, not integration with external tracking systems.
How it compares to PyCaret and to a single CatBoost run
PyCaret is the closest comparison and appears in the related searches for this project. Both wrap preprocessing, model comparison and tuning behind a short API. The difference is in what they hand back. mljar-supervised's stated emphasis is the per-model Markdown report and the explanation surface: decision tree visualizations, linear coefficients, permutation importance and SHAP. PyCaret's emphasis is a compact experiment workflow with its own plotting and deployment helpers. If your review process needs a written artifact per model, mljar-supervised's output is closer to that shape. If you want a tight interactive loop in a notebook, the two are closer in feel and the choice comes down to which output you actually file.
The other alternative is not a framework at all. If you already know CatBoost wins on your data, running CatBoost directly with a small tuning loop gives you full control and one dependency instead of a stack. mljar-supervised earns its keep when you do not yet know which learner wins, or when you need the comparison documented. When the answer is already known, the extra search is overhead.
Maintenance, licence and the cost of upgrading
The repository is not archived, and the last push was on 2026-07-27. Release v1.3.2 landed the same day, following v1.3.1 on 2026-06-11 and v1.3.0 on 2026-05-29. That is a steady cadence across the recent window, and the setup.py version matches the v1.3.2 tag.
The licence is MIT, declared in setup.py and in the repository's LICENSE file. MIT permits commercial use and modification with the copyright notice retained. That is a permission, not legal advice; if you redistribute the package inside a product, have your own counsel read the file.
Upgrade cost is dominated by the dependency floor, not by the package's own API. requirements.txt requires numpy>=2.0.0,<3.0.0, pandas>=2.2.2 and scikit-learn>=1.5.0. Those floors move with the wider ecosystem, and a major scikit-learn or numpy release can force a coordinated upgrade of the whole environment. Pinning the full environment is the practical move. The README does not document a migration path between major versions, so read the release notes for each of v1.3.0 through v1.3.2 before jumping.
Editorial conclusion
Adopt mljar-supervised if you need a documented baseline on tabular data and want to inspect the pipeline rather than trust a black box. Skip it if you need unstructured data, streaming training, or a sub-second prediction service, since the package pulls in xgboost, lightgbm, catboost, shap and optuna-integration and trains many models per run. Before committing, run one dataset in Explain mode and open the generated Markdown report to confirm the model list and explanations match what your team expects.
Frequently asked questions
What is mljar-supervised and who is it for?
It is an Automated Machine Learning Python package for tabular data, described in the README as designed to save time for a data scientist by abstracting preprocessing, model construction and hyperparameter tuning. It suits analysts who have a DataFrame and a target column and need a trained model plus a written record of how it was built.
How do I install mljar-supervised?
Install it from PyPI with pip install mljar-supervised, or from conda-forge. The setup.py declares python_requires of 3.9 or newer and classifies Python 3.9 through 3.12.
What are the available modes in mljar-supervised?
The README lists four: Explain for understanding data with decision trees, linear coefficients, permutation importance and SHAP; Perform for production pipelines; Compete for heavily tuned models with ensembling and stacking; and Optuna for maximum performance when compute time is not limited, available since version 0.10.0.
Can mljar-supervised generate a web app for a trained model?
Yes. The README shows automl.app() to generate the app files, automl.local_app() to start it locally, and automl.publish_app() to publish it, with the app built on Mercury. The generated app can include a single prediction dashboard, batch prediction from CSV, downloadable predictions, and feature importance and feature context plots.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mljar-mljar-supervised)