Library / SDK
rasbt/mlxtend avatar
rasbt/mlxtend

mlxtend: stacking, Apriori and plotting helpers that scikit-learn leaves out

A library of extension and helper modules for Python's data analysis and machine learning libraries.

5,181 stars920 forksPythonNOASSERTION

At a glance

What is it?
mlxtend is a Python library of extensions to scikit-learn and pandas, covering ensemble classifiers, sequential feature selection, association rule mining and plotting utilities. It is a companion toolkit, not a replacement for scikit-learn, and the two overlap in ways worth knowing before you install it.
Who is it for?
Adopt mlxtend if you already work in scikit-learn and need stacking classifiers, sequential feature selection, Apriori or fpgrowth, or decision region plots, and you accept that its classes follow the fit and predict API rather than scikit-learn's own estimators.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap mlxtend fills between scikit-learn and pandas

scikit-learn gives you estimators, pipelines and metrics. It does not ship association rule mining, and its ensemble module does not include a voting classifier with per-model weights in the form mlxtend exposes. The README lists the library's stated purposes: ensemble methods such as stacking and voting classifiers, feature selection and extraction, visualization utilities including decision regions and confusion matrices, plotting helpers, and frequent pattern mining with the Apriori algorithm.

The audience is the working data scientist who already has scikit-learn installed and hits one of those gaps mid-notebook. You are not switching frameworks. You import one class, fit it on the same arrays, and keep the rest of your pipeline. The library also carries a JOSS paper from 2018, which is unusual for a utility package and tells you the author treated it as citable infrastructure rather than a scratchpad.

Where it is weaker is breadth. There is no deep learning, no serving layer, no data versioning. If your problem is model deployment rather than model inspection, mlxtend has nothing for you.

How the ensemble and feature selection classes actually work

The README's example constructs three scikit-learn classifiers, wraps them in EnsembleVoteClassifier with weights [2, 1, 1] and voting='soft', then fits each one separately for plotting. The wrapper is a meta-estimator: it holds references to the base classifiers, and on fit it delegates to each of them in turn. On predict with soft voting it averages the predicted probabilities, weighted, and takes the argmax. That is the mechanism, and it explains the constraint you will hit: soft voting requires every base classifier to expose predict_proba, which is why the README passes probability=True to SVC. Drop that argument and soft voting fails at predict time, not at construction time.

Sequential feature selection in mlxtend follows the same pattern. It is a wrapper that adds or removes features one at a time, refitting a supplied estimator at each step and scoring it with a supplied metric. The cost is linear in the number of candidate features times the number of steps, so on a wide dataset it is slow by construction. The library does not parallelize this for you; the pyproject.toml lists joblib as a dependency, which suggests some parallel paths exist, but the README does not document the selection classes' parallel behaviour, so treat throughput as something you measure yourself.

Association rule mining is a different code path entirely. Apriori operates on a one-hot encoded DataFrame, typically produced by TransactionEncoder, and returns frequent itemsets and rules with support, confidence and lift. It is not a scikit-learn estimator and does not follow fit and predict. That split is visible in the package layout: classifier, feature_selection, frequent_patterns, plotting and data are separate subpackages.

Installing mlxtend and running the README example

The README documents uv as the install path. To add mlxtend to a uv-managed project:

bash
uv add mlxtend

For a one-off check without touching the current project, the README gives this command, which imports the package and prints its version:

bash
uv run --with mlxtend python -c "import mlxtend; print(mlxtend.__version__)"

If you are on plain pip, the package is published on PyPI as mlxtend, so pip install mlxtend resolves the same distribution. The README's own install section only shows uv, so pip is inferred from the PyPI link rather than documented there.

The README also notes that the PyPI version may lag the repository, and gives a dev install:

bash
uv add "mlxtend @ git+https://github.com/rasbt/mlxtend.git"

For a first real use, the README's ensemble example is the shortest path to something visible. It loads iris data through mlxtend.data.iris_data, selects two columns, and plots decision regions for each base classifier plus the ensemble:

python
from sklearn.linear_model import LogisticRegression
from sklearn.svm import SVC
from sklearn.ensemble import RandomForestClassifier
from mlxtend.classifier import EnsembleVoteClassifier
from mlxtend.data import iris_data
from mlxtend.plotting import plot_decision_regions

X, y = iris_data()
X = X[:, [0, 2]]
clf1 = LogisticRegression(random_state=0)
clf2 = RandomForestClassifier(random_state=0)
clf3 = SVC(random_state=0, probability=True)
eclf = EnsembleVoteClassifier(clfs=[clf1, clf2, clf3], weights=[2, 1, 1], voting='soft')

What you should see after fitting and calling plot_decision_regions is a filled background partitioning the two-dimensional feature space by predicted class, with the training points overlaid. If the plot is blank or the axes are unlabelled, the usual cause is a classifier that was not fitted before being passed in, since plot_decision_regions calls predict on the object it receives.

Where mlxtend breaks or is the wrong choice

The dependency floor is the first real constraint. pyproject.toml declares requires-python >= 3.11 and pins scipy>=1.16.3, numpy>=2.3.5, pandas>=2.3.3, scikit-learn>=1.8.0, matplotlib>=3.10.8 and joblib>=1.5.2. On a locked environment running Python 3.10 or numpy 1.x, installing mlxtend will force upgrades across your entire scientific stack. That is a migration, not an install.

Soft voting with an unfitted base classifier is a silent trap in the sense that the failure surfaces late. The README's example passes probability=True to SVC specifically to make soft voting work; nothing in the wrapper checks this for you at construction time.

Apriori on dense transaction data is the other place it disappoints. The algorithm enumerates candidate itemsets level by level, so runtime grows with the number of frequent itemsets rather than the number of rows. On a retail basket dataset with thousands of distinct items and low support thresholds, memory use climbs quickly. The library also exposes fpgrowth, which the search data shows people look for; FP-growth avoids the candidate generation step and is usually the better choice on the same data, but the README does not compare the two, so the choice is left to you.

Finally, mlxtend is not a framework. It does not manage experiments, serialize pipelines for serving, or provide a model registry. Teams expecting a platform will be disappointed by the scope.

mlxtend versus scikit-learn: overlapping ground and different defaults

The comparison people search for is mlxtend versus sklearn, and the honest answer is that they are not competitors. scikit-learn is the estimator library; mlxtend imports from it and adds around it. The overlap is narrow but real. scikit-learn has its own SequentialFeatureSelector, and mlxtend has SequentialFeatureSelector too. Both wrap an estimator and add or remove features greedily. The difference is in defaults and in the surrounding API: scikit-learn's version is built to sit inside a Pipeline and follows the transformer interface with fit and transform, while mlxtend's selection classes are documented as standalone selectors that expose the chosen indices. If your workflow is pipeline-centric, scikit-learn's version composes more cleanly. If you want the selection step to hand you indices and a plot, mlxtend's is more direct.

Voting and stacking are the clearer case. scikit-learn has VotingClassifier and StackingClassifier. mlxtend's EnsembleVoteClassifier predates them and exposes per-classifier weights in the constructor, which is the same idea. There is no strong reason to prefer one over the other on capability alone; the reason to pick mlxtend is that you are already using its plotting or frequent pattern modules and want one dependency instead of two.

For association rules there is no scikit-learn equivalent at all, which is the strongest argument for the library. If Apriori or fpgrowth is what brought you here, mlxtend is not competing with scikit-learn, it is filling a hole scikit-learn left open.

Versioning, maintenance and licence terms

The last push to master was on 2026-09-08, and the most recent release is v0.25.0 from 2026-06-06. Before that, v0.24.0 landed on 2025-12-13 and v0.23.4 on 2025-01-26. The gap between v0.23.4 and v0.24.0 is roughly eleven months, which tells you releases are infrequent and can bundle a wide set of changes. Read the CHANGELOG before bumping a pinned version rather than assuming patch-level compatibility.

The version is dynamic, pulled from mlxtend.__version__ via setuptools, so the installed version always reflects the source tree you built from. For the dev install path shown earlier, that means you get whatever is on master, unreleased and possibly mid-refactor.

On licensing, the README states the project is released under a permissive new BSD open source license, with the file LICENSE-BSD3.txt in the repository, and that it is commercially usable with no warranty, not even for merchantability or fitness for a particular purpose. The pyproject.toml classifier says BSD 3-Clause. Separately, artistic creative works such as figures and images are covered by Creative Commons Attribution 4.0, with LICENSE-CC-BY.txt in the repository, while computer-generated graphics such as matplotlib plots fall under the BSD license. If you redistribute the documentation figures, that distinction matters; if you only use the code, the BSD terms apply. This is a description of what the repository states, not legal advice.

Upgrade cost is mostly the dependency floor. Each release raises the numpy, scipy and scikit-learn minimums, so a version bump in mlxtend can cascade into a coordinated upgrade of the whole stack.

Editorial conclusion

Adopt mlxtend if you already work in scikit-learn and need stacking classifiers, sequential feature selection, Apriori or fpgrowth, or decision region plots, and you accept that its classes follow the fit and predict API rather than scikit-learn's own estimators. Skip it if you need a maintained, single-vendor ML framework, if you cannot tolerate a dependency floor of Python 3.11 with numpy 2.3.5 and scikit-learn 1.8.0, or if the plotting helpers are the only thing you want, since matplotlib alone may cover that. Before committing, install it in a throwaway environment, run the EnsembleVoteClassifier example from the README, and check the CHANGELOG for the version you pin.

Frequently asked questions

Why is there no module named 'mlxtend' in Python?

That error means the package is not installed in the interpreter you are running, not that the module is missing from the project. Install it with uv add mlxtend or pip install mlxtend in the same environment as the notebook or script.

How do I install mlxtend?

The README documents uv add mlxtend for a uv-managed project, and a one-off check via uv run --with mlxtend python -c "import mlxtend; print(mlxtend.__version__)". The package is published on PyPI, so pip install mlxtend resolves the same distribution.

How do I import the plot_decision_regions function in mlxtend plotting?

Import it from the plotting subpackage with from mlxtend.plotting import plot_decision_regions, as the README's ensemble example does. The classifier you pass in must already be fitted, since the function calls predict on it.

What is mlxtend used for?

The README lists ensemble methods such as stacking and voting classifiers, feature selection and extraction, visualization utilities including decision regions and confusion matrices, plotting helpers, and frequent pattern mining with the Apriori algorithm.

How does mlxtend compare with scikit-learn?

mlxtend imports from scikit-learn and extends it rather than replacing it. The overlap is narrow: both have a sequential feature selector and both have voting and stacking classifiers, while association rule mining has no scikit-learn equivalent.

Official sources

  1. Issues
  2. Project website
  3. rasbt/mlxtend on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rasbt-mlxtend.svg)](https://hysenlabs.com/projects/rasbt-mlxtend)