# combo: a Python toolbox for combining ML models and scores

> combo wraps stacking, DCS, DES, EAC and LSCP behind scikit-learn-style classes so you can combine classifiers, clusterers, outlier detectors and raw scores. It suits researchers and competition teams who already have several trained models and want a principled way to merge them.

**yzhao062/combo** — (AAAI' 20) A Python Toolbox for Machine Learning Model Combination

- Repository: https://github.com/yzhao062/combo
- Website: https://pycombo.readthedocs.io
- Stars: 660 · Forks: 105
- Language: Python
- License: BSD-2-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/yzhao062-combo

## The gap combo fills between one model and a pile of models

Most teams do not stop at one estimator. A gradient boosting model handles the tabular signal, a logistic regression gives a stable linear baseline, a k-nearest-neighbours model catches local structure, and someone eventually asks which of them to ship. The usual answer is to pick the best validation score, which throws away the disagreement between models that is often the useful part. The other answer is to hand-write a voting or averaging layer, and that layer tends to grow one ad hoc branch at a time.

combo targets that second answer. The README describes it as a toolbox for combining machine learning models and scores, and frames model combination as a subtask of ensemble learning. The intended audience is visible in the framing: the README points at data science competitions and research use, and the repository ships an examples directory with separate scripts for classifier combination, clustering combination and outlier detector combination. This is a library for people who already have several trained models and want the combination step to be a named, parameterised object rather than a function they wrote on a Friday.

The scope is wider than classifier stacking. The README lists classification, clustering, anomaly detection and raw score combination as covered tasks, and names Stacking, DCS, DES, EAC and LSCP as the implemented approaches. That breadth is the actual pitch: one API surface for four kinds of combination problem.

## How the combination layer is structured

The mechanism that the README and the repository layout make visible is a scikit-learn-shaped wrapper around a set of base estimators. You build a Python list of already-configured estimators, pass that list into a combination class as base_estimators, call fit on training data, then call predict or predict_proba on unseen data. The combination object owns the base models and the rule that merges them, so the caller never touches the merge logic directly.

The README's API demo shows this with Stacking, importing from combo.models.classifier_stacking. The base estimators in that example are scikit-learn classifiers, but the README states that models and scores from scikit-learn, xgboost and LightGBM can be combined. That matters because it means the base models do not have to come from one library, and the examples directory contains a file named classifier_multiple_libs_example.py, which is consistent with that claim.

Under the hood the README credits numba and joblib for what it calls optimized performance with JIT and parallelization when possible. The qualifier is doing real work there. Not every combination algorithm reduces to array loops that numba can compile, and the README does not enumerate which ones do. The requirements.txt pins numba>=0.35 and joblib as hard dependencies, so both are installed whether or not a given algorithm uses them.

The dependency list also includes pyod. That is the same author's outlier detection library, and its presence is the clearest signal of where the anomaly detection combination classes get their detector interfaces from. If you are not doing outlier work, you are still installing pyod.

## Installing combo and running a first stacking experiment

The README recommends pip and explicitly says to make sure the latest version is installed. It gives three install forms. The plain install is the one to start with, and the --pre form is for pre-release versions when you want features that have not been tagged yet.

```bash
pip install combo
pip install --upgrade combo
pip install --pre combo
```

The README also documents a source install, which is what you want if you intend to read the combination code alongside your own experiments.

```bash
git clone https://github.com/yzhao062/combo.git
cd combo
pip install .
```

After installation, the README's API demo is the shortest real use. It builds a list of five scikit-learn classifiers, wraps them in Stacking, fits on training data, and then produces both hard labels and class probabilities from the same fitted object.

```python
from combo.models.classifier_stacking import Stacking

classifiers = [DecisionTreeClassifier(), LogisticRegression(),
               KNeighborsClassifier(), RandomForestClassifier(),
               GradientBoostingClassifier()]

clf = Stacking(base_estimators=classifiers)
clf.fit(X_train, y_train)

y_test_labels = clf.predict(X_test)
y_test_proba = clf.predict_proba(X_test)
```

What you should see is a fitted Stacking object that answers both predict and predict_proba, exactly as a scikit-learn estimator would. That is the point of the wrapper: the combined model drops into code that already expects a scikit-learn classifier. Note that the README's snippet assumes the five classifiers are already imported; it shows the combo import and leaves the scikit-learn imports to the reader. The examples directory has complete runnable versions, including classifier_stacking_example.py.

## Where combo is the wrong tool

The dependency and version story is the first real constraint. The README's installation section lists Python 3.5, 3.6, or 3.7 as supported, and setup.py carries the same classifiers. setup.py also declares Development Status :: 2 - Pre-Alpha, while the only tagged release in the release list is V0.1.0, marked as a stable release, dated 2020-02-19. Those two signals do not agree with each other, and neither one points at a library that has settled its packaging story. The repository's last push was on 2026-09-08, so the code is not abandoned, but a recent commit is not the same as a tested install path on current Python. Anyone deploying on a modern interpreter should treat the first pip install as an experiment, not a step in a runbook.

There is also a conceptual limit. combo combines models that already exist. It does not train them for you, does not tune their hyperparameters, and does not tell you which base models belong in the list. If your problem is that a single model underfits, adding a combination layer over five weak models will not fix it. The library assumes you have done the base-model work and are now deciding how to merge the results.

A third case: if your models are already inside a pipeline framework that has its own ensembling primitives, adding combo means maintaining a second abstraction for the same job. The README positions combo as covering classification, clustering, anomaly detection and raw scores in one API, which is genuinely useful when you need more than one of those. It is less useful when you need exactly one and your existing framework already provides it.

## combo against scikit-learn's own ensemble classes

The obvious comparison is scikit-learn, which already ships VotingClassifier, StackingClassifier and StackingRegressor. The difference is scope rather than quality. scikit-learn's stacking classes are supervised-learning classes: they take estimators, a final estimator and a cross-validation strategy, and they are maintained as part of a library with a published deprecation policy and a current Python floor.

combo's difference is that the same style of wrapper is applied to problems scikit-learn does not cover with an ensemble class. Clustering combination and outlier detector combination have no direct equivalent in scikit-learn's ensemble module, and the README names EAC and LSCP as algorithms in that space. If your task is combining clusterers or combining anomaly scores, scikit-learn is not an alternative at all, and combo is one of the few places those algorithms are packaged behind a consistent API. If your task is classifier stacking on modern Python, scikit-learn is the lower-risk choice and combo's advantage shrinks to the wider algorithm list.

That is the honest split. For classifier stacking alone, combo's value is the additional algorithms and the AAAI 2020 paper documenting them, not the stacking wrapper itself. For clustering and outlier combination, the comparison is against writing the algorithm yourself from the paper.

## Licence and the maintenance cost you are taking on

combo is BSD-2-Clause. That is a permissive licence: it allows use, modification and redistribution in source and binary form provided the copyright notice and the licence text are retained. It does not carry the patent grant language that BSD-3-Clause and Apache-2.0 include, and it does not require derivative works to be open sourced. For most internal and commercial use this is the least restrictive end of the common open source licences. This is a description of the licence text, not legal advice; if the combination layer ends up in a shipped product, have your own counsel read the LICENSE file.

The maintenance cost is the part to weigh. The release list shows one tagged release, V0.1.0, on 2020-02-19. Six years of commits without a second tag means the version you get from pip and the version in the master branch may not be the same code. Upgrading is therefore not a matter of bumping a version pin; it is a matter of deciding whether to track master. The README's own advice to install the latest version, and its --pre install form, both push you toward the moving branch rather than a stable tag.

There is also a dependency cost that is easy to miss. requirements.txt pulls in numba, joblib, pyod, scipy, scikit_learn, numpy and matplotlib. Numba in particular is sensitive to Python and numpy versions, and the README's stated Python range predates several numba releases. A dependency conflict on install is a plausible first experience, and the README does not document a fallback for it.

## Conclusion

Adopt combo if you already train several models and want stacking, DCS, DES, EAC or LSCP without writing the combination layer yourself; the unified API and the examples directory make the first experiment short. Do not adopt it if you need a supported library with a modern Python floor, since setup.py still declares Python 3.5 to 3.7 and the only tagged release is V0.1.0 from 2020-02-19. Before committing, check on PyPI which version pip install combo actually resolves to, and read combo/version.py in the master branch to see what the source tree calls itself.

## FAQ

### How do I install combo?

The README recommends pip and gives three forms: pip install combo for a normal install, pip install --upgrade combo to update, and pip install --pre combo to include a pre-release version. It also documents cloning the repository and running pip install . from the checkout.

### Which Python versions does combo support?

The README's installation section lists Python 3.5, 3.6, or 3.7 as required, and setup.py carries the same Python classifiers. No newer version is listed in either place.

### What algorithms does combo implement for model combination?

The README names Stacking, DCS, DES, EAC and LSCP as the advanced models it covers, and states that combination is supported for classification, clustering, anomaly detection and raw scores.

### Which libraries can the base models come from?

The README states that combo supports combining models and scores from scikit-learn, xgboost and LightGBM. The examples directory includes a file named classifier_multiple_libs_example.py, which is consistent with that claim.

### What licence is combo released under?

The repository's licence badge and LICENSE file point to BSD-2-Clause. That is a permissive licence requiring the copyright notice and licence text to be retained.

## Sources

- [License: BSD-2-Clause](https://github.com/yzhao062/combo/blob/master/LICENSE)
- [Project website](https://pycombo.readthedocs.io)
- [README](https://github.com/yzhao062/combo/blob/master/README.md)
- [Releases](https://github.com/yzhao062/combo/releases)
- [yzhao062/combo on GitHub](https://github.com/yzhao062/combo)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yzhao062-combo
