CLI tool
csinva/imodels avatar
csinva/imodels

csinva/imodels: scikit-learn compatible interpretable models for tabular data

Interpretable ML package 🔍 for concise, transparent, and accurate predictive modeling (sklearn-compatible).

1,621 stars143 forksJupyter NotebookMIT

At a glance

What is it?
imodels packages rule lists, rule sets, greedy trees and sparse linear models behind the familiar fit and predict API. It is a good fit when a stakeholder has to read the model, and a poor fit when you need a differentiable model or a GPU pipeline.
Who is it for?
Adopt imodels when the deliverable is a model a domain expert has to read, and your data is a tabular array you can hand to scikit-learn. Do not adopt it as a general replacement for gradient boosting on wide sparse text or image inputs, and do not expect the Bayesian rule list or rule set to finish quickly on a large dataset.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem imodels addresses, and who feels it

A random forest or a gradient-boosted model usually wins a tabular benchmark, and then someone asks why a particular applicant or patient was scored the way they were. Answering that with post-hoc feature importances or SHAP values gives you an approximation of the model's reasoning, not the model's reasoning. imodels takes the other route: it gives you a library of models whose structure is the explanation. The README describes the package as fitting and using state-of-the-art interpretable models, all compatible with scikit-learn, and states that these models can often replace black-box models such as random forests with simpler models such as rule lists.

The audience is narrow and identifiable. It is the analyst or data scientist who has to put a model in front of a clinician, a credit reviewer, a claims adjuster or a regulator, and who has to answer questions about individual predictions in the room. It is also the researcher who wants a reference implementation of a published interpretable method without reimplementing it from the paper. If nobody will ever read the model, the package's main value does not apply to you.

How imodels works: one API over many interpretable model families

The package is a collection rather than a single algorithm. The README lists rule sets (RuleFit, Skope, Boosted, Slipper, Bayesian), rule lists (Bayesian, Greedy, OneR), fast-and-frugal trees, tree variants (CART, C4.5, TAO), a sparse integer linear model called SLIM, and Tree GAM, among others. Each is exposed as a classifier or regressor class, and the README states that you simply import a classifier or regressor and use the fit and predict methods, same as standard scikit-learn models.

The mechanism differs by family, and the README is explicit about some of them. RuleFit fits a sparse linear model on rules extracted from decision trees. Skope extracts rules from gradient-boosted trees, deduplicates them, then linearly combines them based on their out-of-bag precision. Boosted rule set fits rules sequentially with Adaboost. Slipper learns rules sequentially with the SLIPPER algorithm. Greedy rule list uses CART to fit a list, which is a single path rather than a tree, and OneR restricts the list to a single feature. The Bayesian variants sample a distribution over rules or rule lists, and the README marks both as slow.

The practical consequence of that design is that the output object is not a numeric score you have to explain afterwards. For a tree-family model the repr is the model. The README's example prints a decision tree with hierarchical shrinkage, and the printed form is a set of nested conditions ending in leaf values such as [0.10] and [0.68]. That text is the artifact you show. It also sets the ceiling: a model you can print in a dozen lines cannot represent an interaction that needs thirty.

Installing imodels and fitting a first model

Installation is a single pip command. The README gives it without extras, and pyproject.toml declares the base dependencies as matplotlib, mlxtend, numpy, pandas, requests, scipy, scikit-learn and tqdm, so a plain install pulls in the scientific Python stack.

bash
pip install imodels

The README points to a troubleshooting page in the repository's docs directory for help if that fails. Note that some algorithms need packages outside the base set. pyproject.toml lists an optional dependency group containing cvxpy for slim, statsmodels for bartpy, torch for imodels.util.neural_nets, interpret for imodels.algebraic.gam_multitask, and glmnet for imodels.tree.rf_plus, and it marks the glmnet entry as deprecated and due for removal. If you pick one of those models and skip the extra, expect an import error rather than a silent fallback.

For a first run, the README's own example is the shortest path. It loads a sample clinical dataset through a helper, splits it, fits a tree with hierarchical shrinkage capped at four leaf nodes, and predicts.

python
from imodels import get_clean_dataset, HSTreeClassifierCV
from sklearn.model_selection import train_test_split

X, y, feature_names = get_clean_dataset('csi_pecarn_pred')
X_train, X_test, y_train, y_test = train_test_split(
    X, y, random_state=42)

model = HSTreeClassifierCV(max_leaf_nodes=4)
model.fit(X_train, y_train, feature_names=feature_names)
preds = model.predict(X_test)

Two details in that snippet matter. The fit call takes a feature_names keyword that the standard scikit-learn fit signature does not have; it is what makes the printed tree readable instead of showing column indices. And predict returns discrete predictions with shape (n_test, 1), which is a column vector, not the flat array scikit-learn estimators usually return. If you pass those predictions into code that assumes a one-dimensional array, reshape first. After fitting, printing the model gives the nested-condition text shown in the README, with one leaf value per terminal branch.

Where imodels stops being the right tool

The package is honest about one limitation in its own model table: the Bayesian rule set and Bayesian rule list are both annotated as slow, because they sample. On a dataset with tens of thousands of rows and many candidate rules, that is not a tuning problem, it is a different order of runtime, and no amount of parallelism in the calling code fixes a sampler that has to draw.

The second boundary is the input type. Everything in the supported list is a tabular estimator operating on a feature matrix, and the base dependencies are the classic numeric stack with no text or vision components. The README's own signposting confirms this: it points readers to a separate package, imodelsX, for interpretability in text, and to agentic-imodels for interpretable tools for tabular data with agents. If your input is raw text, images or sequences, you are in the wrong repository.

The third is the accuracy ceiling. The README claims these models can often replace black-box models without sacrificing predictive accuracy. That word, often, is doing real work. A four-leaf tree is a coarse model by construction, and on a dataset with genuine high-order interactions a boosted ensemble will beat it. The right framing is not that imodels matches your existing model, but that you should fit both and decide whether the accuracy gap is one your application can absorb. If it cannot, the interpretable model is the wrong answer and no amount of presentation fixes it.

A fourth, smaller friction point: the optional dependencies are not installed by default, and the glmnet entry that rf_plus needs is flagged in pyproject.toml as deprecated. Treat that model family as the least supported corner of the package.

imodels versus post-hoc explanation with SHAP or LIME

The obvious alternative is not another interpretable-model library but the post-hoc route: keep your gradient-boosted model and explain it with SHAP or LIME. The difference is structural, not a matter of quality. A SHAP value is computed from a surrogate or from sampling around a prediction, so the explanation is an estimate attached to a model that remains opaque. An imodels rule list is the model, so there is no gap between the explanation and the thing being explained, and no question about whether the explainer is faithful.

The cost of that guarantee is flexibility. With SHAP you keep whatever accuracy your original model had and pay in explanation fidelity and runtime. With imodels you accept a restricted hypothesis class and pay in accuracy. Neither is universally correct. If your model is already deployed and the requirement is to annotate its outputs, post-hoc explanation fits the constraint. If the requirement is that a human can audit the decision logic itself, a rule list or a shallow tree is the only thing that satisfies it, and imodels gives you a menu of them behind one interface. The package also sits alongside other interpretable-model implementations; the README links reference code for several of the algorithms, including separate repositories for RuleFit, Skope and fast-and-frugal trees, so if you only need one of those methods you can go directly to its source project.

Maintenance, versioning and licence

The repository is not archived, and its last push was on 2026-08-03, the same date as the v3.0.0 release, which the release notes describe as full-fledged support and compatibility across models. The two releases before it were v2.0.0 in October 2024, described as full compatibility with NumPy 2 and the latest releases of common packages, and v1.4.5 in May 2024, which improved compatibility with recent pandas and scikit-learn versions. The pattern is instructive: most of the release history is compatibility work, not new algorithms. That is normal for a package that wraps a fixed set of published methods, but it means you should not expect the supported-model table to grow quickly.

Upgrade cost is dominated by the dependency floor. pyproject.toml requires Python 3.9 or later and pins scikit-learn only in a comment, noting that version 1.6.0 has an issue with ensemble models. Because the package tracks the wider scientific Python stack, a major NumPy or pandas bump in your environment is the event most likely to force an imodels upgrade, exactly as v2.0.0 and v1.4.5 were. The repository ships a uv.lock and a .python-version file, so the maintainers' tested environment is reproducible if you want to match it.

The licence is MIT, declared in pyproject.toml as license = { text = "MIT" } and marked with the OSI Approved :: MIT License classifier. For most users that means permissive reuse with attribution and no copyleft obligation on your own code. This is a description of what the repository declares, not legal advice; if you are redistributing the package inside a commercial product, read the licence text in the repository yourself.

Editorial conclusion

Adopt imodels when the deliverable is a model a domain expert has to read, and your data is a tabular array you can hand to scikit-learn. Do not adopt it as a general replacement for gradient boosting on wide sparse text or image inputs, and do not expect the Bayesian rule list or rule set to finish quickly on a large dataset. Before committing, verify three things on your own data: that the model you picked is in the supported table, that any optional dependency it needs is installed, and that its printed structure is short enough to show to the person who has to sign off on it.

Frequently asked questions

What is imodels?

imodels is a Python package for concise, transparent and accurate predictive modeling, according to its README. It provides a scikit-learn compatible interface for fitting and using a range of interpretable models, including rule lists, rule sets, trees and sparse linear models.

How do I install imodels?

The README gives a single command, pip install imodels, and points to a troubleshooting page in the repository's docs directory if it fails. Some algorithms need extra packages listed in the optional dependency group in pyproject.toml, such as cvxpy for slim.

Which models does imodels support?

The README's supported-model table lists rule sets (RuleFit, Skope, Boosted, Slipper, Bayesian), rule lists (Bayesian, Greedy, OneR), fast-and-frugal trees, tree variants (CART, C4.5, TAO), the sparse integer linear model SLIM, and Tree GAM, among others. Each is exposed as a classifier or regressor you use with fit and predict.

Does imodels replace scikit-learn?

No. The README states that the models are all compatible with scikit-learn and that you import a classifier or regressor and use the fit and predict methods as with standard scikit-learn models. It is a set of additional estimators, not a replacement for the library.

Why is the Bayesian rule list so slow in imodels?

The README's model table annotates both the Bayesian rule set and the Bayesian rule list as slow, because they fit a distribution over rules or rule lists by Bayesian sampling rather than by a greedy pass over the data.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/csinva-imodels.svg)](https://hysenlabs.com/projects/csinva-imodels)