# ELI5: Feature Importance and Prediction Explanation for scikit-learn, XGBoost, and More

> ELI5 is a Python library for debugging machine learning models by exposing feature weights, feature importances, and per-prediction explanations. It supports scikit-learn, XGBoost, LightGBM, CatBoost, Keras, and several other frameworks, and can explain black-box models using LIME and permutation importance.

**TeamHG-Memex/eli5** — A library for debugging/inspecting machine learning classifiers and explaining their predictions

- Repository: https://github.com/TeamHG-Memex/eli5
- Website: http://eli5.readthedocs.io
- Stars: 2,849 · Forks: 326
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/teamhg-memex-eli5

## What ELI5 Explains and Who It Is For

ELI5 targets machine learning practitioners who need to understand what a trained classifier is doing at the feature level. The name comes from the internet phrase for explaining something simply. The library provides two complementary capabilities: it shows which features a model considers most important globally, and it explains individual predictions by attributing them to specific input features.

The intended users are data scientists and ML engineers who work with tabular data, text, or images, and who need to communicate model behavior to stakeholders, debug classification errors, or identify whether a model is learning relevant patterns.

## Supported Frameworks and What You Get from Each

ELI5 integrates with several frameworks and provides different explanations depending on the model type.

For scikit-learn, the library explains weights and predictions of linear classifiers and regressors, prints decision trees as text or SVG, and shows feature importances for tree-based ensembles. It understands scikit-learn's text processing utilities and can highlight which words in a text input contributed to a classification. Pipeline and FeatureUnion are supported. A specific feature handles HashingVectorizer: ELI5 can undo the hashing to recover the original feature names for debugging.

For Keras, ELI5 generates Grad-CAM visualizations that highlight which regions of an input image drove an image classifier's prediction.

For XGBoost, ELI5 shows feature importances and explains individual predictions for XGBClassifier, XGBRegressor, and xgboost.Booster.

For LightGBM and CatBoost, the library exposes feature importances for their respective classifier and regressor types. CatBoost support covers CatBoostClassifier, CatBoostRegressor, and catboost.CatBoost.

The lightning library (a scikit-learn-compatible set of large-scale linear models) is also supported for weight and prediction explanation.

## Installing ELI5 and Running a First Explanation

The library installs via pip. The requirements.txt lists the core dependencies:

```bash
pip install eli5
```

The core dependencies are numpy, scipy, scikit-learn, attrs, jinja2, graphviz, and tabulate. For specific integrations (XGBoost, LightGBM, Keras, CatBoost), those libraries must also be installed separately.

The primary entry points are `explain_weights`, which returns a model-level explanation of feature importance, and `explain_prediction`, which explains a single prediction. These return an explanation object. To display it in an IPython notebook as an HTML table, pass the object to `eli5.show_weights()` or `eli5.show_prediction()`. The same object can be converted to a pandas DataFrame with `eli5.format_as_dataframe(explanation)` for further processing.

## Black-Box Inspection: LIME and Permutation Importance

For models that do not expose internal structure, ELI5 provides two algorithms that treat the model as a black box.

TextExplainer implements the LIME algorithm (from the 2016 paper by Ribeiro et al.) for text classifiers. LIME fits a locally linear approximation around a specific input by sampling nearby inputs and observing how the model's output changes. TextExplainer wraps this process for the text domain, highlighting which words or phrases most influenced the prediction. The README notes that utilities for non-text data and arbitrary classifiers also exist, though that feature is described as experimental.

Permutation importance is a model-agnostic method that measures how much a model's performance degrades when each feature's values are randomly shuffled. If shuffling feature A causes a large drop in accuracy, A is considered important. ELI5's permutation importance implementation works with any estimator that follows the scikit-learn API.

## Output Formats: Console, Notebook, DataFrame, and JSON

ELI5 separates the explanation computation from its presentation. The same explanation object can be rendered in four formats.

For console output, the text format prints a human-readable table of feature weights or importances. For IPython notebooks and web dashboards, the HTML format renders styled tables that can be embedded directly. The DataFrame format converts the explanation to a pandas DataFrame, allowing further computation, filtering, or charting. The JSON format is intended for cases where the rendering happens on the client side, in a custom visualization or a front-end dashboard.

This separation means the explanation logic runs once and the output is adapted to the environment without rerunning the model.

## Limitations: Python Version Targeting and Framework Age

The setup.py targets Python 2.7 and Python 3.4, and lists `six` as a dependency. The `six` library is a Python 2/3 compatibility shim that became unnecessary once Python 2 reached end of life in 2020. Its presence indicates the core code was written for a Python 2-era environment and has not been fully modernized.

The Keras integration was written for an older Keras API. Keras changed significantly with the TensorFlow 2.x series and again with the Keras 3 release. The README documents Keras support but does not specify which Keras version is compatible, and engineers using current Keras should test compatibility before relying on the Grad-CAM feature.

The library has no GitHub releases and no tagged versions on PyPI visible in the repository. The CHANGES.rst file at the top level is the primary record of changes.

The pyproject.toml-based packaging ecosystem that replaced setup.py is not used here. Adding ELI5 to a project managed with modern packaging tools requires testing for dependency conflicts, particularly the broad version ranges in requirements.txt (e.g., `numpy >= 1.9.0`).

## How ELI5 Compares to SHAP

SHAP (SHapley Additive exPlanations) is a Python library for model explanation based on Shapley values from game theory. It provides a unified measure of feature importance that satisfies several theoretical properties and works with a wide range of model types including tree ensembles, neural networks, and linear models.

ELI5 and SHAP overlap for tree-based models and linear classifiers, where both can produce per-prediction feature attributions. SHAP's TreeExplainer computes exact Shapley values for tree ensembles efficiently. ELI5's black-box support uses LIME rather than Shapley values, which is a local linear approximation rather than a theoretically consistent global measure.

ELI5's advantage is its native understanding of scikit-learn pipelines and text processing utilities, particularly the ability to undo HashingVectorizer to recover readable feature names. SHAP does not provide this pipeline-aware behavior by default. For teams already using scikit-learn pipelines with text features, ELI5's integration is more direct.

## Maintenance and License

The last push to the repository was on 2026-04-08. The repository is not archived. The project classification in setup.py is `Development Status :: 4 - Beta`. The authors listed in setup.py are Mikhail Korobov and Konstantin Lopuhin from TeamHG-Memex.

The documentation is hosted at eli5.readthedocs.io. The license is MIT, located in LICENSE.txt at the repository root.

## Conclusion

ELI5 is a practical choice for teams using scikit-learn, XGBoost, or LightGBM who need to inspect why a classifier produces a particular output. The library's explanation and output formatting layers are separated cleanly, which makes it easy to route results into notebooks, console tools, or dashboards. The dependency list includes `six`, a Python 2/3 compatibility shim, and the setup.py targets Python 2.7 and 3.4, indicating the core code was written for an older Python era. Engineers on Python 3.9+ should verify compatibility before integrating. The last push was on 2026-04-08. The license is MIT.

## FAQ

### Can ELI5 explain predictions from XGBoost and LightGBM models?

Yes. ELI5 supports XGBClassifier, XGBRegressor, and xgboost.Booster for XGBoost, and LGBMClassifier and LGBMRegressor for LightGBM. It can show feature importances and explain individual predictions for both.

### How does ELI5 explain black-box models that do not expose feature weights?

ELI5 provides TextExplainer, which uses the LIME algorithm to explain text classifier predictions by sampling nearby inputs and fitting a local linear approximation. For other model types, permutation importance measures how much a model's performance drops when each feature's values are randomly shuffled.

### What output formats does ELI5 support for displaying explanations?

ELI5 can render explanations as plain text for the console, as HTML for IPython notebooks and web dashboards, as a pandas DataFrame for further processing, and as JSON for custom client-side rendering.

## Sources

- [Issues](https://github.com/TeamHG-Memex/eli5/issues)
- [License: MIT](https://github.com/TeamHG-Memex/eli5/blob/master/LICENSE)
- [Project website](http://eli5.readthedocs.io)
- [README](https://github.com/TeamHG-Memex/eli5/blob/master/README.md)
- [TeamHG-Memex/eli5 on GitHub](https://github.com/TeamHG-Memex/eli5)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/teamhg-memex-eli5
