# SHAP: Shapley Additive exPlanations for Model Outputs

> SHAP assigns each feature a share of a single prediction using Shapley values from cooperative game theory. It is a Python package for engineers who need per-prediction attribution rather than a global importance score.

**shap/shap** — A game theoretic approach to explain the output of any machine learning model.

- Repository: https://github.com/shap/shap
- Website: https://shap.readthedocs.io
- Stars: 25,760 · Forks: 3,752
- Language: Jupyter Notebook
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/shap-shap

## What SHAP computes and who needs it

A model returns a number. It does not tell you which input moved that number. SHAP closes that gap by borrowing a concept from cooperative game theory: the Shapley value, which divides a payout among players according to their marginal contribution across all possible coalitions. Here the players are features and the payout is the difference between a prediction and a baseline.

The README states the project is "a game theoretic approach to explain the output of any machine learning model" and that it connects "optimal credit allocation with local explanations using the classic Shapley values from game theory and their related extensions." The unit of output is therefore local: one row of data, one explanation.

The audience is narrower than the tagline suggests. If you are debugging why a single loan application was rejected, or why one sentence was classified as negative, SHAP is built for that question. If you want a ranked list of which columns matter in general, the same values can be averaged, but that is a derived summary, not the primary output.

## How SHAP turns a prediction into feature attributions

The mechanism is coalition sampling plus credit allocation. For a given instance, SHAP considers subsets of features present and absent, evaluates the model on each configuration, and measures how much the prediction changes when a feature joins. The Shapley value is the weighted average of those marginal contributions over all orderings.

Exact enumeration is exponential in the number of features, so the implementation differs by model type. For trees, the repository ships a fast exact algorithm with C++ implementations, described in the README as supporting XGBoost, LightGBM, CatBoost, scikit-learn and pyspark tree models. For other models, the same explainer API falls back to sampling-based estimation, which is approximate and costs more function evaluations.

There is also a natural language path. The README describes adding "coalitional rules to traditional Shapley values" so that games can explain large transformer models "using very few function evaluations." The repository layout reflects this spread: a shap/ package, a javascript/ directory, a notebooks/ directory, and tests/ alongside a CMakeLists.txt, which is consistent with a mixed Python and C++ build rather than a pure Python package.

## Installing SHAP and explaining your first prediction

Installation is a single command from PyPI, or the conda-forge channel if you prefer conda. The README gives both forms.

```bash
pip install shap
# or
conda install -c conda-forge shap
```

Note the version floor in pyproject.toml: requires-python is ">=3.12". If your environment is older, pip will refuse the install rather than silently downgrade.

The README's tree ensemble example is the shortest path to a real result. It trains an XGBoost regressor on the bundled california dataset, builds an explainer, and plots the first prediction.

```python
import xgboost
import shap

X, y = shap.datasets.california()
model = xgboost.XGBRegressor().fit(X, y)

explainer = shap.Explainer(model)
shap_values = explainer(X)

shap.plots.waterfall(shap_values[0])
```

What you should see is a waterfall chart where each feature pushes the output away from the base value, which the README defines as "the average model output over the training dataset we passed." Red bars push the prediction higher, blue bars push it lower. The same explainer object works for LightGBM, CatBoost, scikit-learn and Spark models according to the README, so the syntax does not change when you swap the model.

For aggregate views, the README shows a beeswarm plot that sorts features by the sum of SHAP value magnitudes across all samples, and a bar plot built from the mean absolute SHAP value per feature, which produces stacked bars for multi-class outputs.

```python
shap.plots.beeswarm(shap_values)
shap.plots.bar(shap_values)
```

GPU acceleration exists but is not a pip flag. The README says to install from source with the CUDA toolkit present and the environment variable set:

```bash
SHAP_ENABLE_CUDA=1 pip install .
```

## Where SHAP gets expensive or misleading

The cost scales with the number of features and the number of rows you explain. Exact tree SHAP is fast, but the sampling-based path for arbitrary models evaluates the model many times per instance. Explaining a full test set of a wide tabular model with a non-tree model can take longer than training it did.

Correlated features are the sharper problem. Shapley values assume players can be added or removed independently. When two features carry nearly the same information, the split between them is arbitrary and can flip between runs with different sampling seeds. The documentation does not offer a correction for this; the values remain internally consistent but are not a causal decomposition.

SHAP is the wrong tool when you need causality. A high SHAP value for a feature means the model leaned on it, not that changing it would change the real-world outcome. It is also the wrong tool when the only question is global feature ranking, since a permutation importance or a simple model-based ranking avoids the per-instance machinery entirely.

The README does not document a rollback or version-pinning procedure for explainer output, and it does not state stability guarantees for sampling-based estimates across releases. If you produce explanations that feed a report, record the SHAP version alongside them.

## SHAP compared with LIME and permutation importance

LIME is the closest alternative in intent. It fits a local surrogate model, usually a sparse linear model, on perturbed samples around the instance and reads the surrogate's coefficients as importance. SHAP instead allocates credit through Shapley values, which come with an axiomatic basis: the README calls the family "the only possible consistent and locally accurate additive feature attribution method based on expectations" in a commented-out passage. The practical difference is that LIME's explanation depends on the surrogate's fit and on the perturbation kernel, while SHAP's depends on coalition sampling and on the background dataset you pass.

Permutation importance answers a different question. It shuffles a column and measures the drop in a global score. It is cheap and model-agnostic, but it produces one number per feature for the whole dataset and cannot tell you why row 42 was scored the way it was. If your deliverable is a chart of the top ten drivers, permutation importance is less work and less to explain to a reviewer.

The honest split is this: reach for SHAP when the per-instance attribution is the deliverable, and reach for permutation importance when the global ranking is. Using SHAP to produce a global ranking works, but you pay the local cost to get it.

## Maintenance, licensing and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-07. Recent releases are v0.53.0rc0 on 2026-09-07, v0.52.0 on 2026-05-28 and v0.51.0 on 2026-03-04. The cadence is roughly quarterly, with release candidates appearing before stable tags, so pinning to a stable version rather than tracking main is the lower-surprise option.

The licence is MIT, declared both in the repository metadata and in pyproject.toml as "MIT License". That permits commercial use and modification with the copyright notice retained. This is not legal advice; if you redistribute SHAP inside a product, have your own counsel review the notice requirements.

Upgrade cost is dominated by the dependency floor. pyproject.toml requires Python 3.12 or newer and pins numpy>=2, scipy, scikit-learn, pandas, tqdm>=4.27.0, slicer==0.0.8 and cloudpickle. Because SHAP follows SPEC 0 for minimum supported dependency versions, the README warns that the project tests against the versions specified there and "may not fix bugs for older versions." Upgrading SHAP on an old environment is therefore a numpy 2 migration, not a patch bump. The mixed C++ and Python build (CMakeLists.txt, scikit-build-core, nanobind) also means source installs need a working compiler toolchain.

## Conclusion

Adopt SHAP when you need to attribute one prediction to its input features and can accept the cost of the explanation step, especially for tree ensembles where the exact algorithm is available. Do not adopt it if you only need a global ranking of features, or if your model is a black box whose internals you cannot sample cheaply. Before committing, verify your Python is at least 3.12 as pyproject.toml requires, and check which explainer class the documentation recommends for your model family, since the choice changes both runtime and the meaning of the output.

## FAQ

### What does SHAP stand for?

SHAP stands for SHapley Additive exPlanations. The README describes it as a game theoretic approach to explaining the output of any machine learning model, built on Shapley values from cooperative game theory.

### What is SHAP in Shapley additive explanations?

It is a Python library that assigns each input feature a share of a single prediction, using Shapley values and their extensions to allocate credit among features. The README presents it as connecting optimal credit allocation with local explanations.

### How do I install SHAP in Python?

The README gives two options: pip install shap, or conda install -c conda-forge shap. pyproject.toml requires Python 3.12 or newer, so an older interpreter will not install the package.

### How do I create a SHAP explainer for a model?

The README's example builds one with shap.Explainer(model) and then calls it on the feature matrix to get shap_values. The same syntax is documented as working for LightGBM, CatBoost, scikit-learn, transformers and Spark models.

### How does SHAP enable GPU-accelerated Tree SHAP?

The README says to install from source with the CUDA toolkit available and the SHAP_ENABLE_CUDA environment variable set, then run SHAP_ENABLE_CUDA=1 pip install . The CUDA toolkit must already be installed on the system.

## Sources

- [License: MIT](https://github.com/shap/shap/blob/main/LICENSE)
- [Project website](https://shap.readthedocs.io)
- [README](https://github.com/shap/shap/blob/main/README.md)
- [Releases](https://github.com/shap/shap/releases)
- [shap/shap on GitHub](https://github.com/shap/shap)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/shap-shap
