Library / SDK
meta-pytorch/captum avatar
meta-pytorch/captum

Captum: Integrated Gradients, Saliency and TCAV for PyTorch Models

Model interpretability and understanding for PyTorch

5,706 stars567 forksPythonBSD-3-Clause

At a glance

What is it?
Captum is a BSD-3-Clause interpretability library from the PyTorch project. It bundles attribution algorithms such as Integrated Gradients, DeepLift and GradientShap behind a consistent API, and the trade-offs show up in baselines, hooks and installation extras.
Who is it for?
Adopt Captum if your models are already nn.Module instances and you want attribution results as tensors you can post-process, or if you are implementing an interpretability method and want a reference implementation to benchmark against.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap Captum fills between a trained PyTorch model and an explanation

A trained PyTorch model returns logits. It does not return a reason. Captum exists to close that distance without asking you to rewrite the model. The README describes it as a model interpretability and understanding library for PyTorch, and the target audience it names is narrow and specific: model developers who need to know which concepts, features or training examples matter, and interpretability researchers who want to implement and compare algorithms. Application engineers using a trained model in production are listed as a secondary audience, mainly for troubleshooting and for producing explanations shown to end users, such as why a particular recommendation appeared.

The library covers three kinds of question. Feature attribution asks which parts of an input drove a prediction. Neuron and layer attribution asks which internal units mattered. Data attribution, through TracIn influence functions, asks which training examples mattered. Captum also ships adversarial attack and minimal perturbation utilities, which the README positions as a way to generate counterfactual explanations as well as adversarial examples. That last point is worth pausing on: the same perturbation machinery serves two very different goals, and the library does not police which one you are pursuing.

If your model is not a PyTorch nn.Module, none of this applies. Captum is not a framework-agnostic explanation service. It is a set of algorithms bound to PyTorch autograd and to the module abstraction.

How attribution actually flows through a Captum call

The mechanism is consistent across the attribution classes. You construct an algorithm object with your model, call attribute on an input tensor, and get back a tensor with the same shape as the input (or, for neuron and layer methods, a tensor shaped according to the target unit). Underneath, the library registers forward and backward hooks on the modules you name, runs the model, and accumulates gradients or modified gradients into an attribution value.

The baseline is the part newcomers underestimate. Algorithms such as IntegratedGradients, DeepLift and GradientShap attribute the change between the input and a baseline, not the prediction in isolation. The README states this directly: baselines belong to the input space and often carry no predictive signal, and a zero tensor can serve as a baseline for many tasks. That is a real modelling decision, not a formality. A zero baseline for an image means a black image; for text embeddings it means something with no natural meaning at all. The library gives you the hook, but it cannot tell you whether your baseline is defensible.

The README's own walkthrough builds a small ToyModel with two linear layers and a ReLU, sets it to eval mode, fixes the random seeds with torch.manual_seed(123) and np.random.seed(123), and then defines input and baseline tensors. That sequence is the whole shape of a Captum workflow: deterministic setup, an explicit baseline, then an attribution call. NoiseTunnel wraps other algorithms to average over noisy samples, which is how SmoothGrad and VarGrad behaviour is expressed rather than as separate classes.

Installing Captum and running a first attribution

Captum installs from PyPI. The README lists Python >= 3.10 and PyTorch >= 2.3 as requirements, and pyproject.toml declares the same floors, so check both before you start. The single command is:

bash
pip install captum

If you want the unreleased code from master, the README gives a manual install. It warns that bleeding-edge features may come with the occasional bug. The extras are worth knowing about: pip install -e .[dev] pulls in testing, linting and docs tooling, and pip install -e .[tutorials] pulls in the packages needed for the tutorial notebooks, which pyproject.toml lists as torchtext and torchvision.

bash
git clone https://github.com/pytorch/captum.git
cd captum
pip install -e .

Once installed, the smallest useful program is an attribution against a baseline. The README's example imports IntegratedGradients and GradientShap from captum.attr alongside LayerConductance and NeuronConductance. The core call shape is this:

python
import torch
from captum.attr import IntegratedGradients

model.eval()
input = torch.randn(1, 3)
baseline = torch.zeros(1, 3)

ig = IntegratedGradients(model)
attributions = ig.attribute(input, baselines=baseline, target=0)

What you get back is a tensor of the same shape as input, holding per-feature attribution values for the chosen target. Read it as a signed map, not a probability. If you want the completeness property that Integrated Gradients is known for, you would sum the attributions and compare against the model output difference between input and baseline, but the README does not spell out that check, so treat it as something to verify yourself rather than a documented guarantee.

Where Captum is the wrong tool

Captum assumes a PyTorch model you can call and hook. That assumption fails in several common situations, and the library does not offer a fallback.

The first is tracing. Layer and neuron attribution depend on hooks and on module structure surviving compilation or export. If your deployment path traces, scripts or compiles the model in a way that changes module boundaries, the layer-level APIs become fragile. The README does not document a supported path for explaining a TorchScript or compiled artifact directly. Test this on your own model before you build a pipeline on top of it.

The second is baseline choice. Because IntegratedGradients, DeepLift and GradientShap all attribute a change from a baseline, an arbitrary baseline produces an arbitrary explanation. For tabular data with a meaningful zero, this is manageable. For images, text and audio, it is a research question the library hands back to you. There is no built-in baseline search in the README's description.

The third is scope. Captum explains a model's behaviour on given inputs. It does not tell you whether the model is fair, whether a feature is causally related to the outcome, or whether a prediction should be trusted. Those are separate analyses. Treating an attribution map as a causal claim is the most common misuse of any library in this category, and Captum's API does nothing to prevent it.

Captum against SHAP and GradCAM

The two comparisons people search for most are Captum vs SHAP and Captum vs GradCAM, and the differences are structural rather than cosmetic.

SHAP is a model-agnostic framework built on Shapley values. It explains any callable by repeatedly evaluating it on perturbed inputs, which means it works outside PyTorch and gives a consistent additive attribution with a well-defined baseline expectation. The cost is evaluation count: model-agnostic Shapley estimation can require many forward passes per explanation. Captum's GradientShap is a gradient-based approximation of Shapley values, so it keeps the Shapley framing but trades exactness for speed by using gradients instead of pure sampling. If your model is not PyTorch, or if you need a single explanation interface across several model types, SHAP is the more natural fit. If your model is PyTorch and you can afford gradients, Captum stays inside the framework you already use.

GradCAM is a specific method, not a library. It produces coarse spatial heatmaps from convolutional feature maps, which is why it is popular for image classification. Captum's LayerConductance and related layer attribution methods answer a related but different question: which layers and neurons contributed, expressed as attribution over units rather than a spatial map. The README's algorithm list does not include a GradCAM implementation, so if a spatial heatmap is the deliverable, GradCAM-style tooling in torchvision or a dedicated package is the more direct route. Captum's value in that comparison is breadth: one API across feature, neuron, layer and training-example attribution, at the cost of not being the sharpest tool for any single one.

Release cadence, licence and the cost of keeping up

The last push to the default branch was on 2026-09-13, and the repository is not archived. Release history is uneven rather than continuous: v0.7.0 in December 2023, v0.8.0 in March 2025, and v0.9.0 in April 2026. That is roughly one minor release a year, which means the pinned version you install today may sit unchanged for a long stretch. Plan upgrades as a deliberate event, not a background drip.

The dependency surface is small. pyproject.toml lists matplotlib, numpy, packaging, torch>=2.3 and tqdm as core dependencies, with optional extras for openai, scikit-learn and annoy that the file says should be lazily imported. That is a reasonable footprint, but the torch>=2.3 floor is the one that will force your hand: upgrading Captum may require upgrading PyTorch, and that is a much larger change than the Captum bump itself.

Licensing is BSD-3-Clause, declared in pyproject.toml and referenced from the README badge. That is a permissive licence, which generally means you can use it in commercial and closed-source products provided you retain the copyright notice and licence text. This is a description of the licence identifier, not legal advice; your counsel should review the actual LICENSE file and any bundled third-party notices before you ship.

Editorial conclusion

Adopt Captum if your models are already nn.Module instances and you want attribution results as tensors you can post-process, or if you are implementing an interpretability method and want a reference implementation to benchmark against. Do not adopt it as a general-purpose explanation layer for scikit-learn pipelines, non-PyTorch runtimes or dataset-level causal claims; the library is scoped to PyTorch models and its data attribution support is narrower than its feature attribution support. Before committing, verify three things: that your model runs under torch.jit tracing or scripting if you rely on the layer and neuron APIs, that you have a defensible baseline tensor for every input modality you plan to explain, and that the version you install satisfies the Python 3.10 and PyTorch 2.3 floors declared in pyproject.toml.

Frequently asked questions

How do I install Captum?

Install the released version with pip install captum. The README requires Python >= 3.10 and PyTorch >= 2.3, and pyproject.toml declares the same floors.

What is Captum used for?

It is a model interpretability and understanding library for PyTorch. It implements algorithms such as Integrated Gradients, TCAV and TracIn that show which features, neurons, layers, concepts or training examples contribute to a model's predictions.

How does Captum compare with SHAP?

SHAP is model-agnostic and estimates Shapley values by evaluating the model on perturbed inputs. Captum is bound to PyTorch and includes GradientShap, a gradient-based approximation of Shapley values, which is faster but relies on gradients rather than pure sampling.

How does Captum compare with GradCAM?

GradCAM is a single method that produces coarse spatial heatmaps from convolutional feature maps. Captum offers layer and neuron attribution such as LayerConductance and NeuronConductance, but the README's algorithm list does not include a GradCAM implementation.

What are alternatives to Captum?

SHAP covers model-agnostic Shapley-value explanations for any callable, and GradCAM-style tooling targets spatial heatmaps for convolutional image models. Captum's distinguishing trait is one API across feature, neuron, layer and training-example attribution inside PyTorch.

Official sources

  1. License: BSD-3-Clause
  2. meta-pytorch/captum on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/meta-pytorch-captum.svg)](https://hysenlabs.com/projects/meta-pytorch-captum)