pytorch-grad-cam: Grad-CAM and Pixel Attribution Methods for PyTorch
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
At a glance
- What is it?
- pytorch-grad-cam is a Python package that implements more than 18 class activation map methods for explaining what drives a neural network's predictions in computer vision. It works with CNNs, Vision Transformers, and tasks beyond classification, including object detection, semantic segmentation, and embedding similarity.
- Who is it for?
- pytorch-grad-cam is the right package for computer vision teams who need to explain model predictions through spatial heatmaps, debug unexpected activations, or benchmark explainability methods against each other. It works across CNNs, Vision Transformers, and multi-task models.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 48 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Why Computer Vision Models Need Spatial Explanations
Neural network classifiers in computer vision produce a prediction but not an explanation. When a model classifies an image as a dog, a practitioner debugging an error or validating the model for deployment wants to know which regions of the image drove that classification. Did the model attend to the dog's face, or to a human in the background? Gradient-based class activation maps (CAMs) answer that question by producing a heatmap over the image that highlights which spatial regions contributed most to the output.
The README positions pytorch-grad-cam as a package for "diagnosing model predictions, either in production or while developing models" and as a benchmark of algorithms for explainability research. Both use cases reflect the same need: understanding whether the model is attending to the right features for the right reasons, rather than relying on spurious correlations in the training data.
The package requires PyTorch 1.7.1 or later, torchvision 0.13 or later, and Python 3.8 or later, as stated in requirements.txt and setup.py. It adds numpy, Pillow, OpenCV (headless), matplotlib, scikit-learn, and scipy as dependencies.
The Method Table: Gradient Strategies from GradCAM to ShapleyCAM
The README lists 18 methods, each with a description of what distinguishes it from GradCAM:
GradCAM weights 2D activations by the average gradient. HiResCAM multiplies activations element-wise with the gradients and offers provable faithfulness guarantees for certain models. GradCAM++ uses second-order gradients. XGradCAM scales gradients by the normalized activations. AblationCAM zeros out activations and measures how the output drops, with a fast batched implementation in this package. ScoreCAM perturbs the image by scaled activations and measures the output drop. EigenCAM takes the first principal component of the 2D activations (no class discrimination). EigenGradCAM adds class discrimination by computing the principal component of Activations multiplied by Grad. LayerCAM weights activations by positive gradients, which the README notes works better in lower layers.
The newer methods address specific limitations. FullGrad computes gradients of biases from all over the network and sums them. KPCA-CAM uses Kernel PCA instead of PCA. FEM binarizes activations by a threshold rule. ShapleyCAM weights activations using gradient and Hessian-vector products. FinerCAM targets fine-grained classification by comparing similar classes. SegEigenCAM adds gradient weighting before SVD for semantic segmentation. RefineCAM combines CAMs from multiple layers for higher resolution. SESS computes CAMs across sliding window patches at multiple scales.
This breadth is a distinguishing feature: the package is a benchmark and comparison tool as much as a production explainability library.
Installing the Package and Running Your First CAM
The README gives the install command at the top of the page:
pip install grad-camNote that the PyPI package name is `grad-cam`, not `pytorch-grad-cam`. After installation, the import path is `pytorch_grad_cam`.
The README shows the basic usage pattern, starting with a ResNet-50 classification model from torchvision:
from pytorch_grad_cam import GradCAM, HiResCAM, ScoreCAM, GradCAMPlusPlus, AblationCAM, XGradCAM, EigenCAM, FullGrad
from pytorch_grad_cam.utils.model_targets import ClassifierOutputTarget
from pytorch_grad_cam.utils.image import show_cam_on_image
from torchvision.models import resnet50, ResNet50_Weights
model = resnet50(weights=ResNet50_Weights.DEFAULT)
target_layers = [model.layer4[-1]]The key step is selecting target layers. For ResNet-50, the last layer of the final residual block (`model.layer4[-1]`) is the conventional choice because it has the richest spatial feature activations before the global pooling that discards spatial information. The README notes that all methods in the package support batches of images.
The `cam.py` file at the repository root contains a more complete usage example with command-line argument handling, which is useful as a starting point for building a custom evaluation or visualization script.
Choosing the Target Layer: Where Architecture Knowledge Matters
The README dedicates a section to choosing target layers, because the quality of a CAM output depends entirely on selecting activations that carry spatial information meaningful to the model's decision.
For CNNs, the last convolutional layer before the classification head is typically the correct choice. For ResNet-50, that is `model.layer4[-1]`. Using earlier convolutional layers produces lower-resolution, less semantically meaningful heatmaps.
For Vision Transformers, the mechanism is different. Vision Transformers use multi-head self-attention and do not have spatial convolutional feature maps in the same sense. The package supports ViT architectures (the README lists Deit Tiny and Swin Transformer Tiny in its example gallery), but the target layer selection requires understanding the transformer block structure and where spatial token relationships are preserved.
The README does not give a universal rule for transformer target layer selection beyond the general guidance. For models not tested in the README's example gallery, practitioners need to inspect the model's architecture, identify where 2D spatial activations exist, and verify the CAM output produces sensible heatmaps on well-understood examples before relying on the explanation.
Beyond Classification: Detection, Segmentation, and Embeddings
pytorch-grad-cam extends beyond single-class classifiers. The README describes four non-classification use cases supported by the package.
For object detection, the package supports producing heatmaps that show which regions of an image contributed to a detected bounding box, rather than a class score. The README includes screenshots comparing detection model explanations.
For semantic segmentation, the SegEigenCAM method is specifically designed for segmentation models, applying gradient weighting before singular value decomposition and handling the sign ambiguity in SVD output. The README shows a 3D medical segmentation example alongside standard 2D segmentation.
For image similarity and embedding models, the package can explain which regions of a query image contributed most to the similarity score against a reference image. This is useful for debugging embedding-based retrieval systems where the model returns unexpected matches.
For CLIP-style vision-language models, the README demonstrates explaining which regions of an image correspond to a given text prompt, showing separate heatmaps for prompts like 'a dog' and 'a cat' on the same image.
Metrics for Evaluating Whether to Trust the Explanation
A CAM that looks plausible is not necessarily a faithful explanation of the model's actual reasoning. pytorch-grad-cam includes metrics for checking whether the generated heatmaps actually correspond to causally important regions. The README mentions this metrics and evaluation section as a component of the package.
AblationCAM and ScoreCAM measure faithfulness by directly testing whether the highlighted regions are causally important: AblationCAM by zeroing out activations and measuring output drops, ScoreCAM by perturbing the image itself. These approaches make the metric and the CAM method the same operation.
For methods like GradCAM and EigenCAM that use gradient signals without direct perturbation, the package provides evaluation tools to verify the heatmap. The tutorials/ directory in the repository contains Jupyter notebooks that demonstrate the metrics.
A limitation of all gradient-based methods is that the gradient signal depends on the model's internal state at inference time. Models with batch normalization may behave differently in train versus eval mode. The README does not describe this as a handled case, so practitioners should ensure consistent eval mode during CAM generation.
AblationCAM Performance, Alternatives, and License
The README explicitly notes that AblationCAM has a fast batched implementation in this package, distinguishing it from the naive serial implementation that made the method slow in other implementations. AblationCAM is gradient-free and potentially more faithful than gradient-based methods for models where backpropagation is slow or unavailable, but it is computationally more expensive per sample than GradCAM-family methods.
Captum is the main alternative for model interpretability in PyTorch. Captum is maintained by Meta and covers a broad range of attribution methods, including integrated gradients, layer conductance, SHAP-inspired approaches, and others that are not spatial. The difference in scope is significant: pytorch-grad-cam focuses exclusively on pixel attribution methods that produce 2D spatial heatmaps for computer vision models, while Captum covers interpretability across model types and data modalities, including tabular and text data. Teams that need only computer vision spatial explanations will find pytorch-grad-cam's method selection and documentation more directly applicable; teams that need interpretability across multiple model types alongside computer vision should evaluate Captum.
The package is MIT-licensed. The last push to the repository was on August 13, 2026. The version in setup.py is 1.5.5. The repository has no GitHub releases, so the PyPI version is the reference for tracking updates.
Editorial conclusion
pytorch-grad-cam is the right package for computer vision teams who need to explain model predictions through spatial heatmaps, debug unexpected activations, or benchmark explainability methods against each other. It works across CNNs, Vision Transformers, and multi-task models. The package does not cover non-spatial interpretability (feature attribution at the input dimension level for tabular or text data), and the quality of an explanation depends on choosing the right target layer, which requires understanding the model architecture. Before deploying explanations to end users, apply the built-in metrics to verify the CAM is actually highlighting causally relevant regions and not producing a plausible-looking but misleading heatmap.
Frequently asked questions
What is Grad-CAM used for in deep learning?
Grad-CAM produces spatial heatmaps that highlight which regions of an input image most influenced a neural network's prediction. The README describes two use cases: diagnosing model predictions in production and comparing explainability methods for research. It is used to verify that a model attends to the expected image regions and to debug unexpected predictions.
How can you use Grad-CAM with a PyTorch model?
Install with pip install grad-cam, then import the desired method from pytorch_grad_cam, specify the target layers in your model, and run the CAM on your input tensor. The README demonstrates this with a ResNet-50 model using model.layer4[-1] as the target layer. The cam.py file in the repository provides a more complete usage example.
How do you install pytorch-grad-cam?
The README gives the install command as pip install grad-cam. Note that the PyPI package name is grad-cam, not pytorch-grad-cam. After installation, the library is imported as pytorch_grad_cam.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/jacobgil-pytorch-grad-cam)