Model or dataset
frgfm/torch-cam avatar
frgfm/torch-cam

TorchCAM's manifest reads 0.6.0.dev0 and the newest tag is 0.5.0

Class activation maps for your PyTorch models (CAM, Grad-CAM, Grad-CAM++, Smooth Grad-CAM++, Score-CAM, SS-CAM, IS-CAM, XGrad-CAM, Layer-CAM, Finer-CAM, LeGrad, RefineCAM)

2,305 stars230 forksPythonApache-2.0

At a glance

What is it?
frgfm/torch-cam is an Apache licensed PyTorch library that explains a prediction with twelve class activation methods behind one API and four runtime dependencies. The hook mechanism, the evidence bundle, and the four dependency floors are all documented with the reason attached. The version numbers are not.
Who is it for?
This fits someone who has a classifier that got one prediction wrong and needs to know which part of the image drove it, rather than someone building a new attribution method. The target layer default and the hook requirements are the two things to understand before you trust a map.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The manifest is a development version and the newest tag is a minor line behind

Three numbers describe this project and they are all different. The published releases are the most recent, a release labelled for vision transformers and prediction debugging, then one the previous autumn for cleaner hooks control, then one in late 2023 for evaluation metrics and PyTorch two. The newest release is dated 2026-10-03, which is after the last commit to the tree on 2026-10-02. And the project file declares its version as a development build of the next minor line. So the tag stream, the manifest, and the commit dates are three separate timelines, and the gap between the 2023 release and the 2025 one is two years. The project file also carries a development status classifier of beta, which is the only statement anywhere about maturity, and it is a self-declared one rather than something the release history supports on its own.

The four dependency floors each have a reason written next to them

The library claims to be lean and the claim is precise: four runtime dependencies, the deep learning framework, a numerical library, an imaging library, and a plotting library. The project file pins all four, and one of those pins carries a two-line comment. The framework floor is 2.4.1, and the comment says that release includes the Windows fix for NumPy two, with a link to the framework's own issue tracker, plus a cross-reference to a discussion in this project. That is a floor set by a Windows bug, not by a feature the library needs, and it is the kind of comment that saves the next maintainer an afternoon. The numerical floor is the two series below three, the plotting floor is just below four, and the imaging library carries only a lower bound. The other toolchain choices are visible in the badges instead: a linter, a separate type checker, and two coverage services.

The explanation is a directory with a manifest, not a picture

The short version of the library is a single function call, and what it returns is worth looking at closely:

python
result = explain(model, weights.transforms()(image).unsqueeze(0), class_names=weights.meta["categories"])
result.save("torchcam-explanation", image)  # CAMs, heatmaps, overlays and manifest.json

The comment is the documentation. The call produces the activation maps, the heatmaps, overlays on the input, and a manifest file, written to a directory. That last artefact is why the project describes itself as agent-ready: a manifest is something a program can read, so the claim is not that a human can look at a heatmap but that a coding agent can be handed the directory and verify it. Two supporting features back the same idea. There is a documented workflow for a predicted-versus-expected comparison, entered by passing the class index you expected so the tool can tell you where the prediction and your expectation diverge, and there is a portable skill file that installs into a coding agent by one command, plus a machine-readable index page for agents to start from.

The default target layer is the last non-reduced convolutional one

This is the setting that decides what every map you see actually means, and the document puts it in a note rather than in the tutorial. By default the layer the map is read from is the last non-reduced convolutional layer, and if you want a different one you pass a target layer argument to the constructor. The extraction itself works by wrapping your model: the library registers PyTorch forward and backward hooks and retrieves what it needs without the user doing anything, so once the extractor is set you run your model as usual and the extra information appears. Two consequences follow. You are reading one layer, and if that layer is not the one carrying the evidence your map will look wrong. And the extractor returns one map per target layer, so asking for more layers gives you more maps rather than a combined one. The documentation also warns that hooking requires a model that supports it, which is what the troubleshooting page is for.

Four failure symptoms are named, and transformers take a different route

The troubleshooting link in the quick tour is specific about what goes wrong. It names a hook that cannot be registered, a gradient requirement error, a not-a-number result, and a blank heatmap. Those are the four failure modes of this mechanism and they are all consequences of the same design: the library reads inside your model's forward pass, so anything that blocks hooks, severs the graph, or produces a degenerate tensor shows up as one of those four. The same design is why transformer support is described differently. For convolutional networks the library resolves a target layer automatically; for vision transformers it uses a different method, a gradient-based approach and reshape transforms, because there is no last non-reduced convolutional layer to stop at. The advanced guide also covers your own model that does not come from the vision library, three dimensional and video data, and batched input.

The readme's code blocks are injected from the documentation source

Five pairs of markers sit around the code in the readme, each one wrapping a python block and carrying a name. The syntax is the snippet-inclusion feature of a documentation generator: a start marker, the block, an end marker. The effect is that the readme's quick tour is not written in the readme, it is pulled from the documentation source at build time so the two cannot drift. The same documentation tree publishes a machine-readable index page, which is one of the two entry points the readme offers a coding agent alongside the skill file. It is a small structural decision with a real payoff, and it is the kind of thing that explains why the quick tour's examples look identical to the guide's. The corollary is that editing the readme directly would be overwritten, and the contributing flow has to go through the documentation source.

Local development runs on the oldest supported Python and a different commit runner

The build file reveals the development toolchain and two details in it are worth knowing. The virtual environment target creates an environment on the oldest Python the project claims to support, which is the same version the dependency floors are written against, so a contributor who upgrades their interpreter is not testing the configuration the project pins. And the pre-commit target does not run pre-commit: it runs a different executable with the same name minus the hyphen, against a configuration file named for pre-commit. So the hook definitions are portable between the two tools, which is the point, but the contributor following the documentation literally will not have the command they expect. Everything else goes through one tool: the environment, the editable install, the quality extra, the linter, the formatter, and the hooks all go through the same runner, and both linter invocations pass an explicit config path pointing back at the project file.

Editorial conclusion

This fits someone who has a classifier that got one prediction wrong and needs to know which part of the image drove it, rather than someone building a new attribution method. The target layer default and the hook requirements are the two things to understand before you trust a map. Check four things. What the map is measuring, because a class activation map explains the model rather than the image, and the library ships faithfulness metrics for that reason. Which layer it is reading, since the default is the last non-reduced convolutional one and that is a choice, not a fact about your network. Whether your model is a convolutional one, because transformers are handled by a different mechanism. And what you save, because the explanation is a directory with overlays and a manifest, which is the part you can hand to someone else to check.

Frequently asked questions

How many class activation methods does TorchCAM implement?

Twelve behind one API, from the original class activation map and its gradient variants through the score-based and layer-wise methods to the two gradient-free approaches and the refinement method. The exhaustive list lives in the reference documentation.

What does explain() save when I call it?

A directory containing the activation maps, the heatmaps, overlays on the input image, and a manifest file. The manifest is what makes the bundle verifiable by a program rather than only viewable by a person. Passing the class index you expected turns the call into a comparison against your own expectation.

Which layer does TorchCAM read by default?

The last non-reduced convolutional layer. If you want a different one you pass a target layer argument to the extractor constructor. The extractor returns one activation map per target layer, and for vision transformers the library resolves the layer differently since there is no convolutional layer to stop at.

What are TorchCAM's runtime dependencies?

Four: the deep learning framework, a numerical library, an imaging library, and a plotting library. The framework floor is set to a release that includes a Windows fix for NumPy two, with the reason and a tracker link written into the project file.

Does TorchCAM work with vision transformers and 3D data?

Yes, through a separate mechanism from the convolutional path, and the advanced usage guide covers both, along with your own model that does not come from the vision library and batched input. The troubleshooting guide names the four failure modes to expect if hooking is blocked or the graph is severed.

Official sources

  1. frgfm/torch-cam on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/frgfm-torch-cam.svg)](https://hysenlabs.com/projects/frgfm-torch-cam)