# Quantus: a metrics harness for judging XAI explanations

> Quantus collects explanation evaluation metrics under one API so attribution methods get scored rather than eyeballed. Six metric categories, a citation back to the original paper for every metric, and a release line whose newest tag is v0.6.0 from 2025-07-21.

**understandable-machine-intelligence-lab/Quantus** — [JMLR 2023] Quantus is an eXplainable AI toolkit for responsible evaluation of neural network explanations

- Repository: https://github.com/understandable-machine-intelligence-lab/Quantus
- Website: https://quantus.readthedocs.io/
- Stars: 677 · Forks: 94
- Language: Jupyter Notebook
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/understandable-machine-intelligence-lab-quantus

## The newest tag is v0.6.0 from 2025-07-21 while main keeps moving

The release record is uneven. v0.5.2 and v0.5.3 both landed in early December 2023, four days apart, and then nothing until v0.6.0 on 2025-07-21. The repository itself was last pushed on 2026-09-05, so the branch has moved well past that tag without a new release.

The project is explicit about this. A line near the top of the README says Quantus is under active development and asks you to carefully note the release version to ensure reproducibility of your work. That request makes sense against the numbers: installing from PyPI gets you 0.6.0 from July 2025, while the repository at 676 stars, 94 forks and 65 open issues contains months of unreleased commits. Anyone reproducing a paper has to record which of the two they used.

## Two metrics in the faithfulness list share one name

Metrics are grouped into six categories: faithfulness, robustness, localisation, complexity, randomisation for sensitivity, and axiomatic metrics. Faithfulness measures to what extent an explanation follows the model's predictive behaviour, that more important features play a larger role in model outcomes.

Two entries in that list are both called Monotonicity Metric. The first, after Arya et al. 2019, starts from a reference baseline and incrementally replaces each feature in a sorted attribution vector, measuring the effect on model performance. The second, after Nguyen et al. 2020, measures the Spearman rank correlation between the absolute values of the attribution and the uncertainty in the probability estimation. Different measurements, identical names, so an import path or a result table has to disambiguate them by paper rather than by label.

## The catalogue is counted twice and the two counts disagree

The library overview says the toolbox includes 30 or more different metrics. The highlights list above it says more than 35 metrics in six categories. Neither number is reconciled anywhere, and the highlights phrasing stacks a comparison word on top of a plus sign, which makes the figure harder to cite than it should be.

The same section has other small repairs waiting. The main heading and the CI workflow badges are wrapped in HTML comments, so they do not render, and the note about dropping Python 3.7 support sits inside a comment too. The codecov badge points at `branch=master` while the default branch is `main`. None of this affects the code, and all of it is the kind of drift that tells a reader which parts of a README were last touched.

## Every metric carries the paper it came from

The citation section asks for the JMLR paper, Quantus: An Explainable AI Toolkit for Responsible Evaluation of Neural Network Explanations and Beyond, volume 24, number 34, pages 1 to 11, from 2023, authored by Anna Hedström, Leander Weber, Daniel Krakowczyk, Dilyara Bareeva, Franz Motzkus, Wojciech Samek, Sebastian Lapuschkin and Marina M.-C. Höhne. It then adds the rule that matters most in practice: when you apply an individual metric, cite the original author's work too.

The metric list follows that rule literally, since each entry links its source paper. Faithfulness Correlation points at Bhatt et al. 2020, Faithfulness Estimate at Alvarez-Melis et al. 2018, Pixel Flipping at Bach et al. 2015, and the two newer randomisation metrics, EfficientMPRT and SmoothMPRT, are credited to Hedström et al. 2023 through OpenReview.

## Four gradient methods motivate the whole toolkit

The argument for Quantus starts with a comparison that does not work. The README takes four gradient based attribution methods, Saliency, Integrated Gradients, GradientShap and FusionGrad, and observes that lining their heatmaps up side by side is not enough to decide which one works best. Visual inspection is still the common practice when there is no ground truth for the explanation itself.

What is offered instead has two parts. The first quantifies each method across several evaluation criteria at once. The second is the sharper idea: take one parameter, such as the pixel replacement strategy inside a faithfulness test, vary it, and watch the ranking of the methods move. A metric whose result survives that sweep says more about the model than a metric whose result depends on one hand picked constant.

## The licence classifier says GPLv3 while the metadata says nothing

The licence story needs reading from three places. `pyproject.toml` sets `license = { file = "LICENSE" }`, which defers to a file instead of naming a licence expression. The PyPI classifier list says `License :: OSI Approved :: GNU General Public License v3 (GPLv3)`. And the repository root carries three licence files at once: `LICENSE`, `COPYING` and `COPYING.LESSER`.

On top of that, the API metadata for this project returns no licence value at all, so tools that read that field rather than the repository see a blank. If you are vendoring Quantus into a product, the classifier and the COPYING pair are the only statements to go on, and the LGPL text sitting next to the GPL text suggests the exception was deliberate.

## Dependency floors, a Python 3.8 minimum and flit versioning

`requires-python = ">=3.8"`, and every dependency is a lower bound with no ceiling: numpy at 1.19.5, pandas at 1.5.3, opencv-python at 4.5.5.62, scikit-image at 0.19.3, scikit-learn at 0.24.2, scipy at 1.7.3, tqdm at 4.62.3, matplotlib at 3.3.4, plus cachetools with no version at all. Framework support is optional, split into extras, and the file shows a `torch` group whose comment mentions Mac, Windows or Linux and a commented example of `pip install quantus[tensorflow]`.

Version numbers are dynamic rather than written in the manifest, the pattern used by flit, so the tag and the installed version can disagree if you build from a checkout rather than from PyPI. Maintenance of the project sits with one person: the authors and maintainers lists both name Anna Hedstrom and the same address.

## A metrics library whose main language is a notebook

The repository is recorded as primarily Jupyter Notebook, which is unusual for a package that people import. The `quantus/` directory holds the library, `tests/` and `pytest.ini` and `tox.ini` hold the tests, and `mypy.ini` sits beside them for typing. Then there is `tutorials/`, a Binder link that launches the tutorial folder, and a Colab notebook that walks through an ImageNet example with all metrics applied to one model.

The ecosystem claims line up with that shape: more than 35 metrics over image, time series and tabular data, with NLP marked as next up, PyTorch and TensorFlow supported, and built-in hooks for captum, tf-explain and zennit as explanation methods. For training data attribution the README points somewhere else entirely, at a separate toolkit called quanda.

## Conclusion

Quantus fits teams that have to choose between attribution methods, because a side by side plot is not a result and the sensitivity analysis is the strongest part of the package. It fits badly if you want one opinionated score, since two entries share the name Monotonicity Metric and the README counts its own catalogue twice. Pin the version either way: the newest tag is v0.6.0 from 2025-07-21 while the repository was pushed on 2026-09-05, and the project itself asks you to note the release version for reproducibility.

## FAQ

### What is Quantus in explainable AI?

A metrics toolkit to evaluate neural network explanations, from the Understandable Machine Intelligence Lab and accepted to the Journal of Machine Learning Research in 2023. It covers image, time-series and tabular data for PyTorch and TensorFlow models.

### How many explanation metrics does Quantus have?

Two figures appear: the library overview says 30 or more different metrics, the highlights say more than 35. The six categories are faithfulness, robustness, localisation, complexity, randomisation for sensitivity, and axiomatic metrics.

### Which explanation methods does Quantus work with?

Built-in support is listed for captum, tf-explain and zennit, over PyTorch and TensorFlow models. Training data attribution evaluation is pointed at a separate toolkit, quanda.

### What licence is Quantus released under?

The PyPI classifier says GNU General Public License v3 and the repository carries LICENSE, COPYING and COPYING.LESSER, while pyproject.toml points at LICENSE through `license = { file = "LICENSE" }` and the project API metadata returns no licence value.

## Sources

- [Issues](https://github.com/understandable-machine-intelligence-lab/Quantus/issues)
- [Project website](https://quantus.readthedocs.io/)
- [README](https://github.com/understandable-machine-intelligence-lab/Quantus/blob/main/README.md)
- [Releases](https://github.com/understandable-machine-intelligence-lab/Quantus/releases)
- [understandable-machine-intelligence-lab/Quantus on GitHub](https://github.com/understandable-machine-intelligence-lab/Quantus)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/understandable-machine-intelligence-lab-quantus
