# tensorflow/tcav: testing which concepts a classifier actually relies on

> TCAV turns a trained network into something you can question in human terms: how much did a concept like color or gender matter for a class prediction? This is a walkthrough of what the library does, how to install it, and where its assumptions break.

**tensorflow/tcav** — Code for the TCAV ML interpretability project

- Repository: https://github.com/tensorflow/tcav
- Stars: 654 · Forks: 148
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/tensorflow-tcav

## The gap TCAV fills: class-level concept importance instead of pixel weights

Most interpretability tooling answers a narrow question. It takes one input, one image or one row, and returns a map of which input features moved the prediction. The README draws the contrast directly: typical methods show importance weights per input feature such as a pixel, while TCAV reports the importance of high level concepts such as color, gender or race for a prediction class. The unit of explanation changes from one example to a whole class.

The intended reader is not necessarily a machine learning engineer. The README states that TCAV is designed to make sense to everyone as long as they can understand the high level concept. That is a deliberate design constraint, and it explains why the output is a number attached to a word like "striped" rather than a heatmap.

The method also claims something narrower than causality. It tells you how sensitive a class prediction is to a direction in activation space, not why the model was trained that way. If you need to audit a single misclassified image, this is the wrong tool.

## How TCAV works: concept examples become a direction, then a directional derivative

The mechanism has three moving parts. First, you supply a set of examples that share a concept and a set that does not. The README says TCAV learns concepts from examples, and gives the example of needing a couple of images of female and something not female to learn a gender concept. Second, the library trains a linear classifier on the activations of those two sets at a chosen layer, the bottleneck. The normal vector of that separating hyperplane is the Concept Activation Vector, or CAV. Third, TCAV computes the directional derivative of the class logit with respect to that CAV direction, for many inputs of the target class.

The result is aggregated into a score per concept per class. A random experiment count is part of the configuration, which is how the method calibrates against concepts learned from random example sets. In the README's usage snippet, num_random_exp is set to 2, which is a small number and worth raising for anything you intend to report.

Two properties follow from this design. The network does not need to be changed or retrained, as the README states. And the concept is defined entirely by the examples you choose, so a badly chosen negative set produces a CAV that measures something other than what you named it.

## Installing tcav and running a first TCAV object

TensorFlow is deliberately excluded from the package dependencies. The README explains that a user may want either tensorflow or tensorflow-gpu, so you install one of those yourself alongside the tcav package. The repository's requirements.txt pins tensorflow==2.5.1, numpy==1.19.2, scikit-learn==0.20.3 and others, which is a useful signal about the version range the project was tested against.

Install the package and a TensorFlow build:

```bash
pip install tcav
pip install tensorflow
```

The README points to Run_TCAV.ipynb for a step by step guide after installation. The constructor takes a session, a target, concepts, bottlenecks, an activation generator, a list of alphas and a CAV directory:

```python
mytcav = tcav.TCAV(sess,
                   target,
                   concepts,
                   bottlenecks,
                   act_gen,
                   alphas,
                   cav_dir=cav_dir,
                   num_random_exp=2)

results = mytcav.run()
```

What you should expect is a results object keyed by concept and target, with the per-concept scores inside. The repository also ships FetchDataAndModels.sh, which the README does not describe in detail, so check that file before assuming the notebooks run without network access. For non-image data, the README points to tcav/tcav_examples/discrete/ and a KDD99 notebook at tcav/tcav_examples/discrete/kdd99_discrete_example.ipynb.

## Where TCAV breaks down: concept quality, layer choice and the random baseline

The largest limitation is that the method inherits every flaw in your concept sets. If the positive examples for "gender" correlate with image source, camera model or annotation style, the CAV will point along whatever separates those sets in activation space, and the score will look meaningful. The README does not describe any procedure for validating that a learned CAV corresponds to the intended concept.

The bottleneck layer is a second unforced choice. Concepts are only linearly separable at some layers and not others, so the same concept can produce very different scores depending on where you read activations. Nothing in the README prescribes how to pick one.

The random experiment count is a third. The README example uses num_random_exp=2, which is enough to demonstrate the API and not enough to establish that a concept score exceeds chance. Treat any published number with a small random experiment count as provisional.

Finally, TCAV is not a debugging tool for a single prediction. It gives a class-level statement, and the README is explicit that this is a global explanation beyond one image. If you need to explain why this particular image was classified as a zebra, use a feature attribution method instead.

## Alternatives: LIME, SHAP and Grad-CAM answer a different question

LIME and SHAP both produce local explanations: for one input, which features pushed the prediction. Grad-CAM produces a coarse localization map over the input for a chosen class, which is closer to TCAV in that it is class-conditional, but it still lives in input space rather than concept space.

The practical difference is what you can say afterwards. With LIME or SHAP you can say that this image's prediction depended on these pixels or these tabular columns. With TCAV you can say that the class prediction is sensitive to the direction that separates striped from unstriped activations, and you can say it without retraining. If your stakeholders ask about concepts, TCAV is the one that speaks their language. If they ask about a single decision, it is not.

There is also a cost difference. LIME and SHAP need no concept examples, only the model and an input. TCAV needs a curated concept dataset before it can return anything, and that curation is the expensive part.

## Maintenance, releases and what the Apache 2.0 licence means here

The repository is not archived, and the last push was on 2026-07-22. The release history is thin: 0.1 on 2018-11-14 and 0.2, labelled API Redesign, on 2018-11-21. setup.py declares version 0.2.2, so the package version and the tagged releases have drifted apart. Anyone planning to depend on tcav should read the source tree rather than assume the release notes describe the current API.

The upgrade surface is small because the API is small: a TCAV constructor, a run method, and a set of model wrappers. The real upgrade risk sits in the dependency pins. requirements.txt fixes scikit-learn==0.20.3 and numpy==1.19.2, both old, and a modern environment will fight those constraints. Expect to unpin and test rather than to upgrade cleanly.

Licensing is Apache 2.0, declared both in the repository LICENSE file and in setup.py. That permits commercial use and modification with the usual notice and patent terms. It is a permissive licence, not a copyleft one, so it does not force you to publish your own model code. This is a description of the licence text, not legal advice; have counsel review anything you ship.

## Conclusion

Use tensorflow/tcav when you already have a frozen TensorFlow classifier and a set of concept examples you can label, and when you need a class-level answer rather than a per-image saliency map. Skip it if you cannot assemble positive and negative examples for every concept, because the method has nothing to learn from otherwise. Before committing, verify that your TensorFlow version matches the pinned requirements, check that your model exposes a bottleneck layer that TCAV can read, and confirm the CAV directory is writable. The last push to the repository was on 2026-07-22, and the newest release is 0.2 from 2018-11-21, so treat the API as settled rather than evolving.

## FAQ

### What are concept activation vectors in tensorflow/tcav?

A CAV is the normal vector of a linear classifier trained on the activations of a concept's positive and negative example sets at a chosen bottleneck layer. TCAV then measures how sensitive a class prediction is to that direction.

### Does tensorflow/tcav require me to retrain my model?

No. The README states that you do not need to change or retrain your network to use TCAV. You supply the trained model plus concept example sets, and the library reads activations from a bottleneck layer.

### Can tensorflow/tcav be used on non-image models?

Yes. The README provides an example for models trained on discrete, non-image data under tcav/tcav_examples/discrete/, including a KDD99 notebook.

### Why is TensorFlow not installed automatically with the tcav package?

The README explains that TensorFlow is left out of install_requires because a user may want either the tensorflow or tensorflow-gpu package, so you install one of them yourself.

## Sources

- [Issues](https://github.com/tensorflow/tcav/issues)
- [License: Apache-2.0](https://github.com/tensorflow/tcav/blob/master/LICENSE)
- [README](https://github.com/tensorflow/tcav/blob/master/README.md)
- [Releases](https://github.com/tensorflow/tcav/releases)
- [tensorflow/tcav on GitHub](https://github.com/tensorflow/tcav)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tensorflow-tcav
