# Five years of commits, no release since October 2021, and a toolchain from 2019

> A visual interface for probing a trained classifier or regressor, served through TensorBoard or embedded in a notebook, with fairness and attribution tooling. The repository is still being worked on, its newest tag is from 2021, and the only package manifest inside it belongs to another project.

**PAIR-code/what-if-tool** — Source code/webpage/demos for the What-If Tool

- Repository: https://github.com/PAIR-code/what-if-tool
- Website: https://pair-code.github.io/what-if-tool
- Stars: 1,018 · Forks: 184
- Language: HTML
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pair-code-what-if-tool

## Five years of commits, no release since October 2021, and a 2019 toolchain

The version history and the commit history have been living in different decades.

The three newest tags are from 2021, 2020 and 2020, with the most recent a patch release in October 2021. The default branch was pushed in September 2026. The repository is not marked archived.

So the tree has moved roughly five years past its newest release, and there is no release note that covers any of it. For a tool that lives inside another product and is versioned against that product, that is a plausible arrangement, but it means a user has two different version numbers to reason about and this file does not help with either.

The dependency set tells you which era the recent commits live in. The only package manifest in the repository pins Angular at 8, TypeScript at roughly 3.4, Bazel between 0.23 and 0.34, a formatter at 1.18, and type definitions for a Node major version in the teens. Those are all 2019 releases.

That is the finding worth sitting with. A branch receiving commits in 2026 is being built with a frontend toolchain from 2019, which means the maintenance is real and confined to keeping a 2019 stack working. It also means the newest thing in the repository is old, and a reader who assumes otherwise from the push date will be wrong about almost everything except the fact that someone cared recently.

## The only package manifest in the repository belongs to TensorBoard

There is a package manifest at the root of this repository, and it is not this project's.

It declares the name tensorboard, a version string that reads zero point zero point zero unused, a description naming a visualisation toolkit, an author listed as the authors of the framework, a licence, and both a repository address and a homepage pointing at a different repository entirely.

Its scripts explain why it is there. Installing it compiles an Angular metadata configuration. Building it runs a Bazel build over everything. Testing it uses the watch mode variant of Bazel. Linting and fixing lint run a formatter over a path scoped to one of the top-level directories.

None of that publishes anything. The manifest exists so that a frontend build, for a user interface hosted inside another product, has the node tooling it needs, and the version string says so explicitly.

This is worth naming because it is the kind of detail that makes a repository confusing to enter. There is no manifest for the project itself, so there is nothing to check a version against, nothing to publish, and no dependency list you can read to learn what this tool requires. The build system here is Bazel, configured through three files at the root, and the JavaScript story is a vendored fragment of someone else's.

The practical consequence is that the tool's compatibility surface is defined by the hosting product's requirements rather than by anything declared here.

## The serving route speaks gRPC, so the port you want is not the port you expect

In the TensorBoard integration there are only two things to provide: where the model server is, and which port. Everything else depends on which of three serving interfaces the model was exported through.

Two of them are described as the simplest case, needing no configuration in the tool beyond setting the model type. The third is where the specification gets exact.

For a model served through the general predict interface, the input must be serialized in one of two example proto formats, and the output shape is prescribed. A classification model must return a two-dimensional float tensor holding class probabilities for every class index, for each example. A regression model must return a float tensor holding one score per example. Those are shapes, not suggestions, and a model that exports a differently shaped tensor does not qualify.

Then the port. The tool queries the served model over the gRPC interface and not the RESTful one. The container image for the serving framework uses one specific port for that interface, and the document points out that if you are using the container approach, that is the port you should give the tool.

That is a genuinely costly detail to discover the hard way, because every piece of serving documentation a person is likely to have read will name the REST port first. The interface looks like it should connect to the same endpoint the rest of your tooling uses, and the failure would appear as a connection problem rather than as a configuration mistake.

So: two values to supply, one interface to choose, one output shape to satisfy, and one port that is not the obvious one.

## The escape hatch is a runtime flag with unsafe written into its name

There is an alternative to serving a model at all, and the project names its own risk in the flag.

Instead of pointing the tool at a model server, you can hand it a Python function that takes examples and returns predictions, passed through a runtime argument whose name contains the word unsafe. Whatever the argument's full spelling, the naming is a deliberate signal: this path bypasses the verification that a served model went through an export pipeline.

What the path buys you is model independence. The document says you can load any model, including models from frameworks other than the one this tool is built around, including ones that do not take example protos as input at all, as long as the input and output specifications of your function are correct.

Read that clause carefully, because it carries the entire correctness burden. The tool does not check your function. If it returns the wrong shape, or the wrong order, or a probability vector where a score was expected, the interface will render whatever it receives. Nothing in the description mentions validation, and the flag name is the only warning offered.

There is a real trade here. The unsafe path is the only way to use this on a model that was not exported through the serving stack, and for a lot of models that is the only option that exists. Used carelessly it is a way to generate a confident visualisation of a function that does not do what you think.

Attribution display is available on both of the non-served paths, which is the one place the unsafe route is strictly more capable rather than merely more permissive.

## No code required holds on the served path and not on the notebook path

The stated purpose is a simple and intuitive way to play with a trained model through a visual interface with absolutely no code required.

That claim is true for one of the two integration routes and false for the other.

On the TensorBoard route, a person who has already exported a model through one of the two simpler serving interfaces supplies a host and a port and a model type. Nothing else is required. That is genuinely no code.

On the notebook route, the requirements are stated in the same document and they are substantial. A model of a particular estimator family has to take one of two example proto formats. A hosted machine-learning platform model has to take example protos, sequence protos, or raw JSON objects. Or you write a prediction function yourself, which is the unsafe path again.

So the no-code promise is scoped to a model that was already exported, and the document does not draw that line explicitly. It reads as a general property of the tool when it is a property of one deployment route.

The intended user is still clear from the rest of it. A notebook that starts from a spreadsheet, converts the rows to the proto format, trains a classifier and then opens the interface on it is the documented path for someone who has a dataset and a model to train. That is code, and the document says so by showing it.

## Four demos share port 6006, which is where the plugin itself lives

The demo section is the fastest way to see what the tool does, and it is built as four servers with one notable property.

There is a binary classifier on a census dataset predicting whether someone earns more or less than a stated threshold, a binary image classifier for smile detection on a face dataset, a multiclass classifier splitting a flower dataset into three classes from four measurements, and a regression model predicting a person's age. Each has its own build target under the dashboard package, and each is launched with the same build tool.

The property is the port. All four are served on the same address, and it is the address the tensor board application itself uses by default. So the demos occupy the port the tool normally lives on, which is convenient for a demo and means you cannot run a demo and a live session at once.

The fourth demo is the interesting one. It returns attribution values in addition to predictions, and it does so using plain gradients rather than a dedicated attribution library, specifically so the interface can be shown displaying attributions that came from a real model rather than from a toy example.

The demos are not a separate example package. Their build targets sit under the same directory as the dashboard itself, which is why they inherit its dependencies and its port.

There is also a page on the project's site collecting both web and hosted-notebook demos, so the browser versions and the notebook versions are presented as the same set in two forms.

## Eight notebooks at the root describe what people do with it

The repository's worked examples are at its top level rather than in a directory, and their filenames are the best description of how the tool gets used.

There are notebooks for age regression, for the smile detector, for comparing two models, for comparing two toxicity text models, for one fairness dataset comparison, for the same comparison again with a different attribution method, a templated issue notebook, and the getting-started notebook that walks from a spreadsheet through conversion and training to the interface.

Two things follow from those names. Fairness and bias auditing is a first-class use, not a side feature: a whole notebook is about one public fairness dataset, and another compares text toxicity models, which is a harm-category question rather than an accuracy question. And attribution has more than one route into the interface: the age demo computes gradients, while one notebook does the same analysis with a separate explanation library.

The repeated-notebook pattern is also the repository's testing story. There are no unit tests for the interface in the visible root; instead the demonstrations are analyses you can run, and a change that breaks a plot shows up when one of them stops rendering. That is a weaker guarantee than a test suite and a better sample of real usage.

One more root document is a workshop write-up, which suggests the tool has been taught in a course setting as well as used in production.

## Conclusion

This is still worth using, and the reason is the entry points rather than the code. If you already export a model through the serving APIs that accept example protos, you can point the interface at it and look at performance and fairness over subsets without writing a visualization. If your model is anything else, the custom prediction function is the route, and the flag that enables it carries the word unsafe in its own name, which is a reasonable hint about where the responsibility sits. Three things to check first. Nothing has been released in five years, so treat the branch rather than a tag as the artefact and expect no changelog to explain what changed. The serving route speaks gRPC and not the REST interface, so the port you give it is not the port most serving documentation mentions. And the repository ships notebooks rather than tests as its worked examples, which means the demonstrations of what this tool does are analyses of particular datasets rather than guarantees about its behaviour.

## FAQ

### What is the What-If Tool?

A visual interface for exploring a trained classification or regression model. You run inference over a large set of examples and see the results plotted, you can edit individual examples by hand or programmatically and re-run them to see what changed, and there is tooling for looking at model performance and fairness over subsets of a dataset.

### How do I use the What-If Tool in TensorBoard?

You supply a model server host and port served through the machine learning serving stack, and the model type. Models exported through the classification or regression interfaces need no further configuration. Models exported through the general predict interface must take serialized example protos and return the prescribed tensor shapes.

### Which port does the What-If Tool connect to?

It queries the served model over the gRPC interface rather than the RESTful one. The serving container image uses port 8500 for that interface, so when you use the container approach that is the port to give the tool, not the REST port that most serving documentation names first.

### Can the What-If Tool work with a model that is not from TensorFlow?

Yes, through a custom prediction function passed as a runtime argument, which is the flag whose name contains the word unsafe. The document says you can load any model, including ones that do not take example proto inputs, as long as your function's input and output specifications are correct, and it does not describe any validation of that function.

### Is the What-If Tool still maintained, and what version should I use?

The repository is not marked archived and its default branch was pushed in September 2026, but the newest release tag is from October 2021, so the branch is roughly five years ahead of the newest release. The frontend toolchain in the only package manifest pins 2019 versions of the build and language tooling.

## Sources

- [License: Apache-2.0](https://github.com/PAIR-code/what-if-tool/blob/master/LICENSE)
- [PAIR-code/what-if-tool on GitHub](https://github.com/PAIR-code/what-if-tool)
- [Project website](https://pair-code.github.io/what-if-tool)
- [README](https://github.com/PAIR-code/what-if-tool/blob/master/README.md)
- [Releases](https://github.com/PAIR-code/what-if-tool/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pair-code-what-if-tool
