# PAIR-code/lit: a browser UI for inspecting what your model actually learned

> LIT is Google's Learning Interpretability Tool, a Python server plus a TypeScript front end for salience maps, counterfactuals and side-by-side model comparison. It installs from PyPI as lit-nlp, but the web app itself has to be built from source.

**PAIR-code/lit** — The Learning Interpretability Tool: Interactively analyze ML models to understand their behavior in an extensible and framework agnostic interface.

- Repository: https://github.com/PAIR-code/lit
- Website: https://pair-code.github.io/lit
- Stars: 3,667 · Forks: 366
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/pair-code-lit

## What LIT is for, and who ends up using it

LIT answers three questions the README states directly: which examples a model performs poorly on, why a given prediction was made, and whether the model behaves consistently when surface features like verb tense or pronoun gender change. That framing matters, because it puts LIT in the debugging category rather than the reporting category. It is not a dashboard you point at production traffic. It is a workbench for a model you already trained and a dataset you already have on disk.

The intended user is an ML engineer or researcher who is comfortable writing Python wrappers. Two APIs do the work: a Dataset loader and a Model wrapper. If you can express your data as records and your model as a callable that returns predictions, LIT will render those predictions in a browser. The tool covers classification, regression, span labeling, seq2seq and language modeling, and it supports multi-head models and multiple input features. Text, image and tabular data are all in scope, which is wider than most interpretability tools that pick one modality and stay there.

## The architecture: a Python server, a TypeScript client, and a small API surface

The repository is split cleanly. lit_nlp/ holds the Python package, including the server, the built-in interpretation components and the examples. The TypeScript front end lives alongside it and is compiled by yarn. The top-level pyproject.toml declares the package as lit-nlp with requires-python >= 3.9, and the dependency list is where the actual interpretability machinery comes from: saliency, shap, scikit-learn, annoy, Levenshtein, rouge-score, sacrebleu, matplotlib, pandas and numpy pinned below 2.0.0.

Data flows in one direction at startup and then on demand. You construct a server object, hand it a dict of datasets and a dict of models, and it serves a UI on a port. When you select an example in the browser, the client asks the server for predictions and explanations, and the server calls into your Model wrapper. That means the compute cost of an explanation is paid per interaction, not in a batch job. The annoy dependency is what backs embedding-space views; shap and saliency back the local explanation panels.

The design is deliberately framework agnostic. Nothing in the dependency list forces TensorFlow or PyTorch on you. Your wrapper does the framework-specific work and returns plain arrays. The cost of that freedom is that LIT cannot introspect your model graph. It only knows what your wrapper tells it.

## Install LIT with pip and run the GLUE quickstart

The fastest path is PyPI. The default install covers the Python API, the built-in interpretability components and the web application.

```bash
pip install lit-nlp
```

The README notes that dependencies for the bundled demos are optional extras. If you want the discriminative examples such as GLUE and Penguin, or the generative prompt-debugging examples, install the matching extra.

```bash
pip install 'lit-nlp[examples-discriminative-ai]'
pip install 'lit-nlp[examples-generative-ai]'
```

With the discriminative extra in place, the quickstart launches a small BERT-based model fine-tuned on the Stanford Sentiment Treebank. The port is passed explicitly.

```bash
python -m lit_nlp.examples.glue.demo --port=5432 --quickstart
```

Navigate to http://localhost:5432 and the UI loads with that model selected. The README says you can switch to STS-B or MultiNLI from the toolbar or the gear icon in the upper right. That is the whole first-run experience: no config file, no dataset path, no model checkpoint to download by hand. Other examples follow the same shape, `python -m lit_nlp.examples.<example_name>.demo --port=5432` with optional arguments.

One caveat belongs here rather than in a footnote. The README states plainly that the repository does not include a distributable version of the LIT app and that you must build it from source. The pip package covers the Python side; if you are working from a clone and want the front end, you need Node tooling and a build step.

```bash
(cd lit_nlp; yarn && yarn build)
```

The README also points at a known yarn issue on Ubuntu and Debian and links to the yarn install instructions, so a failed yarn run on those distributions is a version problem, not a LIT problem.

## Wrapping your own model is the real integration cost

The README's own instructions for custom models are three bullets: write a data loader following the Dataset API, write a model wrapper following the Model API, and pass models, datasets and any additional interpretation components to the LIT server class. It links to a full walkthrough under the API documentation.

That is a small surface, and it is also the entire cost of adoption. LIT has no auto-discovery. It will not read a Hugging Face checkpoint or a SavedModel and figure out the rest. Every prediction type you want in the UI, whether that is a classification score, a regression value, a span label or a generated sequence, has to be produced by code you write and shaped the way the API expects.

The upside is that the wrapper is where you put the awkward parts of your setup: tokenization, batching, device placement, feature extraction. LIT never sees them. The downside is that a bug in your wrapper looks like a bug in LIT's UI, and the two are hard to tell apart when a salience map comes back empty.

## Where LIT stops being the right tool

LIT is a local, interactive tool. The README describes it as runnable as a standalone server or inside notebook environments such as Colab, Jupyter and Google Cloud Vertex AI notebooks. Nothing in the README describes authentication, multi-tenancy, or a hosted deployment model. If your requirement is a shared explainability endpoint that other teams call, LIT is not that, and the container guide points at Docker and Podman for local use rather than at a service topology.

The second boundary is scale. Explanations are computed when you interact with an example, and the aggregate views operate on the dataset you loaded. The README's feature list mentions slicing, binning and custom metrics, which are the tools you use to find patterns, but the workflow is still a person looking at a browser. Teams that want automated regression checks on explanation quality will not find them here.

The third is the build split. Anyone who installs lit-nlp from PyPI and expects the full web application to be present should read the source-install section carefully. The Python package and the compiled front end are separate concerns, and only one of them comes from pip.

## How LIT differs from SHAP, Captum and notebook-only explainability

SHAP is the closest comparison, and it is already inside LIT: pyproject.toml pins shap>=0.42.0,<0.46.0 as a dependency. The difference is the layer each one occupies. SHAP is a library that returns numbers, and you decide how to plot them. LIT is a server and a UI that calls explanation code, including SHAP-style attribution, and renders the result next to the input, the prediction and a set of counterfactuals you can edit by hand.

That distinction has practical consequences. With SHAP alone, comparing two models on the same example means writing a loop and producing two figures. In LIT, the README lists side-by-side mode as a first-class feature for comparing two or more models, or one model on a pair of examples. Counterfactual generation is also built in, either through manual edits or generator plug-ins, and those generated examples go back through the model for evaluation.

The trade-off runs the other way too. A SHAP script runs in CI, produces a file and diffs cleanly. A LIT session produces a browser state that a human has to interpret. If your goal is a number in a report, reach for the library. If your goal is to find the twenty examples your model gets wrong for a reason you did not anticipate, the interactive loop is the point.

## Licence, maintenance and what an upgrade costs you

LIT is Apache-2.0, declared in the LICENSE file and in the pyproject.toml classifier as OSI Approved :: Apache Software License. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on your own code. It does not settle the question of what you may do with model outputs or datasets you load into the UI, which is a separate matter.

The last push to the default branch was on 2026-09-09, and the repository is not archived. The most recent tagged release is v1.3.1 from 2024-12-20, following v1.3 in 2024-10-22 and v1.2 in 2024-06-26. So the release cadence and the commit activity are not the same story, and anyone pinning a version should plan around the tags rather than the branch.

Upgrade cost is dominated by the dependency pins, not by LIT's own API. numpy is held below 2.0.0, matplotlib is held below 3.9.0 in requirements.txt, and shap is held below 0.46.0. In a shared environment those ceilings will collide with other packages, and resolving the conflict is your problem. The pyproject.toml and requirements.txt files carry LINT.IfChange markers around the version and dependency blocks, which tells you the two are kept in sync by tooling and that editing one by hand is the wrong move. There is also a build_package.sh at the repository root and a RELEASE.md, so the release process is documented rather than improvised.

## Conclusion

Adopt LIT if you already have a trained model and a dataset you can wrap in the Dataset and Model APIs, and you want a browser UI for salience, slicing and counterfactuals rather than a script that prints numbers. Skip it if you need a hosted service, a one-line explainability call, or a tool that ships a prebuilt web bundle. Before committing, verify that your Python is 3.9 or newer, that your model wrapper returns the prediction types your task needs, and that you can run yarn build in lit_nlp, because the repository does not include a distributable copy of the app.

## FAQ

### What is a lit code?

In this context LIT stands for the Learning Interpretability Tool, formerly the Language Interpretability Tool, a visual and interactive model-understanding tool from PAIR-code that supports text, image and tabular data. It is distributed as the Python package lit-nlp.

### How do I install LIT?

The README gives pip install lit-nlp for the Python API, built-in interpretability components and web application. Dependencies for the bundled demos are optional extras, for example lit-nlp[examples-discriminative-ai] or lit-nlp[examples-generative-ai].

### Does LIT work with PyTorch as well as TensorFlow?

The README describes LIT as framework-agnostic and compatible with TensorFlow, PyTorch and more. Neither framework appears in the core dependency list, because your Model wrapper does the framework-specific work and returns predictions to the server.

### Can I run LIT inside a Jupyter or Colab notebook?

Yes. The README states that LIT can be run as a standalone server or inside notebook environments such as Colab, Jupyter and Google Cloud Vertex AI notebooks, and it points to notebooks under lit_nlp/examples/notebooks plus a Colab demo for a sentiment classifier.

### Why is there no prebuilt LIT web app in the repository?

The README states that the LIT repo does not include a distributable version of the LIT app and that you must build it from source. The build step given is (cd lit_nlp; yarn && yarn build), and the source install requires Python 3.9 or newer.

## Sources

- [License: Apache-2.0](https://github.com/PAIR-code/lit/blob/main/LICENSE)
- [PAIR-code/lit on GitHub](https://github.com/PAIR-code/lit)
- [Project website](https://pair-code.github.io/lit)
- [README](https://github.com/PAIR-code/lit/blob/main/README.md)
- [Releases](https://github.com/PAIR-code/lit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/pair-code-lit
