Model or dataset
TransformerLensOrg/TransformerLens avatar
TransformerLensOrg/TransformerLens

TransformerLens 4.0: What the TransformerBridge Rewrite Changes for Mechanistic Interpretability

A library for mechanistic interpretability of GPT-style language models

3,926 stars699 forksPythonMIT

At a glance

What is it?
TransformerLens exposes the internal activations of GPT-style models so you can cache, edit and ablate them. Version 4.0 replaces HookedTransformer with a TransformerBridge that preserves raw HuggingFace numerics, and that change is the whole story for existing users.
Who is it for?
Adopt TransformerLens if you need to cache and intervene on internal activations across many architecture families, and you are starting fresh on 4.x with TransformerBridge. Do not adopt it if your existing research code calls HookedTransformer.from_pretrained, because that entry point was removed in 4.0 and only bridge.enable_compatibility_mode() restores the old numerics.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TransformerLens Is For, and Who It Is Actually For

The problem TransformerLens addresses is that a trained language model is a matrix of weights with no labels on the parts. Mechanistic interpretability treats the model as a program to reverse engineer, and that requires reading intermediate tensors while the forward pass runs. A stock HuggingFace model gives you logits. TransformerLens gives you the residual stream, attention patterns, MLP outputs and per-head values, and lets you overwrite them mid-pass.

The audience is narrow and specific. It is researchers who already know what an induction head is and want to ablate one. The README points newcomers at the ARENA mechanistic interpretability tutorials and Neel Nanda's getting-started guide rather than pretending the library is self-teaching. The gallery is a list of papers (grokking progress measures, sparse probing, automated circuit discovery, the tuned lens) that used the library, which tells you the intended consumer is someone producing that kind of artifact, not someone doing prompt engineering.

If you want to inspect a model's behaviour through its outputs, you do not need this. If you want to know which attention head writes the token that flips a prediction, you do.

TransformerBridge: The 4.0 Architecture and the Numerics Trap

The recommended entry point is TransformerBridge. The README's own example boots a model and returns logits plus a cache in two lines. Under the hood the bridge loads through HuggingFace and, by default, preserves raw HuggingFace weights, so logits and activations match what HF produces.

That default is a deliberate break from the past. Legacy HookedTransformer folded LayerNorm and centered weights, which meant the same model produced different internal numbers depending on which library you loaded it with. Anyone comparing a TransformerLens activation against a HuggingFace activation had to account for that transformation. TransformerLens 4.0 removes the discrepancy by not applying it, and the README states plainly that logits and activations match HF, not legacy HookedTransformer.

The escape hatch exists: calling bridge.enable_compatibility_mode() after booting gives HookedTransformer-equivalent numerics. That is a real migration aid, but it is opt-in, so any notebook or script written before 4.0 that silently relied on the old folding will now produce different activations without raising an error. Numerics changes that do not throw are the expensive kind in this field, because a circuit that looked clean under one normalization can look different under another. Treat the migration guide as required reading, not optional.

The scale claim is concrete: 15,000+ open source models across 140+ architecture families, with the full inventory in transformer_lens/tools/model_registry/data/supported_models.json. That file, not the README prose, is where you confirm whether your model is covered.

Installing TransformerLens and Running a First Cached Forward Pass

Installation is a single pip command. The README gives a separate pin for Python 3.8 or 3.9, but pyproject.toml sets requires-python to >=3.10,<4.0, so on a current interpreter the plain install is what you want.

bash
pip install transformer_lens

On Python 3.8 or 3.9 the README instructs installing the 2.x line instead:

bash
pip install 'transformer_lens~=2.0'

For a first real use, boot GPT-2 Small on CPU and run one forward pass that returns both logits and the activation cache. This is the README's quick-start example verbatim.

python
from transformer_lens.model_bridge import TransformerBridge

# Load a model (eg GPT-2 Small)
bridge = TransformerBridge.boot_transformers("gpt2", device="cpu")

# Run the model and get logits and activations
logits, activations = bridge.run_with_cache("Hello World")

What you should see is a logits tensor with a vocabulary dimension and an activations object holding the cached tensors for that pass. The bridge defaults to raw HuggingFace numerics, so cross-checking those logits against the same model loaded through transformers should agree. If you need the old HookedTransformer numbers, call bridge.enable_compatibility_mode() after booting and before running.

Gated models (Llama, Mistral, Gemma and others) need a HuggingFace token in the environment. The repository ships a .env.example that documents the variable and how to source it.

bash
set -a; source .env; set +a

The same file documents two optional switches: TRANSFORMERLENS_HF_RETRY=1 to retry HuggingFace Hub 429s in ad-hoc scripts, and TRANSFORMERLENS_ALLOW_MPS=1 to permit the MPS device on macOS, which is off by default to avoid divergence from CUDA and CPU. If your first run fails on a gated checkpoint, the missing HF_TOKEN is the first thing to check, not the model name.

Where TransformerLens Breaks Down

The 4.0 migration is the sharpest limitation. HookedTransformer.from_pretrained was removed, and the README says so directly. Every tutorial, Colab notebook and blog post written against 3.x that calls that method is now broken until it is rewritten around TransformerBridge. The compatibility mode restores numerics, not the old API surface, so it does not rescue code that still imports HookedTransformer.

The second constraint is numerical comparability across versions. Because the default path changed from folded and centered weights to raw HuggingFace weights, results computed under 3.x are not directly comparable to results computed under 4.x unless you opt into compatibility mode. For a field where a published circuit diagram is the deliverable, that is a reproducibility hazard, and the documentation does not promise a per-layer diff of what changed.

The third is hardware. MPS on macOS is disabled unless you set TRANSFORMERLENS_ALLOW_MPS=1, and the .env.example explains the reason: divergence from CUDA and CPU is hard to debug. That is an honest trade-off, but it means an Apple laptop is not a first-class target for anything you intend to publish.

Finally, this is the wrong tool when the model you care about is not a GPT-style transformer covered by the registry, or when your question is about training dynamics rather than the learned algorithm. The README frames the goal as taking a trained model and reverse engineering it from its weights. If you are still training, or if you need a behavioural evaluation rather than an internal one, the activation cache is not the abstraction you want.

TransformerLens Against Generic HuggingFace Access

The obvious alternative is using transformers directly and registering forward hooks yourself. The difference is not capability, because PyTorch hooks can capture any tensor. The difference is the abstraction you maintain. Doing it by hand means writing per-architecture logic for naming conventions, residual stream locations, per-head slicing and attention pattern extraction, and redoing that work every time you move to a new model family. TransformerLens centralizes that in the model registry and the bridge, and the 140+ architecture family count is the payoff for accepting its naming scheme.

A second alternative is training a proxy model you fully control, the approach behind the toy model of universality paper listed in the gallery. That gives you ground truth and cheap iteration, at the cost of the question you actually care about, which is what a real large model learned.

The honest framing is that TransformerLens buys you a shared vocabulary for activations across many architectures. You pay for it with a dependency on that vocabulary staying stable, and 4.0 is the moment where it did not.

Maintenance, Licensing and Upgrade Cost

The repository is not archived and the last push was on 2026-09-23. The release history shows v4.0.0 on 2026-09-21, v3.9.0 on 2026-09-11 and v4.0.0b2 on 2026-09-02, so the 4.x line is days old at the time of writing and the 3.x line was still receiving releases two weeks before it. The README names Bryce Meyer and Jonah Larson as maintainers and Neel Nanda as the creator.

The licence is MIT, declared both in the repository and in pyproject.toml. MIT is permissive: it allows commercial and academic use, modification and redistribution with the licence text preserved. That is a statement about the licence terms, not legal advice about your situation.

The upgrade cost is the part to budget for. Moving from 3.x to 4.x means rewriting model loading around TransformerBridge, deciding per project whether to call enable_compatibility_mode(), and re-validating any cached activation results you intend to compare against older work. The dependency floor is also worth noting: torch>=2.6 and transformers>=5.9.0 are required, so an environment pinned to older versions of either will need to move before TransformerLens 4.x installs cleanly.

Editorial conclusion

Adopt TransformerLens if you need to cache and intervene on internal activations across many architecture families, and you are starting fresh on 4.x with TransformerBridge. Do not adopt it if your existing research code calls HookedTransformer.from_pretrained, because that entry point was removed in 4.0 and only bridge.enable_compatibility_mode() restores the old numerics. Verify first that the specific model you care about appears in transformer_lens/tools/model_registry/data/supported_models.json, and that your code paths do not depend on LayerNorm folding, which the default bridge path no longer applies.

Frequently asked questions

What is TransformerLens?

It is a Python library for mechanistic interpretability of GPT-style language models, described in the README as a library for reverse engineering the algorithms a trained model learned from its weights. It exposes internal activations so you can cache, edit, remove or replace them as the model runs.

What does TransformerLens do?

It loads open source language models and lets you cache any internal activation during a forward pass, then intervene on those activations. The README states it covers 15,000+ models across 140+ architecture families, with the inventory in transformer_lens/tools/model_registry/data/supported_models.json.

What are the alternatives to TransformerLens?

The main practical alternative is using HuggingFace transformers directly with your own forward hooks, which can capture any tensor but requires you to write and maintain per-architecture logic for activation naming and per-head slicing. A second option is training a small proxy model you fully control, which gives ground truth at the cost of not studying a real large model.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. TransformerLensOrg/TransformerLens on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/transformerlensorg-transformerlens.svg)](https://hysenlabs.com/projects/transformerlensorg-transformerlens)