TransformerLens vs nnsight is a choice between two ways of looking inside a language model. TransformerLens gives you a standard, interpretability-friendly view of a model, with consistent hook names on every layer and a one-line way to cache all activations; it is the library most mechanistic interpretability tutorials teach. nnsight wraps the model you already have, any PyTorch model, and lets you read and edit its internals inside a with model.trace() block, and it can run the same code remotely on very large open models through the free NDIF service. Choose TransformerLens to learn and to follow existing research recipes; choose nnsight for big models, unusual architectures, or when you cannot host the model yourself.
The big 2026 change: TransformerLens 4.0
If you learned TransformerLens from older tutorials, note that version 4.0, released on PyPI in September 2026, removed the classic HookedTransformer class. According to the official migration guide, TransformerBridge, introduced in 3.0, is now the only path. The old 3.x branch still gets bug fixes, but no new features or models.
# TransformerLens 4.0
from transformer_lens.model_bridge import TransformerBridge
model = TransformerBridge.boot_transformers("gpt2")
model.enable_compatibility_mode() # HookedTransformer-style weight processing
logits, cache = model.run_with_cache("Hello World")
By default the bridge keeps the raw Hugging Face weights, so its outputs match Transformers. Calling enable_compatibility_mode() reproduces the old defaults, such as folding LayerNorm into the weights, so results from older notebooks line up. The project's README says the bridge supports more than 15,000 models across over 140 architecture families.
TransformerLens vs nnsight at a glance
| TransformerLens | nnsight | |
|---|---|---|
| Approach | Re-exposes the model with standard hook points | Traces the original PyTorch model in place |
| Latest PyPI release (checked October 5, 2026) | 4.0.0 | 0.7.0 |
| Model support | Hugging Face models through TransformerBridge | Any PyTorch model, with wrappers for Hugging Face, diffusers and vLLM |
| Remote execution | No | Yes, through NDIF with an API key |
| Best known for | Teaching and classic circuit-style research | Large models and flexible interventions |
| Licence | MIT | MIT |
How TransformerLens works
TransformerLens was created by Neel Nanda and is now maintained by a community team. Its goal, in the README's words, is to reverse engineer the algorithms a model learned from its weights. You load a model, run it with a cache, and every internal activation, from attention patterns to residual stream values, is available under a predictable name. You can also attach hook functions to edit, remove or replace activations as the model runs, which is how techniques like activation patching are done.
The strengths are consistency and community. Hook names follow the same scheme across supported models, which makes notebooks far easier to move from one model to another. Much of the published interpretability curriculum, including the ARENA course linked from the README, is built on it, and SAELens offers deep integration for working with sparse autoencoders.
The trade-off is that TransformerLens needs an adapter for each architecture, and very new models may lag. It also runs everything locally, so you need enough memory to hold the model; our VRAM guide helps you estimate that.
How nnsight works
nnsight comes from the NDIF team, described in the paper NNsight and NDIF. Instead of re-implementing the model, it records what you want to do and runs it alongside the real forward pass. The current release's quick start looks like this:
from nnsight import LanguageModel
model = LanguageModel("openai-community/gpt2", device_map="auto", dispatch=True)
with model.trace("The Eiffel Tower is in the city of"):
model.transformer.h[0].output[0][:] = 0 # edit a layer's output
hidden = model.transformer.h[-1].output[0].save() # read a later layer
You write the intervention in the order the model executes it, and mark what you want to keep with .save(). The nnsight README also shows generation, batching several prompts, gradients with respect to internal values, and applying modules out of order, as in the logit lens. Its main branch is moving to a new TransformersModel class, so check the docs for the version you install.
The headline feature is remote execution. Add remote=True and an NDIF API key, and the same trace runs on NDIF's servers. The NDIF website describes it as a National Science Foundation project offering free remote access to large open models for research. Available models change, so check NDIF's status page.
The trade-off is that you work with each model's own module names, which differ between architectures, so code is less portable than TransformerLens notebooks.
Side-by-side: common tasks
- Cache every activation: TransformerLens's
run_with_cachedoes it in one call. In nnsight, you save the specific values you need. - Activation patching: both handle it well. TransformerLens uses hook functions; nnsight lets you assign values from one trace into another.
- Steering vectors: both can add a vector to the residual stream; nnsight's in-place edits read naturally.
- Sparse autoencoders: SAELens integrates most deeply with TransformerLens, but its README says SAEs work with Transformers, nnsight or any PyTorch model if you extract activations yourself. You can explore published SAE features on Neuronpedia.
- Toy models: TransformerLens 4.0 can build a model from a config with
TransformerBridge.boot_native, handy if you train a small LLM from scratch to study. - Models too big for your GPU: nnsight with NDIF.
Pros and cons
TransformerLens
- Pros: consistent hook names across models; one-line activation caching; the most tutorials and research recipes; strong sparse autoencoder tooling.
- Cons: version 4.0 breaks older
HookedTransformercode; depends on architecture adapters; local only.
nnsight
- Pros: works with any PyTorch model; very flexible interventions; free remote access to large models through NDIF; supports vLLM for faster generation.
- Cons: code depends on each model's module layout; the deferred-execution style takes some getting used to; remote use requires an account and API key.
Which should you choose?
- You are learning mechanistic interpretability: TransformerLens, following the ARENA curriculum.
- You are reproducing a paper from the last few years: probably TransformerLens, with compatibility mode on.
- You study a model with tens of billions of parameters and lack the hardware: nnsight with NDIF.
- You work on a brand-new or non-standard architecture: nnsight.
- You do sparse autoencoder work: start with SAELens and TransformerLens, and move to nnsight if your model is not supported.
Both libraries depend on open weights; our explainer on open-weight vs open-source models covers which models allow this kind of research.
Who this is for
Researchers, students and engineers getting started with mechanistic interpretability, and anyone whose old TransformerLens code broke after the 4.0 release.
FAQ
What is the difference between TransformerLens and nnsight?
TransformerLens exposes a model through standard, consistent hook points and makes it easy to cache all activations. nnsight traces the original PyTorch model in place, works with any architecture, and can run experiments remotely on large models through NDIF.
Was HookedTransformer removed from TransformerLens?
Yes. TransformerLens 4.0 removed the HookedTransformer stack. Use TransformerBridge.boot_transformers instead, and call enable_compatibility_mode() to reproduce the old weight processing.
Is NDIF free?
NDIF describes itself as providing free remote access to large AI models for research, funded by the US National Science Foundation. You need an API key, and the list of available models changes.
Does SAELens work with nnsight?
Yes, for inference. SAELens integrates most deeply with TransformerLens, but its README says SAEs can be used with nnsight or any PyTorch model by extracting activations and passing them to the SAE.
Which is better for beginners?
TransformerLens, because most interpretability tutorials and courses use it, and its consistent hook names make examples easy to follow.
Do these libraries work with Llama and Qwen models?
Generally yes. TransformerLens supports thousands of Hugging Face models through TransformerBridge, and nnsight works with any PyTorch model. Gated models such as Llama need a Hugging Face access token.