SAELens
github.com · Interpretability
Toolkit for training and analysing sparse autoencoders that decompose model activations into interpretable features.
About the Interpretability category
Libraries and platforms for opening the black box — hooking activations, training sparse autoencoders, attributing outputs and publishing circuit-level findings.
SAELens is one of 8 interpretability tools indexed on Axiomi. Facts on this page come from the tool's official site and public APIs; prices and features change, so confirm details on the official website before you commit.