REVIEW 15 cited by
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Disentangling model activations into meaningful features is a central problem in interpretability. However, the absence of ground-truth for these features in realistic scenarios makes validating recent approaches, such as sparse dictionary learning, elusive. To address this challenge, we propose a framework for evaluating feature dictionaries in the context of specific tasks, by comparing them against \emph{supervised} feature dictionaries. First, we demonstrate that supervised dictionaries achieve excellent approximation, control, and interpretability of model computations on the task. Second, we use the supervised dictionaries to develop and contextualize evaluations of unsupervised dictionaries along the same three axes. We apply this framework to the indirect object identification (IOI) task using GPT-2 Small, with sparse autoencoders (SAEs) trained on either the IOI or OpenWebText datasets. We find that these SAEs capture interpretable features for the IOI task, but they are less successful than supervised features in controlling the model. Finally, we observe two qualitative phenomena in SAE training: feature occlusion (where a causally relevant concept is robustly overshadowed by even slightly higher-magnitude ones in the learned features), and feature over-splitting (where binary features split into many smaller, less interpretable features). We hope that our framework will provide a useful step towards more objective and grounded evaluations of sparse dictionary learning methods.
Forward citations
Cited by 15 Pith papers
-
Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation
In ablation-based SAE evaluation, the measurement token is selected by the dictionary under test, and holding that token fixed collapses most of the variance that is usually attributed to differences between dictionaries.
-
Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations
SA-GSAE with Bi-Jump-ReLU enables one latent to encode both polarities of anticorrelated features, Pareto-dominating or matching full-width gated SAEs while reducing dead latents by up to 500x on some LLM hookpoints.
-
Sparse Autoencoders Do Not Find Canonical Units of Analysis
Across different dictionary sizes, SAE latents are neither complete nor atomic, so SAEs do not learn a canonical set of features.
-
Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
A single-token feature's causal necessity under zero-ablation depends on which SAE family found it: GemmaScope and BatchTopK features stay causally anchored while LlamaScope features are locally redundant.
-
Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models
Persistent SAEs learn per-feature persistence coefficients from reconstruction, splitting features into fast local detectors and slow topic-tracking states that retain prompt-injection signals over long contexts.
-
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
PS-Eval, a Word-in-Context-based benchmark, shows that SAEs optimized for MSE-L0 do not necessarily extract better word-meaning features, and that separation improves in deeper layers and attention outputs.
-
Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders
Sparse autoencoders with single-layer encoders provably incur a non-zero amortisation gap, and more expressive encoders (MLPs, iterative optimization) improve both feature recovery and interpretability.
-
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.
-
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
CRA trains sparse autoencoders on PRM activations and applies backdoor adjustment to estimate true rewards, reducing reward hacking in math reasoning.
-
On the transferability of Sparse Autoencoders for interpreting compressed models
Pruning a pretrained sparse autoencoder can produce an interpretability tool for a WANDA-pruned LLM that is roughly comparable to retraining an SAE on the pruned model, though with notable caveats in the reported metrics.
-
Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders
Sparse autoencoders trained on Mammo-CLIP features expose a small set of concept-aligned and confounding latent neurons in breast cancer predictions.
-
InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
A new conditional-generation architecture plus a feature-consistency metric make activation inversion practical for LLMs up to 7B parameters, with experiments on IOI, RAVEL, and in-context learning.
-
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
A new survey organizes LLM interpretation methods by workflow stage and connects them to safety enhancement strategies and tools, covering around 70 works.
-
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
Nested 'Matryoshka' sparse autoencoders outperform pruned vanilla sparse autoencoders on reconstruction and recaptured language-model loss, but pruned vanilla features remain more interpretable.
-
Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning
A scoping review surveying circuit analysis, sparse autoencoders, activation steering, and neurosymbolic frameworks for interpreting and controlling Transformer-based neural networks.
Discussion (0). Continue with ORCID to comment.