Pith. sign in

REVIEW 15 cited by

Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.08366 v3 pith:NKIOL3VW submitted 2024-05-14 cs.LG

classification cs.LG
keywords featuresdictionariesfeaturesparsesupervisedevaluationsframeworkinterpretability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Disentangling model activations into meaningful features is a central problem in interpretability. However, the absence of ground-truth for these features in realistic scenarios makes validating recent approaches, such as sparse dictionary learning, elusive. To address this challenge, we propose a framework for evaluating feature dictionaries in the context of specific tasks, by comparing them against \emph{supervised} feature dictionaries. First, we demonstrate that supervised dictionaries achieve excellent approximation, control, and interpretability of model computations on the task. Second, we use the supervised dictionaries to develop and contextualize evaluations of unsupervised dictionaries along the same three axes. We apply this framework to the indirect object identification (IOI) task using GPT-2 Small, with sparse autoencoders (SAEs) trained on either the IOI or OpenWebText datasets. We find that these SAEs capture interpretable features for the IOI task, but they are less successful than supervised features in controlling the model. Finally, we observe two qualitative phenomena in SAE training: feature occlusion (where a causally relevant concept is robustly overshadowed by even slightly higher-magnitude ones in the learned features), and feature over-splitting (where binary features split into many smaller, less interpretable features). We hope that our framework will provide a useful step towards more objective and grounded evaluations of sparse dictionary learning methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

    cs.LG 2026-08 accept novelty 7.0 of 10

    In ablation-based SAE evaluation, the measurement token is selected by the dictionary under test, and holding that token fixed collapses most of the variance that is usually attributed to differences between dictionaries.

  2. Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

    cs.LG 2026-05 conditional novelty 7.0 of 10

    SA-GSAE with Bi-Jump-ReLU enables one latent to encode both polarities of anticorrelated features, Pareto-dominating or matching full-width gated SAEs while reducing dead latents by up to 500x on some LLM hookpoints.

  3. Sparse Autoencoders Do Not Find Canonical Units of Analysis

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Across different dictionary sizes, SAE latents are neither complete nor atomic, so SAEs do not learn a canonical set of features.

  4. Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single-token feature's causal necessity under zero-ablation depends on which SAE family found it: GemmaScope and BatchTopK features stay causally anchored while LlamaScope features are locally redundant.

  5. Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Persistent SAEs learn per-feature persistence coefficients from reconstruction, splitting features into fast local detectors and slow topic-tracking states that retain prompt-injection signals over long contexts.

  6. Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words

    cs.CL 2025-01 conditional novelty 6.0 of 10

    PS-Eval, a Word-in-Context-based benchmark, shows that SAEs optimized for MSE-L0 do not necessarily extract better word-meaning features, and that separation improves in deeper layers and attention outputs.

  7. Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Sparse autoencoders with single-layer encoders provably incur a non-zero amortisation gap, and more expressive encoders (MLPs, iterative optimization) improve both feature recovery and interpretability.

  8. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  9. Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    CRA trains sparse autoencoders on PRM activations and applies backdoor adjustment to estimate true rewards, reducing reward hacking in math reasoning.

  10. On the transferability of Sparse Autoencoders for interpreting compressed models

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Pruning a pretrained sparse autoencoder can produce an interpretability tool for a WANDA-pruned LLM that is roughly comparable to retraining an SAE on the pruned model, though with notable caveats in the reported metrics.

  11. Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Sparse autoencoders trained on Mammo-CLIP features expose a small set of concept-aligned and confounding latent neurons in breast cancer predictions.

  12. InverseScope: Scalable Activation Inversion for Interpreting Large Language Models

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A new conditional-generation architecture plus a feature-consistency metric make activation inversion practical for LLMs up to 7B parameters, with experiments on IOI, RAVEL, and in-context learning.

  13. Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

    cs.SE 2025-06 accept novelty 5.0 of 10

    A new survey organizes LLM interpretation methods by workflow stage and connects them to safety enhancement strategies and tools, covering around 70 works.

  14. Empirical Evaluation of Progressive Coding for Sparse Autoencoders

    cs.LG 2025-04 conditional novelty 4.0 of 10

    Nested 'Matryoshka' sparse autoencoders outperform pruned vanilla sparse autoencoders on reconstruction and recaptured language-model loss, but pruned vanilla features remain more interpretable.

  15. Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

    cs.LG 2026-07 accept

    A scoping review surveying circuit analysis, sparse autoencoders, activation steering, and neurosymbolic frameworks for interpreting and controlling Transformer-based neural networks.

Pith tools