Pith. sign in

REVIEW 31 cited by

Automatically Interpreting Millions of Features in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13928 v3 pith:3MYT5KLZ submitted 2024-10-17 cs.LG cs.CL

classification cs.LGcs.CL
keywords explanationssaesfeaturesneuronstechniquestrainedactivationsavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

While the activations of neurons in deep neural networks usually do not have a simple human-understandable interpretation, sparse autoencoders (SAEs) can be used to transform these activations into a higher-dimensional latent space which may be more easily interpretable. However, these SAEs can have millions of distinct latent features, making it infeasible for humans to manually interpret each one. In this work, we build an open-source automated pipeline to generate and evaluate natural language explanations for SAE features using LLMs. We test our framework on SAEs of varying sizes, activation functions, and losses, trained on two different open-weight LLMs. We introduce five new techniques to score the quality of explanations that are cheaper to run than the previous state of the art. One of these techniques, intervention scoring, evaluates the interpretability of the effects of intervening on a feature, which we find explains features that are not recalled by existing methods. We propose guidelines for generating better explanations that remain valid for a broader set of activating contexts, and discuss pitfalls with existing scoring techniques. We use our explanations to measure the semantic similarity of independently trained SAEs, and find that SAEs trained on nearby layers of the residual stream are highly similar. Our large-scale analysis confirms that SAE latents are indeed much more interpretable than neurons, even when neurons are sparsified using top-$k$ postprocessing. Our code is available at https://github.com/EleutherAI/sae-auto-interp, and our explanations are available at https://huggingface.co/datasets/EleutherAI/auto_interp_explanations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features

    cs.LG 2026-05 accept novelty 8.0 of 10

    Many distinct SAE features share identical explanations, with the average annotation resolving only 70% of feature identity in a large annotated dataset.

  2. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  3. MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents

    cs.LG 2026-04 conditional novelty 7.0 of 10

    Joint training of a primary SAE with a meta SAE that applies a decomposability penalty on decoder directions produces more atomic latents, shown by 7.5% lower mean absolute phi and 7.6% higher fuzzing scores on GPT-2.

  4. Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

    econ.EM 2025-11 conditional novelty 7.0 of 10

    A framework that uses LLM sparse autoencoder features as high-dimensional outcome variables, with new k-FWER controlling bootstrap tests and formalized autointerp evaluation, to make interpretable discoveries from uns...

  5. Training, Reading, and Editing Legible Transformers

    cs.LG 2026-07 conditional novelty 6.5 of 10

    A variance-floor objective plus learned operator gates produce an end-to-end legible transformer whose crisp units are 50–184× more local to edit and can be reshaped from fan-out to fan-in circuits without quality loss.

  6. ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

    cs.CL 2026-08 conditional novelty 6.0 of 10

    ChronoLens uses feature-aligned crosscoders to show that historical language change has comparable magnitude across linguistic levels within a language, but divergent timing and direction across five parliamentary languages.

  7. Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Sparse autoencoders provide a basis for sensible concept hierarchies on visual data but are undermined by hard and soft feature absorption.

  8. Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    A new SAE-based framework extracts visual, textual, and multimodal concepts from VLMs and reports up to 45% better visual concept quality on a VQA dataset while identifying multimodal concepts.

  9. ICA Lens: Interpreting Language Models Without Training Another Dictionary

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    ICALens applies an optimized ICA workflow to LLM activations and recovers compact interpretable directions that match or exceed public SAEs on SAEBench probing and perturbation tasks without per-layer dictionary training.

  10. Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Sparse autoencoders on a TTS language model yield interpretable features that causally control attributes such as laughter, gender, and speech rate via targeted interventions.

  11. Are Sparse Autoencoder Benchmarks Reliable?

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    An audit of SAEBench reveals that Targeted Probe Perturbation and Spurious Correlation Removal metrics fail reliability tests and should not be used to evaluate sparse autoencoders.

  12. The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    SAEs exhibit a rate-distortion-polysemanticity tradeoff where monosemanticity increases rate and distortion, with optimal polysemanticity set by feature co-occurrence probabilities in the data.

  13. Why Retrieval-Augmented Generation Fails: A Graph Perspective

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Attribution graphs reveal that RAG failures arise from shallow fragmented evidence flow in LLMs, enabling topology-based detection and targeted interventions that reinforce question-guided routing.

  14. Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    LLMs perform in-context learning as trajectories through a structured low-dimensional conceptual belief space, with the structure visible in both behavior and internal representations and causally manipulable via inte...

  15. Domain Restriction via Multi SAE Layer Transitions

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Multi-layer SAE transitions capture domain-specific signatures that distinguish OOD texts in Gemma-2 models.

  16. Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Tree SAE learns hierarchical feature structures by combining activation coverage with a new reconstruction condition, outperforming prior SAEs on hierarchical pair detection while matching state-of-the-art benchmark p...

  17. Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Tree SAE learns hierarchical feature pairs in sparse autoencoders by combining activation coverage with a new reconstruction condition, outperforming prior methods on hierarchy detection while remaining competitive on...

  18. From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features

    cs.AI 2026-05 conditional novelty 6.0 of 10

    Graph-motif clustering of SAE features via a frequency-binned WL kernel recovers structural families not captured by decoder cosine similarity or token histograms.

  19. Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    DACO curates a 15,000-concept dictionary from 400K image-caption pairs and uses it to initialize an SAE that enables granular, concept-specific steering of MLLM activations, raising safety scores on MM-SafetyBench and...

  20. LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

    cs.CL 2026-04 conditional novelty 6.0 of 10

    LangFIR isolates sparse language-specific SAE features via random-token filtering on monolingual data and uses them to steer multilingual LLM generation better than parallel-data methods.

  21. LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Language-specific sparse autoencoder features can be identified from monolingual data alone by filtering out features that also activate on random-token sequences, and steering with these features improves language control.

  22. Visual Persuasion: What Influences Decisions of Vision-Language Models?

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Naturalistic edits to background, lighting, and staging systematically shift the choices of frontier vision-language models, and a new optimization-plus-interpretation framework (CVPO) discovers and explains the visua...

  23. Prototype Transformer: Towards Language Model Architectures Interpretable by Design

    cs.AI 2026-02 conditional novelty 6.0 of 10

    ProtoT is an autoregressive language model whose attention is replaced by learned prototype channels that are claimed to capture nameable concepts and allow targeted edits, at linear sequence cost but with slightly lo...

  24. Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

    econ.EM 2025-11 unverdicted novelty 6.0 of 10

    A new framework combines AI-derived concept embeddings with high-dimensional selective inference to enable statistically principled, interpretable discovery from unstructured data in empirical economics.

  25. Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders

    cs.CL 2025-07 conditional novelty 6.0 of 10

    MTMs replace word counts with SAE features, producing more abstract topics and enabling topic steering, and outperform word-based baselines in LLM and human pairwise evaluations.

  26. Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    A new pipeline uses interpretability to characterize concepts in preference data and shape rewards via feature or data interventions during LM post-training.

  27. SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    SAEExplainer applies activation-guided preference optimization in two iterative rounds to improve explanations of SAE features and reduce hallucinations.

  28. Features have life history. And we should care

    q-bio.NC 2026-05 unverdicted novelty 5.0 of 10

    Language model features form an early stable carrier scaffold of about 50 sparse features that is load-bearing, predictable from onset firing, and recruits most later features.

  29. Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

    cs.LG 2025-09 unverdicted novelty 5.0 of 10

    Safe-SAIL supplies a pre-explanation metric and segment-level simulation to interpret 1758 safety SAE features across pornography, politics, violence, and terror, with public models and tools released.

  30. Distribution-Aware Feature Selection for SAEs

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Sampled-SAE pre-selects a candidate pool of features using batch-level norms or entropy before batch top-K, creating a tunable family that trades reconstruction fidelity for improved probing and reduced absorption on ...

  31. CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs

    cs.LG 2026-03 conditional novelty 4.0 of 10

    CLT-Forge packages feature-sharded CLT training, compressed activation caching, automated interpretability, Circuit-Tracer attribution graphs, and a Dash visualizer into a single open-source library.

Pith tools