REVIEW 14 cited by
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Identifying the features learned by neural networks is a core challenge in mechanistic interpretability. Sparse autoencoders (SAEs), which learn a sparse, overcomplete dictionary that reconstructs a network's internal activations, have been used to identify these features. However, SAEs may learn more about the structure of the datatset than the computational structure of the network. There is therefore only indirect reason to believe that the directions found in these dictionaries are functionally important to the network. We propose end-to-end (e2e) sparse dictionary learning, a method for training SAEs that ensures the features learned are functionally important by minimizing the KL divergence between the output distributions of the original model and the model with SAE activations inserted. Compared to standard SAEs, e2e SAEs offer a Pareto improvement: They explain more network performance, require fewer total features, and require fewer simultaneously active features per datapoint, all with no cost to interpretability. We explore geometric and qualitative differences between e2e SAE features and standard SAE features. E2e dictionary learning brings us closer to methods that can explain network behavior concisely and accurately. We release our library for training e2e SAEs and reproducing our analysis at https://github.com/ApolloResearch/e2e_sae
Forward citations
Cited by 14 Pith papers
-
Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation
In ablation-based SAE evaluation, the measurement token is selected by the dictionary under test, and holding that token fixed collapses most of the variance that is usually attributed to differences between dictionaries.
-
Sparse Autoencoders Trained on the Same Data Learn Different Features
Seed variation alone causes sparse autoencoders to learn different, often equally interpretable feature sets, so SAE features are pragmatic decompositions rather than ground-truth units.
-
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
Attribution-based Parameter Decomposition splits a network's parameters into faithful, minimal, and simple components and recovers ground-truth mechanisms in toy models of superposition and compressed computation.
-
Training, Reading, and Editing Legible Transformers
A variance-floor objective plus learned operator gates produce an end-to-end legible transformer whose crisp units are 50–184× more local to edit and can be reshaped from fan-out to fan-in circuits without quality loss.
-
Scaling sparse feature circuit finding for in-context learning
Task vectors for in-context learning decompose into sparse SAE features that detect and execute tasks, linked by an attention/MLP circuit in Gemma-1 2B.
-
Low-Rank Adapting Models for Sparse Autoencoders
LoRA fine-tuning of the language model around a fixed SAE reduces the SAE-insertion loss gap by 30-55% and matches end-to-end SAEs 2-20x faster.
-
Can Input Attributions Explain Inductive Reasoning in In-Context Learning?
Using synthetic inductive reasoning tasks with a single 'aha' example, the paper shows simple gradient-norm attribution often beats integrated gradients for identifying the crucial example, while interpretability wors...
-
Obfuscated Activations Bypass LLM Latent-Space Defenses
Obfuscation attacks that jointly optimize for target behavior and for low monitor scores bypass sparse autoencoders, probes, and OOD detectors on LLMs, while performance degrades mainly on hard tasks like writing correct SQL.
-
Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning
Gradient-weighted TopK selection in sparse autoencoders improves downstream loss fidelity and steering efficacy of learned features without reducing human-rated interpretability.
-
Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders
Procrustes-conditioned joint End-to-end Top-K SAEs recover more cross-seed universal features (r≥0.70) from independent BERT seeds than post-hoc alignment baselines on three NLP datasets.
-
Stable and Steerable Sparse Autoencoders with Weight Regularization
L2 weight regularization in TopK SAEs increases cross-seed feature overlap and roughly doubles measured steering success on Pythia-70M, at the cost of collapsing most latents to zero.
-
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
The paper introduces an Explanatory Virtues Framework and argues, via a qualitative rubric, that Compact Proofs are the most promising method for mechanistic interpretability.
-
Discovering Chunks in Neural Embeddings for Interpretability
Recurring 'chunks' in neural embeddings can be extracted, predict input patterns, and be perturbed to steer a model's outputs.
-
Empirical Evaluation of Progressive Coding for Sparse Autoencoders
Nested 'Matryoshka' sparse autoencoders outperform pruned vanilla sparse autoencoders on reconstruction and recaptured language-model loss, but pruned vanilla features remain more interpretable.
Discussion (0). Continue with ORCID to comment.