Individual transformer dimensions encode concepts via signs alone, enabling training-free detection and steering without learned feature dictionaries.
URL https://distill
10 Pith papers cite this work, alongside 823 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Toy models demonstrate that polysemanticity arises when neural networks store more sparse features than neurons via superposition, producing a phase transition tied to polytope geometry and increased adversarial vulnerability.
The paper introduces compositional interpretability as a category-theoretic framework that casts mechanistic explanations as commuting syntactic-semantic mappings optimized under faithfulness and complexity constraints derived from minimum description length.
Sparse autoencoders resolve superposition in image-based neuron representations, recovering geometric fidelity and enabling scRNA-seq adaptation plus GW-map alignment to reconstruct pathology pathways without spatial transcriptomics.
JSAE jointly factorizes pooled vision and language activations in VLMs into aligned interpretable features, revealing layer-dependent asymmetry in additive steering versus suppression on three models.
Feature visualization on TRIBE v2 brain encoders recovers the known ventral visual hierarchy from V1 to V4 and produces distinctive patterns for MT, FFA, and PPA, with optimized stimuli driving ~4x higher activation than natural images.
HOLE applies persistent homology to latent embeddings in neural networks and uses visualizations such as cluster flow diagrams to reveal patterns of class separation, feature disentanglement, and robustness.
Reasoning in large output spaces proceeds via shortlisting then fine-grained reasoning; this characterization enables a mechanistic distillation strategy that outperforms standard distillation.
NeuroViz offers interactive real-time visualization of neural network forward and backward passes, achieving top usability scores in a study with 31 participants compared to existing tools.
A review paper that organizes conceptual, practical, and socio-technical open problems in mechanistic interpretability.
citing papers explorer
-
The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions
Individual transformer dimensions encode concepts via signs alone, enabling training-free detection and steering without learned feature dictionaries.
-
Toy Models of Superposition
Toy models demonstrate that polysemanticity arises when neural networks store more sparse features than neurons via superposition, producing a phase transition tied to polytope geometry and increased adversarial vulnerability.
-
From Mechanistic to Compositional Interpretability
The paper introduces compositional interpretability as a category-theoretic framework that casts mechanistic explanations as commuting syntactic-semantic mappings optimized under faithfulness and complexity constraints derived from minimum description length.
-
Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images
Sparse autoencoders resolve superposition in image-based neuron representations, recovering geometric fidelity and enabling scRNA-seq adaptation plus GW-map alignment to reconstruct pathology pathways without spatial transcriptomics.
-
Steering Vision-Language Models with Joint Sparse Autoencoders
JSAE jointly factorizes pooled vision and language activations in VLMs into aligned interpretable features, revealing layer-dependent asymmetry in additive steering versus suppression on three models.
-
Feature Visualization Recovers Known Cortical Selectivity from TRIBE v2
Feature visualization on TRIBE v2 brain encoders recovers the known ventral visual hierarchy from V1 to V4 and produces distinctive patterns for MT, FFA, and PPA, with optimized stimuli driving ~4x higher activation than natural images.
-
HOLE: Homological Observation of Latent Embeddings for Neural Network Interpretability
HOLE applies persistent homology to latent embeddings in neural networks and uses visualizations such as cluster flow diagrams to reveal patterns of class separation, feature disentanglement, and robustness.
-
Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces
Reasoning in large output spaces proceeds via shortlisting then fine-grained reasoning; this characterization enables a mechanistic distillation strategy that outperforms standard distillation.
-
NeuroViz: Real-time Interactive Visualization of Forward and Backward Passes in Neural Network Training
NeuroViz offers interactive real-time visualization of neural network forward and backward passes, achieving top usability scores in a study with 31 participants compared to existing tools.
-
Open Problems in Mechanistic Interpretability
A review paper that organizes conceptual, practical, and socio-technical open problems in mechanistic interpretability.