REVIEW 18 cited by
Towards Automated Circuit Discovery for Mechanistic Interpretability
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Through considerable effort and intuition, several recent works have reverse-engineered nontrivial behaviors of transformer models. This paper systematizes the mechanistic interpretability process they followed. First, researchers choose a metric and dataset that elicit the desired model behavior. Then, they apply activation patching to find which abstract neural network units are involved in the behavior. By varying the dataset, metric, and units under investigation, researchers can understand the functionality of each component. We automate one of the process' steps: to identify the circuit that implements the specified behavior in the model's computational graph. We propose several algorithms and reproduce previous interpretability results to validate them. For example, the ACDC algorithm rediscovered 5/5 of the component types in a circuit in GPT-2 Small that computes the Greater-Than operation. ACDC selected 68 of the 32,000 edges in GPT-2 Small, all of which were manually found by previous work. Our code is available at https://github.com/ArthurConmy/Automatic-Circuit-Discovery.
Forward citations
Cited by 18 Pith papers
-
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
LLMs' source-attribution ability is not fixed: it flips with conversational memory structure, and corrective feedback can invert judgments or sever confidence from accuracy.
-
LAWFUL: Law-Aligned Witness for Faithful Use of Latents
LAWFUL defines coverage-aware physical-consistency scores and circuit tests, reporting that a MoCap-to-Radar transformer's 9-component temporal circuit carries Doppler-law consistency via attention patterns.
-
IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning
IFCLoRA allocates LoRA ranks before fine-tuning using information-flow centrality from a calibration-set interaction graph, improving GSM8K by 1.36–1.82 points over LoRA at low ranks.
-
Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures
Denoising diffusion models encode visual illusions in internal layers, yet these representations do not influence the generated image.
-
Targeted Recovery of Weight-Space Mechanisms From Neural Networks
A targeted decomposition method recovers the weight-space mechanisms behind specific inputs at low FLOPs, enabling focused ablation and rewiring of a 12-block transformer.
-
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
Removing a single fitted activation direction at a mid-depth layer reduces the evaluation-vs-deployment behavioral gap on held-out prompts in 10 of 12 fine-tuned LLM settings, with matched controls staying flat.
-
Language Model Circuits Are Sparse in the Neuron Basis
MLP neuron activations are shown to be as sparse and faithful a basis for circuit tracing as sparse autoencoder features, enabling simplified interpretability pipelines.
-
Circuit Stability Characterizes Language Model Generalization
Circuit stability, measured as rank correlation between soft circuits across subtasks, is proposed as a predictor of language model generalization.
-
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
A two-layer transformer solving an in-context meta-learning task acquires skill in three abrupt phases, each corresponding to a distinct attention circuit: bigram, label attention, then chunking plus label attention.
-
Transcoders Beat Sparse Autoencoders for Interpretability
Skip transcoders beat sparse autoencoders on both reconstruction fidelity and automated interpretability scores for transformer MLP layers.
-
MechELK: A Mechanistic Interpretability Framework for Eliciting Latent Knowledge in Large Language Models
MechELK combines SAE features, causal patching, and representation engineering to elicit latent LLM knowledge at 84.7% average accuracy, beating CCS by 6.2%.
-
Enhancing Multi-Robot Exploration Using Probabilistic Frontier Prioritization with Dirichlet Process Gaussian Mixtures
DP-GMM-based probabilistic frontier prioritization improves two multi-agent frontier explorers by roughly 10–14% across clutter, team size, and communication settings.
-
From Features to Actions: Explainability in Traditional and Agentic AI Systems
Attribution explanations that work for static classifiers do not diagnose failures in multi-step AI agents; trace-grounded rubric evaluation does, with state-tracking inconsistency 2.7x more common in failed agent runs.
-
From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits
GPT-2 small performs syllogisms through truth-copying attention heads and a suppression-plus-MLP pathway that can output a negated truth value.
-
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content
Zeroing attack-vulnerable attention heads improves BERT/RoBERTa toxicity classifier accuracy on PGD-adversarial inputs, with distinct heads implicated per demographic group.
-
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
A perspective paper that advocates direct, post hoc interpretability for multi-agent deep reinforcement learning and offers a taxonomy of where those methods might apply.
-
AI Governance through Markets
Market governance mechanisms, supported by standardized AI disclosures, can create financial incentives for responsible AI development, according to this policy paper.
-
Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning
A scoping review surveying circuit analysis, sparse autoencoders, activation steering, and neurosymbolic frameworks for interpreting and controlling Transformer-based neural networks.
Discussion (0). Continue with ORCID to comment.