Pith. sign in

REVIEW 12 cited by

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.04709 v4 pith:HR6OHQBW submitted 2023-01-11 cs.AI

classification cs.AI
keywords causalabstractioninterpretabilitymechanisticanalysisfoundationmechanismmechanisms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models. Our contributions are (1) generalizing the theory of causal abstraction from mechanism replacement (i.e., hard and soft interventions) to arbitrary mechanism transformation (i.e., functionals from old mechanisms to new mechanisms), (2) providing a flexible, yet precise formalization for the core concepts of polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and (3) unifying a variety of mechanistic interpretability methods in the common language of causal abstraction, namely, activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Safety from Honesty in a Disinterested AI Predictor

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.

  2. Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Structural probes on UD-invariant wh-movement stimuli reveal phase-count gradients and phase-internal cohesion effects in 12-13 of 13 LLMs, indicating syntactic abstractions beyond UD annotations.

  3. Temporal Preference Concepts and their Functions in a Large Language Model

    cs.LG 2026-05 unverdicted novelty 6.5 of 10

    Temporal preference in Qwen3-4B-Instruct-2507 localizes to layers 17–35 (especially L24 attention), has curved residual-stream geometry, is behaviorally unstable, and can be bidirectionally steered.

  4. Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A single-token feature's causal necessity under zero-ablation depends on which SAE family found it: GemmaScope and BatchTopK features stay causally anchored while LlamaScope features are locally redundant.

  5. Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies

    stat.ML 2025-07 conditional novelty 6.0 of 10

    A nonlinear causal model reduction framework learns interpretable high-level causes of reward in RL policies, with uniqueness guarantees for a class of additive noise models.

  6. Can Interpretation Predict Behavior on Unseen Data?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Presence of hierarchical attention heads on in-distribution data predicts hierarchical out-of-distribution generalization across 270 small transformers, independent of causal support.

  7. How Do Transformers Learn Variable Binding in Symbolic Programs?

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A Transformer trained from scratch on symbolic variable-assignment programs develops a systematic dereferencing mechanism through three phases, building on early line-based heuristics rather than replacing them.

  8. Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A position paper proposing feature consistency, measured by PW-MCC, as a core SAE evaluation criterion, with evidence that TopK SAEs achieve high consistency on LLM activations.

  9. Explaining Neural Networks with Reasons

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A new interpretability method computes 'reasons vectors' from neuron activations and measures how strongly each neuron supports propositions about the input, with experiments on MNIST, Adult, and SST2.

  10. Identifiability in Causal Abstractions: A Hierarchy of Criteria

    cs.AI 2025-07 conditional novelty 5.0 of 10

    The paper formalizes distinct notions of identifiability over collections of causal diagrams and proves a hierarchy among them, leaving an open conjecture on the gap between two central notions.

  11. Identifying a Circuit for Verb Conjugation in GPT-2

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A candidate 12-head circuit for subject-verb agreement in GPT-2 Small is identified, but it generalizes poorly to more complex agreement conditions.

  12. Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

    cs.AI 2025-05 reject novelty 5.0 of 10

    SafetyNet is an ensemble of standard outlier detectors for LLM backdoor monitoring, but its key mechanistic claim and headline numbers are contradicted by inconsistent tables and a mismatched abstract.

Pith tools