Pith. sign in

REVIEW 5 cited by

Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17030 v2 pith:JT33FSWV submitted 2023-11-28 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords subspaceinterpretabilitymodelactivationfactpatchingsubspacesaims
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activations. Specifically, recent studies have explored subspace interventions (such as activation patching) as a way to simultaneously manipulate model behavior and attribute the features behind it to given subspaces. In this work, we demonstrate that these two aims diverge, potentially leading to an illusory sense of interpretability. Counterintuitively, even if a subspace intervention makes the model's output behave as if the value of a feature was changed, this effect may be achieved by activating a dormant parallel pathway leveraging another subspace that is causally disconnected from model outputs. We demonstrate this phenomenon in a distilled mathematical example, in two real-world domains (the indirect object identification task and factual recall), and present evidence for its prevalence in practice. In the context of factual recall, we further show a link to rank-1 fact editing, providing a mechanistic explanation for previous work observing an inconsistency between fact editing performance and fact localization. However, this does not imply that activation patching of subspaces is intrinsically unfit for interpretability. To contextualize our findings, we also show what a success case looks like in a task (indirect object identification) where prior manual circuit analysis informs an understanding of the location of a feature. We explore the additional evidence needed to argue that a patched subspace is faithful.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LAWFUL: Law-Aligned Witness for Faithful Use of Latents

    cs.LG 2026-07 conditional novelty 7.0 of 10

    LAWFUL defines coverage-aware physical-consistency scores and circuit tests, reporting that a MoCap-to-Radar transformer's 9-component temporal circuit carries Doppler-law consistency via attention patterns.

  2. Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

    q-bio.GN 2026-07 conditional novelty 6.0 of 10

    Sparse dictionaries extracted from two genomic language models contain features whose ablation shifts masked-token predictions specifically at ChIP-seq bound versus unbound sites, after GC-composition controls.

  3. Transformers converge to invariant algorithmic cores

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Trained transformers contain low-dimensional causal subspaces — algorithmic cores — that recur across runs and scales and can be extracted, characterized, and steered.

  4. Can Interpretation Predict Behavior on Unseen Data?

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Presence of hierarchical attention heads on in-distribution data predicts hierarchical out-of-distribution generalization across 270 small transformers, independent of causal support.

  5. Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A position paper proposing feature consistency, measured by PW-MCC, as a core SAE evaluation criterion, with evidence that TopK SAEs achieve high consistency on LLM activations.

Pith tools