REVIEW 2 cited by
Transformer Circuit Faithfulness Metrics are not Robust
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Mechanistic interpretability work attempts to reverse engineer the learned algorithms present inside neural networks. One focus of this work has been to discover 'circuits' -- subgraphs of the full model that explain behaviour on specific tasks. But how do we measure the performance of such circuits? Prior work has attempted to measure circuit 'faithfulness' -- the degree to which the circuit replicates the performance of the full model. In this work, we survey many considerations for designing experiments that measure circuit faithfulness by ablating portions of the model's computation. Concerningly, we find existing methods are highly sensitive to seemingly insignificant changes in the ablation methodology. We conclude that existing circuit faithfulness scores reflect both the methodological choices of researchers as well as the actual components of the circuit - the task a circuit is required to perform depends on the ablation used to test it. The ultimate goal of mechanistic interpretability work is to understand neural networks, so we emphasize the need for more clarity in the precise claims being made about circuits. We open source a library at https://github.com/UFO-101/auto-circuit that includes highly efficient implementations of a wide range of ablation methodologies and circuit discovery algorithms.
Forward citations
Cited by 2 Pith papers
-
Certified Circuits: Stability Guarantees for Mechanistic Circuits
Certified Circuits uses deletion-based randomized smoothing to guarantee that circuit components stay included or excluded under bounded edits to the concept dataset, yielding more compact and more accurate circuits.
-
Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning
SICAF traces per-token self-influence inside extracted circuits to map GPT-2's reasoning on the IOI task.
Discussion (0). Continue with ORCID to comment.