Pith. sign in

REVIEW 5 cited by

Mechanistic?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.09087 v1 pith:RJQFVFW7 submitted 2024-10-07 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords interpretabilitymechanisticcommunitydefinitionculturaltermhowevermodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models. However, this jargon has also led to a fair amount of confusion. So, what does it mean to be "mechanistic"? We describe four uses of the term in interpretability research. The most narrow technical definition requires a claim of causality, while a broader technical definition allows for any exploration of a model's internals. However, the term also has a narrow cultural definition describing a cultural movement. To understand this semantic drift, we present a history of the NLP interpretability community and the formation of the separate, parallel "mechanistic" interpretability community. Finally, we discuss the broad cultural definition -- encompassing the entire field of interpretability -- and why the traditional NLP interpretability community has come to embrace it. We argue that the polysemy of "mechanistic" is the product of a critical divide within the interpretability community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Propositional Interpretability in Artificial Intelligence

    cs.AI 2025-01 conditional novelty 7.0 of 10

    Chalmers proposes propositional interpretability, interpreting AI in terms of beliefs, desires, and credences, and sets the challenge of thought logging all such attitudes over time.

  2. Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An unsupervised method, EDDP, jointly learns encoding-decoding direction pairs for concepts in CNN latent spaces, recovering interpretable and influential concepts without labels, validated on synthetic and real data.

  3. Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Circuits extracted from inflectionally varied Polish sentences are more robust to adversarial word attacks than circuits from syncretic or English variants, identifying layer-0 attention heads as inflection-specific.

  4. Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A position paper unifying feature, data, and component attribution under three shared techniques, perturbation, gradient, and linear approximation, and proposing cross-attribution research directions.

  5. Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The paper introduces an Explanatory Virtues Framework and argues, via a qualitative rubric, that Compact Proofs are the most promising method for mechanistic interpretability.

Pith tools