Pith. sign in

Mechanistic Interpretability Needs Philosophy

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it
abstract

Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important to examine not just models themselves, but the assumptions, concepts and explanatory strategies implicit in MI research. We argue that mechanistic interpretability needs philosophy as an ongoing partner in clarifying its concepts, refining its methods, and navigating the epistemic and ethical complexities of interpreting AI systems. There is significant unrealised potential for progress in MI to be gained through deeper engagement with philosophers and philosophical frameworks. Taking three open problems from the MI literature as examples, this paper illustrates the value philosophy can add to MI research, and outlines a path toward deeper interdisciplinary dialogue.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1 cs.CV 1

years

2026 2

verdicts

UNVERDICTED 2

roles

background 1

polarities

background 1

representative citing papers

Vision-Language Asymmetry in Bistable Image Captioning

cs.CV · 2026-06-06 · unverdicted · novelty 6.0

Behavioral tests and SAE probing on 83 bistable images show simultaneous vision-tower activation of both aspects in 72% of cases, with causal steering succeeding on default-dominant but not force-balanced stimuli, locating the commitment bottleneck downstream of the vision tower.

citing papers explorer

Showing 2 of 2 citing papers.

  • Vision-Language Asymmetry in Bistable Image Captioning cs.CV · 2026-06-06 · unverdicted · none · ref 15 · internal anchor

    Behavioral tests and SAE probing on 83 bistable images show simultaneous vision-tower activation of both aspects in 72% of cases, with causal steering succeeding on default-dominant but not force-balanced stimuli, locating the commitment bottleneck downstream of the vision tower.

  • Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space cs.CL · 2026-04-06 · unverdicted · none · ref 74 · internal anchor

    PAM, a complex-valued associative memory model, exhibits steeper power-law scaling in loss and perplexity than a matched real-valued baseline when trained on WikiText-103 from 5M to 100M parameters.