Behavioral tests and SAE probing on 83 bistable images show simultaneous vision-tower activation of both aspects in 72% of cases, with causal steering succeeding on default-dominant but not force-balanced stimuli, locating the commitment bottleneck downstream of the vision tower.
Mechanistic Interpretability Needs Philosophy
2 Pith papers cite this work. Polarity classification is still indexing.
abstract
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important to examine not just models themselves, but the assumptions, concepts and explanatory strategies implicit in MI research. We argue that mechanistic interpretability needs philosophy as an ongoing partner in clarifying its concepts, refining its methods, and navigating the epistemic and ethical complexities of interpreting AI systems. There is significant unrealised potential for progress in MI to be gained through deeper engagement with philosophers and philosophical frameworks. Taking three open problems from the MI literature as examples, this paper illustrates the value philosophy can add to MI research, and outlines a path toward deeper interdisciplinary dialogue.
citation-role summary
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
PAM, a complex-valued associative memory model, exhibits steeper power-law scaling in loss and perplexity than a matched real-valued baseline when trained on WikiText-103 from 5M to 100M parameters.
citing papers explorer
-
Vision-Language Asymmetry in Bistable Image Captioning
Behavioral tests and SAE probing on 83 bistable images show simultaneous vision-tower activation of both aspects in 72% of cases, with causal steering succeeding on default-dominant but not force-balanced stimuli, locating the commitment bottleneck downstream of the vision tower.
-
Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space
PAM, a complex-valued associative memory model, exhibits steeper power-law scaling in loss and perplexity than a matched real-valued baseline when trained on WikiText-103 from 5M to 100M parameters.