REVIEW 3 major objections 4 minor 15 references
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A trained neural network contains an implicit explanation of its own behavior, so mechanistic interpretability can be a principled science with a well-defined target: explanatory faithfulness.
desk verdict Programmatic philosophy of MI with a useful demarcation and an honest conjecture, but the well-definedness claim rests on an unproven uniqueness assertion. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing concept is the ur-explanation: the idealised explanation of a model's behavior over a data distribution, given in terms of the model's learned internal structures and computations over representations. It carries the argument by supplying a target for interpretability; with it, explanatory faithfulness can be defined as layerwise activation matching, and without it, claims about circuit equivalence reduce to behavioral statistics. A second mechanism is a two-step No Miracles Argument: first, a model's extraordinary predictive success justifies realism about its internal features; second, an interpretation that successfully predicts activations at each layer under causal interventions is likely to be approximately faithful to the unique ur-explanation. Representations themselves are characterised by three criteria: information, use, and the possibility of misrepresentation.
What would settle it
For one concrete check, take a trained, generalising model and a fixed data distribution, then search over candidate explanations of increasing complexity for the best layerwise activation match; if the minimum mismatch stays bounded away from zero even as explanation complexity grows, explanatory faithfulness as defined has no attainable target.
Extended reading notes
Core claim
The paper's central claim is that a trained, generalising neural network already contains an implicit explanation of its own behavior, called the ur-explanation, as computations over learned representations, so understanding is not imposed on the model from outside but extracted from it. Explanatory faithfulness is then defined operationally: an explanation E matches model M over data distribution D to the degree that the intermediate activations predicted by E at each layer match the model's actual intermediate activations. The paper claims this definition is only possible once ur-explanations are accepted, and it uses the No Miracles Argument, transferred to neural networks, to support the existence and uniqueness of ur-explanations. It further claims that mechanistic interpretability is demarcated as the study of Model-level, Ontic, Causal-Mechanistic, and Falsifiable explanations and that Explanatory Optimism, the conjecture that important machine concepts are human-understandable, is a necessary precondition for the project to succeed.
Load-bearing premise
The load-bearing premise is that the No Miracles Argument transfers from natural science to neural networks, so a generalising model's predictive success justifies believing that its internal representations track real entities and that its ur-explanation exists and is unique.
Editorial extensions
If this is right
- Mechanistic interpretability explanations can be objectively evaluated as faithful or unfaithful by comparing layerwise activations, not just matching input-output behavior.
- The field gains a demarcation criterion: only explanations that are Model-level, Ontic, Causal-Mechanistic, and Falsifiable count as mechanistic interpretability.
- If Explanatory Optimism is false, mechanistic interpretability cannot achieve its stated goal of understanding generalising neural networks, so the conjecture's truth value is load-bearing for the research agenda.
- Models can be trusted for content-based reasons, because extracting an explanation faithful to the model's internal mechanisms provides reasons for a prediction rather than just a track record.
- The success of neural networks as predictors becomes evidence that their internal representations correspond to real structure in the world, making interpretability a route to scientific discovery.
Reading between the lines
- A natural extension is to turn explanatory faithfulness into a quantitative benchmark: average layerwise activation divergence between model and explanation over a held-out distribution; the paper gestures at this through hypothesis testing but does not propose a specific metric.
- The uniqueness claim for ur-explanations could be sharpened as a testable prediction: independently trained networks that solve the same task and generalise equally well should converge on the same high-level causal structure, and robust divergence across such networks would put pressure on uniqueness.
- Explanatory Optimism could be probed empirically by measuring whether increasing model scale expands the residue of concepts that resist translation into human-comprehensible abstractions; if that residue grows with capability, weak optimism would fail.
- If explanatory faithfulness is accepted as the goal, then explanation quality should be evaluated by intervention-based activation prediction rather than by human ratings, which would shift interpretability evaluation toward causal and mechanistic criteria.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues for an Explanatory View Hypothesis: trained neural networks contain implicit explanations of their own behaviour (ur-explanations), and therefore mechanistic interpretability (MI) has a well-defined target. Explanatory Faithfulness is defined as the degree to which an algorithmic explanation's intermediate activations match the model's intermediate activations. The paper proposes a demarcation of MI as the production of Model-level, Ontic, Causal-Mechanistic, and Falsifiable explanations, discusses limits of MI arising from value-ladenness, theory-ladenness, model-level versus system-level analysis, and low-abstraction explanations, and formulates the Principle of Explanatory Optimism as a conjecture about human-understandability of machine concepts. Appendix B attempts to ground the existence and uniqueness of ur-explanations via the No Miracles Argument.
Significance. If the central claim were established, the paper would give MI a principled target and would usefully separate explanatory faithfulness from behavioural faithfulness, which is a substantive contribution. The paper also provides a clear and useful demarcation of MI, and it explicitly formulates Explanatory Optimism as a testable conjecture, which is valuable for focusing future work. The paper is a serious synthesis of philosophical and technical themes, and it cites relevant literature, including works that challenge its own assumptions. However, the central claim that Explanatory Faithfulness is well-defined is not currently supported: the uniqueness of ur-explanations is asserted rather than derived, and the operational activation-matching criterion is basis-dependent.
major comments (3)
- [Appendix B.3 and Section 3.2.2] The uniqueness of the ur-explanation is load-bearing for the claim that Explanatory Faithfulness is well-defined, but it is asserted rather than established. Appendix B.3 concludes that 'ur-explanations are to be unique', yet no derivation is given; the No Miracles Argument, even if it transferred successfully to neural networks, would at most support the approximate truth of some representation or internal structure, not the uniqueness of a privileged decomposition. Appendix F provides a forward trace, which gives existence of at least one implementation-level explanation, but existence does not imply uniqueness. Without uniqueness, the definition of explanatory faithfulness as matching 'the' ur-explanation is not well-defined.
- [Section 3.2.2] The operational criterion for Explanatory Faithfulness is basis-dependent, so it fails to define a unique target even if the No Miracles Argument is granted. The definition compares intermediate activations s_i from the explanation E with the model's activations x_i. However, for any invertible matrix U, reparameterizing the layer-i weights as W_i' = U W_i and the downstream weights accordingly leaves the input-output map unchanged while changing the activations x_i to U x_i. Two behaviorally identical parameterizations of the same model would therefore assign different faithfulness scores to the same explanation E. The paper does not supply a canonical decomposition, and its own discussion of theory-ladenness in Section 5.1.2, including Dmitry's Koan, reinforces that no scale-invariant or interpreter-independent decomposition is provided. This is a failure of well-definedness internal to the formalism, independent of the status of the No Miracles Argument.
- [Section 7.2 and Contributions] The paper claims that the Principle of Explanatory Optimism is a necessary precondition for the success of MI, but this necessity is not established. Section 4 defines MI without any reference to human understanding, and Section 7.1 argues only that 'it is not clear how MI research can proceed' if alien concepts are prevalent, which is weaker than the claim that EO is necessary. The contribution statement says 'We show that without the Principle of Explanatory Optimism the project of MI is intractable', but the body of the paper presents EO as a conjecture and explicitly says no arguments for its truth are provided. The necessity claim needs either a direct argument or a reframing as a working assumption rather than a demonstrated precondition.
minor comments (4)
- [Appendix F] There is a typo in Appendix F: 'Implementation leval' should be 'Implementation level'.
- [References] The reference list contains a typo: 'Thomasn D. Aquinas' should be 'Thomas D. Aquinas'.
- [Section 3.2.1] The relation between 'learned internal structures' in the definition of ur-explanations and 'intermediate activations' in the operational definition of explanatory faithfulness could be stated more explicitly; as written, the two notions may not coincide, and the paper does not clarify whether the ur-explanation is the full computational trace or a compressed abstraction over it.
- [Section 7.1] The naming of 'Dmitry's Koan++' as 'Nora's Koan' is informal; if retained, a brief justification would help readers unfamiliar with the provenance.
Circularity Check
The claimed NMA2 result that activation-matching explanations correspond to the ur-explanation restates the definition of explanatory faithfulness; uniqueness of the ur-explanation is asserted, not derived.
-
self definitional
[Section 3.2.1–3.2.2; Appendix B.3 (NMA2)]
"These network computations, outputs, and intermediate activations together constitute not only a prediction of some answer, but also an explanation of the process by which the model came to such a result. ... An explanation E is explanatorily faithful to the model to the extent that it matches the model’s ur-explanation. ... if (1) the explanation E highly successfully predicts the relevant information in the neural network activations at each layer of the network ... then we can conclude (C) we have reason to believe that E in fact does approximately correspond to the ur-explanation U."
The §3.2.2 definition makes 'explanatorily faithful' mean matching the ur-explanation and immediately operationalizes it as matching the model's per-layer intermediate activations. Since §3.2.1 already identifies the ur-explanation with the network computations, outputs, and intermediate activations, B.3's NMA2 inference—activation-predicting E corresponds to U—is just the definition of explanatory faithfulness in new clothing. No NMA-style probabilistic step is needed; the conclusion is true by construction. Thus the alleged 'result' is the input definition.
-
self definitional
[Appendix B.3 (NMA2 summary)]
"In summary, through theNMA2 argument, we believe that the following two conclusions follow: 1. We have reason to believe that ML models that are extraordinarily successful at making accurate novel predictions contain explanatory knowledge. That is, their ur-explanations are well-defined and we ought to be realists about features, as variables of computation. Further, ur-explanations are to be unique."
The uniqueness of U is the load-bearing condition for 'Explanatory Faithfulness is well-defined' (Section 3.2.2 and abstract). But the NMA2 premises already presume a single U: they speak of 'the ur-explanation U' and define faithfulness as matching U. The conclusion that ur-explanations 'are to be unique' is not derived from any premise about predictive success; it is the definite article and the co-definition restated. So the paper's own conclusion 1 imports uniqueness from the definition rather than establishing it.
full rationale
The paper has independent philosophical content: the four-part demarcation of MI (Model-level, Ontic, Causal-Mechanistic, Falsifiable), the treatment of theory-ladenness, and Explanatory Optimism are framed as conjectures and not derived from the circular step. But the central claim—that Explanatory Faithfulness is well-defined—is not independently grounded. Its operational criterion compares intermediate activations, and its ur-explanation is characterized by those very computations and activations, so the NMA2 conclusion reduces to the definition. The uniqueness assertion in Appendix B.3 is a further definitional placeholder, not a proof. Separately (a correctness concern, not itself a circularity), the activation-matching criterion is basis-dependent under invertible reparameterizations, so even the intended definition requires a choice of scale/precision; the paper's own Section 5.1 Dmitry's Koan++ concedes this. Because the target 'prediction' reduces by construction while the surrounding taxonomy remains substantive, the score is 6.
Assumptions & free parameters
assumptions (4)
- domain assumption The No Miracles Argument is valid and transfers from science to neural networks: predictive success licenses realism about internal representations.
- ad hoc to paper The ur-explanation of a model exists and is unique.
- domain assumption Features are real, discoverable entities (representational realism about internal activations).
- domain assumption The goal of MI includes human understanding of model behavior.
invented entities (4)
-
ur-explanation
-
Alien Concepts (CA)
-
explanatory complexity class
-
features (as real, non-neuron entities of computation)
Cite this review
Pith. "Pith review of A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i." pith.science (2026). https://pith.science/paper/NFLIUKE3
@misc{pith2026250500808,
author = {Pith},
title = {Pith review of: A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i},
year = {2026},
howpublished = {\url{https://pith.science/paper/NFLIUKE3}},
note = {Machine review of arXiv:2505.00808}
}
read the original abstract
Mechanistic Interpretability aims to understand neural networks through causal explanations. We argue for the Explanatory View Hypothesis: that Mechanistic Interpretability research is a principled approach to understanding models because neural networks contain implicit explanations which can be extracted and understood. We hence show that Explanatory Faithfulness, an assessment of how well an explanation fits a model, is well-defined. We propose a definition of Mechanistic Interpretability (MI) as the practice of producing Model-level, Ontic, Causal-Mechanistic, and Falsifiable explanations of neural networks, allowing us to distinguish MI from other interpretability paradigms and detail MI's inherent limits. We formulate the Principle of Explanatory Optimism, a conjecture which we argue is a necessary precondition for the success of Mechanistic Interpretability.
Figures
Reference graph
Works this paper leans on
-
[1]
Scientific theories are (extraordinarily) successful in the sense that they make accu- rate novel empirical predictions about phenomena of interest
-
[2]
If our scientific theories are very far from the truth, then it would be miraculous that they are so successful
-
[3]
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy
URL https://arxiv.org/abs/2504.04072. Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014. Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. Large language model based multi-agents: A survey of progre...
arXiv 2014
-
[4]
Note that the NMA has two distinct conclusions
Therefore, two conclusions follow: (a) Firstly, that our best scientific theories are approximately true (or approximately correctly describe mind-independent laws) (b) Secondly, that the entities posited by such scientific theories are real (or approx- imately characterize mind-independent entities). Note that the NMA has two distinct conclusions. The fi...
work page 2016
-
[6]
Literaturverzeichnis: Seite 266-270 ; Hier auch später erschienene, unveränderte Nachdrucke. Larry Laudan. The demise of the demarcation problem. In Robert S. Cohen and Larry Laudan (eds.), Physics, Philosophy and Psychoanalysis: Essays in Honor of Adolf Grünbaum, pp. 111–127. D. Reidel, 1983. Matthew L. Leavitt and Ari Morcos. Towards falsifiable interpr...
arXiv 1983
-
[7]
Forthcoming. James Myers. Cognitive styles in two cognitive sciences. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 34, 2012. Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. arXiv preprint arXiv:2301.05217, 2023. 22 Chris Olah, Nick ...
arXiv 2012
-
[8]
URL https://www.alignmentforum.org/posts/64MizJXzyvrYpeKqm/ sparsify-a-mechanistic-interpretability-research-agenda . Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adria Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbel...
arXiv 2023
-
[11]
Given the choice between a straightforward reason for the success of scientific theories and a seemingly miraculous sense in which all our theories are just co- incidentally producing accurate novel predictions, one should clearly prefer the former. 28
Show all 15 references
-
[13]
That is, their ur-explanations are well-defined and we ought to be realists about features, as variables of computation
We have reason to believe that ML models that are extraordinarily successful at making accurate novel predictions contain explanatory knowledge. That is, their ur-explanations are well-defined and we ought to be realists about features, as variables of computation. Further, ur...
-
[14]
Mechanistic
We have reason to believe that MI explanations that provide successful algorithmic predictions of how the neural activations are structured at each sequential layer of the model, are likely to be approximately explanatorily faithful to the model’s ur-explanation and contribute...
2023
-
[15]
even if humans and machines do not share all concepts, AI concepts are understandable by humans given the right explanation
and the theory of Natural Kinds (Khalidi, 2023) all argue for Universality of Learned Concepts in some form. If Universality of Learned Concepts is true, then we should expect that the problem of non-shared concepts becomes less significant for interpretability with scale.43 I...
2023
-
[1987]
ISBN 978-0-674-79291-3 and 0-674-79290-4 and 0-674-79291-2 and 978-0-674-79290-
-
[1990]
Raphaël Millière and Cameron Buckner
doi:10.1007/bf00361036. Raphaël Millière and Cameron Buckner. Interventionist methods for interpreting deep neu- ral networks. In Gualtiero Piccinini (ed.), Neurocognitive Foundations of Mind. Routledge,
-
[2024]
arXiv:2406.04093 [cs] version: 1
URL http://arxiv.org/abs/2406.04093. arXiv:2406.04093 [cs] version: 1. Atticus Geiger, Chris Potts, and Thomas Icard. Causal abstraction for faithful model interpretation. arXiv preprint arXiv:2301.04709, 2023. Thomas F. Gieryn. Cultural Boundaries of Science: Credibility on t...
2023 arXiv
-
[2025]
James Chua and Owain Evans
URL https://assets.anthropic.com/m/71876fabef0f0ed4/original/reasoning_ models_paper.pdf. James Chua and Owain Evans. Inference-time-compute: More faithful? a research note. arXiv preprint arXiv:2501.08156, 2025. Bilal Chughtai, Lawrence Chan, and Neel Nanda. A toy model of un...
2025 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.