REVIEW 4 major objections 5 minor 35 references
A Brief Summary of Explanatory Virtues
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This report argues that explanatory virtues from philosophy, psychology, and cognitive science can be organized into an expanded taxonomy and then applied to local AI explanations by treating the explanation as a theory and the model's…
desk verdict A clean taxonomic survey with a definitional drift in the XAI section that a referee should catch. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the explanatory-virtue taxonomy, an expanded version of a four-category classification from the philosophy of science (evidential, coherential, aesthetic, and diachronic virtues) to which the report adds a fifth category, coverage, for virtues about how widely a theory applies. The taxonomy carries the argument: Section 4 runs each virtue against local XAI explanations, translating it into a concrete property of an explanation of a single prediction, and records which virtues survive the translation and which do not. Coverage is the load-bearing new piece, because it is where the report places unification and where it locates the under-researched gap in XAI.
What would settle it
A user study that keeps one model, one instance, and one feature-attribution explanation fixed, then adds only an applicability statement (e.g., 70 of 80 similar cars are acceptable), and asks users to rate understanding, trust, and action would test the coverage claim; if the added statement changes none of the ratings after word count is controlled, the paper's claim that coverage virtues warrant further investigation is not supported.
Extended reading notes
Core claim
The central claim is that the virtues long used to judge scientific theories can double as criteria for judging local XAI explanations. The report gathers virtues from the literature into an expanded five-category taxonomy—evidential, coherential, coverage, aesthetic, and diachronic—and then proposes the theory-explanation mapping: the explanation of an ML model's prediction is the theory, and the model's behavior on the instance is the observation. Under this mapping, evidential accuracy, counterfactual-style explanatory depth, optimality, internal and universal consistency, conservativeness, simplicity, modesty, completeness, scope, consilience, and unification apply directly; causal adequacy and refutability drop out because a local explanation neither causes the model's outcome nor faces imaginable refuting instances; and diachronic virtues survive only as properties of the procedure that generates explanations. The report also singles out unification and the other coverage virtues as the least studied XAI dimension, supported by a user study in which unifying explanations were about as well liked as conservative explanations.
Load-bearing premise
The central mapping rests on treating a local explanation of a machine-learning prediction as structurally analogous to a scientific theory, with the model's behavior as the observations; if that analogy is invalid, the transfer of virtues such as optimality, coherence, and unification to explainable AI has no foundation.
Editorial extensions
If this is right
- Local XAI evaluation can be organized by the five-category taxonomy, with each virtue translated into a concrete property of an explanation of a single prediction.
- Several classical virtues drop out of local evaluation: causal adequacy and refutability do not apply, and diachronic virtues apply to explanation-generation procedures rather than to the explanation itself.
- Conversational maxims give a practical route to operationalizing coherence and depth, so explanations can be scored on clarity, order, quality, and quantity of information.
- Coverage virtues—especially unification—emerge as the under-researched but promising direction for XAI, with one user study showing unifying explanations are about as well liked as conservative ones.
Reading between the lines
- If the theory-explanation analogy is taken literally, explanation quality becomes a comparative, dataset-dependent property rather than a property of the explanation text alone, because coverage virtues depend on how many other instances the explanation generalizes to.
- The taxonomy makes trade-offs explicit: a simple explanation can be incomplete, and a complete one can be ad hoc, so applying the virtues to XAI means deciding how to weight competing criteria, much as theory choice does in philosophy.
- A natural next experiment is to turn each transferred virtue into a rating item and test whether the virtues predict user trust and action beyond explanation length; the report cites adjacent evidence but does not run such a study itself.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a brief literature summary of 'explanatory virtues' (EVs) drawn from philosophy, psychology, and cognitive science, organized mainly around Keas's (2018) taxonomy with an added 'coverage' category. It then attempts to link these virtues to explainable AI (XAI), focusing on local explanations, by viewing each local explanation as a theory and the ML model's behavior on an instance as observations. Section 3 provides definitions and two summary tables, and Section 4 gives an item-by-item mapping of virtues to XAI, discussing which are applicable, adaptable, or not relevant.
Significance. If the summary and the XAI mapping were sound, the paper would offer a useful interdisciplinary bridge, particularly by directing attention to coverage virtues such as unification, which are underrepresented in the XAI literature. The paper's organizational scheme, with its two tables and explicit distinction between epistemic/pragmatic and ontological dimensions, is a practical reference for researchers entering this area. However, the paper's usefulness is weakened by several internal inconsistencies: a concrete example of 'unification' in Section 4 does not satisfy the paper's own definition, one stated implication among virtues is logically false, and the analogy between local explanations and scientific theories is asserted without justification. These issues affect the central claimed contribution of linking EVs to XAI, but they are local enough to be fixable in a revision.
major comments (4)
- [Section 4] The unification example does not instantiate the definition of unification given in Section 3. Section 3 defines unification as 'a theory T explains more kinds of facts than rival explanations with the same amount of theoretical content,' and the text glosses 'theoretical content' as concepts or rules mentioned in an explanation. In Section 4, however, the example adds the italicized sentence 'Indeed, 70 of 80 cars with four doors and a large boot are deemed acceptable by the model' to the initial attribution. That sentence is a statistical generalization, i.e., an additional rule, not a feature value. Therefore E2 has more theoretical content than E1, and under Section 3's definition it is not more unifying. If the author intended 'theoretical content' to mean only feature values, then the added sentence constitutes empirical evidence, which would make it an evidential rather than a coverage virtue. Either way, the example conflates two different notions. The section should be revised so that the example genuinely keeps theoretical content fixed while expanding the explained observations, or the definition of unification must be explicitly revised and justified.
- [Section 3] The claim that 'Since Modesty demands a subset relation between the requirements of rival explanations, it implies Simplicity' is not correct under the definitions given. Modesty is defined as logical weakness: theory T is more modest than another if T is implied by the other. This does not require that T and the rival explain the same facts; T may explain strictly fewer facts. Simplicity, by contrast, requires that T 'explains the same facts as rival explanations, but with fewer theoretical requirements.' A subset relation between theoretical requirements does not by itself guarantee identical explanatory scope. For example, a theory with requirements {A} is implied by a rival with requirements {A, B}, but the former explains only facts about A while the latter also explains facts about B. Thus Modesty does not imply Simplicity without an additional assumption. The implication should be removed or the definitions amended to make the relationship valid.
- [Section 3] The entailment relations among Completeness, Unification, and Consilience are stated without proof and do not obviously follow from the definitions. The paper says 'Completeness is entailed by Unification' and 'Since Unification poses no requirements regarding what a theory cannot explain, it is entailed by Consilience.' Unification requires 'the same amount of theoretical content' and 'more kinds of facts,' whereas Completeness requires 'more observations' without a content constraint; the first entailment may hold if 'more kinds of facts' implies 'more observations,' but this is not argued. The second entailment is even less immediate: Consilience does not mention theoretical content, so a consilient theory could have more theoretical content than a rival and therefore not be unifying under the stated definition. These 'Hence' inferences should either be proved from the definitions or explicitly weakened to statements of compatibility rather than entailment.
- [Sections 1 and 4] The central analogy on which the XAI mapping rests—that a local explanation of an ML model is structurally analogous to a scientific theory, with model behavior as observations—is asserted in the introduction and used throughout Section 4 but never defended. This matters because the transfer of virtues such as refutability, optimality, and diachronic virtues (or their dismissal as 'not applicable') depends on the strength of this analogy. The paper should either provide a substantive argument for why a single-instance explanation can be treated as a theory, or explicitly frame the mapping as an exploratory hypothesis and discuss how it could be empirically tested. As written, the section gives the impression that the analogy is a settled fact, which is not supported by the cited literature.
minor comments (5)
- [Abstract] There is a typo in the abstract: 'thes e concepts' should read 'these concepts.'
- [Section 4] The phrase 'As mention above' should be 'As mentioned above.'
- [Section 4] The paper repeatedly gives examples only for regression, logistic regression, and decision trees, but the applicability of many virtues (e.g., internal coherence, conservativeness) to other model classes is not addressed. A brief note on the generality of the mapping would help.
- [Section 3] Footnote 2 explains that Keas (2018) categorized Unification as an aesthetic virtue, but the paper provides no discussion of why it should be moved to the new 'coverage' category. Since this is a structural choice of the paper, a sentence of justification would be appropriate.
- [Section 3] The two tables are useful but not cross-referenced in the text; for instance, Table 2 lists 'Depth (van Cleave, 2016; Lipton, 1991)' under coherential virtues, while Table 1 assigns Lipton's Unification to a separate note. This could be confusing to readers and should be harmonized or annotated.
Circularity Check
No significant circularity: the paper is a literature summary; its Section 4 mapping is a definitional proposal, not a derivation from fitted inputs.
full rationale
The paper is an expository survey of explanatory virtues and a proposal for how they might apply to local XAI explanations. It does not derive a quantitative result, fit a parameter, or rename a fitted value as a prediction. The only self-citations are to Maruf et al. (2023, 2024); these are used as empirical evidence that conservative and unifying explanations are perceived similarly by users, and the cited 2024 study is an externally published user study, so it is independent support rather than a load-bearing self-citation chain. The central claim—that the expanded Keas taxonomy can be linked to XAI—does not reduce to the author's own prior work. Section 4's operationalization of Unification as 'the same feature values' is a definitional choice; even if the 70-of-80 example adds content beyond the feature values and thus sits uneasily with Section 3's definition, that is an internal-consistency or conceptual concern, not a circular derivation. No equation or construction equates the output with the input, so no circular step can be exhibited.
Assumptions & free parameters
assumptions (3)
- domain assumption Local XAI explanations can be treated as 'theories' and the behavior of an ML model as the 'observations' those theories account for.
- domain assumption Keas's (2018) four-category ontology (evidential, coherential, aesthetic, diachronic) is an acceptable base taxonomy for explanatory virtues.
- domain assumption The survey assumes that the cited authors' intended meanings of terms such as 'Unification' and 'Consilience' are correctly captured by the grouped definitions in Tables 1 and 2.
Cite this review
Pith. "Pith review of A Brief Summary of Explanatory Virtues." pith.science (2026). https://pith.science/paper/XMXQTX4Y
@misc{pith2026241116709,
author = {Pith},
title = {Pith review of: A Brief Summary of Explanatory Virtues},
year = {2026},
howpublished = {\url{https://pith.science/paper/XMXQTX4Y}},
note = {Machine review of arXiv:2411.16709}
}
read the original abstract
In this report, I provide a brief summary of the literature in philosophy, psychology and cognitive science about Explanatory Virtues, and link these concepts to eXplainable AI.
Reference graph
Works this paper leans on
-
[1]
Ausubel, D. P. (1962). A subsumption theory of meaningful verbal learning and retention. Journal of General Psychology , 6:213--224
work page 1962
-
[2]
Bastani, O., Kim, C., and Bastani, H. (2017). Interpreting blackbox models via model extraction. arXiv preprint arXiv:1705.08504
arXiv 2017
-
[3]
Biran, O. and McKeown, K. R. (2017). Human-centric justification of Machine Learning predictions. In International Joint Conference on Artificial Intelligence , IJCAI'17, pages 1461--1467, Melbourne, Australia
work page 2017
-
[4]
Bu c inca, Z., Malaya, M. B., and Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI -assisted decision-making. Proceedings of the ACM on Human-Computer Interaction , 5(CSCW1):1--21
work page 2021
-
[5]
Glymour, C. (2015). Probability and the explanatory virtues. British Journal for the Philosophy of Science , 66(3):591--604
work page 2015
-
[6]
Grice, H. P. (1975). Logic and conversation. In Speech Acts [Syntax and Semantics 3] , pages 41--58
work page 1975
-
[7]
Guidotti, R., Monreale, A., Matwin, S., and Pedreschi, D. (2019). Black box explanation by learning image exemplars in the latent feature space. In ECML PKDD , pages 189--205. Springer
work page 2019
-
[8]
He, G., Balayn, A., Buijsman, S., Yang, J., and Gadiraju, U. (2024). Opening the analogical portal to explainability: Can analogies help laypeople in ai-assisted decision making? Journal of Artificial Intelligence Research , 81
work page 2024
Show all 35 references
-
[9]
Hempel, C. G. and Oppenheim, P. (1948). Studies in the logic of explanation. Philosophy of science , 15(2):135--175
1948
-
[10]
and Woodward, J
Hitchcock, C. and Woodward, J. (2003). Explanatory generalizations, Part II : Plumbing explanatory depth. No\^ u s , 37(2):181--199
2003
-
[11]
R., Mueller, S
Hoffman, R. R., Mueller, S. T., Klein, G., and Litman, J. (2018). Metrics for explainable AI : Challenges and prospects. arXiv preprint arXiv:1812.04608
2018 arXiv
-
[12]
Keas, M. N. (2018). Systematizing the theoretical virtues. Synthese , 195(6):2761--2793
2018
-
[13]
Kuhn, T. (1977). Objectivity, value judgment, and theory choice. In The Essential Tension . Chicago University Press
1977
-
[14]
Lakkaraju, H., Kamar, E., Caruana, R., and Leskovec, J. (2017). Interpretable and explorable approximations of black box models. In SIGKDD2017 Workshop on Fairness, Accountability, and Transparency in Machine Learning , Halifax, Canada
2017
-
[15]
Lipton, P. (1991). Inference to the best explanation . Routledge, London and New York, 1st edition
1991
-
[16]
Lipton, P. (2004). Inference to the best explanation . Routledge, London and New York, 2nd edition
2004
-
[17]
Mackonis, A. (2013). Inference to the best explanation, coherence and other explanatory virtues. Synthese , 190(6):975--995
2013
-
[18]
Maruf, S., Zukerman, I., Reiter, E., and Haffari, G. (2023). Influence of context on users’ views about explanations for decision-tree predictions. Computer Speech & Language , 81:101483
2023
-
[19]
Maruf, S., Zukerman, I., Situ, X., Paris, C., and Haffari, G. (2024). Generating simple, conservative and unifying explanations for logistic regression models. In Proceedings of the 17th International Conference on Natural Language Generation , INLG 2024, pages 103--120, Tokyo, Japan
2024
-
[20]
McMullin, E. (2014). The virtues of a good theory. In Curd, M. and Psillos, S., editors, The Routledge companion to Philosophy of Science , pages 561--571. Routledge, New York
2014
-
[21]
Miller, T. (2019). Explanation in Artificial Intelligence : Insights from the social sciences. Artificial Intelligence , 267:1--38
2019
-
[22]
Newton-Smith, W. (1981). The rationality of science . Routlege
1981
-
[23]
Psillos, S. (2002). Simply the best: A case for abduction. In Kakas, A. and Sadri, F., editors, Computational Logic (Kowalski Festschrift) , LNAI 2408, pages 605--625. Springer-Verlag Berlin Heidelberg
2002
-
[24]
and Ullian, J
Quine, W. and Ullian, J. (1978). The web of belief . Random House
1978
-
[25]
T., Singh, S., and Guestrin, C
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). `` W hy should I trust you?": Explaining the predictions of any classifier. In Proceedings of the ACM/SIGKDD Conference on Knowledge Discovery and Data Mining , KDD'16, pages 1135--1144, San Francisco, California
2016
-
[26]
and Morton, A
Rosales, A. and Morton, A. (2021). Scientific explanation and trade-offs between explanatory virtues. Foundations of Science , 26:1075--1087
2021
-
[27]
Schupbach, J. N. and Sprenger, J. (2011). The logic of explanatory power. Philosophy of Science , 78(1):105--127
2011
-
[28]
and Flach, P
Sokol, K. and Flach, P. (2020). One explanation does not fit all: The promise of interactive explanations for Machine Learning transparency. K \"u nstliche Intelligenz , 34:235--250
2020
-
[29]
M., Gatala, A., and Pereira-Fari\ n a, M
Stepin, I., Alonso, J. M., Gatala, A., and Pereira-Fari\ n a, M. (2020). Generation and evaluation of factual and counterfactual explanations for decision trees and fuzzy rule-based classifiers. In 2020 IEEE International Conference on Fuzzy Systems , FUZZ-IEEE, page 1–8, Glas...
2020
-
[30]
Thagard, P. (1978). The best explanation: Criteria for theory choice. The Journal of Philosophy , 75(2):76--92
1978
-
[31]
Thagard, P. (1989). Explanatory coherence. Behavioral and Brain Sciences , pages 435--502
1989
-
[32]
van Cleave, M. (2016). Introduction to Logic and Critical Thinking . Lansing Community College
2016
-
[33]
van der Waa, J., Robeer, M., van Diggelen, J., Brinkhuis, M., and Neerincx, M. (2018). Contrastive explanations with local foil trees. In Proceedings of the ICML-18 Workshop on Human Interpretability in Machine Learning , WHI'18, pages 41--46, Stockholm, Sweden
2018
-
[34]
van Fraassen , B. (1980). The scientific image . Oxford University Press
1980
-
[35]
C., Sloman, S., Bechlivanidis, C., and Lagnado, D
Zemla, J. C., Sloman, S., Bechlivanidis, C., and Lagnado, D. A. (2017). Evaluating everyday explanations. Psychonomic Bulletin & Review , 24(5):1488--1500
2017
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.