{"id":"94f67961-4787-4c93-b36c-b9ee5678891a","arxiv_id":"2411.16709","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A concise literature survey that organizes explanatory virtues into Keas's four categories plus a new coverage category, and relates them to explainable AI.","lead":"This report surveys how philosophers and psychologists define the qualities of a good explanation and maps those qualities to AI systems that explain their decisions. It gives XAI researchers a shared vocabulary for evaluating and improving machine-generated explanations.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4's Unification example adds theoretical content rather than holding it fixed, so the taxonomy is redefined rather than applied to local XAI explanations.","rationale":"The paper is a useful and largely well-organized review: Tables 1 and 2 are clear, the definitions in Section 3 are sourced, and Section 4 explicitly acknowledges that several virtues do not transfer to local explanations. The author is not claiming a formal theorem; this is a conceptual mapping. However, the mapping is the paper's only non-expository contribution, and Section 4's flagship example for coverage virtues is self-undermining. If Unification is meant to be the same virtue as in Section 3, adding a statistical sentence to an attribution violates the 'same theoretical content' condition. If the virtue is meant to be adapted, the paper does not say what the new definition is or why it preserves the normative force of Section 3's virtue. This is a real soft spot in the central claim, not merely a disagreement with the philosophical literature. It is also more specific than the reader's weakest assumption: the problem is not just that the theory/observation analogy is unproven; it is that Section 4's operationalization breaks the internal definitions the paper itself sets up. Consequently, the link to XAI is not achieved by application but by redefinition, and the strongest experimental support cited for the central claim (Maruf et al. 2024) is an empirical study of user preferences, not evidence that the Section 3 virtue has been transferred. The paper remains UNVERDICTED because it is a review with no testable research claim; the concern is grounds for a revision of Section 4, not a rejection of the summary as a whole. I therefore keep the reader's verdict unchanged while flagging the Section 4 definitions for correction.","tokens_in":8362,"tokens_out":7148,"duration_ms":69900,"concrete_test":"Take the two explanations from §4's car example and represent them in the Section 3 framework: E1 = (C1={four doors, large boot}, O1={the instance}); E2 = E1 plus '70 of 80 cars with four doors and a large boot are deemed acceptable' (C2=C1∪{70/80 statistic}, O2=O1∪{80 model outputs}). Check whether E2 is more unifying by Section 3's definition: it requires C2⊆C1 (same theoretical content) and O2⊃O1. Since C2 contains an additional statistical claim, C2⊄C1, so E2 is not more unifying. If one instead declares C2=C1 on the grounds that 'theoretical content' means only feature values, then the statistical sentence adds no content and the example reduces to reporting model accuracy. Either branch shows that §4 does not transfer the Section 3 definition unchanged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's stated contribution (abstract; §1, §4) is to link explanatory virtues to local XAI explanations by treating each local explanation as a theory and model behavior as observations. The load-bearing step is Section 4's concrete operationalization of coverage virtues, especially Unification. Section 3 defines Unification as: T explains more kinds of facts than a rival 'with the same amount of theoretical content,' and says Completeness is entailed by Unification. Section 4 then gives as its example of a Unifying explanation: the attribution 'this car is acceptable because it has four doors and a large boot' augmented with 'Indeed, 70 of 80 cars with four doors and a large boot are deemed acceptable by the model.' The italicized sentence is new theoretical content (a statistical generalization about 80 cars), so E2 does not have 'the same amount of theoretical content' as E1; under §3 it is not more unifying. If the author intends 'theoretical content' to mean only feature values, then the 70/80 sentence is not part of the theory and supplies empirical evidence for the explanation, which is an evidential virtue (Evidential accuracy or Conservativeness), not a coverage virtue. Either way, the transfer changes the meaning of Unification. The same pattern recurs for Completeness and Scope: §4 defines them by how many additional model applications an explanation mentions ('more instances,' 'additional types of situations'), whereas §3 defines them by how many observations the theory explains. A local explanation for one prediction can mention the model's performance elsewhere, but it does not thereby explain those outputs.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a brief literature summary of 'explanatory virtues' (EVs) drawn from philosophy, psychology, and cognitive science, organized mainly around Keas's (2018) taxonomy with an added 'coverage' category. It then attempts to link these virtues to explainable AI (XAI), focusing on local explanations, by viewing each local explanation as a theory and the ML model's behavior on an instance as observations. Section 3 provides definitions and two summary tables, and Section 4 gives an item-by-item mapping of virtues to XAI, discussing which are applicable, adaptable, or not relevant.","tokens_in":8634,"tokens_out":6677,"duration_ms":61932,"significance":"If the summary and the XAI mapping were sound, the paper would offer a useful interdisciplinary bridge, particularly by directing attention to coverage virtues such as unification, which are underrepresented in the XAI literature. The paper's organizational scheme, with its two tables and explicit distinction between epistemic/pragmatic and ontological dimensions, is a practical reference for researchers entering this area. However, the paper's usefulness is weakened by several internal inconsistencies: a concrete example of 'unification' in Section 4 does not satisfy the paper's own definition, one stated implication among virtues is logically false, and the analogy between local explanations and scientific theories is asserted without justification. These issues affect the central claimed contribution of linking EVs to XAI, but they are local enough to be fixable in a revision.","major_comments":[{"comment":"The unification example does not instantiate the definition of unification given in Section 3. Section 3 defines unification as 'a theory T explains more kinds of facts than rival explanations with the same amount of theoretical content,' and the text glosses 'theoretical content' as concepts or rules mentioned in an explanation. In Section 4, however, the example adds the italicized sentence 'Indeed, 70 of 80 cars with four doors and a large boot are deemed acceptable by the model' to the initial attribution. That sentence is a statistical generalization, i.e., an additional rule, not a feature value. Therefore E2 has more theoretical content than E1, and under Section 3's definition it is not more unifying. If the author intended 'theoretical content' to mean only feature values, then the added sentence constitutes empirical evidence, which would make it an evidential rather than a coverage virtue. Either way, the example conflates two different notions. The section should be revised so that the example genuinely keeps theoretical content fixed while expanding the explained observations, or the definition of unification must be explicitly revised and justified.","section":"Section 4"},{"comment":"The claim that 'Since Modesty demands a subset relation between the requirements of rival explanations, it implies Simplicity' is not correct under the definitions given. Modesty is defined as logical weakness: theory T is more modest than another if T is implied by the other. This does not require that T and the rival explain the same facts; T may explain strictly fewer facts. Simplicity, by contrast, requires that T 'explains the same facts as rival explanations, but with fewer theoretical requirements.' A subset relation between theoretical requirements does not by itself guarantee identical explanatory scope. For example, a theory with requirements {A} is implied by a rival with requirements {A, B}, but the former explains only facts about A while the latter also explains facts about B. Thus Modesty does not imply Simplicity without an additional assumption. The implication should be removed or the definitions amended to make the relationship valid.","section":"Section 3"},{"comment":"The entailment relations among Completeness, Unification, and Consilience are stated without proof and do not obviously follow from the definitions. The paper says 'Completeness is entailed by Unification' and 'Since Unification poses no requirements regarding what a theory cannot explain, it is entailed by Consilience.' Unification requires 'the same amount of theoretical content' and 'more kinds of facts,' whereas Completeness requires 'more observations' without a content constraint; the first entailment may hold if 'more kinds of facts' implies 'more observations,' but this is not argued. The second entailment is even less immediate: Consilience does not mention theoretical content, so a consilient theory could have more theoretical content than a rival and therefore not be unifying under the stated definition. These 'Hence' inferences should either be proved from the definitions or explicitly weakened to statements of compatibility rather than entailment.","section":"Section 3"},{"comment":"The central analogy on which the XAI mapping rests—that a local explanation of an ML model is structurally analogous to a scientific theory, with model behavior as observations—is asserted in the introduction and used throughout Section 4 but never defended. This matters because the transfer of virtues such as refutability, optimality, and diachronic virtues (or their dismissal as 'not applicable') depends on the strength of this analogy. The paper should either provide a substantive argument for why a single-instance explanation can be treated as a theory, or explicitly frame the mapping as an exploratory hypothesis and discuss how it could be empirically tested. As written, the section gives the impression that the analogy is a settled fact, which is not supported by the cited literature.","section":"Sections 1 and 4"}],"minor_comments":[{"comment":"There is a typo in the abstract: 'thes e concepts' should read 'these concepts.'","section":"Abstract"},{"comment":"The phrase 'As mention above' should be 'As mentioned above.'","section":"Section 4"},{"comment":"The paper repeatedly gives examples only for regression, logistic regression, and decision trees, but the applicability of many virtues (e.g., internal coherence, conservativeness) to other model classes is not addressed. A brief note on the generality of the mapping would help.","section":"Section 4"},{"comment":"Footnote 2 explains that Keas (2018) categorized Unification as an aesthetic virtue, but the paper provides no discussion of why it should be moved to the new 'coverage' category. Since this is a structural choice of the paper, a sentence of justification would be appropriate.","section":"Section 3"},{"comment":"The two tables are useful but not cross-referenced in the text; for instance, Table 2 lists 'Depth (van Cleave, 2016; Lipton, 1991)' under coherential virtues, while Table 1 assigns Lipton's Unification to a separate note. This could be confusing to readers and should be harmonized or annotated.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is more of a position paper or literature survey than a technical research contribution. Its main value lies in the organized compilation of virtues and the explicit mapping to XAI, but the logical slips and the problematic unification example would need to be corrected before publication. The author's self-citations (Maruf et al., 2023, 2024) are used to support the claim that coverage virtues warrant further investigation; while not inappropriate, the novelty of that claim relative to those prior works should be clarified. The fit with a cs.AI venue is plausible if the paper is positioned as a provocation for further study rather than a definitive framework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a tidy, accurately-labeled literature summary with a plausible taxonomic tweak. The problem is in Section 4, where the definitions drift when applied to XAI, so the 'coverage' virtues end up meaning something slightly different than in Section 3.\n\nWhat's actually new: the coverage category (Completeness, Scope, Consilience, Unification, Importance) and the two tables that map terms to authors. The tables are genuinely handy; grouping Unification under coverage rather than Keas's aesthetic category is defensible. The XAI section also does something useful: it goes through each virtue and says which ones transfer to local explanations, which ones apply only to explanation-generating procedures, and which ones are irrelevant. That's a reasonable constraint on an often hand-wavy discussion. The citation pattern looks fair; the self-citations to Maruf et al. are relevant, since the 2024 paper is the experiment behind the car example.\n\nWhere it gets soft: the worked example of Unification. Section 3 defines Unification as explaining more kinds of facts than a rival 'with the same amount of theoretical content.' Section 4 reinterprets that as 'the same feature values' and gives the car example, where the second sentence ('70 of 80 cars...') is added. That sentence isn't a feature value — it's a statistical generalization. If it counts as theoretical content, the explanation has more theoretical content, so it doesn't fit §3. If it doesn't count, then the virtue doing the work is evidential accuracy, not coverage. Either way the example doesn't instantiate the definition. The same drift appears with Completeness and Scope: §4 defines them by how many model applications the explanation mentions, whereas §3 defines them by observations the theory explains. Mentioning other instances isn't the same as explaining them.\n\nAlso the theory-observation analogy for local XAI is asserted, not argued. For a brief summary that's acceptable, but it's worth noting that the 'coverage' vocabulary carries a lot of weight in Section 4 on a thin base.\n\nThe paper is honest about being a summary. I wouldn't fault it for not having a new theorem or experiment. The reader's UNVERDICTED verdict is right; the significance is modest but real.\n\nFor a serious venue I'd send it to review — the taxonomic organization is solid enough to deserve feedback, and a referee can pin down the Section 4 drift. For a workshop, accept with caveats. I'd cite the tables if I needed to index the literature, though I'd probably cite Keas for the core taxonomy.","headline":"A clean taxonomic survey with a definitional drift in the XAI section that a referee should catch.","tokens_in":9159,"tokens_out":4382,"would_cite":false,"duration_ms":36925,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This report argues that explanatory virtues from philosophy, psychology, and cognitive science can be organized into an expanded taxonomy and then applied to local AI explanations by treating the explanation as a theory and the model's…","keywords":["explanatory virtues","explainable AI","local explanations","taxonomy","coverage virtues","unification","philosophy of science"],"falsifier":"A user study that keeps one model, one instance, and one feature-attribution explanation fixed, then adds only an applicability statement (e.g., 70 of 80 similar cars are acceptable), and asks users to rate understanding, trust, and action would test the coverage claim; if the added statement changes none of the ratings after word count is controlled, the paper's claim that coverage virtues warrant further investigation is not supported.","tokens_in":8164,"feed_emoji":"🔍","tokens_out":11673,"duration_ms":103431,"temperature":0.7,"pith_summary":"This report attempts to establish that the explanatory virtues studied in philosophy, psychology, and cognitive science can be organized into a single taxonomy and then carried over to explainable AI (XAI). The carry-over works by treating a local explanation of a machine-learning prediction as a theory and the model's behavior on that instance as the observations the theory must fit. If the mapping holds, XAI gains a ready-made vocabulary for judging explanations: accuracy, coherence, simplicity, unification, and the rest, each with a concrete meaning for a single prediction. The report's most forward-looking claim is that the coverage virtues, especially unification, are under-studied in XAI and deserve more investigation.","feed_headline":"XAI explanations inherit philosophy's explanatory virtues","feed_subtitle":"Local model explanations become theories judged by philosophical virtues; coverage is flagged as the research gap.","key_machinery":"The central object is the explanatory-virtue taxonomy, an expanded version of a four-category classification from the philosophy of science (evidential, coherential, aesthetic, and diachronic virtues) to which the report adds a fifth category, coverage, for virtues about how widely a theory applies. The taxonomy carries the argument: Section 4 runs each virtue against local XAI explanations, translating it into a concrete property of an explanation of a single prediction, and records which virtues survive the translation and which do not. Coverage is the load-bearing new piece, because it is where the report places unification and where it locates the under-researched gap in XAI.","core_discovery":"The central claim is that the virtues long used to judge scientific theories can double as criteria for judging local XAI explanations. The report gathers virtues from the literature into an expanded five-category taxonomy—evidential, coherential, coverage, aesthetic, and diachronic—and then proposes the theory-explanation mapping: the explanation of an ML model's prediction is the theory, and the model's behavior on the instance is the observation. Under this mapping, evidential accuracy, counterfactual-style explanatory depth, optimality, internal and universal consistency, conservativeness, simplicity, modesty, completeness, scope, consilience, and unification apply directly; causal adequacy and refutability drop out because a local explanation neither causes the model's outcome nor faces imaginable refuting instances; and diachronic virtues survive only as properties of the procedure that generates explanations. The report also singles out unification and the other coverage virtues as the least studied XAI dimension, supported by a user study in which unifying explanations were about as well liked as conservative explanations.","pith_inferences":["If the theory-explanation analogy is taken literally, explanation quality becomes a comparative, dataset-dependent property rather than a property of the explanation text alone, because coverage virtues depend on how many other instances the explanation generalizes to.","The taxonomy makes trade-offs explicit: a simple explanation can be incomplete, and a complete one can be ad hoc, so applying the virtues to XAI means deciding how to weight competing criteria, much as theory choice does in philosophy.","A natural next experiment is to turn each transferred virtue into a rating item and test whether the virtues predict user trust and action beyond explanation length; the report cites adjacent evidence but does not run such a study itself."],"forward_implications":["Local XAI evaluation can be organized by the five-category taxonomy, with each virtue translated into a concrete property of an explanation of a single prediction.","Several classical virtues drop out of local evaluation: causal adequacy and refutability do not apply, and diachronic virtues apply to explanation-generation procedures rather than to the explanation itself.","Conversational maxims give a practical route to operationalizing coherence and depth, so explanations can be scored on clarity, order, quality, and quantity of information.","Coverage virtues—especially unification—emerge as the under-researched but promising direction for XAI, with one user study showing unifying explanations are about as well liked as conservative ones."],"supporting_citations":[{"why":"Provides the base four-category taxonomy of theoretical virtues that the report extends with a coverage category.","marker":"Keas, 2018"},{"why":"Supplies the epistemic/pragmatic distinction and core virtue definitions, including unification.","marker":"van Fraassen, 1980"},{"why":"Provides the conversational maxims used to operationalize coherence and depth for local explanations.","marker":"Grice, 1975"},{"why":"Defines the four explanatory attributes used to compare explanations in the cited user study.","marker":"Hoffman et al., 2018"},{"why":"User study showing unifying explanations rival conservative ones, grounding the claim that coverage virtues warrant investigation.","marker":"Maruf et al., 2024"},{"why":"One of the two studies cited as having examined coverage-style explanations in XAI, supporting the under-studied claim.","marker":"Bucinca et al., 2021"},{"why":"Canonical local-explanation method that the report's mapping is intended to cover.","marker":"Ribeiro et al., 2016"}],"fun_headline_variants":["Coverage virtues: XAI's untapped evaluation criteria","Theory virtues judge local XAI explanations","Local explanations become theories for virtue checks","Five virtue categories map to XAI explanation quality","Philosophy's theory virtues apply to XAI explanations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central mapping rests on treating a local explanation of a machine-learning prediction as structurally analogous to a scientific theory, with the model's behavior as the observations; if that analogy is invalid, the transfer of virtues such as optimality, coherence, and unification to explainable AI has no foundation.","fun_headline_variants_meta":{"raw":{"variants":["Coverage virtues: XAI's untapped evaluation criteria","Theory virtues judge local XAI explanations","Local explanations become theories for virtue checks","Five virtue categories map to XAI explanation quality","Philosophy's theory virtues apply to XAI explanations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000978,"raw_usage":{"total_tokens":4052,"prompt_tokens":743,"completion_tokens":3309,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":359,"completion_tokens_details":{"reasoning_tokens":3239}},"tokens_in":359,"tokens_out":3309,"duration_ms":22351,"temperature":1.0,"reasoning_tokens":3239,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:41:28.825875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A user study that keeps one model, one instance, and one feature-attribution explanation fixed, then adds only an applicability statement (e.g., 70 of 80 similar cars are acceptable), and asks users to rate understanding, trust, and action would test the coverage claim; if the added statement changes none of the ratings after word count is controlled, the paper's claim that coverage virtues warrant further investigation is not supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the base four-category taxonomy of theoretical virtues that the report extends with a coverage category."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the epistemic/pragmatic distinction and core virtue definitions, including unification."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the conversational maxims used to operationalize coherence and depth for local explanations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"User study showing unifying explanations rival conservative ones, grounding the claim that coverage virtues warrant investigation."},{"cited_title":"T., Singh, S., and Guestrin, C","cited_arxiv_id":null,"evidence_quote":"Canonical local-explanation method that the report's mapping is intended to cover."}],"review_version":1}