{"id":"b47bc3f0-ebee-4dfa-ac20-277c3f255244","arxiv_id":"2505.10188","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In a study with eight physicians, AI explanations that list supporting features were rated more comprehensible and plausible than counterfactual or exclusion-based explanations, yet doctors were reluctant to use any AI explanation with colleagues.","lead":"This paper reports a study in which eight epilepsy doctors rated three types of AI-generated explanations for diagnostic support systems. Attribution-based explanations were preferred, but the doctors said they would not use any of them to explain a diagnosis to a colleague.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No inferential test supports the 'significantly better' wording in Section 6; the comparative claim rests on visual boxplot differences with n=8 raters.","rationale":"I focused on the comparative claim because it is the paper's main contribution and the phrase 'significantly better' appears in the conclusion. The qualitative finding about missing uncertainty/reliability information is better supported: it comes from follow-up interviews with direct quotes, and the paper honestly acknowledges the small sample size in its limitations. The descriptive patterns are plausible, and the paper does not hide the data limitations. But the central quantitative claim requires inferential support that is absent. The reader's identified weakest assumption (template instantiations confounding type) is real and related; my concern is narrower and easier to check: no statistical test is reported. This does not move the reader's CONDITIONAL verdict. If the requested significance test fails, the condition should be enforced by removing 'significantly' and describing the result as a descriptive finding. The paper is sound as a descriptive qualitative study, but it cannot support the 'significantly' wording as written.","tokens_in":14735,"tokens_out":5375,"duration_ms":50999,"concrete_test":"Ask the authors to release the per-rater, per-case Likert ratings and rerun the comparison with a repeated-measures ordinal model (e.g., cumulative-link mixed model with random intercepts for rater and case) or, failing that, a Friedman test with post-hoc Wilcoxon signed-rank tests and appropriate multiple-comparison correction for each dimension (comprehensibility, plausibility, completeness, colleague-use). Report adjusted p-values and 95% confidence intervals for the attribution-vs-counterfactual and attribution-vs-exclusion contrasts. If these contrasts do not meet the chosen significance threshold, the Section 6 conclusion should be revised to say 'descriptively rated higher' rather than 'significantly better'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline comparative conclusion (Section 6) is: 'In terms of plausibility and comprehensibility, attribution-based arguments were evaluated significantly better than counterfactual or exclusion-based arguments.' The only quantitative evidence offered is the boxplots in Figures 2–5; no test statistic, p-value, effect size, confidence interval, or repeated-measures model is reported anywhere in Section 5. With eight raters, ten cases, and Likert-scale items, descriptive medians can easily be driven by a few participants or by case-level variation. Therefore, the word 'significantly' makes a statistical claim that the reported analysis does not support. This is load-bearing independently of the template-instantiation concern: even if the three argument types were instantiated in perfectly matched text, the data as presented do not establish a significant difference. The reader's confound concern (Section 4.3/4.4) is a further threat, but it is secondary; the immediate defect is that the central claim overstates what the analysis can show.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a user study with eight physicians evaluating three template-based natural-language explanations for a diagnostic decision-support system in transient loss of consciousness (TLOC). The three explanation types are attribution-based (derived from LIME), counterfactual, and exclusion-based. Participants rated the explanations on comprehensibility, plausibility, completeness, and willingness to use them when explaining a case to a colleague, and later took part in semi-structured follow-up interviews. The paper reports that attribution-based explanations received the best ratings on plausibility and comprehensibility, that willingness to use any AI-generated explanation was low, and that interview participants attributed this reluctance to missing information about prediction quality, reliability, uncertainty, and to the lack of case-specific tailoring.","tokens_in":14771,"tokens_out":4182,"duration_ms":43143,"significance":"The study addresses a relevant and under-explored question in XAI and medical decision support: how clinicians perceive different argumentative explanation types. Its strengths include the use of real patient cases, domain-expert participants, a concrete diagnostic task, and the combination of quantitative ratings with qualitative interview data. The descriptive findings and the interview quotes about uncertainty and reliability signaling are valuable for designers of explanation interfaces. However, the central comparative claim is not supported by the statistical analysis as reported, and the explanation-type comparison may be confounded by the specific feature instantiations used for each case. If the authors can provide appropriate inferential statistics and address the confound, the contribution would be a useful empirical data point for argumentative XAI; in its current form the main conclusion overstates what the data show.","major_comments":[{"comment":"The sentence \"attribution-based arguments were evaluated significantly better than counterfactual or exclusion-based arguments\" is not supported by any inferential statistical analysis in the manuscript. Section 5 reports only visual boxplot comparisons and descriptive statements; no test statistic, p-value, confidence interval, or effect size is provided, and no repeated-measures model is fitted. With eight raters and repeated measurements on ten cases, apparent median differences can easily be driven by a few participants or by case-level variation. Please add an appropriate analysis (e.g., Friedman test with post-hoc comparisons, or a mixed-effects model with by-participant and by-case random effects) or remove the word \"significantly\" and frame the finding as a descriptive trend.","section":"Section 6, with evidence in Section 5 (Figures 2-5)"},{"comment":"The comparison across explanation types is potentially confounded by the specific instantiation of each template. For each case, participants saw one realization of each template filled with particular LIME or counterfactual features, and the attribution template even permits replacing the top LIME features with other, \"less important\" features when the top features are not among the expert-defined relevant variables. Ratings may therefore reflect the suitability, wording, or case-fit of the particular features rather than the argument type itself. In addition, the number of feature mentions differs (two in the attribution template versus four in the counterfactual and exclusion templates) and the latter two include alternative diagnoses. Please provide evidence that the instantiations are matched on relevant properties, or soften the cross-type claims and explicitly discuss this confound as a limitation.","section":"Sections 4.3 and 4.4"}],"minor_comments":[{"comment":"The sentence \"For both CFs and attribution-based explanations, the median value is significantly lower\" appears to contain a typo: the context suggests it should refer to \"CFs and exclusion-based explanations\"; \"significantly\" should also be removed unless a statistical test is reported.","section":"Section 5, paragraph after Figure 2"},{"comment":"The feature selection for exclusion-based explanations is not described: the text explains how LIME features are chosen for attribution-based explanations and how counterfactuals are generated, but it does not state how the features in the exclusion-principle template are selected.","section":"Section 4.3"},{"comment":"The manuscript does not include a statement on ethics approval or informed consent for the user study with physicians; please add this information, as is standard for empirical studies with human participants.","section":"Section 4.4 and Appendix"},{"comment":"Several boxplots have medians that coincide with quartiles, making the plots hard to read; adding jittered individual data points would improve interpretability.","section":"Figures 2 and 3"},{"comment":"References [25] and [26] appear to be the same paper (Krause, Ambler, Elvang-Gøransson, and Fox, \"A logic of argumentation for reasoning under uncertainty\", Computational Intelligence 11:113-131, 1995) and should be consolidated.","section":"References"},{"comment":"The term \"prototype\" in the explanation patterns is not defined in the text; please explain what constitutes a prototype-based explanation in this context.","section":"Table 2"},{"comment":"The conclusion states that \"all explanations score reasonably well along the first three dimensions,\" but Section 5 reports that completeness is perceived as \"rather low, or moderate in the case of attribution-based arguments\"; please align this wording with the reported results.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The core descriptive findings are potentially useful, but the paper currently hinges on a comparative claim that the reported analysis cannot support. If the authors cannot provide a proper repeated-measures analysis with the existing data, they should remove the word \"significantly\" and reframe the conclusion as a descriptive observation; that alone would substantially improve the manuscript. The template-instantiation confound also needs to be addressed honestly, even if only as a clearly stated limitation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a small, honestly reported user study that would be a solid workshop paper or a revised journal paper, but the conclusion claims a statistical difference that the data analysis doesn't support.\n\nWhat is actually new: a side-by-side evaluation of three explanation templates—attribution via LIME, counterfactual, and exclusion-principle—in a specific clinical diagnostic task (transient loss of consciousness), with eight practicing epileptologists. The qualitative follow-up is the most valuable part: doctors said they wouldn't use the AI explanations because they lacked information about prediction quality, reliability, and uncertainty. That's a concrete, actionable design requirement.\n\nThe paper does several things well. The templates are clearly specified, the case reports are realistic, and the authors are upfront about the small sample and its limitations. The boxplots and interview quotes are consistent with the descriptive claims.\n\nNow the soft spots. The word 'significantly' appears in the conclusion: attribution-based arguments were evaluated significantly better than the other two on plausibility and comprehensibility. No significance test, CI, or effect size is reported. With eight raters and ten cases, descriptive medians can shift with one or two respondents. The boxplots do show a visible difference, so the descriptive claim is plausible; the inferential claim is not. The authors should either run a proper repeated-measures analysis (which will likely be underpowered) or soften the language to 'visibly better in this sample' or 'trending better.'\n\nThe second concern is that the three explanation types are each instantiated with case-specific features, so the comparison could reflect the particular LIME outputs or the template wording rather than the type of argument itself. The authors do mention a feature-replacement rule, but the rule is ad hoc and could introduce bias. This is a secondary threat; it doesn't negate the study, but it means the effect sizes are likely inflated.\n\nAlso, no code or data are released, which makes it hard to reanalyze. That's a reproducibility minus.\n\nWho this is for: researchers working on XAI evaluation, human-centered AI, or clinical decision support. The qualitative finding about uncertainty is worth taking seriously, but the paper should be treated as an exploratory pilot, not a confirmatory test.\n\nRecommendation: Yes, engage with this. Send it to peer review, but the authors should be asked to fix the statistical language, add tests or clearly label as exploratory, and ideally release a reproducible analysis script. If they do that, it's a useful addition to the XAI user-study literature.","headline":"Honest small pilot study with a useful qualitative insight about uncertainty, but the 'significantly better' conclusion is not backed by any inferential test.","tokens_in":15432,"tokens_out":2790,"would_cite":true,"duration_ms":26366,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Attribution-based explanations beat counterfactual and exclusion arguments in physician ratings, yet doctors would not use any AI-generated explanation to explain a case to a colleague.","keywords":["explainable AI","argumentative explanations","diagnostic decision support","user study","counterfactual explanations","feature attribution","transient loss of consciousness","Bayesian network"],"falsifier":"Run a follow-up study in which each case is paired with several independently generated instantiations of each explanation template. If attribution-based arguments do not consistently outrank counterfactual and exclusion arguments once the concrete features and wording are varied, the reported advantage is an artifact of the specific instantiations rather than of the argument type.","tokens_in":14400,"feed_emoji":"🩺","tokens_out":5985,"duration_ms":52141,"temperature":0.7,"pith_summary":"The paper reports a user study in which eight neurologists rated three kinds of AI-generated explanations produced for a diagnostic-support system that classifies transient loss of consciousness. The central claim is that attribution-based explanations, which state that a feature supports a given diagnosis, were rated significantly higher in plausibility and comprehensibility than counterfactual or exclusion-based explanations. The paper also claims that, despite moderate-to-good scores on those dimensions, the doctors would not use any of the AI explanations to explain a case to a colleague, because the explanations gave no information about prediction quality, reliability, or uncertainty. If this is right, explanation type is a real lever on clinician acceptance, but no verbalization format can replace explicit uncertainty signaling.","feed_headline":"Doctors rate simple AI explanations best yet won't use them","feed_subtitle":"Neurologists judged feature-attribution arguments most plausible, then said no AI explanation was fit for a colleague.","key_machinery":"The central object is a set of three natural-language argument templates: attribution ('Because feature is A, the patient most likely has diagnosis X'), counterfactual ('If feature were A, the diagnosis would most likely change to Y'), and exclusion ('Because feature is A, diagnosis Y can most likely be ruled out'). These templates are instantiated with feature-value pairs coming from LIME and from a counterfactual generator, using a Bayesian network (a causal probabilistic model with layers of risk factors, diseases, and symptoms) trained on 300 patient records as the underlying predictor. The templates are what allow the comparison of XAI methods to proceed over explicit arguments instead of raw feature lists.","core_discovery":"On the paper's own terms, the discovery is that when the outputs of standard XAI methods are turned into natural-language arguments, physicians evaluating real TLOC cases rate the simple attribution argument as the most plausible and comprehensible, significantly above counterfactual and exclusion-based arguments, while ratings for completeness and 'would I use this with a colleague' are low for all three. The follow-up interviews identify the missing information about the quality and reliability of the prediction as the main cause of the doctors' reluctance, together with the lack of case-specific detail in the templated explanations.","pith_inferences":["If the specific features selected for the counterfactual and exclusion templates in the chosen ten cases happened to be less clinically salient than those used for attribution templates, the reported ranking could partly reflect content rather than argument type; a re-run with multiple independent instantiations per case would separate the two.","The doctors' own explanations frequently combined a prototype statement with ruled-out alternatives, which suggests a 'typical case, and here is what it is not' template might match clinical communication more closely than any single current template.","A direct design fix suggested by the interview data is to append a confidence or reliability statement to each explanation; the study predicts this would raise the 'use with a colleague' rating.","The results favor evaluating XAI by use-oriented acceptance tests, such as whether a clinician would relay the explanation to a peer, rather than by perceived quality ratings alone."],"forward_implications":["Explanation format is not presentationally neutral: converting the same model prediction into attribution, counterfactual, or exclusion wording changes how clinicians judge the explanation.","For this diagnostic task, simpler feature-attribution arguments appear to be a safer default than counterfactual or exclusion arguments, at least when judged on plausibility and comprehensibility.","A diagnostic-support explanation that omits prediction confidence, reliability, and uncertainty will be rejected for inter-physician communication regardless of which XAI method generated it.","Tailoring the explanation to the complexity of the case and the experience of the clinician, rather than generating one template for every case, is a necessary next step for acceptance."],"supporting_citations":[{"why":"Supplies the LIME feature-attribution method whose outputs are verbalized into attribution-based arguments.","marker":"[41]"},{"why":"Supplies the counterfactual-explanation algorithm used to generate the counterfactual arguments being tested.","marker":"[24]"},{"why":"Provides the 300-patient dataset from which the Bayesian-network model's parameters are learned.","marker":"[53]"},{"why":"Provides the causal Bayesian-network modeling approach used to construct the disease model.","marker":"[42]"},{"why":"Motivates counterfactual explanations by their resemblance to human explanation and defines their expected form.","marker":"[49]"},{"why":"Establishes human evaluation as the gold standard for explanation assessment, justifying the user study.","marker":"[8]"},{"why":"Supports the choice of two features per explanation as the complexity sweet spot.","marker":"[29]"},{"why":"Supplies the operative definition of counterfactual explanations as 'if X had been different, the prediction would have changed from P to Q'.","marker":"[21]"}],"fun_headline_variants":["Doctors favor simple AI arguments, then decline to use them","Why top-rated AI explanations still fail doctors' trust","Simple AI explanations win on plausibility, lose on use","Neurologists rate attribution best, yet find it unfit","Doctors like plain AI reasons but want reliability proof"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that each doctor's score reflects the explanation type, not the particular features and wording chosen for that type in each specific case; if the instantiations differed in quality for reasons unrelated to the method, the ranking could be an artifact of the examples.","fun_headline_variants_meta":{"raw":{"variants":["Doctors favor simple AI arguments, then decline to use them","Why top-rated AI explanations still fail doctors' trust","Simple AI explanations win on plausibility, lose on use","Neurologists rate attribution best, yet find it unfit","Doctors like plain AI reasons but want reliability proof"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000202,"raw_usage":{"total_tokens":1327,"prompt_tokens":838,"completion_tokens":489,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":409}},"tokens_in":454,"tokens_out":489,"duration_ms":5441,"temperature":1.0,"reasoning_tokens":409,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:14:15.255354+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a follow-up study in which each case is paired with several independently generated instantiations of each explanation template. If attribution-based arguments do not consistently outrank counterfactual and exclusion arguments once the concrete features and wording are varied, the reported advantage is an artifact of the specific instantiations rather than of the argument type.","supporting_citations":[{"cited_title":"why should i trust you?","cited_arxiv_id":null,"evidence_quote":"Supplies the LIME feature-attribution method whose outputs are verbalized into attribution-based arguments."},{"cited_title":"Klaise, A","cited_arxiv_id":null,"evidence_quote":"Supplies the counterfactual-explanation algorithm used to generate the counterfactual arguments being tested."},{"cited_title":"Wardrope, J","cited_arxiv_id":null,"evidence_quote":"Provides the 300-patient dataset from which the Bayesian-network model's parameters are learned."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the causal Bayesian-network modeling approach used to construct the disease model."},{"cited_title":"Wachter, B","cited_arxiv_id":null,"evidence_quote":"Motivates counterfactual explanations by their resemblance to human explanation and defines their expected form."},{"cited_title":"Liedeker, C","cited_arxiv_id":null,"evidence_quote":"Supports the choice of two features per explanation as the complexity sweet spot."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the operative definition of counterfactual explanations as 'if X had been different, the prediction would have changed from P to Q'."}],"review_version":1}