{"id":"820c9406-8348-4495-8d15-78feea93ba6d","arxiv_id":"2505.06620","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An MHRA-convened expert group and an eight-clinician pilot found that local explainable AI outputs improve trust and accuracy in heart attack risk predictions, while also exposing automation bias risks.","lead":"UK regulators, doctors, and data scientists met to test whether explainable AI can help doctors trust heart attack risk tools. A small pilot found explanations raised confidence and accuracy but also created over-reliance, which matters for how medical AI devices are approved and trained.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pilot cannot support the abstract's causal claim: clinicians saw AI diagnosis, confidence, and all XAI outputs together, with no control or local-explanation-only arm, and Appendix A records the study as not generalisable.","rationale":"Reader's weakest assumption is that eight clinicians, ten selected cases, no control, and no statistics cannot support a general claim. I agree, and I would sharpen it: even within the pilot, the post-AI condition bundles the AI prediction, confidence, and global/LIME/counterfactual explanations, so no causal role for local explanations is identifiable. The reported 5-of-6 shift is also not statistically significant under an exact binomial test (one-sided p≈0.11), and the acknowledged automation-bias case cuts against an unqualified accuracy gain. Appendix A's 'not generalisable' response is an explicit self-limitation. The qualitative lessons (global explanations for regulators, local explanations 'essential' for trust, counterfactuals must be actionable) remain coherent and useful, so conditional acceptance with a reframed empirical claim is the right call. Hence I do not move the reader's verdict.","tokens_in":17926,"tokens_out":6736,"duration_ms":70223,"concrete_test":"Apply an exact binomial test to the six disagreement instances reported in Section 3.2: under the null hypothesis that aligning with AI is equally likely to improve or worsen accuracy, 5 or more improvements out of 6 has a one-sided p-value of 7/64 ≈ 0.11 (two-sided p ≈ 0.22). If the full 10×8 before/after decision matrix is supplied, run a paired McNemar test on clinician-level accuracy; either way, a non-significant result would require the abstract to be reframed. The attribution to local explanations would additionally require an arm showing only LIME explanations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract asserts that local explanations increase clinicians' trust and diagnostic accuracy. The pilot described in Section 3 does not support that causal attribution. In Section 3.1, each clinician first diagnoses from raw records, then is shown 'the corresponding AI diagnosis, along with explanations from the XAI models and the confidence level'; the XAI outputs include global explanations, LIME local explanations, and ExMatrix counterfactuals from Section 2.2. There is no arm that presents local explanations alone, nor an arm with the AI prediction and confidence but no explanation, so any change in behaviour cannot be attributed to local explanations specifically. Section 3.2 reports 'five out of six instances' where clinicians adjusted toward the AI prediction and 'an overall improvement in diagnostic accuracy', but no per-case or per-clinician accuracy table, no statistical test, and no control condition are provided. The acknowledged exception, where clinicians adopted an incorrect AI output, is labelled automation bias; whether the net effect is an increase in accuracy requires a paired analysis that is not reported. Finally, the HRA decision tool in Appendix A explicitly records 'No' to 'Are your findings going to be generalisable?', which is the paper's own admission that the empirical basis is not generalisable. The qualitative workshop insights may stand, but the abstract's empirical claim is overreach.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a multidisciplinary expert working group convened by the UK MHRA to evaluate AI/XAI-based clinical decision support for heart attack risk prediction. It presents three models (logistic regression, random forest, neural network) on a synthetic CPRD dataset, global, local, and counterfactual explanation methods, qualitative insights from regulators and clinicians, and a pilot study in which eight clinicians diagnosed ten synthetic records before and after seeing AI outputs. The headline empirical claim is that local explanations increase clinicians' trust and diagnostic accuracy, with an automation-bias caveat. The paper concludes with recommendations on model selection, explanation design, stakeholder training, and regulatory safeguards.","tokens_in":18114,"tokens_out":3931,"duration_ms":38374,"significance":"If the empirical claim were properly supported, the paper would offer useful evidence for AI-CDSS design and regulation. The workshop-based qualitative insights and the concrete recommendations around model selection, explanation presentation, and trust calibration are valuable, and the use of a public synthetic dataset is a strength. However, the central quantitative claim about local explanations improving accuracy and trust is currently not supported by the pilot design and analysis, and the paper's own HRA tool output says the findings are not generalisable. The paper would still be a useful multidisciplinary position piece if the empirical claims were appropriately downgraded and clearly labelled as exploratory.","major_comments":[{"comment":"The abstract states that 'the study reveals an overall increase in clinicians' trust and diagnostic accuracy when using local explanations,' but the pilot described in §3.1 and §3.2 cannot support a causal attribution to local explanations. Each clinician first diagnosed from raw records and was then shown the AI diagnosis, confidence, and all explanation types (global, LIME local, and ExMatrix counterfactual) together, with no control condition and no arm isolating local explanations. Appendix A explicitly records 'No' to generalisability, which is the paper's own admission that the empirical basis is not generalisable. The sentence should be revised to describe an exploratory observation or be supported by a properly controlled study.","section":"Abstract and §3.2"},{"comment":"The claim that 'five out of six instances' of clinician adjustment produced 'an overall improvement in diagnostic accuracy' is unsupported by the reported data. No per-case or per-clinician before/after accuracy table, no paired comparison, and no statistical test are provided, and the acknowledged automation-bias case could offset any accuracy gain. The authors should report the full contingency of clinician diagnoses before and after AI/XAI exposure and, if appropriate, a paired analysis; otherwise the conclusion should be limited to a qualitative observation.","section":"§3.2"},{"comment":"The comparison that 'random forest outperformed the others' is made from point estimates of sensitivity, specificity, precision, and AUC without confidence intervals, bootstrap resampling, or cross-validation variability. Because the workshop discussion and the pilot selection depend on model performance, the authors should either provide uncertainty measures for these metrics or soften the comparison to descriptive reporting.","section":"§2.2.1 and Table 1"}],"minor_comments":[{"comment":"The text and Table 2 use 'odd ratios' where 'odds ratios' is the correct term; this should be corrected throughout.","section":"§2.2.2 and Table 2"},{"comment":"There are typographical errors such as 'ans it considered the simplest model' and 'the complexiest'; these should read 'as it is considered' and 'the most complex.'","section":"§2.2.1"},{"comment":"The lesson numbering jumps from Lesson 5 to Lesson 7; the lessons should be renumbered or Lesson 6 should be added.","section":"§4"},{"comment":"The caption 'T able 3' and the phrase 'as in shown in Table 1' are formatting/grammatical issues that should be cleaned up.","section":"§2.3 and Table 3"},{"comment":"The sentence 'ensuring that they were provided with exactly the same data that has been used by the AI/ML models' is grammatically awkward and should be rephrased.","section":"§3.1"},{"comment":"References [23] and [27] appear to cite the same Herm et al. paper; this duplicate should be merged or disambiguated.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's own HRA decision tool output in Appendix A states that the findings are not generalisable, which directly undercuts the headline empirical claim in the abstract. I would ask the authors to align the abstract, results, and conclusions with the exploratory status of the pilot, or to provide a properly controlled study. Also, the same team built the models, produced the explanations, ran the workshops, and wrote the report; this role overlap does not by itself invalidate the qualitative insights, but the editor may want to ensure that the stakeholder views are reported transparently and not over-interpreted. There is no construction-circularity concern of the kind that would require rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, Read this one for the workshop material, not for the pilot's empirical claims. The genuinely new thing here is the joint framing: regulators, clinicians, and data scientists looking at the same model outputs and explanation visualizations together, in structured workshops, with a small pilot feeding their discussion. The regulatory insights — that global explanations matter most for model-level assessment, that variation in feature rankings across models is informative rather than disqualifying, that counterfactuals must be actionable — are sensible and clearly reported. That part is worth a serious referee. The paper does several things well. The ML setup is transparent at the level of models and data: logistic regression, random forest, ANN on a synthetic CVD dataset derived from CPRD, with 10-fold CV. The performance table and odds ratio confidence intervals are there. The tutorials given to the expert group are a good idea, and the authors honestly report that clinicians found LIME visualizations time-consuming and that the original ExMatrix was overwhelming. The HRA decision tool output is included verbatim in Appendix A, which is unusually candid. The soft spot is the abstract's sentence about increased trust and diagnostic accuracy. The pilot does not support a causal claim about local explanations. Eight clinicians, ten deliberately selected synthetic cases, no control arm, no explanation-only arm, no statistical analysis, and the paper's own HRA appendix says the findings are not generalisable. The stress-test note is right: the design shows participants the AI diagnosis, confidence, and all XAI outputs together, so any change in behavior cannot be attributed to local explanations specifically. The claim \"overall increase in diagnostic accuracy\" rests on five out of six instances of clinicians adjusting toward the AI, which is a small, uncontrolled observation, not an effect estimate. The authors should either reframe this as exploratory and prominently state the limitation in the abstract, or the abstract's conclusion should be weakened. The same pilot produced useful qualitative material about automation bias and feature alignment, and that material is fine as hypothesis-generating. Minor: no code or hyperparameters are released, and the ANN appears to underperform the logistic regression on this tabular task, which is worth a sentence of discussion since it bears on the simplicity preference they endorse. Citation pattern looks fine; the relevant XAI and trust calibration literature is represented. Bottom line: conditional accept if the empirical claim is reframed as exploratory pilot evidence. The workshop insights and recommendations are a legitimate contribution to the medical-device XAI discussion, especially given the MHRA involvement. I would send it to peer review and would cite it for the stakeholder-perspective material, not for the causal claim. Recommendation: accept peer review, with the abstract as the main revision target.","headline":"A useful multidisciplinary workshop report whose pilot evidence is being oversold in the abstract; the qualitative regulatory-clinical insights are the real contribution, not the causal claim about explanations improving accuracy.","tokens_in":724,"tokens_out":1261,"would_cite":true,"duration_ms":19950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that adding local explanations from LIME to a black-box clinical decision support system increases clinicians' trust and diagnostic accuracy while also producing over-reliance on AI suggestions, a safety concern for…","keywords":["explainable AI","clinical decision support systems","medical device regulation","trust calibration","LIME","counterfactual explanations","automation bias","pilot study"],"falsifier":"A randomized trial with a representative clinician sample and real or realistic patient cases, comparing a group that sees AI predictions with LIME explanations against a group that sees only the predictions, could settle the claim: if the explanation group shows no gain in diagnostic accuracy or a higher rate of accepting incorrect AI recommendations, the paper's central conclusion fails.","tokens_in":17676,"feed_emoji":"🩺","tokens_out":3697,"duration_ms":39211,"temperature":0.7,"pith_summary":"This paper tries to establish that local explainability, specifically LIME explanations attached to a black-box neural network, improves clinicians' trust and diagnostic accuracy when they use a clinical decision support system, but that the same explanations create measurable over-reliance on AI suggestions. The evidence comes from an expert working group of regulators, clinicians, and data scientists, plus a pilot study in which eight clinicians diagnosed ten synthetic patient records before and after seeing AI predictions with explanations. The paper argues that regulators and clinicians need different kinds of transparency, and it offers recommendations for safely adopting AI decision support, including preferring simpler models when performance is comparable and using global and counterfactual explanations carefully.","feed_headline":"Local AI explanations boost accuracy but fuel over-reliance","feed_subtitle":"A pilot with eight clinicians shows trust rises with explanations, yet wrong AI calls were adopted too.","key_machinery":"The central object is the pilot study design: eight practicing clinicians each reviewed ten synthetic patient records, made an initial diagnosis, then saw the neural network's prediction with LIME local explanations and confidence scores, and were allowed to revise their diagnosis. LIME serves as the local explainer, generating feature-importance bars by fitting an interpretable surrogate model around each instance, while ExMatrix provides counterfactual visualizations from the random forest. The before-and-after comparison of clinician diagnoses, interpreted by the expert group, is the mechanism that carries the argument.","core_discovery":"On its own terms, the paper's central claim is that local explanations increase clinicians' confidence and diagnostic accuracy, especially for low-risk cases, but that this benefit is accompanied by automation bias: in five of six instances where AI and clinician diagnoses differed, clinicians shifted their diagnosis toward the AI, including one case where the AI was wrong. The discovery is that trust calibration, rather than explanation fidelity alone, is the key safety variable in AI-based clinical decision support, and that transparency needs diverge between regulators, who value global feature importance for model selection, and clinicians, who value local and counterfactual explanations for individual cases.","pith_inferences":["The pilot's lack of a control condition and its small self-selected sample mean the observed accuracy gain could plausibly come from task familiarity or regression to the mean rather than from explanations; a controlled replication is needed before acting on the recommendation.","If explanations increase agreement with AI, they may amplify systematic model errors in biased or under-represented subgroups; regulators might therefore require subgroup-specific interaction testing, not just overall accuracy.","Local explanations such as LIME may be better suited as training tools for junior clinicians than as real-time decision aids for experts, since time pressure was reported as a barrier to using them in practice.","The observed divergence between clinician and model feature rankings suggests a testable extension: showing clinicians the model's top features before their own diagnosis could either reduce anchoring bias or increase it, depending on how the information is framed."],"forward_implications":["If local explanations improve accuracy but induce over-reliance, then AI clinical decision support should include safeguards such as override prompts, uncertainty display, and training in when to discard AI recommendations.","Regulators should evaluate human-AI interaction evidence, not just model performance metrics, when approving AI-based medical devices.","Divergent global feature rankings across models mean model selection should involve clinical review of top-ranked features, not accuracy metrics alone.","Counterfactual explanations that suggest changing immutable features like age or medical history are clinically unusable and should be filtered before presentation.","Simpler, interpretable models with comparable performance are preferable to black-box models in high-risk clinical settings."],"supporting_citations":[{"why":"Supplies the LIME method used to produce local explanations in the pilot study.","marker":"[32]"},{"why":"Grounds the claim that different explanation classes impact trust calibration in clinical decision support.","marker":"[17]"},{"why":"Documents over-reliance and errors in physician-AI collaboration, which the paper's automation-bias finding extends.","marker":"[5]"},{"why":"Provides the ExMatrix counterfactual visualization method evaluated by the expert group.","marker":"[52]"},{"why":"Supplies the high-fidelity synthetic dataset used for model training and the pilot study cases.","marker":"[41]"},{"why":"Frames the UK regulatory context for AI-enabled medical devices that the recommendations address.","marker":"[10]"}],"fun_headline_variants":["AI explanations aid diagnosis but skew trust","Local AI explanations improve diagnosis but cause automation bias","Explainable AI: accuracy up, over-reliance down? Not quite","Trust calibration key for clinical AI, not just explanation fidelity","MHRA group: transparency alone won't fix AI over-reliance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the behavior of eight self-selected clinicians on ten deliberately chosen synthetic patient records, with no control condition and no statistical analysis, represents what clinicians in general would do when given local explanations.","fun_headline_variants_meta":{"raw":{"variants":["AI explanations aid diagnosis but skew trust","Local AI explanations improve diagnosis but cause automation bias","Explainable AI: accuracy up, over-reliance down? Not quite","Trust calibration key for clinical AI, not just explanation fidelity","MHRA group: transparency alone won't fix AI over-reliance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000569,"raw_usage":{"total_tokens":2640,"prompt_tokens":837,"completion_tokens":1803,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":1723}},"tokens_in":453,"tokens_out":1803,"duration_ms":12154,"temperature":1.0,"reasoning_tokens":1723,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:37:48.300207+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A randomized trial with a representative clinician sample and real or realistic patient cases, comparing a group that sees AI predictions with LIME explanations against a group that sees only the predictions, could settle the claim: if the explanation group shows no gain in diagnostic accuracy or a higher rate of accepting incorrect AI recommendations, the paper's central conclusion fails.","supporting_citations":[{"cited_title":"In: Third International Workshop on Learning with Imbalanced Domains: Theory and Applications, pp","cited_arxiv_id":null,"evidence_quote":"Supplies the high-fidelity synthetic dataset used for model training and the pilot study cases."},{"cited_title":"International Journal of Human-Computer Studies 169, 102941 (2023)","cited_arxiv_id":null,"evidence_quote":"Grounds the claim that different explanation classes impact trust calibration in clinical decision support."},{"cited_title":"Available at SSRN 4773350 (2024)","cited_arxiv_id":null,"evidence_quote":"Documents over-reliance and errors in physician-AI collaboration, which the paper's automation-bias finding extends."},{"cited_title":"IEEE Transactions on Visualization and Computer Graphics 27(2), 1427–1437 (2020) 46 Appendix A: Ethical Approval Requirements for the Pilot Study) Go straight to content","cited_arxiv_id":null,"evidence_quote":"Provides the ExMatrix counterfactual visualization method evaluated by the expert group."},{"cited_title":"In: Proceedings of the First International Symposium on Trustworthy Autonomous Systems","cited_arxiv_id":null,"evidence_quote":"Frames the UK regulatory context for AI-enabled medical devices that the recommendations address."}],"review_version":1}