{"id":"8135c189-4883-4b40-8864-6617846ac961","arxiv_id":"2412.01829","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review that categorizes XAI techniques and surveys medical applications across vision, audio, and multimodal data, with an analysis of current challenges.","lead":"This paper reviews recent research on explainable artificial intelligence in medicine, covering visual, audio, and multimodal applications. It organizes the field with a taxonomy of 19 XAI techniques and points out where more work is needed.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed scope of 'over 100 papers' is contradicted by the tables: counting Tables 2–4 yields ~87 unique studies, so the review's comprehensiveness claim is unsupported as stated.","rationale":"The reader's verdict of CONDITIONAL is appropriate, and the identified concerns (the 'over 100' mismatch and selection bias) are both real. However, I judge the most load-bearing concern to be the concrete numerical discrepancy, because it is a directly falsifiable internal inconsistency: the text claims a scope that the tables do not support. Selection bias is a legitimate but more speculative concern, as it depends on what a systematic review would have found; the count issue is demonstrable from the manuscript alone. This discrepancy undermines the central claim's comprehensiveness component and should be corrected by either revising the number or expanding the survey. The taxonomy framework and the qualitative trends may still hold, so a conditional acceptance requiring this correction and a documented selection protocol is the right outcome. My view does not change the reader's verdict, hence UNCHANGED; I partially agree with the reader's choice of weakest assumption because both issues were flagged in the reader's rationale, but I prioritize the count discrepancy over selection bias as the single most load-bearing point.","tokens_in":34550,"tokens_out":7335,"duration_ms":67294,"concrete_test":"Collate all reference numbers listed in Table 2, Table 3, and Table 4 (including starred variants), deduplicate them, and count the total. If the total is less than 100, scan Section 4's text for any additional analyzed papers that are not assigned a table row; if none exist, the 'over 100 papers' claim is false. Then recompute the descriptive statistics from Section 5 (e.g., the majority perception-based, post-hoc, model-specific trends) using the corrected dataset to confirm the qualitative conclusions remain unchanged.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The introduction states 'we analyse over 100 papers published in the past five years' and Section 7 reiterates that 'a multitude of XAI application cases... were gathered.' This quantitative claim is directly contradicted by the paper's own tables. Counting the rows in Table 2 (visual) gives 71 unique entries, Table 3 (audio) gives 5, and Table 4 (multimodal) gives 11, for a total of 87 unique references. Section 4's narrative discussion adds no analyzed studies beyond those appearing in the tables; background works cited in the audio subsection (e.g., [1, 42, 6, 35, 67, 100, 118, 153, 160]) are not XAI application papers and are not analyzed. Thus the central claim of analyzing 'over 100 papers' is factually unsupported. This is not a stylistic vagueness but an internal inconsistency between the stated scope and the evidence presented. The discrepancy also signals that the selection process described in Section 4 (subjective criteria: anatomical location, technique diversity, timeframe) may not have been systematically tracked against a defined inclusion count, which reinforces the reader's representativeness concern. While the taxonomy framework may be useful, the review's comprehensiveness, a key part of its claimed contribution, is overstated as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a narrative review of explainable artificial intelligence (XAI) in medical applications. It consolidates four existing taxonomy criteria (perception/mathematics, ante-hoc/post-hoc, model-agnostic/model-specific, local/global) into a framework of 19 XAI techniques, assigns those techniques to the criteria in Table 1, and organizes recent medical XAI studies into three application domains: visual (Table 2), audio (Table 3), and multimodal (Table 4). The authors state in Section 1 that they analyze over 100 papers from the past five years, and in Section 7 they claim to have gathered a 'multitude' of application cases spanning visual, audio, and multimodal domains. The paper also discusses challenges and future directions, including standardization, quantitative evaluation, formal XAI, and multimodal/personalized explanations.","tokens_in":34811,"tokens_out":5817,"duration_ms":53382,"significance":"If the central claims were fully supported, the consolidated taxonomy and cross-modality synthesis would be a useful entry point for researchers and medical practitioners navigating the XAI literature. The paper's Table 1 and Figure 3 provide a compact comparison of 19 XAI techniques, and the division into visual, audio, and multimodal applications is a sensible organizing principle. However, the review's quantitative claim of analyzing over 100 papers is contradicted by its own tables, which enumerate 87 unique studies, and the selection of studies is explicitly subjective as described in Section 4. These issues directly affect the claimed comprehensiveness and the generality of the empirical trends reported in Section 5. The taxonomy itself is a reasonable contribution, but the review's scope claims and the evidentiary support for its observations need to be reconciled with the presented data.","major_comments":[{"comment":"Section 1 states 'we analyse over 100 papers published in the past five years', but Tables 2, 3, and 4 enumerate 71, 5, and 11 unique studies, respectively, totalling 87; no additional analyzed studies appear outside these tables, and the background citations in §4.2 (e.g., [1], [42], [6], [35], [67], [100], [118], [153], [160]) are not XAI application papers. The quantitative claim is therefore unsupported by the evidence presented, and the authors should either expand the included studies to exceed 100 or revise the claim to match the actual count.","section":"§1, Introduction"},{"comment":"In Section 4 the authors write that they 'selected the representative XAI medical applications based on medical anatomical locations, diversity of XAI techniques, and publication timeframe' and that this selection rests on 'our subjective assessment'. No inclusion/exclusion criteria, search results, or screening decisions are reported, so the study set is not reproducible. Because the cross-domain trends in §5.2 (e.g., that audio and multimodal XAI are 'less abundant') are derived from this selected sample rather than from a systematic census, the observed imbalances may be artifacts of the selection rather than robust properties of the literature. The authors should either supply a reproducible selection protocol with a defined inclusion count or moderate the generality of the empirical observations.","section":"§4, XAI Applications in Medicine Review"},{"comment":"Section 5.1 makes unquantified frequency claims such as 'the majority of XAI techniques used in the medical field are perception-based' and 'The majority of XAI techniques currently used in medical AI applications are model-specific' without providing an aggregation of the per-study technique entries in Tables 2-4. Table 1 is a taxonomy of techniques rather than an application survey and cannot by itself substantiate these frequency statements. The authors should add explicit counts or a summary table cross-tabulating the reviewed studies by the four taxonomy criteria, or replace these claims with hedged observations that reflect the small, subjectively selected sample.","section":"§5.1, Observation on Criteria Taxonomy and Techniques"}],"minor_comments":[{"comment":"The phrase 'past five years' is inconsistent with the search window '2018 and 2024' described in Section 4, which spans six calendar years; please align the wording.","section":"§1, Introduction"},{"comment":"Several typographical errors need correction: 'according approaches' (§1), 'explainablility' (§2.2), 'we are not reiterate' (§3.1), 'a implementation-based approaches' (§7), 'Data Sacarcity' (Table 4), and 'In 2022, 10% of patients assumed they did not feel involved' (§2.1).","section":"Multiple sections"},{"comment":"In Table 1, multiple techniques receive checks in both members of a binary criterion (e.g., CAM is checked for both Perception and Mathematics, and SM for both Local and Global); because §3.3 introduces each pair as an alternative, please add a note explaining that the criteria are treated as multi-label axes.","section":"Table 1"},{"comment":"The 'Cons' column in Tables 2-4 is not defined; please spell it out in the table footnotes as, for example, 'shortcomings identified by the authors from the original paper's limitations section and their own assessment'.","section":"Tables 2-4"},{"comment":"The Boolean search string in §4 ('XAI' OR 'explainable artificial intelligence' AND 'healthcare' OR 'medicine') is missing parentheses and therefore has an ambiguous parse; adding explicit grouping would clarify the intended query.","section":"§4, Selection methodology"}],"recommendation":"major_revision","confidential_remarks":"The paper's material contribution lies in its consolidated taxonomy and the organized tables; the empirical framing as a systematic survey of 'over 100 papers' is not supported by the presented data. In my view the issues are correctable within the scope of a revision: the count should be reconciled, the selection procedure should be made reproducible or the generality claims moderated, and the Section 5 observations should be supported with explicit aggregations of the tables. Because these points bear on the review's claimed comprehensiveness, I would not accept the manuscript in its current form, but I do not see a fundamental barrier to revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this one before you cite it: the review's central quantity claim is off. The intro and conclusion say they analyse over 100 recent papers, but Tables 2–4 list 71 visual, 5 audio, and 11 multimodal studies — 87 unique entries, and the narrative doesn't add any analyzed studies beyond the tables. So \"over 100\" is simply not supported by the evidence in front of us. That's a real inconsistency, and it matters because comprehensiveness is part of what the authors claim as their contribution. I'd want it fixed before publication, but it's not load-bearing: the taxonomy and the survey are still usable without that number.\n\nWhat's actually good: the paper consolidates four taxonomy criteria (perception, ante-hoc/post-hoc, model-agnostic/specific, local/global) to classify 19 XAI techniques, and then organizes recent medical applications by modality — vision, audio, multimodal. The tables are informative and the coverage of audio and multimodal work is a genuine gap-filler next to earlier reviews from Loh, Band, Singh, and Chaddad. The authors also explicitly acknowledge that the classification in Table 1 reflects their subjective reading, which is honest. The discussion of challenges (data quality, generalizability, evaluation without standardized metrics) is sensible and well-grounded in the cited literature.\n\nSoft spots beyond the overcount: the study selection criteria are described as \"representative\" and based on anatomical location, technique diversity, and timeframe, but there's no systematic protocol or a PRISMA-style flow. That makes the observed trends (e.g., dominance of post-hoc gradient methods) suggestive rather than generalizable. The authors do flag this as subjective, so it's a transparency issue more than a fatal one. Some table entries use starred variants (Grad-CAM*) without a clear note on which variant each study used, but that's minor.\n\nBottom line: this is a useful orientation piece for researchers entering medical XAI, and the taxonomy alone is worth having. It deserves a proper peer review — not a desk reject — but the referee should insist on reconciling the \"over 100\" claim with the tables and on documenting the selection logic. If that gets cleaned up, I'd cite it. I'd bring it to reading group as a maybe: useful for context, not a must-read.","headline":"Useful survey with a solid taxonomy, but the 'over 100 papers' claim doesn't match the paper's own tables — a fixable overstatement, not a fatal flaw.","tokens_in":35297,"tokens_out":1624,"would_cite":true,"duration_ms":17622,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that medical AI can be made accountable with a single four-criteria taxonomy grid spanning 19 explanation techniques, and that current practice skews toward visual, post-hoc, model-specific methods.","keywords":["Explainable Artificial Intelligence (XAI)","Machine Learning","Literature Review","Medical Information System","medical imaging","audio XAI","multimodal XAI","post-hoc explanation"],"falsifier":"A formal systematic review of medical XAI published from 2018 to 2024 with a pre-registered protocol that found audio and multimodal explanation studies about as numerous as visual ones, or that found most techniques ante-hoc rather than post-hoc, would directly contradict the trends the paper reports.","tokens_in":34376,"feed_emoji":"🩺","tokens_out":4018,"duration_ms":36920,"temperature":0.7,"pith_summary":"This review argues that medical AI's accountability gap can be addressed by a structured map of explanation techniques, and that the field currently over-relies on visual, post-hoc, model-specific methods. It consolidates four existing taxonomy criteria into one framework covering 19 XAI techniques. It then uses that framework to organize over 100 recent medical XAI studies across vision, audio, and multimodal data, drawing out trends and limitations. The value of the work is giving future researchers a shared vocabulary and grid for comparing explanation methods in clinical settings.","feed_headline":"Review maps 19 XAI techniques across 100+ medical studies","feed_subtitle":"Four taxonomies merge into one grid, exposing medicine's reliance on visual, post-hoc explanations.","key_machinery":"The central organizing device is a four-criteria taxonomy grid, adapted from prior frameworks: perceptive interpretability versus interpretability by mathematical structures, ante-hoc versus post-hoc, model-agnostic versus model-specific, and local versus global. The paper maps 19 XAI techniques onto this grid and groups them by implementation principle into perturbation-based, backpropagation-based, gradient-based, instance-based, and other approaches. This grid carries the review's arguments by letting the authors compare disparate medical studies on the same terms and extract trends about what the field is and is not doing.","core_discovery":"On its own terms, the paper's discovery is that the scattered XAI literature for medicine can be unified: the four criteria (perceptive versus mathematical, ante-hoc versus post-hoc, model-agnostic versus model-specific, local versus global) are jointly enough to describe 19 commonly used techniques such as LIME, SHAP, Grad-CAM, LRP, and counterfactuals. Classifying those techniques by implementation principle into perturbation-based, backpropagation-based, gradient-based, instance-based, and other approaches reveals that most medical deployments are perceptive, post-hoc, and model-specific. Applied to over 100 studies from 2018 to 2024, the framework shows that visual explanations dominate, audio and multimodal XAI are underdeveloped, and no single technique covers all four criteria.","pith_inferences":["The same four-criteria grid could be applied outside medicine to high-stakes domains such as credit or criminal justice, where the local/global distinction maps onto individual decisions versus policy-level audits; the authors do not draw this comparison.","A testable extension of the paper's scarcity claim is that audio XAI would grow faster if explainability shifted from pixel-level heatmaps to time-frequency segmentations and listenable sonified explanations, a direction the authors mention only as outlook.","The paper's own selection logic could be stress-tested by a formal systematic review with pre-registered inclusion criteria; the authors explicitly state that their choices were based on anatomical location, technique diversity, and publication timeframe rather than a systematic protocol."],"forward_implications":["Researchers entering medical XAI can use the 19-technique grid as a checklist, making it easier to see which explanatory perspective a new method covers and which it leaves out.","The observed dominance of visual, post-hoc, gradient-based explanations implies that audio and multimodal explanations are an open, comparatively empty niche for new work.","The finding that explanations are rarely combined suggests that multi-technique and multi-criteria evaluation, rather than any single saliency map, is the route to clinically usable explanations.","The paper's outlook implies that standardizing terminology and adopting quantitative evaluation frameworks would make medical XAI comparisons more meaningful."],"supporting_citations":[{"why":"Supplies the perceptive-versus-mathematical interpretability taxonomy criterion.","marker":"[183]"},{"why":"Supplies the ante-hoc versus post-hoc distinction for explanation timing.","marker":"[70]"},{"why":"Supplies the model-agnostic versus model-specific classification criterion.","marker":"[34]"},{"why":"Supplies the local versus global explanation criterion and grounds tree-based global explanation with SHAP.","marker":"[106]"},{"why":"Provides the DARPA definition of XAI that the review adopts for medical contexts.","marker":"[61]"},{"why":"Defines LIME, a prominent perturbation-based technique in the framework.","marker":"[152]"},{"why":"Defines SHAP, a prominent game-theoretic technique in the framework.","marker":"[107]"}],"fun_headline_variants":["Medical XAI: one grid unifies 19 techniques, 100+ studies","Visual explanations dominate medical AI; audio lags behind","No single XAI method covers all needs, review finds","Four criteria classify 19 medical AI explanation tools","Explainable AI review maps the field, spots gaps in audio"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's trend claims rest on the assumption that its chosen studies, picked for anatomical coverage, technique variety, and recency rather than through a formal systematic protocol, represent the whole body of medical XAI work.","fun_headline_variants_meta":{"raw":{"variants":["Medical XAI: one grid unifies 19 techniques, 100+ studies","Visual explanations dominate medical AI; audio lags behind","No single XAI method covers all needs, review finds","Four criteria classify 19 medical AI explanation tools","Explainable AI review maps the field, spots gaps in audio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000601,"raw_usage":{"total_tokens":2788,"prompt_tokens":909,"completion_tokens":1879,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1795}},"tokens_in":525,"tokens_out":1879,"duration_ms":14776,"temperature":1.0,"reasoning_tokens":1795,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:56:35.208942+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A formal systematic review of medical XAI published from 2018 to 2024 with a pre-registered protocol that found audio and multimodal explanation studies about as numerous as visual ones, or that found most techniques ante-hoc rather than post-hoc, would directly contradict the trends the paper reports.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the perceptive-versus-mathematical interpretability taxonomy criterion."}],"review_version":1}