{"id":"bab48906-1220-46f1-a78b-cb6e325bde33","arxiv_id":"2508.10806","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of 79 XAI evaluation studies shows almost no involvement of visually impaired users, and a three-person prototype study suggests simplified multimodal explanations work best for non-visual users.","lead":"An analysis of 79 explainable AI studies finds that almost none evaluate explanations with blind or visually impaired users, and most explanations are visual. A small prototype with three sight-loss users suggests simplified, multimodal explanations may be easier to understand, but the evidence is only preliminary.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"§2 search string may miss accessible-XAI studies using different vocabulary, so the 1-in-79 gap could be overstated.","rationale":"The central claim is that XAI evaluations rarely include disabled users and mostly rely on visual explanations. The empirical basis is the literature review's 79-study corpus, selected by a single search string. The reader identified this search string as the weakest assumption, and I agree. The search string is the gate through which all evidence flows; if it is too narrow, both the inclusion statistic and the visual-format claim are undermined. The paper does not report the kind of methodological safeguards (per-database strings, protocol registration, dual coding) that would make the retrieval trustworthy. My concrete test—an expanded search with accessibility-related terms—directly probes whether studies were missed. If the expanded search finds such studies, the paper's headline gap is overstated; if not, the claim stands. Because this concern is exactly what the reader flagged, and because it warrants the already-assigned CONDITIONAL verdict, I do not change the verdict.","tokens_in":25322,"tokens_out":8085,"duration_ms":85452,"concrete_test":"Re-run the §2 search over the same seven databases and 2019–2024 window using an expanded block: (XAI OR 'explainable AI' OR 'interpretable machine learning' OR 'model explanation' OR 'explainability') AND (accessib* OR disability OR blind OR 'visual impairment' OR 'screen reader' OR 'non-visual' OR 'assistive technology' OR 'low vision') AND (user OR evaluation OR usability OR co-design). Compare the retrieved set with the original 79. If additional studies that evaluate XAI with visually impaired or disabled participants emerge, the claim that such evaluations are nearly absent is weakened; if no new studies appear beyond the original set, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The literature review in §2 uses the single string (XAI OR “AI explanation”) AND (end-user OR user) AND (evaluation OR validation) over 2019–2024. This presumes that all relevant studies self-identify as “XAI” or “AI explanation” and appear under “end-user/user” plus “evaluation/validation.” Research on accessible AI explanations may instead use vocabulary such as “blind,” “screen reader,” “non-visual,” “assistive technology,” “low vision,” or “interpretable machine learning,” and may describe a “user study” or “co-design” rather than “evaluation”/“validation.” The paper reports no per-database search variants, truncation, or Boolean expansions, and provides no protocol registration or inter-coder reliability. If even a small number of accessible-XAI evaluation studies were missed, the headline statistic (1 of 79 included studies; 3 of 79 mentioning accessibility) would overstate the field's neglect. The load-bearing premise is that the search string defines the population of XAI evaluation studies; that premise is unverified and could be wrong.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses accessibility in explainable AI (XAI), focusing on users with vision impairments. It reports a PRISMA-style literature review of 79 studies (2019–2024) that evaluate XAI techniques with end users, concluding that only one included study involves participants with visual impairment, three mention accessibility concerns, and most explanations rely on inherently visual formats. The authors then present a four-part methodological proof of concept: a categorization of AI systems, a persona (Caroline) created with CNIB input, a functional prototype implementing simplified and detailed LIME/SHAP explanations with screen-reader support, and an evaluation with accessibility experts and users with lived experience of sight loss. Based on three active participants, the user evaluation suggests that simplified explanations are more comprehensible than detailed ones for non-visual users and that multimodal presentation is needed. The authors explicitly label the user study preliminary.","tokens_in":25531,"tokens_out":5486,"duration_ms":62897,"significance":"If the prevalence claim is robust, the finding that only 1 of 79 XAI evaluation studies includes participants with visual impairment is important and actionable for the XAI and HCI communities. The proof of concept is concrete, involves collaboration with accessibility experts from CNIB, follows ARIA practices, and was tested with several screen readers; this provides a useful template for future inclusive XAI design. The paper is honest about the small sample and preliminary nature of the user results. Its main weakness is that the literature-review methodology and the modality-based conclusions are not reported in sufficient detail to fully support the central prevalence claim.","major_comments":[{"comment":"The single search string (XAI OR “AI explanation”) AND (end-user OR user) AND (evaluation OR validation) may systematically miss accessible-XAI studies that use different vocabulary, such as “screen reader”, “non-visual”, “blind users”, “assistive technology”, “low vision”, or “interpretable machine learning” combined with “user study” or “co-design”. Because the headline statistic (79 studies; 1 with visual-impairment participants) is the main evidence for the claimed gap, this is load-bearing. The paper does not report per-database search strings, truncation or date filters, inclusion/exclusion decision details, or screening reliability. The authors should expand and validate the search, report a protocol, and ideally add a citation/screening check against known accessible-XAI work; otherwise the 1/79 figure may overstate the field’s neglect.","section":"§2, search string and PRISMA flow"},{"comment":"The claim that “most explanations rely on inherently visual formats” is not supported by systematic modality coding. Table S1 records XAI techniques, evaluation methods, metrics, and evaluators, but it does not code whether each study’s explanations were visual-only, textual, audio, haptic, or multimodal. Inferring visual dependence from technique names (SHAP, LIME, Grad-CAM) conflates the default implementation with the actual presentation used in the study; LIME and SHAP can be rendered as text or speech. A separate analysis or an additional column in Table S1 that operationalizes “inherently visual” is needed before this general conclusion is stated.","section":"§2 and Table S1"},{"comment":"The secondary claim—simplified explanations are more comprehensible for non-visual users than detailed ones—rests on three active participants and qualitative agreement, with no coding scheme, inter-rater reliability, or quantitative comparison. The authors correctly call this preliminary in the Conclusion, but the abstract presents it as a finding. Either report the evidence with the caveat directly in the abstract, or frame the user component as an illustrative co-design session rather than an empirical evaluation result. This does not invalidate the proof of concept, but it must not be read as a measured comparison of explanation formats.","section":"§3.4, user evaluation"}],"minor_comments":[{"comment":"The “randomly selected combination” of AI categories is not reproducible. Specify the random selection procedure or state explicitly that it was an arbitrary illustrative choice.","section":"§3.1"},{"comment":"The pseudocode contains informal lines such as “Import lib imports”. Clean up the pseudocode and use consistent notation for functions and data structures.","section":"§3.3, Algorithm 1"},{"comment":"The row for Gunning and Aha [2019] lacks a concrete evaluation method and metrics; use “N/A” or clarify what was extracted. Also, some evaluator counts are missing (“quantity not provided”); mark these consistently.","section":"Table S1"},{"comment":"The expert evaluation with six CNIB experts is described only narratively. While useful, a short summary of how expert feedback was recorded and aggregated would improve transparency.","section":"§3.4"},{"comment":"Some URLs appear as footnotes and some as reference entries; format consistently according to the venue style.","section":"References and footnotes"}],"recommendation":"major_revision","confidential_remarks":"The core directional claim—that XAI evaluation rarely includes users with visual impairments—is plausible and important, but the review methodology currently does not fully support the precise 1/79 prevalence figure. The user-study evidence is appropriately labeled preliminary, but the abstract slightly overstates it. I recommend major revision with emphasis on expanding the search strategy, adding modality coding, and reporting screening reliability. The self-citation to the authors’ prior survey is acceptable, but the new review should be presented as an independent data collection with its own protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper gives us a concrete number for a gap we all suspected—XAI evaluation is heavily visual and almost no one includes users with sight loss—and it reports its own work honestly. But the number (1 of 79 studies) rests on a search string that is probably too narrow, and the user study is three people. Don't take the number as gospel, but do take the paper seriously as a proof of concept.\n\nWhat's genuinely new: the PRISMA review of 79 end-user XAI evaluation studies, with a table breaking down techniques, metrics, and participant groups. As far as I know, no one has tabulated that specific slice before. The finding that only one study includes participants with a visual impairment (and even that one doesn't adapt the explanations) is a useful, citable piece of evidence. The prototype itself is also a real contribution: a working traffic-flow dashboard with LIME/SHAP in simplified and detailed forms, built with screen-reader compatibility in mind and refined with input from CNIB experts. That's a good template for future work.\n\nSoft spots, in order of size. First, the search string: (XAI OR \"AI explanation\") AND (end-user OR user) AND (evaluation OR validation). That presumes any accessible-XAI study would self-identify with those terms. A paper about non-visual explanations might say 'screen reader' or 'assistive technology' rather than 'XAI', and 'user study' rather than 'evaluation' or 'validation'. The authors don't report per-database search strings, truncation, or a protocol, and no coding reliability is given. So the 1-in-79 statistic is a lower bound on what the field is doing, not a precise measure. The qualitative claim—'rarely'—probably survives a broader search, but the specific number is fragile. This is the one part I'd want fixed before using the number in a paper.\n\nSecond, the co-design had three active participants. The authors are honest about this, and they call the result preliminary. That's fair. But the abstract still says 'preliminary findings suggest simplified explanations are more comprehensible,' which reads as a finding. I'd rather see that phrased as 'anecdotal' or 'hypothesis-generating.' The evaluation has no baseline, no quantitative measure, just qualitative feedback. That's fine for a proof of concept, not for a claim.\n\nMinor: the prior survey from the same group (Nwokoye et al. 2024) is cited to motivate the gap, and the new review is independent but overlapping. It would have been tidy to state the relationship explicitly.\n\nSo: this is a paper for the XAI and accessibility communities. It deserves a serious peer review—the gap measurement is worth having, and the prototype is a useful design example. The literature review methodology needs to be tightened, and the user-study conclusions need to be softened or supported with a larger sample. If you're looking at whether to engage, I'd say yes—as a citable proof of concept, not as a definitive empirical claim.","headline":"A useful, honestly-reported gap analysis and prototype for accessible XAI, but the headline statistic rests on a search string that may miss relevant work and the user study has three participants.","tokens_in":26078,"tokens_out":4026,"would_cite":true,"duration_ms":39598,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the XAI evaluation literature is built around sighted users: only 1 of 79 end-user studies includes participants with visual impairment, and the dominant techniques default to visual formats.","keywords":["explainable AI","accessibility","visual impairment","non-visual explanations","SHAP","LIME","user evaluation","inclusive design"],"falsifier":"Run the same seven-database review with search terms expanded to include 'screen reader,' 'assistive technology,' 'visual impairment,' 'accessible explanation,' and 'non-visual'; if this surfaces more than a handful of end-user XAI evaluations that include blind or low-vision participants, the claimed 1-in-79 gap is overstated. Separately, a comprehension experiment with a larger sample comparing simplified and detailed screen-reader LIME and SHAP explanations would test the preliminary finding.","tokens_in":25177,"feed_emoji":"♿","tokens_out":8551,"duration_ms":83103,"temperature":0.7,"pith_summary":"The paper sets out to show that explainable AI has an accessibility gap: across 79 end-user evaluation studies published from 2019 to 2024, only one includes participants with any visual impairment, and only three mention accessibility at all. The dominant techniques—SHAP, LIME, Grad-CAM, decision trees—are evaluated and displayed in visual formats such as bar charts, heatmaps, and graphs, so the field's evidence about 'understandable AI' is built almost entirely on sighted users. To show the gap can be closed, the authors build a screen-reader-friendly traffic-prediction prototype offering simplified and detailed LIME and SHAP explanations, then assess it with accessibility experts and three users with lived experience of sight loss. Their preliminary finding is that simplified explanations are understood better than detailed ones by non-visual users, and that multimodal presentation is needed for equitable interpretability.","feed_headline":"Only 1 in 79 AI explanation studies includes disabled users","feed_subtitle":"A seven-database review finds XAI relies on charts and graphs, locking out the 2.2 billion people with vision loss.","key_machinery":"The argument runs on two machines. First, a systematic review following a PRISMA flow diagram, with search string (XAI OR \"AI explanation\") AND (end-user OR user) AND (evaluation OR validation) across seven databases, whose exclusion criteria leave 79 studies; this supplies the 1-in-79 and 3-in-79 counts that define the gap. Second, a four-part methodological proof of concept—categorization of AI systems, persona definition, prototype implementation, expert and user assessment—that turns the gap into a testable design: a web prototype with a machine-learning traffic-flow model and simplified and detailed LIME and SHAP explanations, made accessible through screen readers, keyboard navigation,","core_discovery":"The central claim is that accessibility, specifically for people with vision impairments, is missing from the empirical XAI literature and from the default design of XAI techniques. In a review of 79 end-user evaluation studies, only three mention accessibility concerns and only one includes participants who report visual impairment; 33 of the 79 evaluate SHAP and 30 evaluate LIME, whose typical outputs are visual. The paper argues that visual-only explanations block users from contesting biased AI decisions and exclude them from meaningful participation in AI governance. Based on a small co-design session, it further argues that simplified explanations outperform detailed ones for non-visua","pith_inferences":["A testable extension: re-run a standard task-based XAI evaluation, such as decision accuracy with SHAP versus LIME, using screen-reader output and compare outcomes to published sighted results; if accuracy gaps appear, existing XAI metrics may be measuring visual literacy as much as understanding.","The preliminary finding implies that adding alt-text or image descriptions to existing charts may not be enough; non-visual comprehension may require a fundamentally different representation, such as sonified feature weights or narrative text.","If the review's gap is real, it also bears on transparency regulation: an explanation that cannot be perceived is arguably no explanation at all, so accessibility standards may need to be read into AI transparency duties.","Because the review used one narrow search string, a broader replication using terms like 'screen reader,' 'assistive technology,' and 'accessible explanation' would show whether the gap is in the literature or partly in the search vocabulary."],"forward_implications":["If the review's counts hold, published claims about how well users understand SHAP, LIME, and other XAI techniques apply to sighted users only; generalizing them to the roughly 2.2 billion people with near or distance vision impairment is unsupported.","XAI technique choice would have to treat accessibility as a first-class criterion alongside model-agnosticism and fidelity, since the default outputs of the most popular methods are visual.","Explanation design for non-visual users should offer simplified overviews first and detailed content as an option; the preliminary result says details without an accessible overview hurt comprehension.","Multimodal presentation—text, audio, simplified point form, linear charts over tables—becomes a requirement for equitable access, not an enhancement.","The four-part proof of concept offers a reusable template for evaluating any AI system's explanations with disabled users before deployment."],"supporting_citations":[{"why":"Prior survey cited for the claim that accessible XAI research is nearly nonexistent; it frames the motivating gap the review investigates.","marker":"[Nwokoye et al., 2024]"},{"why":"The one study in the 79-paper sample whose participants report visual impairment; it serves as the exception that defines the gap.","marker":"[Labarta et al., 2024]"},{"why":"Supplies the PRISMA flow methodology used to identify and screen the 523 candidate records down to 79 studies.","marker":"[Haddaway et al., 2022]"},{"why":"Defines LIME, one of the two XAI techniques implemented in the prototype and one of the most evaluated techniques in the review.","marker":"[Ribeiro et al., 2016]"},{"why":"Defines SHAP, the other prototype technique and the most frequently evaluated method in the 79 studies.","marker":"[Lundberg and Lee, 2017]"}],"fun_headline_variants":["Just 1 of 79 AI explanation studies includes disabled users","AI explainability research barely includes disabled users","Visual AI explanations exclude vision-impaired users","Simpler AI explanations work better for non-visual users","For vision-impaired users, AI explanations need multimodal design"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The review's count of only one accessibility-aware study in 79 rests on the assumption that the search string (XAI OR \"AI explanation\") AND (end-user OR user) AND (evaluation OR validation) captures the whole field of end-user XAI evaluation studies, so a gap in vocabulary would not masquerade as a gap in research.","fun_headline_variants_meta":{"raw":{"variants":["Just 1 of 79 AI explanation studies includes disabled users","AI explainability research barely includes disabled users","Visual AI explanations exclude vision-impaired users","Simpler AI explanations work better for non-visual users","For vision-impaired users, AI explanations need multimodal design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000321,"raw_usage":{"total_tokens":1623,"prompt_tokens":703,"completion_tokens":920,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":447,"completion_tokens_details":{"reasoning_tokens":844}},"tokens_in":447,"tokens_out":920,"duration_ms":10486,"temperature":1.0,"reasoning_tokens":844,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:13:24.904232+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same seven-database review with search terms expanded to include 'screen reader,' 'assistive technology,' 'visual impairment,' 'accessible explanation,' and 'non-visual'; if this surfaces more than a handful of end-user XAI evaluations that include blind or low-vision participants, the claimed 1-in-79 gap is overstated. Separately, a comprehension experiment with a larger sample comparing simplified and detailed screen-reader LIME and SHAP explanations would test the preliminary finding.","supporting_citations":[],"review_version":1}