{"id":"153fd95a-de92-450b-8d54-e9428c6f6ce0","arxiv_id":"2506.13904","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 82 healthcare XAI user studies produces an updated property framework and context-sensitive guidelines for evaluation design.","lead":"This paper reviews 82 user studies of explainable AI in healthcare and proposes an updated framework of evaluation properties plus guidelines for choosing what to measure. It is a synthesis and framework paper, not an experimental study.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No inter-rater reliability is reported for the property coding that the framework and guidelines are built on; consensus coding in §3.3 could embed the coders' prior framework, so the framework's empirical grounding is unverified.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing methodological gap: the property coding is the empirical foundation for the framework and guidelines, yet no inter-rater reliability is reported for it. My reading of the full text confirms this. The paper is otherwise transparent about screening, includes a PRISMA flow, provides clear coding criteria, and honestly labels the framework as provisional and in need of empirical validation. The absence of coding reliability and the lack of a shared coded dataset are real weaknesses, but they are addressable and do not force rejection of the framework claim. They do, however, mean the current evidence base for the central claim is conditional: the framework's content and the guideline recommendations could shift if the coding is not reproducible. The proposed concrete test would settle the concern by quantifying coding agreement on a sample. I therefore agree with the CONDITIONAL verdict and recommend no change to it.","tokens_in":51500,"tokens_out":2051,"duration_ms":26259,"concrete_test":"Recode a random sample of 20 of the 82 included studies using the published property definitions and coding rules, with two or three independent coders who were not involved in the original coding and who are blind to the original property assignments. Compute per-property Cohen's kappa (or Fleiss kappa) for property presence and for relation identification. If mean agreement falls below 0.6 for property assignments, the frequency tables in §4.12 and the relation network in Fig. 9 are not demonstrably robust, and the framework changes in §5 should be re-examined on the basis of a more reliable coded dataset.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the proposed framework and context-sensitive guidelines are empirically grounded in a systematic review of 82 user studies. That grounding depends on the property coding described in §3.3: three researchers deductively coded each study with the framework properties from Donoso-Guzmán et al. [15], plus inductively added new codes, and 'discussed in regular sessions to reach consensus'. No inter-rater reliability is reported for this coding. The only agreement statistic in the paper (Fleiss Kappa > 0.8, §3.2.1) concerns title/abstract screening, not the assignment of properties to study measurements. The consensus procedure means disagreements were resolved through discussion, so the reported frequencies (§4.12), relations (§4.11, Fig. 9), and the framework updates (§5) reflect the shared interpretation of the coders rather than a demonstrated stable coding scheme. Because the coders started from the authors' own prior framework, the coding could systematically favour that framework, making the 'updated framework' partially an artefact of the coding lens. The paper's Limitations section (§8) acknowledges other limitations but does not address coding reliability or provide the coded dataset. This does not invalidate the review, but it means the empirical support for the central claim is weaker than reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a systematic review of 82 user studies evaluating explainable AI (XAI) systems in healthcare, following a PRISMA-style workflow over 13,860 records from five databases. The authors code each study along dimensions such as user knowledge, usage context, data type, AI model, XAI method, explanation characteristics, study type, and measured properties, using a property framework from their earlier work [15] together with inductively developed codes. The main contributions are: (1) a synthesis of current evaluation practices in healthcare XAI, (2) a map of reported relations among explanation properties, and (3) an updated user-centred evaluation framework with context-sensitive guidelines for selecting evaluation properties. The paper also proposes a layered model linking domain context, AI context, explanation design, and evaluation design, and gives recommendations for reporting results.","tokens_in":51723,"tokens_out":7046,"duration_ms":77618,"significance":"If the empirical grounding is sound, this would be a valuable resource for the XAI evaluation community: it is the largest domain-specific review of user-centred XAI evaluation I am aware of, and it makes a concrete attempt to move from property taxonomies to actionable, context-dependent guidelines. The strengths are substantial: the screening process is documented in detail, the inclusion/exclusion criteria are explicit, the coding scheme is described, and Appendix B provides an unusually rich mapping from individual measurement items to framework properties. The transparent reuse of the authors' prior framework [15] is also a strength, since it allows the field to see how the framework evolved. The main weakness is that the empirical foundation for the framework and guidelines rests on a property-coding exercise whose reliability is not reported. If the property assignments are unstable or systematically biased by the coders' prior framework, the updated framework, the property frequencies, and the relation counts would all shift.","major_comments":[{"comment":"The paper reports Fleiss kappa above 0.8 only for title and abstract screening, not for the property coding that is the core of the analysis. Section 3.3 states that three researchers coded studies in ATLAS.ti and 'discussed in regular sessions to reach consensus', but no inter-rater reliability is reported for the assignment of studies to properties, for the inductive codes, or for the recorded relations. Because the framework updates (§5), the property frequencies (§4.12), and the relation counts (§4.11, Fig. 9) are all derived from this coding, the abstract's claim that the guidelines are 'based on' the 82 studies is not yet fully supported. Please report per-property agreement statistics on a random subsample, provide the full codebook and the coded dataset as supplementary material, and state explicitly how many studies were double-coded and how disagreements were resolved.","section":"§3.3, §3.2.1, §4.11, §4.12, §5"},{"comment":"The guidelines in Section 6 are the second main contribution, but the evidence linking each recommendation to the coded studies is not shown. For example, Fig. 12 maps 'lay user / patient' to properties such as SSA information expectedness, SSA forms of cognitive chunks, UX satisfaction, and UX trust, and Fig. 13 maps usage contexts to properties, but no table or count indicates how many studies support each mapping, what the direction of the evidence is, or whether these rules were derived from the coding or from the authors' expert judgment. Without this traceability, the context-sensitive claim of the guidelines is difficult to audit. Please add an evidence table linking each recommendation to the supporting studies, or clearly label which recommendations are expert proposals rather than direct findings of the review.","section":"§6, Figs. 12–14"},{"comment":"The paper presents the relations between properties as a major outcome ('We found 85 relations between properties'), but it aggregates counts of studies reporting significant correlations or qualitative causal statements without assessing the quality, effect size, or consistency of the underlying studies, and no risk-of-bias assessment of the included studies is reported. Since the Limitations section (Section 8) does not mention this omission, the reader may over-interpret the relation network in Fig. 9. Please either add a study-quality assessment and discuss how it affects the relation counts, or explicitly frame the relation map as an unweighted narrative synthesis.","section":"§4.11, §8"}],"minor_comments":[{"comment":"The PRISMA numbers do not fully reconcile: 13,860 identified records minus 5,633 duplicates equals 8,227, but the text and figure report 8,226 records screened; please correct or explain the discrepancy.","section":"§3.1, Fig. 1"},{"comment":"The phrase 'Capability A' appears truncated; it should read 'Capability Assessment'.","section":"§4.4"},{"comment":"The section begins with 'As shown in section 4.8', which is a self-referential cross-reference; please point to the relevant table or figure instead.","section":"§4.8"},{"comment":"The term 'Orthopedagogy' appears in the table and text; if 'orthopedics' or 'orthopaedics' is intended, please correct it.","section":"Table 8, §4.10"},{"comment":"In the Trust paragraph, the numbers 14 mixed + 10 quantitative + 5 qualitative sum to 29, not 31; and in the Confidence paragraph, 'Case Difficulty' is listed twice among the related properties. Please check these counts.","section":"§4.12.6"},{"comment":"Several labels in these figures are fragmentary, e.g., 'lay user / patientif is' and 'Criterion usage context'; please clean up the format so each branch is readable.","section":"Figs. 12–14"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and makes a real contribution if the coding reliability concern is addressed. I do not see a need for rejection: the issue is fixable by adding a reliability analysis, a codebook, and an evidence trace for the guidelines. The heavy use of the authors' own prior framework [15] is transparent, but the absence of independent reliability data means the framework update is not yet convincingly validated. I would encourage the editor to request the supplementary coded dataset as part of the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a careful, useful systematic review of user-centred XAI evaluation in healthcare, and the updated framework is a genuine synthesis, not a repackaging. The load-bearing weakness is that the property coding is not assessed for reliability, so the empirical grounding is weaker than the paper suggests. That is a fixable problem, not a fatal one.\n\nWhat is actually new: the healthcare-specific lens. Reviews like Kim et al. and Lopes et al. are general; this one stays inside clinical domains and derives context-sensitive guidance from 82 user studies. The framework changes are concrete and justified: Understanding splits into explanation understanding and model behaviour understanding; Continuity and Representativeness merge; and seven properties are added, including Confidence, Prediction Expectedness, and Variance User Decision. The layered model in Section 6 is a practical contribution that connects usage context and system characteristics to property selection.\n\nWhat the paper does well: the screening is transparent and thorough (13,860 records down to 82), inclusion/exclusion criteria are explicit, and the authors decouple properties from measurement methods, which many reviews fail to do. They also give credit where due: they start from their own prior framework [15] and say so. The Appendix B measurement list is a useful resource.\n\nThe soft spot is exactly the one flagged in the stress-test note: Section 3.3 says three researchers coded in ATLAS.ti and 'discussed in regular sessions to reach consensus,' but the only agreement statistic is for title/abstract screening, not for the property coding that the framework depends on. Without an inter-rater reliability figure or a shared coded dataset, the frequencies in §4.12 and the relations in §4.11 are not independently checkable. Since the coding started from the authors' own framework, there is a real—if transparent—risk of confirmation bias. The paper itself acknowledges it offers no validated instruments and that the guidelines need empirical testing, which is honest. But the central claim that the framework is empirically grounded in the review is somewhat stronger than the evidence supports.\n\nWho should read this: anyone designing or evaluating healthcare XAI systems. It deserves a serious referee; the missing reliability data should be addressed in revision by reporting agreement on the property coding and making the coded data available. I'd send it to review.","headline":"A careful systematic review that builds a genuinely useful healthcare-specific XAI evaluation framework, but the missing inter-rater reliability for the property coding leaves the empirical grounding weaker than claimed.","tokens_in":52263,"tokens_out":2220,"would_cite":true,"duration_ms":24176,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A review of 82 healthcare user studies yields an updated framework of atomic explanation properties and a layered model that tells evaluators which properties to measure, and when.","keywords":["explainable AI","XAI evaluation","explainability framework","evaluation XAI in healthcare","user-centred evaluation","systematic review","user studies","healthcare"],"falsifier":"Have two independent researchers, using only the paper's property definitions, re-code a random sample of about 20 of the 82 studies; if their agreement on which properties each study measured falls below roughly 0.6 on a standard inter-rater agreement statistic, the frequency counts and the 85 documented relations that ground the updated framework would not reproduce.","tokens_in":51323,"feed_emoji":"🩺","tokens_out":11235,"duration_ms":98949,"temperature":0.7,"pith_summary":"This paper tries to turn a messy research area into a usable design tool: when an AI system in healthcare explains itself, what exactly should a study measure, and how should that choice depend on the situation? The authors reviewed 82 user studies of explainable AI in healthcare and coded each one against a pre-existing framework of atomic explanation properties, adding new properties that surfaced during coding. The result is an updated framework of well-defined properties organized into seven components, plus a layered model that guides evaluators from the medical domain and usage context down to the choice of study type, properties, and measurements. If the framework and the guidelines are right, interdisciplinary teams would have a concrete, defensible way to decide what to evaluate, instead of defaulting to the usual trio of trust, understanding, and performance.","feed_headline":"82 healthcare AI studies become an evaluation guide","feed_subtitle":"Distills 82 healthcare user studies into a framework that says which explanation qualities to measure, and when.","key_machinery":"The carrying object is the updated User Centric Evaluation framework: a set of atomic, non-overlapping property definitions (for example Necessity, Sufficiency, Information Expectedness, Relevance to the Task, and Reliance) grouped into seven conceptual components — Personal Characteristics, Situational Characteristics, Objective System Aspects, Explanation Aspects, Subjective System Aspects, User Experience, and Interaction — whose definitions were taken from the authors' earlier framework [15] and applied deductively, with inductive codes added while coding. The second mechanism is the layered evaluation-design model, which chains Domain Context, AI Context, Explanation Design, and Evaluation Design, together with property-selection maps that connect criteria such as usage context, user type, XAI method, scope, and interactivity to recommended properties. These two mechanisms convert 82 heterogeneous studies into a single consistent vocabulary and a decision procedure for choosing what to measure.","core_discovery":"The paper's central claim is that the user experience of AI explanations in healthcare can be decomposed into a coherent set of atomic, non-overlapping properties, and that the choice of which properties to evaluate should be derived from the system's context rather than from habit. Analysing 82 user studies, the authors find which properties are actually measured in practice — Understanding (33 papers) and Trust (30) dominate, while Satisfaction (11) is less central than general frameworks suggest — and they document 85 significant relations among properties. From this evidence they update an earlier framework: seven new properties are added (including Confidence, Information Correctness, Prediction Expectedness, and Variance User Decision), Continuity and Representativeness are merged, and Understanding is split into Understanding Explanation and Understanding Model Behaviour. They then propose a layered model — domain context, AI context, explanation design, evaluation design — with explicit recommendations for which properties to measure given the usage context, user type, data, scope, and interactivity of the system. The claim is that this closes the gap left by prior taxonomies, which organise evaluation aspects but never state when to measure them.","pith_inferences":["The 85 recorded property relations could seed a quantitative meta-analysis or a predictive model of clinician experience, weighting each relation by study quality and effect size — a synthesis the paper does not attempt.","As generative-AI explanations enter clinical workflows, Information Correctness — whether the explanation's factual content is right, as opposed to merely expected — is likely to become the most load-bearing of the new properties; the paper flags the trend but does not develop it.","The total absence of the Adapting Control usage context may reflect the review's exclusion of Wizard-of-Oz studies rather than that context's rarity in healthcare; a review that includes simulated-AI studies could decide between those two readings.","The appendix's list of measurements could grow into a shared bank of validated items, one per atomic property; if the field adopted common items, cross-study comparison would become routine, a step the paper calls for but does not take."],"forward_implications":["A team designing a healthcare XAI evaluation can start from the usage context — decision support, capability assessment, or model auditing — and read off which properties are worth measuring.","Understanding can no longer be treated as one construct: the split between Understanding Explanation and Understanding Model Behaviour implies that a question about the one does not measure the other.","The documented relations among properties (for example, Case Difficulty driving Confidence, and Information Expectedness feeding Trust, Usefulness, and Intention to Use) provide ready-made hypotheses for confirmatory studies.","Because more than half of the quantitative studies had 16 or fewer participants, the review implies that small-sample quantitative designs should give way to qualitative or mixed designs when recruiting clinicians is hard.","Consistent reuse of the framework's property names, and reporting results per experience group, would make evaluation results comparable across studies."],"supporting_citations":[{"why":"The authors' earlier framework whose property definitions are reused as the deductive coding scheme; the updated framework is a direct revision of these definitions.","marker":"[15]"},{"why":"Supplies the six usage contexts (decision support, capability assessment, model auditing, and the rest) used to code the AI context layer and to drive the property-selection guidelines.","marker":"[26]"},{"why":"Provides the conceptual component structure (Objective System Aspects, Subjective System Aspects, User Experience) that organises the framework's components.","marker":"[48]"},{"why":"The closest prior evaluation survey; the paper positions its contribution as closing the gap this survey leaves open.","marker":"[17]"},{"why":"One of the taxonomies the paper contrasts with its own, as it organises evaluation methods but does not say which aspects to measure.","marker":"[32]"},{"why":"A prior user-study survey that offers general guidelines but no property-selection step; a direct comparison target for the proposed guidelines.","marker":"[16]"},{"why":"Defines the formal, instrumental, and personal knowledge types used to classify participants in the coded studies.","marker":"[43]"},{"why":"Provides the explanation abstraction types (feature importance, counterfactual, example-based, rule-based) used to code explanation content.","marker":"[44]"},{"why":"Supplies the data, model, and post-hoc taxonomy used to categorise the XAI methods found in the reviewed studies.","marker":"[36]"}],"fun_headline_variants":["82 studies yield AI explanation evaluation playbook","When to measure AI trust? New guide from 82 healthcare studies","Healthcare XAI evaluation gets atomic property framework","From 82 studies: guidelines for evaluating AI explanations","AI explanation testing: new framework from 82 user studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the property coding of the 82 studies is consistent and unbiased: the review reports an inter-rater agreement above 0.8 only for the title-and-abstract screening step, not for the property coding itself, which relied on three coders reaching consensus.","fun_headline_variants_meta":{"raw":{"variants":["82 studies yield AI explanation evaluation playbook","When to measure AI trust? New guide from 82 healthcare studies","Healthcare XAI evaluation gets atomic property framework","From 82 studies: guidelines for evaluating AI explanations","AI explanation testing: new framework from 82 user studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1236,"prompt_tokens":1000,"completion_tokens":236,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":160}},"tokens_in":616,"tokens_out":236,"duration_ms":3380,"temperature":1.0,"reasoning_tokens":160,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:26:07.886073+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent researchers, using only the paper's property definitions, re-code a random sample of about 20 of the 82 studies; if their agreement on which properties each study measured falls below roughly 0.6 on a standard inter-rater agreement statistic, the frequency counts and the 85 documented relations that ground the updated framework would not reproduce.","supporting_citations":[{"cited_title":"Donoso-Guzmán, J","cited_arxiv_id":null,"evidence_quote":"The authors' earlier framework whose property definitions are reused as the deductive coding scheme; the updated framework is a direct revision of these definitions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the six usage contexts (decision support, capability assessment, model auditing, and the rest) used to code the AI context layer and to drive the property-selection guidelines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the conceptual component structure (Objective System Aspects, Subjective System Aspects, User Experience) that organises the framework's components."}],"review_version":1}