Pith. sign in

REVIEW 3 major objections 6 minor 279 references

A Systematic Review of User-Centred Evaluation of Explainable AI in Healthcare

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A review of 82 healthcare user studies yields an updated framework of atomic explanation properties and a layered model that tells evaluators which properties to measure, and when.

desk verdict A careful systematic review that builds a genuinely useful healthcare-specific XAI evaluation framework, but the missing inter-rater reliability for the property coding leaves the empirical grounding weaker than claimed. read the letter →

arxiv 2506.13904 v1 pith:OZNEJONI submitted 2025-06-16 cs.HC cs.AIcs.LG

classification cs.HCcs.AIcs.LG
keywords explainableAIXAIevaluationexplainabilityframeworkinhealthcareuser-centredsystematicreviewuserstudies
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to turn a messy research area into a usable design tool: when an AI system in healthcare explains itself, what exactly should a study measure, and how should that choice depend on the situation? The authors reviewed 82 user studies of explainable AI in healthcare and coded each one against a pre-existing framework of atomic explanation properties, adding new properties that surfaced during coding. The result is an updated framework of well-defined properties organized into seven components, plus a layered model that guides evaluators from the medical domain and usage context down to the choice of study type, properties, and measurements. If the framework and the guidelines are right, interdisciplinary teams would have a concrete, defensible way to decide what to evaluate, instead of defaulting to the usual trio of trust, understanding, and performance.

What carries the argument

The carrying object is the updated User Centric Evaluation framework: a set of atomic, non-overlapping property definitions (for example Necessity, Sufficiency, Information Expectedness, Relevance to the Task, and Reliance) grouped into seven conceptual components — Personal Characteristics, Situational Characteristics, Objective System Aspects, Explanation Aspects, Subjective System Aspects, User Experience, and Interaction — whose definitions were taken from the authors' earlier framework [15] and applied deductively, with inductive codes added while coding. The second mechanism is the layered evaluation-design model, which chains Domain Context, AI Context, Explanation Design, and Evaluation Design, together with property-selection maps that connect criteria such as usage context, user type, XAI method, scope, and interactivity to recommended properties. These two mechanisms convert 82 heterogeneous studies into a single consistent vocabulary and a decision procedure for choosing what to measure.

What would settle it

Have two independent researchers, using only the paper's property definitions, re-code a random sample of about 20 of the 82 studies; if their agreement on which properties each study measured falls below roughly 0.6 on a standard inter-rater agreement statistic, the frequency counts and the 85 documented relations that ground the updated framework would not reproduce.

Watch

Extended reading notes

Core claim

The paper's central claim is that the user experience of AI explanations in healthcare can be decomposed into a coherent set of atomic, non-overlapping properties, and that the choice of which properties to evaluate should be derived from the system's context rather than from habit. Analysing 82 user studies, the authors find which properties are actually measured in practice — Understanding (33 papers) and Trust (30) dominate, while Satisfaction (11) is less central than general frameworks suggest — and they document 85 significant relations among properties. From this evidence they update an earlier framework: seven new properties are added (including Confidence, Information Correctness, Prediction Expectedness, and Variance User Decision), Continuity and Representativeness are merged, and Understanding is split into Understanding Explanation and Understanding Model Behaviour. They then propose a layered model — domain context, AI context, explanation design, evaluation design — with explicit recommendations for which properties to measure given the usage context, user type, data, scope, and interactivity of the system. The claim is that this closes the gap left by prior taxonomies, which organise evaluation aspects but never state when to measure them.

Load-bearing premise

The load-bearing premise is that the property coding of the 82 studies is consistent and unbiased: the review reports an inter-rater agreement above 0.8 only for the title-and-abstract screening step, not for the property coding itself, which relied on three coders reaching consensus.

Editorial extensions

If this is right

  • A team designing a healthcare XAI evaluation can start from the usage context — decision support, capability assessment, or model auditing — and read off which properties are worth measuring.
  • Understanding can no longer be treated as one construct: the split between Understanding Explanation and Understanding Model Behaviour implies that a question about the one does not measure the other.
  • The documented relations among properties (for example, Case Difficulty driving Confidence, and Information Expectedness feeding Trust, Usefulness, and Intention to Use) provide ready-made hypotheses for confirmatory studies.
  • Because more than half of the quantitative studies had 16 or fewer participants, the review implies that small-sample quantitative designs should give way to qualitative or mixed designs when recruiting clinicians is hard.
  • Consistent reuse of the framework's property names, and reporting results per experience group, would make evaluation results comparable across studies.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 85 recorded property relations could seed a quantitative meta-analysis or a predictive model of clinician experience, weighting each relation by study quality and effect size — a synthesis the paper does not attempt.
  • As generative-AI explanations enter clinical workflows, Information Correctness — whether the explanation's factual content is right, as opposed to merely expected — is likely to become the most load-bearing of the new properties; the paper flags the trend but does not develop it.
  • The total absence of the Adapting Control usage context may reflect the review's exclusion of Wizard-of-Oz studies rather than that context's rarity in healthcare; a review that includes simulated-AI studies could decide between those two readings.
  • The appendix's list of measurements could grow into a shared bank of validated items, one per atomic property; if the field adopted common items, cross-study comparison would become routine, a step the paper calls for but does not take.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This manuscript reports a systematic review of 82 user studies evaluating explainable AI (XAI) systems in healthcare, following a PRISMA-style workflow over 13,860 records from five databases. The authors code each study along dimensions such as user knowledge, usage context, data type, AI model, XAI method, explanation characteristics, study type, and measured properties, using a property framework from their earlier work [15] together with inductively developed codes. The main contributions are: (1) a synthesis of current evaluation practices in healthcare XAI, (2) a map of reported relations among explanation properties, and (3) an updated user-centred evaluation framework with context-sensitive guidelines for selecting evaluation properties. The paper also proposes a layered model linking domain context, AI context, explanation design, and evaluation design, and gives recommendations for reporting results.

Significance. If the empirical grounding is sound, this would be a valuable resource for the XAI evaluation community: it is the largest domain-specific review of user-centred XAI evaluation I am aware of, and it makes a concrete attempt to move from property taxonomies to actionable, context-dependent guidelines. The strengths are substantial: the screening process is documented in detail, the inclusion/exclusion criteria are explicit, the coding scheme is described, and Appendix B provides an unusually rich mapping from individual measurement items to framework properties. The transparent reuse of the authors' prior framework [15] is also a strength, since it allows the field to see how the framework evolved. The main weakness is that the empirical foundation for the framework and guidelines rests on a property-coding exercise whose reliability is not reported. If the property assignments are unstable or systematically biased by the coders' prior framework, the updated framework, the property frequencies, and the relation counts would all shift.

major comments (3)
  1. [§3.3, §3.2.1, §4.11, §4.12, §5] The paper reports Fleiss kappa above 0.8 only for title and abstract screening, not for the property coding that is the core of the analysis. Section 3.3 states that three researchers coded studies in ATLAS.ti and 'discussed in regular sessions to reach consensus', but no inter-rater reliability is reported for the assignment of studies to properties, for the inductive codes, or for the recorded relations. Because the framework updates (§5), the property frequencies (§4.12), and the relation counts (§4.11, Fig. 9) are all derived from this coding, the abstract's claim that the guidelines are 'based on' the 82 studies is not yet fully supported. Please report per-property agreement statistics on a random subsample, provide the full codebook and the coded dataset as supplementary material, and state explicitly how many studies were double-coded and how disagreements were resolved.
  2. [§6, Figs. 12–14] The guidelines in Section 6 are the second main contribution, but the evidence linking each recommendation to the coded studies is not shown. For example, Fig. 12 maps 'lay user / patient' to properties such as SSA information expectedness, SSA forms of cognitive chunks, UX satisfaction, and UX trust, and Fig. 13 maps usage contexts to properties, but no table or count indicates how many studies support each mapping, what the direction of the evidence is, or whether these rules were derived from the coding or from the authors' expert judgment. Without this traceability, the context-sensitive claim of the guidelines is difficult to audit. Please add an evidence table linking each recommendation to the supporting studies, or clearly label which recommendations are expert proposals rather than direct findings of the review.
  3. [§4.11, §8] The paper presents the relations between properties as a major outcome ('We found 85 relations between properties'), but it aggregates counts of studies reporting significant correlations or qualitative causal statements without assessing the quality, effect size, or consistency of the underlying studies, and no risk-of-bias assessment of the included studies is reported. Since the Limitations section (Section 8) does not mention this omission, the reader may over-interpret the relation network in Fig. 9. Please either add a study-quality assessment and discuss how it affects the relation counts, or explicitly frame the relation map as an unweighted narrative synthesis.
minor comments (6)
  1. [§3.1, Fig. 1] The PRISMA numbers do not fully reconcile: 13,860 identified records minus 5,633 duplicates equals 8,227, but the text and figure report 8,226 records screened; please correct or explain the discrepancy.
  2. [§4.4] The phrase 'Capability A' appears truncated; it should read 'Capability Assessment'.
  3. [§4.8] The section begins with 'As shown in section 4.8', which is a self-referential cross-reference; please point to the relevant table or figure instead.
  4. [Table 8, §4.10] The term 'Orthopedagogy' appears in the table and text; if 'orthopedics' or 'orthopaedics' is intended, please correct it.
  5. [§4.12.6] In the Trust paragraph, the numbers 14 mixed + 10 quantitative + 5 qualitative sum to 29, not 31; and in the Confidence paragraph, 'Case Difficulty' is listed twice among the related properties. Please check these counts.
  6. [Figs. 12–14] Several labels in these figures are fragmentary, e.g., 'lay user / patientif is' and 'Criterion usage context'; please clean up the format so each branch is readable.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the updated framework is an iterative synthesis of 82 coded studies, and the prior framework [15] is a transparent coding lens, not a fitted target.

full rationale

The paper's central claim is a synthesis and update, not a prediction derived from its own inputs. Section 3.3.8 states that the authors 'use the framework of properties defined in [15] to deductively code the properties measured in the studies,' so the authors' own prior framework serves as the coding lens. This is self-citation, but it is disclosed and is not invoked as an external proof or uniqueness theorem. The output is not forced to equal the input: Section 5 reports seven new inductively discovered properties, a merge of Continuity and Representativeness, a split of Understanding into two properties, and a redefinition of Explanation Power, each motivated by coded study evidence and by re-analysis of cited literature. The context-sensitive guidelines in Section 6 are grounded in observed usage contexts (from Liao et al. [26]), user/task/domain characteristics, and reported property relations, rather than in equations or fits that reproduce the input framework. No quantity is fitted to a subset and then relabelled as a prediction; the reported frequencies, relations, and framework updates are descriptive codings of the 82 included studies. The absence of inter-rater reliability for the property coding, while a legitimate methodological limitation, is a reliability concern rather than a circularity concern. The paper also acknowledges in Section 8 that the proposed framework requires further empirical validation, so it does not claim that the framework is proven by its own construction. No specific reduction of a claimed result to its own definition or to a self-citation chain can be exhibited.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

This is a systematic review, so there are no free parameters or invented entities. The axioms listed are the domain assumptions inherent in using a predefined coding scheme and in treating the selected studies as representative of the literature.

assumptions (2)
  • domain assumption The framework of 30 properties from Donoso-Guzmán et al. [15] is an appropriate starting point for coding, and its definitions are atomic and non-overlapping.
    Invoked in Section 3.3.8 where properties are defined and used as the deductive coding scheme; the review later modifies several definitions, indicating they are not fully stable.
  • domain assumption The 82 included studies are representative of user-centred XAI evaluation in healthcare despite exclusions of non-peer-reviewed work, Wizard-of-Oz studies, and studies with unsuitable participants.
    This underlies all frequency claims in Section 4; the authors acknowledge in Limitations that excluding Wizard-of-Oz may have narrowed the usage contexts observed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Systematic Review of User-Centred Evaluation of Explainable AI in Healthcare." pith.science (2026). https://pith.science/paper/OZNEJONI

@misc{pith2026250613904,
  author       = {Pith},
  title        = {Pith review of: A Systematic Review of User-Centred Evaluation of Explainable AI in Healthcare},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OZNEJONI}},
  note         = {Machine review of arXiv:2506.13904}
}
read the original abstract

Despite promising developments in Explainable Artificial Intelligence, the practical value of XAI methods remains under-explored and insufficiently validated in real-world settings. Robust and context-aware evaluation is essential, not only to produce understandable explanations but also to ensure their trustworthiness and usability for intended users, but tends to be overlooked because of no clear guidelines on how to design an evaluation with users. This study addresses this gap with two main goals: (1) to develop a framework of well-defined, atomic properties that characterise the user experience of XAI in healthcare; and (2) to provide clear, context-sensitive guidelines for defining evaluation strategies based on system characteristics. We conducted a systematic review of 82 user studies, sourced from five databases, all situated within healthcare settings and focused on evaluating AI-generated explanations. The analysis was guided by a predefined coding scheme informed by an existing evaluation framework, complemented by inductive codes developed iteratively. The review yields three key contributions: (1) a synthesis of current evaluation practices, highlighting a growing focus on human-centred approaches in healthcare XAI; (2) insights into the interrelations among explanation properties; and (3) an updated framework and a set of actionable guidelines to support interdisciplinary teams in designing and implementing effective evaluation strategies for XAI systems tailored to specific application contexts.

Figures

Figures reproduced from arXiv: 2506.13904 by the authors.

Figure 1
Figure 1. Prisma schema of the survey process 8 [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Summary of criteria used to code the selected papers. [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Explanation elements as defined by [15]. These elements help to identify characteristics of explanations that are defined at the different steps of the process and within the explanation product. on the data features. This type of explanation has proven to be relevant in healthcare domain settings [45, 46]. For the format level, we focused on what the user would interact with and what type of cognitive effort they w… view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Publication year of selected papers. Most of them ( [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: Groups of participants in the selected papers, according to their knowledge of [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]
Figure 6
Figure 6. Figure 6: Number of properties studied by papers and type of study. Triangles represent [PITH_FULL_IMAGE:figures/full_fig_p027_6.png]
Figure 7
Figure 7. Figure 7: Measurement method for each property depending on the study type. Quantitative [PITH_FULL_IMAGE:figures/full_fig_p029_7.png]
Figure 8
Figure 8. Figure 8: Qualitative methods for mixed and quantitative studies. By far, the most used [PITH_FULL_IMAGE:figures/full_fig_p031_8.png]
Figure 9
Figure 9. Figure 9: Relations between the properties. Each square represents a relation between two properties. Darker squares mean that more papers found that relation. Most of the relations are found between the Subjective System Aspects and User experience aspects. 33 [PITH_FULL_IMAGE…
Figure 10
Figure 10. Figure 10: User Centric Evaluation framework updated. New properties are marked with * [PITH_FULL_IMAGE:figures/full_fig_p050_10.png]
Figure 11
Figure 11. Figure 11: Layered model to design evaluations. During a [PITH_FULL_IMAGE:figures/full_fig_p051_11.png]
Figure 12
Figure 12. Figure 12: Recommendation and suggestion to select properties according to the [PITH_FULL_IMAGE:figures/full_fig_p052_12.png]
Figure 13
Figure 13. Figure 13: Recommendation and suggestion to select properties according to the [PITH_FULL_IMAGE:figures/full_fig_p053_13.png]
Figure 14
Figure 14. Figure 14: Recommendation and suggestion to select properties according to the [PITH_FULL_IMAGE:figures/full_fig_p053_14.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

279 extracted references · 41 canonical work pages

  1. [15]

    Donoso-Guzmán, J

    I. Donoso-Guzmán, J. Ooge, D. Parra, K. Verbert, Towards a Compre- hensive Human-Centred Evaluation Framework for Explainable AI, in: Communications in Computer and Information Science, volume 1903 CCIS, Springer Science and Business Media Deutschland GmbH, 2023, pp. 183–204. doi:10.1007/978-3-031-44070-0{\_}10. 62

  2. [1]

    Čartolovni, A

    A. Čartolovni, A. Tomičić, E. Lazić Mosler, Ethical, legal, and social considerations of AI-based medical decision-support tools: A scop- ing review, International Journal of Medical Informatics 161 (2022) 104738. URL: https://www.sciencedirect.com/science/article/ pii/S1386505622000521?via=ihub. doi: 10.1016/J.IJMEDINF.2022. 104738

  3. [2]

    Tonekaboni, S

    S. Tonekaboni, S. Joshi, M. D. McCradden, A. Goldenberg, What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use, Proceedings of Machine Learning Research (2019). URL: http://arxiv.org/abs/1905.05134. 60

  4. [3]

    Bussone, S

    A. Bussone, S. Stumpf, D. O’Sullivan, The Role of Explanations on Trust and Reliance in Clinical Decision Support Systems, in: 2015 In- ternational Conference on Healthcare Informatics, IEEE, 2015, pp. 160–169. URL: http://ieeexplore.ieee.org/document/7349687/. doi:10.1109/ICHI.2015.26

  5. [4]

    Reyes, R

    M. Reyes, R. Meier, S. Pereira, C. A. Silva, F.-M. Dahlweid, H. v. Tengg- Kobligk, R. M. Summers, R. Wiest, On the Interpretability of Artificial Intelligence in Radiology: Challenges and Opportunities, Radiology: Artificial Intelligence 2 (2020) e190043. URL:http://pubs.rsna.org/ doi/10.1148/ryai.2020190043. doi:10.1148/ryai.2020190043

  6. [5]

    Amann, D

    J. Amann, D. Vetter, S. N. Blomberg, H. C. Christensen, M. Coffee, S. Gerke, T. K. Gilbert, T. Hagendorff, S. Holm, M. Livne, A. Spezzatti, I. Strümke, R. V. Zicari, V. I. Madai, o. b. o. t. Z.-I. initiative, To explain or not to explain?—Artificial intelligence explainability in clini- cal decision support systems, PLOS Digital Health 1 (2022) e0000016. ...

  7. [6]

    C. M. Cutillo, K. R. Sharma, L. Foschini, S. Kundu, M. Mackintosh, K. D. Mandl, T. Beck, E. Collier, C. Colvis, K. Gersing, V. Gordon, R. Jensen, B. Shabestari, N. Southall, Machine intelligence in health- care—perspectives on trustworthiness, explainability, usability, and transparency, 2020. doi:10.1038/s41746-020-0254-2

  8. [7]

    Wiens, S

    J. Wiens, S. Saria, M. Sendak, M. Ghassemi, V. X. Liu, F. Doshi-Velez, K. Jung, K. Heller, D. Kale, M. Saeed, P. N. Ossorio, S. Thadaney- Israni, A. Goldenberg, Do no harm: a roadmap for responsible machine learning for health care, Nature Medicine 25 (2019) 1337–1340. doi:10. 1038/s41591-019-0548-6

Show all 279 references
  1. [8]

    European Institute of Innovation and Technology Health, McKinsey & Company, Transforming Healthcare with AI: The Impact on the Workforce and Organisations, Technical Report, European Institute of Innovation and Technology Health, 2020

  2. [9]

    H. W. Loh, C. P. Ooi, S. Seoni, P. D. Barua, F. Molinari, U. R. Acharya, Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011–2022), Computer Methods 61 and Programs in Biomedicine 226 (2022) 107161. doi:10.1016/J.CMPB...

  3. [10]

    Chaddad, J

    A. Chaddad, J. Peng, J. Xu, A. Bouridane, Survey of Explainable AI Techniques in Healthcare, Sensors (Basel, Switzerland) 23 (2023). doi:10.3390/s23020634

  4. [11]

    A. F. Markus, J. A. Kors, P. R. Rijnbeek, The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies, Journal of Biomedical Informatics 113 (2020) 103655. URL:...

  5. [12]

    K. S. Kacafírková, S. Polak, M. S. Smitt, S. Elprama, A. Jacobs, Trustworthy enough? Evaluation of an AI decision support system for healthcare professionals, in: L. Longo (Ed.), Joint Proceedings of the xAI-2023 Late-breaking Work, Demos and Doctoral Consortium co-located wit...

  6. [13]

    Miller, Explanation in artificial intelligence: Insights from the social sciences, Artificial Intelligence 267 (2019) 1–38

    T. Miller, Explanation in artificial intelligence: Insights from the social sciences, Artificial Intelligence 267 (2019) 1–38. URL:https: //linkinghub.elsevier.com/retrieve/pii/S0004370218305988. doi:10.1016/j.artint.2018.07.007

  7. [14]

    Velmurugan, C

    M. Velmurugan, C. Ouyang, Y. Xu, R. Sindhgatta, B. Wickra- manayake, C. Moreira, Developing guidelines for functionally- grounded evaluation of explainable artificial intelligence using tab- ular data, Engineering Applications of Artificial Intelligence 141 (2025) 109772. URL:...

  8. [16]

    Y. Rong, T. Leemann, T. T. Nguyen, L. Fiedler, P. Qian, V. Unhelkar, T. Seidel, G. Kasneci, E. Kasneci, Towards Human-Centered Explain- able AI: A Survey of User Studies for Model Explanations, IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (2024) 2104–2122....

  9. [17]

    J. Kim, H. Maathuis, D. Sent, Human-centered evaluation of explainable AI applications: a systematic review, Frontiers in Artificial Intelligence 7 (2024) 1456486. doi:10.3389/FRAI.2024.1456486/PDF

  10. [18]

    Naveed, G

    S. Naveed, G. Stevens, D. Robin-Kern, An Overview of the Empirical Evaluation of Explainable AI (XAI): A Comprehensive Guideline for User-Centered Evaluation in XAI, Applied Sciences (Switzerland) 14 (2024). doi:10.3390/APP142311288

  11. [19]

    B. Y. Lim, Q. Yang, A. Abdul, D. Wang, Why these Explanations? Selecting Intelligibility Types for Explanation Goals, IUI Workshops (2019)

  12. [20]

    S. S. Y. Kim, E. A. Watkins, O. Russakovsky, R. Fong, A. Monroy- Hernández, ”Help Me Help the AI”: Understanding How Explain- ability Can Support Human-AI Interaction, in: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2023, pp. 1–17. URL: http:/...

  13. [21]

    A. Suh, I. Hurley, N. Smith, H. C. Siu, Fewer Than 1% of Explainable AI Papers Validate Explainability with Humans, in: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, ACM, New York, NY, USA, 2025, pp. 1–7. URL: https://dl.acm...

  14. [22]

    Nauta, J

    M. Nauta, J. Trienes, S. Pathak, E. Nguyen, M. Peters, Y. Schmitt, J. Schlötterer, M. Van Keulen, C. Seifert, From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI, ACM Computing Surveys 55 (2023). doi:10.1145/ 3583558. 63

  15. [23]

    Vilone, L

    G. Vilone, L. Longo, Notions of explainability and evaluation approaches for explainable artificial intelligence, Information Fusion 76 (2021) 89–106. doi:10.1016/J.INFFUS.2021.05.009

  16. [24]

    Beckh, S

    K. Beckh, S. Müller, S. Rüping, A Quantitative Human-Grounded Evaluation Process for Explainable Machine Learning, in: LWDA’22: Lernen, Wissen, Daten, Analysen, 2022. URL:http://ceur-ws.org

  17. [25]

    Sokol, P

    K. Sokol, P. Flach, Explainability fact sheets: A framework for sys- tematic assessment of explainable approaches, in: FAT* 2020 - Pro- ceedings of the 2020 Conference on Fairness, Accountability, and Trans- parency, Association for Computing Machinery, Inc, 2020, pp. 56–67. d...

  18. [26]

    Q. V. Liao, Y. Zhang, R. Luss, F. Doshi-Velez, A. Dhurandhar, Con- necting Algorithmic Research and Usage Contexts: A Perspective of Contextualized Evaluation for Explainable AI, Proceedings of the AAAI Conference on Human Computation and Crowdsourcing 10 (2022) 147–159. URL: ...

  19. [27]

    Löfström, K

    H. Löfström, K. Hammar, U. Johansson, A Meta Survey of Quality Evaluation Criteria in Explanation Methods (2022). URL:http:// arxiv.org/abs/2203.13929

  20. [28]

    D. V. Carvalho, E. M. Pereira, J. S. Cardoso, Machine Learn- ing Interpretability: A Survey on Methods and Metrics, Electron- ics 8 (2019) 832. URL:https://www.mdpi.com/2079-9292/8/8/832. doi:10.3390/electronics8080832

  21. [29]

    J. H.-w. Hsiao, H. H. T. Ngai, L. Qiu, Y. Yang, C. C. Cao, Roadmap of Designing Cognitive Metrics for Explainable Artificial Intelligence (XAI) (2021). URL:https://arxiv.org/abs/2108.01737v1. doi:10. 48550/arxiv.2108.01737

  22. [31]

    Saeed, C

    W. Saeed, C. Omlin, Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities, Knowledge-Based Systems 263 (2023) 110273. doi:10.1016/j.knosys.2023.110273

  23. [32]

    Lopes, E

    P. Lopes, E. Silva, C. Braga, T. Oliveira, L. Rosado, XAI Systems Eval- uation: A Review of Human and Computer-Centred Methods, Applied Sciences 12 (2022) 9423. URL:https://www.mdpi.com/2076-3417/ 12/19/9423. doi:10.3390/app12199423

  24. [33]

    Guidotti, A

    R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F. Giannotti, D. Pe- dreschi, A survey of methods for explaining black box models, ACM Computing Surveys 51 (2018) 1–45. doi:10.1145/3236009

  25. [34]

    Barredo Arrieta, N

    A. Barredo Arrieta, N. Díaz-Rodríguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. Garcia, S. Gil-Lopez, D. Molina, R. Ben- jamins, R. Chatila, F. Herrera, Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI...

  26. [35]

    Sahakyan, Z

    M. Sahakyan, Z. Aung, T. Rahwan, Explainable Artificial Intelligence for Tabular Data: A Survey, IEEE Access 9 (2021) 135392–135422. doi:10.1109/ACCESS.2021.3116481

  27. [36]

    S. Ali, T. Abuhmed, S. El-Sappagh, K. Muhammad, J. M. Alonso-Moral, R. Confalonieri, R. Guidotti, J. Del Ser, N. Díaz-Rodríguez, F. Herrera, Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence, Information Fusi...

  28. [37]

    J. Ooge, G. Stiglic, K. Verbert, Explaining artificial intelligence with visual analytics in healthcare, WIREs Data Mining and Knowledge Discovery 12 (2022). doi:10.1002/widm.1427

  29. [38]

    A. M. Antoniadi, Y. Du, Y. Guendouz, L. Wei, C. Mazo, B. A. Becker, C. Mooney, Current Challenges and Future Opportunities for XAI in Machine Learning-Based Clinical Decision Support Systems: A System- atic Review, Applied Sciences 2021, Vol. 11, Page 5088 11 (2021) 5088. URL:...

  30. [39]

    I. D. Mienye, G. Obaido, N. Jere, E. Mienye, K. Aruleba, I. D. Emmanuel, B. Ogbuokiri, A survey of explain- able artificial intelligence in healthcare: Concepts, applications, and challenges, Informatics in Medicine Unlocked 51 (2024) 101587. URL: https://linkinghub.elsevier.c...

  31. [40]

    Mohseni, N

    S. Mohseni, N. Zarei, E. D. Ragan, A Multidisciplinary Survey and Framework for Design and Evaluation of Explainable AI Systems, ACM Transactions on Interactive Intelligent Systems 1 (2021) 1–45. URL: http://arxiv.org/abs/1811.11839. doi:10.1145/3387166

  32. [41]

    Vilone, L

    G. Vilone, L. Longo, Classification of Explainable Artificial Intelli- gence Methods through Their Output Formats, Machine Learning and Knowledge Extraction 2021, Vol. 3, Pages 615-661 3 (2021) 615–661. URL: https://www.mdpi.com/2504-4990/3/3/32/htmhttps://www. mdpi.com/2504-4...

  33. [42]

    L. E. Holmquist, Intelligence on Tap: Artificial Intelligence as a New Design Material, Interactions 24 (2017) 28–33. URL: https://dl.acm.org/doi/10.1145/3085571. doi:10.1145/3085571/ ASSET/7D59872F-1989-49EA-8E74-B7228D403DB5/ASSETS/3085571. FP.PNG

  34. [43]

    Suresh, S

    H. Suresh, S. R. Gomez, K. K. Nam, A. Satyanarayan, Beyond Ex- pertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs, in: Proceedings of the 2021 CHI Conference on Human Factors in Computing Sys- tems, volume 16, ACM,...

  35. [44]

    D. Wang, Q. Yang, A. Abdul, B. Y. Lim, Designing Theory-Driven User- Centric Explainable AI, in: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, ACM, New York, NY, USA, 2019, pp. 1–15. URL: https://dl.acm.org/doi/10.1145/3290605. 3300831. doi:10.1...

  36. [45]

    Szymanski, V

    M. Szymanski, V. Vanden Abeele, K. Verbert, Designing and Eval- uating Explanations for a Predictive Health Dashboard: A User- 66 Centred Case Study, Conference on Human Factors in Computing Sys- tems - Proceedings (2024). URL:https://dl.acm.org/doi/10.1145/ 3613905.3637140. d...

  37. [47]

    Szymanski, C

    M. Szymanski, C. Conati, V. V. Abeele, K. Verbert, Designing and personalising hybrid multi-modal health explanations for lay users, in: Proceedings of the 10th Joint Workshop on Interfaces and Human Decision Making for Recommender Systems (IntRS 2023), volume 3534, CEUR-WS.or...

  38. [48]

    B. P. Knijnenburg, M. C. Willemsen, Evaluating Recommender Systems with User Experiments, in: Recommender Systems Hand- book, Springer US, Boston, MA, 2015, pp. 309–352. URL:https:// link.springer.com/10.1007/978-1-4899-7637-6_9. doi: 10.1007/ 978-1-4899-7637-6{\_}9

  39. [49]

    J. Qu, J. Arguello, Y. Wang, Understanding the cognitive influences of interpretability features on how users scrutinize machine-predicted categories, in: CHIIR 2023 - Proceedings of the 2023 Conference on Hu- man Information Interaction and Retrieval, Association for Computin...

  40. [50]

    Hwang, T

    J. Hwang, T. Lee, H. Lee, S. Byun, A clinical decision support system for sleep staging tasks with explanations from artificial intelligence: User-centered design and evaluation study, 2022. doi:10.2196/28659

  41. [51]

    Taylor, J

    P. Taylor, J. Fox, A. T. Pokropek, The development and evaluation of cadmium: A prototype system to assist in the interpretation of mammograms, Medical Image Analysis 3 (1999) 321–337. doi:10.1016/ S1361-8415(99)80027-9. 67

  42. [52]

    Julià-Sapé, C

    M. Julià-Sapé, C. Majós, Àngels Camins, A. Samitier, M. Baquero, M. Serrallonga, S. Doménech, E. Grivé, F. A. Howe, K. Opstad, J. Cal- var, C. Aguilera, C. Arús, Multicentre evaluation of the interpret decision support system 2.0 for brain tumour classification, NMR in Biomedi...

  43. [53]

    Srivastava, M

    S. Srivastava, M. Theune, A. Catala, The role of lexical alignment in human understanding of explanations by conversational agents, in: International Conference on Intelligent User Interfaces, Proceedings IUI, Association for Computing Machinery, 2023, pp. 423–435. doi:10. 114...

  44. [54]

    F. Liu, J. Zhou, M. Zuo, Y. Li, Dual-process theory-driven transparent approach for seniors to accept health misinformation detection results, Information Processing and Management 61 (2024). doi:10.1016/j. ipm.2024.103751

  45. [55]

    Sivaprasad, E

    A. Sivaprasad, E. Reiter, N. Tintarev, N. Oren, Evaluation of human- understandability of global model explanations using decision tree, in: Communications in Computer and Information Science, volume 1947, Springer Science and Business Media Deutschland GmbH, 2024, pp. 43–65. ...

  46. [56]

    Sovrano, F

    F. Sovrano, F. Vitali, An objective metric for explainable ai: How and why to estimate the degree of explainability, Knowledge-Based Systems 278 (2023). doi:10.1016/j.knosys.2023.110866

  47. [57]

    Upadhyay, P

    R. Upadhyay, P. Knoth, G. Pasi, M. Viviani, Explainable online health information truthfulness in consumer health search, Frontiers in Artificial Intelligence 6 (2023). doi:10.3389/frai.2023.1184851

  48. [58]

    Oberste, F

    L. Oberste, F. Rüffer, O. Aydingül, J. Rink, A. Heinzl, Designing user- centricexplanationsformedicalimagingwithinformedmachinelearning, in: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vo...

  49. [59]

    Katzmann, O

    A. Katzmann, O. Taubmann, S. Ahmad, A. Mühlberg, M. Sühling, H. M. Groß, Explaining clinical decision support systems in medical imaging 68 using cycle-consistent activation maximization, Neurocomputing 458 (2021) 141–156. doi:10.1016/j.neucom.2021.05.081

  50. [60]

    Hegselmann, C

    S. Hegselmann, C. Ertmer, T. Volkert, A. Gottschalk, M. Dugas, J. Varghese, Development and validation of an interpretable 3 day intensive care unit readmission prediction model using explainable boosting machines, Frontiers in Medicine 9 (2022). doi:10.3389/fmed. 2022.960296

  51. [61]

    Nauta, M

    M. Nauta, M. J. van Putten, M. C. Tjepkema-Cloostermans, J. P. Bos, M. van Keulen, C. Seifert, Interactive explanations of internal representations of neural network layers: An ex- ploratory study on outcome prediction of comatose patients,

  52. [62]

    Röhrl, H

    S. Röhrl, H. Maier, M. Lengl, C. Klenk, D. Heim, M. Knopp, S. Schumann, O. Hayden, K. Diepold, Explainable artificial intelli- gence for cytological image analysis, in: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture ...

  53. [63]

    Guezennec, J.Bouaud, B.Séroussi, Explainable artificial intelligence for breast cancer: A visual case-based reasoning approach, Artificial Intelligence in Medicine 94 (2019) 42–53

    J.B.Lamy, B.Sekar, G. Guezennec, J.Bouaud, B.Séroussi, Explainable artificial intelligence for breast cancer: A visual case-based reasoning approach, Artificial Intelligence in Medicine 94 (2019) 42–53. doi:10. 1016/j.artmed.2019.01.001

  54. [64]

    Deperlioglu, U

    O. Deperlioglu, U. Kose, D. Gupta, A. Khanna, F. Giampaolo, G. Fortino, Explainable framework for glaucoma diagnosis by im- age processing and convolutional neural network synergy: Analysis with doctor evaluation, Future Generation Computer Systems 129 (2022) 152–169. doi:10.1...

  55. [65]

    N. Das, S. Happaerts, I. Gyselinck, M. Staes, E. Derom, G. Brus- selle, F. Burgos, M. Contoli, A. T. Dinh-Xuan, F. M. Franssen, S.Gonem, N.Greening, C.Haenebalcke, W.D.Man, J.Moisés, R.Peché, V. Poberezhets, J. K. Quint, M. C. Steiner, E. Vanderhelst, M. Abdo, M. Topalovic, W....

  56. [66]

    van der Waa, S

    J. van der Waa, S. Verdult, K. van den Bosch, J. van Diggelen, T. Haije, B. van der Stigchel, I. Cocu, Moral decision making in human-agent teams: Human control and the role of explanations, Frontiers in Robotics and AI 8 (2021). doi:10.3389/frobt.2021.640647

  57. [67]

    W. Jin, M. Fatehi, R. Guo, G. Hamarneh, Evaluating the clinical utility of artificial intelligence assistance and its explanation on the glioma grading task, Artificial Intelligence in Medicine 148 (2024). doi:10.1016/j.artmed.2023.102751

  58. [68]

    Chanda, K

    T. Chanda, K. Hauser, S. Hobelsberger, T. C. Bucher, C. N. Garcia, C. Wies, H. Kittler, P. Tschandl, C. Navarrete-Dechent, S. Podlipnik, E. Chousakos, I. Crnaric, J. Majstorovic, L. Alhajwan, T. Foreman, S. Peternel, S. Sarap, İrem Özdemir, R. L. Barnhill, M. Llamas-Velasco, G...

  59. [69]

    D. Song, J. Yao, Y. Jiang, S. Shi, C. Cui, L. Wang, L. Wang, H. Wu, H. Tian, X. Ye, D. Ou, W. Li, N. Feng, W. Pan, M. Song, J. Xu, D. Xu, L. Wu, F. Dong, A new xai framework with feature explainability for tumors decision-making in ultrasound data: comparing with grad- cam, Co...

  60. [70]

    M. H. Lee, C. J. Chew, Understanding the effect of counterfactual explanations on trust and reliance on ai for human-ai collaborative clinical decision making, Proceedings of the ACM on Human-Computer Interaction 7 (2023). doi:10.1145/3610218

  61. [71]

    A. A. Tutul, E. H. Nirjhar, T. Chaspari, Investigating trust in human-ai collaboration for a speech-based data analytics task, International Jour- nal of Human-Computer Interaction (2024). doi:10.1080/10447318. 2024.2328910

  62. [72]

    Kyrimi, S

    E. Kyrimi, S. Mossadegh, N. Tai, W. Marsh, An incremental explana- tion of inference in bayesian networks for increasing model trustworthi- ness and supporting clinical decision making, Artificial Intelligence in Medicine 103 (2020). doi:10.1016/j.artmed.2020.101812

  63. [73]

    Stork, I

    L. Stork, I. Tiddi, R. Spijker, A. ten Teije, Explainable drug repur- posing in context via deep reinforcement learning, in: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), volume 13870 LNCS,...

  64. [74]

    S. G. Anjara, A. Janik, A. Dunford-Stenger, K. M. Kenzie, A. Collazo- Lorduy, M. Torrente, L. Costabello, M. Provencio, Examining explain- able clinical decision support systems with think aloud protocols, PLoS ONE 18 (2023). doi:10.1371/journal.pone.0291443

  65. [75]

    L. Xiao, H. Zhou, J. Fox, Towards a systematic approach for argu- mentation, recommendation, and explanation in clinical decision sup- port, Mathematical Biosciences and Engineering 19 (2022) 10445–10473. doi:10.3934/mbe.2022489. 71

  66. [76]

    Q. Wang, K. Huang, P. Chandak, M. Zitnik, N. Gehlenborg, Extending the nested model for user-centric xai: A design study on gnn-based drug repurposing, IEEE Transactions on Visualization and Computer Graphics 29 (2023) 1266–1276. doi:10.1109/TVCG.2022.3209435

  67. [77]

    Cabitza, C

    F. Cabitza, C. Natali, L. Famiglini, A. Campagner, V. Caccavella, E. Gallazzi, Never tell me the odds: Investigating pro-hoc explanations in medical decision making, Artificial Intelligence in Medicine 150 (2024). doi:10.1016/j.artmed.2024.102819

  68. [78]

    Metta, A

    C. Metta, A. Beretta, R. Guidotti, Y. Yin, P. Gallinari, S. Rinzivillo, F. Giannotti, Improving trust and confidence in medical skin lesion diagnosis through explainable deep learning, International Journal of Data Science and Analytics (2023). doi:10.1007/s41060-023-00401-z

  69. [79]

    Sabol, P

    P. Sabol, P. Sinčák, P. Hartono, P. Kočan, Z. Benetinová, A. Blichárová, Ľudmila Verbóová, E. Štammová, A. Sabolová-Fabianová, A. Jašková, Explainable classifier for improving the accountability in decision- making for colorectal cancer diagnosis from histopathological images,...

  70. [80]

    Sarkar, M

    N. Sarkar, M. Kumagai, S. Meyr, S. Pothapragada, M. Unberath, G. Li, S. R. Ahmed, E. B. Smith, M. A. Davis, G. D. Khatri, A. Agrawal, Z. S. Delproposto, H. Chen, C. G. Caballero, D. Dreizin, An aser ai/ml expert panel formative user research study for an interpretable interact...

  71. [81]

    F. M. Calisto, C. Santiago, N. Nunes, J. C. Nascimento, Breastscreening- ai: Evaluating medical intelligent agents for human-ai interactions, Artificial Intelligence in Medicine 127 (2022). doi:10.1016/j.artmed. 2022.102285

  72. [82]

    Cabitza, A

    F. Cabitza, A. Campagner, L. Famiglini, E. Gallazzi, G. A. L. Maida, Color shadows (part i): Exploratory usability evalua- tion of activation maps in radiological machine learning, Lec- ture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligenc...

  73. [83]

    Kumar, R

    A. Kumar, R. Manikandan, U. Kose, D. Gupta, S. C. Satapathy, Doc- tor’s dilemma: Evaluating an explainable subtractive spatial lightweight convolutional neural network for brain tumor diagnosis, ACM Transac- tions on Multimedia Computing, Communications and Applications 17 (20...

  74. [84]

    Singla, M

    S. Singla, M. Eslami, B. Pollack, S. Wallace, K. Batmanghelich, Ex- plaining the black-box smoothly—a counterfactual approach, Medical Image Analysis 84 (2023). doi:10.1016/j.media.2022.102721

  75. [85]

    Glick, M

    A. Glick, M. Clayton, N. Angelov, J. Chang, Impact of explainable artificial intelligence assistance on clinical decision-making of novice dental clinicians, JAMIA Open 5 (2022). doi:10.1093/jamiaopen/ ooac031

  76. [86]

    F. M. Calisto, C. Santiago, N. Nunes, J. C. Nascimento, Introduction of human-centric ai assistant to aid radiologists for multimodal breast image classification, International Journal of Human Computer Studies 150 (2021). doi:10.1016/j.ijhcs.2021.102607

  77. [87]

    Pumplun, F

    L. Pumplun, F. Peters, J. F. Gawlitza, P. Buxmann, Bringing machine learning systems into clinical practice: A design science approach to explainable machine learning-based clinical decision support systems, Journal of the Association for Information Systems 24 (2023) 953–979....

  78. [88]

    J. S. Chen, S. L. Baxter, A. V. D. Brandt, A. Lieu, A. S. Camp, J. L. Do, D. S. Welsbie, S. Moghimi, M. Christopher, R. N. Weinreb, L. M. Zangwill, Usability and clinician acceptance of a deep learning-based clinical decision support tool for predicting glaucomatous visual fie...

  79. [89]

    Metta, A

    C. Metta, A. Beretta, R. Guidotti, Y. Yin, P. Gallinari, S. Rinzivillo, F. Giannotti, Advancing dermatological diagnostics: Interpretable ai for enhanced skin lesion classification, Diagnostics 14 (2024). doi:10. 3390/diagnostics14070753. 73

  80. [90]

    Natali, L

    C. Natali, L. Famiglini, A. Campagner, G. A. L. Maida, E. Gallazzi, F. Cabitza, Color shadows 2: Assessing the impact of xai on diagnostic decision-making, in: Communications in Computer and Information Sci- ence, volume 1901 CCIS, Springer Science and Business Media Deutsch- ...

  81. [91]

    Gamage, U

    L. Gamage, U. Isuranga, D. Meedeniya, S. D. Silva, P. Yogarajah, Melanoma skin cancer identification with explainability utilizing mask guided technique, Electronics (Switzerland) 13 (2024). doi:10.3390/ electronics13040680

  82. [92]

    Cabitza, A

    F. Cabitza, A. Campagner, L. Famiglini, C. Natali, V. Caccavella, E. Gallazzi, Let me think! investigating the effect of explanations feeding doubts about the ai advice, in: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lec...

  83. [93]

    Derathé, F

    A. Derathé, F. Reche, P. Jannin, A. Moreau-Gaudry, B. Gibaud, S. Voros, Explaining a model predicting quality of surgical practice: a first presentation to and review by clinical experts, International Jour- nal of Computer Assisted Radiology and Surgery 16 (2021) 2009–2019. d...

  84. [94]

    Domínguez-Rodríguez, H

    S. Domínguez-Rodríguez, H. Liz-López, A. Panizo-LLedot, Álvaro Ballesteros, R. Dagan, D. Greenberg, L. Gutiérrez, P. Rojo, E. Otheo, J. C. Galán, S. Villanueva, S. García, P. Mosquera, A. Tagarro, C. Moraleda, D. Camacho, Testing the performance, adequacy, and applicability of...

  85. [95]

    J. E. Huh, J. H. Lee, E. J. Hwang, C. M. Park, Effects of expert- determined reference standards in evaluating the diagnostic performance of a deep learning model: A malignant lung nodule detection task on chest radiographs, Korean Journal of Radiology 24 (2023) 155–165. doi:1...

  86. [96]

    A. M. Antoniadi, M. Galvin, M. Heverin, L. Wei, O. Hardiman, C. Mooney, A clinical decision support system for the prediction of quality of life in als, Journal of Personalized Medicine 12 (2022). doi:10.3390/jpm12030435

  87. [97]

    Hendawi, J

    R. Hendawi, J. Li, S. Roy, A mobile app that addresses interpretability challenges in machine learning–based diabetes predictions: Survey-based user study, JMIR Formative Research 7 (2023). doi:10.2196/50328

  88. [98]

    A. J. Barda, C. M. Horvat, H. Hochheiser, A qualitative research frame- work for the design of user-centered displays of explanations for machine learning model predictions in healthcare, BMC Medical Informatics and Decision Making 20 (2020). doi:10.1186/s12911-020-01276-x

  89. [99]

    Panigutti, A

    C. Panigutti, A. Beretta, D. Fadda, F. Giannotti, D. Pedreschi, A. Per- otti, S. Rinzivillo, Co-design of human-centered, explainable ai for clinical decision support, ACM Transactions on Interactive Intelligent Systems 13 (2023). doi:10.1145/3587271

  90. [100]

    Kaczmarek-Majer, G

    K. Kaczmarek-Majer, G. Casalino, G. Castellano, M. Dominiak, O. Hryniewicz, O. Kamińska, G. Vessio, N. Díaz-Rodríguez, Ple- nary: Explaining black-box models in natural language through fuzzy linguistic summaries, Information Sciences 614 (2022) 374–399. doi:10.1016/j.ins.2022.10.010

  91. [101]

    Wysocki, J

    O. Wysocki, J. K. Davies, M. Vigo, A. C. Armstrong, D. Landers, R. Lee, A. Freitas, Assessing the communication gap between ai models and healthcare professionals: Explainability, utility and trust in ai-driven clinical decision-making, Artificial Intelligence 316 (2023). doi:...

  92. [102]

    Z. Jin, S. Cui, S. Guo, D. Gotz, J. Sun, N. Cao, Carepre, ACM Trans- actions on Computing for Healthcare 1 (2020). doi:10.1145/3344258

  93. [103]

    Gutiérrez, N

    F. Gutiérrez, N. N. Htun, V. V. Abeele, R. D. Croon, K. Verbert, Explaining call recommendations in nursing homes: a user-centered design approach for interacting with knowledge-based health decision support systems, in: International Conference on Intelligent User Interfaces,...

  94. [104]

    N. C. Rajashekar, Y. E. Shin, Y. Pu, S. Chung, K. You, M. Giuffre, C. E. Chan, T. Saarinen, A. Hsiao, J. Sekhon, A. H. Wong, L. V. Evans, R. F. Kizilcec, L. Laine, T. McCall, D. Shung, Human-algorithmic interaction using a large language model-augmented artificial intelligence...

  95. [105]

    S. S. Samuel, N. N. B. Abdullah, A. Raj, Interpretation of svm using data mining technique to extract syllogistic rules: Exploring the notion of explainable ai in diagnosing cad, in: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligenc...

  96. [106]

    W. J. She, C. S. Ang, R. A. Neimeyer, L. A. Burke, Y. Zhang, A. Jatowt, Y. Kawai, J. Hu, M. Rauterberg, H. G. Prigerson, P. Siriaraya, Investigation of a web-based explainable ai screening for prolonged grief disorder, IEEE Access 10 (2022) 41164–41185. doi:10.1109/ACCESS.2022.3163311

  97. [107]

    D. Gu, K. Su, H. Zhao, A case-based ensemble learning system for explainable breast cancer recurrence prediction, Artificial Intelligence in Medicine 107 (2020). doi:10.1016/j.artmed.2020.101858

  98. [108]

    Pisirir, J

    E. Pisirir, J. M. Wohlgemut, E. Kyrimi, R. S. Stoner, Z. B. Perkins, N. R. Tai, D. W. R. Marsh, A process for evaluating explanations for transparent and trustworthy ai prediction models, in: Proceedings - 2023 IEEE 11th International Conference on Healthcare Informatics, ICHI...

  99. [109]

    J. Hu, Y. Liang, W. Zhao, K. McAreavey, W. Liu, An interactive xai interface with application in healthcare for non-experts, in: Commu- nications in Computer and Information Science, volume 1901 CCIS, Springer Science and Business Media Deutschland GmbH, 2023, pp. 649–670. doi...

  100. [110]

    F. L. D. Morais, A. C. B. Garcia, P. S. M. D. Santos, L. A. P. Ribeiro, Do explainable ai techniques effectively explain their rationale? a case study 76 from the domain expert’s perspective, in: Proceedings of the 2023 26th International Conference on Computer Supported Coope...

  101. [111]

    Aliyeva, N

    K. Aliyeva, N. Mehdiyev, Uncertainty-aware multi-criteria decision analysis for evaluation of explainable artificial intelligence methods: A use case from the healthcare domain, Information Sciences 657 (2024). doi:10.1016/j.ins.2023.119987

  102. [112]

    Z. Wang, I. Samsten, V. Kougia, P. Papapetrou, Style-transfer counterfactual explanations: An application to mortality preven- tion of icu patients, Artificial Intelligence in Medicine 135 (2023). doi:10.1016/j.artmed.2022.102457

  103. [113]

    G. T. Berge, O. C. Granmo, T. O. Tveit, B. E. Munkvold, A. L. Ruthjersen, J. Sharma, Machine learning-driven clinical decision support system for concept-based searching: a field trial in a norwegian hospital, BMC Medical Informatics and Decision Making 23 (2023). doi:10.1186/...

  104. [114]

    Jaber, H

    D. Jaber, H. Hajj, F. Maalouf, W. El-Hajj, Medically-oriented design for explainable ai for stress prediction from physiological measurements, BMC Medical Informatics and Decision Making 22 (2022). doi:10.1186/ s12911-022-01772-2

  105. [115]

    Nagendran, P

    M. Nagendran, P. Festor, M. Komorowski, A. C. Gordon, A. A. Faisal, Quantifying the impact of ai recommendations with explanations on prescription decision making, npj Digital Medicine 6 (2023). doi:10. 1038/s41746-023-00955-z

  106. [116]

    L. V. Herm, K. Heinrich, J. Wanner, C. Janiesch, Stop ordering machine learning algorithms by their explainability! a user-centered investigation of performance and explainability, International Journal of Information Management 69 (2023). doi:10.1016/j.ijinfomgt.2022.102538

  107. [117]

    S. V. Kovalchuk, G. D. Kopanitsa, I. V. Derevitskii, G. A. Matveev, D. A. Savitskaya, Three-stage intelligent support of clinical decision making for higher trust, validity, and explainability, Journal of Biomedical Informatics 127 (2022). doi:10.1016/j.jbi.2022.104013. 77

  108. [118]

    A. Rind, D. Slijepčević, M. Zeppelzauer, F. Unglaube, A. Kranzl, B. Horsak, Trustworthy visual analytics in clinical gait analysis: A case study for patients with cerebral palsy, in: Proceedings - 2022 IEEE Workshop on TRust and EXpertise in Visual Analytics, TREX 2022, Instit...

  109. [119]

    Cheng, D

    F. Cheng, D. Liu, F. Du, Y. Lin, A. Zytek, H. Li, H. Qu, K. Veeramacha- neni, Vbridge: Connecting the dots between features and data to explain healthcare models, IEEE Transactions on Visualization and Computer Graphics 28 (2022) 378–388. doi:10.1109/TVCG.2021.3114836

  110. [120]

    Y. Du, A. M. Antoniadi, C. McNestry, F. M. McAuliffe, C. Mooney, The role of xai in advice-taking from a clinical decision support system: A comparative user study of feature contribution-based and example- based explanations, Applied Sciences (Switzerland) 12 (2022). doi:10. ...

  111. [121]

    C. He, V. Raj, H. Moen, T. Gröhn, C. Wang, L. M. Peltonen, S. Koivusalo, P. Marttinen, G. Jacucci, Vms: Interactive visualization to support the sensemaking and selection of predictive models, in: ACM International Conference Proceeding Series, Association for Computing Machin...

  112. [122]

    Khodabandehloo, D

    E. Khodabandehloo, D. Riboni, A. Alimohammadi, Healthxai: Collab- orative and explainable ai for supporting early diagnosis of cognitive decline, Future Generation Computer Systems 116 (2021) 168–189. doi:10.1016/j.future.2020.10.030

  113. [123]

    Panigutti, A

    C. Panigutti, A. Beretta, F. Giannotti, D. Pedreschi, Understanding the impact of explanations on advice-taking: a user study for ai-based clinical decision support systems, in: Conference on Human Factors in Computing Systems - Proceedings, Association for Computing Machin- e...

  114. [124]

    M. H. Lee, D. P. Siewiorek, A. Smailagic, A human-ai collaborative approach for clinical decision making on rehabilitation assessment, in: Conference on Human Factors in Computing Systems - Proceedings, Association for Computing Machinery, 2021. doi:10.1145/3411764. 3445472. 78

  115. [125]

    W. Jin, X. Li, M. Fatehi, G. Hamarneh, Guidelines and evaluation of clinical explainable ai in medical image analysis, Medical Image Analysis 84 (2023). doi:10.1016/j.media.2022.102684

  116. [126]

    M. H. Lee, D. P. Siewiorek, A. Smailagic, A. Bernardino, S. B. I. Badia, Co-design and evaluation of an intelligent decision support system for stroke rehabilitation assessment, Proceedings of the ACM on Human-Computer Interaction 4 (2020). doi:10.1145/3415227

  117. [127]

    W. Jin, X. Li, G. Hamarneh, Evaluating explainable ai on a multi- modal medical imaging task: Can existing algorithms fulfill clinical requirements?, AAAI Conference on Artificial Intelligence 36 (2022) 11945–11953. doi:10.1609/AAAI.V36I11.21452

  118. [128]

    M. H. Lee, D. P. Siewiorek, A. Smailagic, A. Bernardino, S. B. I. Badia, Interactive hybrid approach to combine machine and human intelligence for personalized rehabilitation assessment, in: ACM CHIL 2020 - Proceedings of the 2020 ACM Conference on Health, Inference, and Learn...

  119. [129]

    Huang, Z

    G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, Densely Con- nected Convolutional Networks, in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2017, pp. 2261–2269. URL: https://ieeexplore.ieee.org/document/8099726/. doi: 10. 1109/CVPR.2017.243

  120. [130]

    K. He, X. Zhang, S. Ren, J. Sun, Deep Residual Learning for Image Recognition, in: 2016 IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), IEEE, 2016, pp. 770–778. URL: http://ieeexplore.ieee.org/document/7780459/. doi:10.1109/CVPR.2016.90

  121. [131]

    Geurts, D

    P. Geurts, D. Ernst, L. Wehenkel, Extremely randomized trees, Machine Learning 63 (2006) 3–42. URL: https://link. springer.com/article/10.1007/s10994-006-6226-1. doi:10.1007/ S10994-006-6226-1/METRICS

  122. [132]

    J. H. Friedman, Greedy function approximation: A gra- dient boosting machine, The Annals of Statistics 29 79 (2001) 1189–1232. URL: https://projecteuclid.org/ journals/annals-of-statistics/volume-29/issue-5/ Greedy-function-approximation-A-gradient-boosting-machine/ 10.1214/ao...

  123. [133]

    S. M. Lundberg, S.-I. Lee, A Unified Approach to Interpreting Model Predictions, in: Advances in Neural Information Processing Systems, 2017, pp. 4766–4775. URL:https://dl.acm.org/doi/pdf/10.5555/ 3295222.3295230. doi:10.5555/3295222.3295230

  124. [134]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-CAM: Visual Explanations from Deep Networks via Gradient- Based Localization, Proceedings of the IEEE International Conference on Computer Vision 2017-October (2017) 618–626. doi:10.1109/ICCV. 2017.74

  125. [135]

    Venkatesh, M

    V. Venkatesh, M. G. Morris, G. B. Davis, F. D. Davis, User acceptance of information technology: Toward a unified view, MIS Quarterly: Management Information Systems 27 (2003) 425–478. doi:10.2307/ 30036540

  126. [136]

    Jeyaraj, Y

    A. Jeyaraj, Y. K. Dwivedi, V. Venkatesh, Intention in information systems adoption and use: Current state and research directions, In- ternational Journal of Information Management 73 (2023) 102680. doi:10.1016/J.IJINFOMGT.2023.102680

  127. [137]

    Lins de Holanda Coelho, P

    G. Lins de Holanda Coelho, P. H. P. Hanel, L. J. Wolf, The Very Efficient Assessment of Need for Cognition: De- veloping a Six-Item Version*, Assessment 27 (2020) 1870–1885. URL: https://journals.sagepub.com/doi/10. 1177/1073191118793208. doi: 10.1177/1073191118793208/ASSET/ 9...

  128. [138]

    O. P. John, S. Srivastava, The Big-Five trait taxonomy: History, mea- surement, and theoretical perspectives, 1999

  129. [139]

    Kouki, J

    P. Kouki, J. Schaffer, J. Pujara, J. O’Donovan, L. Getoor, Generating and Understanding Personalized Explanations in Hybrid Recommender Systems, ACM Transactions on Interactive Intelligent Systems (TiiS) 80 10 (2020). URL:https://dl.acm.org/doi/10.1145/3365843. doi:10. 1145/3365843

  130. [140]

    Coroama, A

    L. Coroama, A. Groza, Evaluation Metrics in Explainable Artificial In- telligence (XAI), in: Communications in Computer and Information Sci- ence, volume 1675 CCIS, Springer Science and Business Media Deutsch- land GmbH, 2022, pp. 401–413. doi:10.1007/978-3-031-20319-0{\_ }30

  131. [141]

    R. R. Hoffman, S. T. Mueller, G. Klein, J. Litman, Metrics for Explainable AI: Challenges and Prospects (2018) 1–50. URL:http: //arxiv.org/abs/1812.04608

  132. [142]

    S. G. Hart, L. E. Staveland, Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Re- search, in: P. A. Hancock, N. Meshkati (Eds.), Human Mental Workload, volume 52 of Advances in Psychology , North-Holland, 1988, pp. 139–183. URL:https://www.scienc...

  133. [143]

    Doshi-Velez, B

    F. Doshi-Velez, B. Kim, Towards A Rigorous Science of Interpretable Machine Learning, Arxiv (2017) 1–13. URL:http://arxiv.org/abs/ 1702.08608

  134. [144]

    Szymanski, M

    M. Szymanski, M. Millecamp, K. Verbert, Visual, textual or hybrid: The effect of user expertise on different explanations, International Con- ference on Intelligent User Interfaces, Proceedings IUI (2021) 109–119. URL: /doi/pdf/10.1145/3397481.3450662?download=true. doi:10. 11...

  135. [145]

    O. Asan, A. E. Bayrak, A. Choudhury, Artificial Intelligence and Human Trust in Healthcare: Focus on Clinicians, Journal of Medical Internet Research 22 (2020) e15154. URL:https://www.jmir.org/ 2020/6/e15154. doi:10.2196/15154

  136. [146]

    P. K. Kahr, G. Rooks, M. C. Willemsen, C. C. P. Snijders, Under- standing Trust and Reliance Development in AI Advice: Assessing Model Accuracy, Model Explanations, and Experiences from Previous Interactions, ACM Transactions on Interactive Intelligent Systems 81 14 (2024) 1–3...

  137. [147]

    Ullah, J

    N. Ullah, J. A. Khan, I. De Falco, G. Sannino, Explainable Artificial Intelligence: Importance, Use Domains, Stages, Output Shapes, and Challenges, ACM Computing Surveys 57 (2024). URL: https://dl.acm.org/doi/10.1145/3705724. doi:10.1145/3705724/ ASSET/357B3925-47D2-415F-B28F-...

  138. [148]

    Re- trieved June 6, 2025, from https://www.merriam-webster.com/dictio- nary/satisfaction

    Satisfaction, in: Merriam-Webster.com, Merriam-Webster, 2025. Re- trieved June 6, 2025, from https://www.merriam-webster.com/dictio- nary/satisfaction

  139. [149]

    Brooke, SUS – a quick and dirty usability scale, 1996, pp

    J. Brooke, SUS – a quick and dirty usability scale, 1996, pp. 189–194

  140. [150]

    F. D. Davis, Perceived usefulness, perceived ease of use, and user acceptance of information technology, MIS Quarterly: Management Information Systems 13 (1989) 319–339. doi:10.2307/249008

  141. [151]

    Velmurugan, C

    M. Velmurugan, C. Ouyang, C. Moreira, R. Sindhgatta, Developing a Fidelity Evaluation Approach for Interpretable Machine Learning (2021). URL: http://arxiv.org/abs/2106.08492

  142. [152]

    M. A. Webb, J. P. Tangney, Too Good to Be True: Bots and Bad Data From Mechanical Turk, Perspectives on Psychological Science (2022). URL: https://scholar.google.com/scholar_url?url=https:// journals.sagepub.com/doi/pdf/10.1177/17456916221120027&hl= es&sa=T&oi=ucasa&ct=usl&ei=...

  143. [153]

    P. Pu, L. Chen, R. Hu, A user-centric evaluation framework for rec- ommender systems, in: Proceedings of the fifth ACM conference on Recommender systems - RecSys ’11, ACM Press, New York, New York, USA, 2011, p. 157. URL:http://dl.acm.org/citation.cfm?doid= 2043932.2043962. do...

  144. [154]

    Y. Jin, L. Chen, W. Cai, X. Zhao, Y. Jin, L. Chen, W. Cai, X. Zhao, CRS-Que: A User-centric Evaluation Framework for Conversational 82 Recommender Systems, ACM Transactions on Recommender Systems 2 (2024) 1–34. URL:https://dl.acm.org/doi/pdf/10.1145/3631534. doi:10.1145/363153...

  145. [156]

    It is important for me to know how uncer- tain (in %) the model is about its recom- mendation AI model performance Metric [116] Accuracy

  146. [157]

    Accuracy, Precision, Recall, F1

  147. [158]

    Sensitivity, Specificity, AUC

  148. [159]

    Absolute error, Percentage of error

  149. [160]

    Accuracy Alignment with situational context Closed Ques- tions

  150. [161]

    Did you feel that the AI assistance acceler- ated or slowed the workflow? Open Questions

  151. [162]

    We gathered feedback from the clinicians on the usefulness of information integrated into the conceptual prototype (MR1) and the correspondence to the diag- nostic work- flow (MR4)

  152. [163]

    What suggestions would you have for mak- ing AI and XAI usable in the medical do- main?

  153. [164]

    In these interviews we discussed the usabil- ity of the visual interface and how such a visual analytics tool can be integrated into the clinical workflow

  154. [165]

    Adoption strategy for the tool

  155. [166]

    How did you use each explanation strat- egy during the experiment?

  156. [167]

    Wasthereanynotablestrategyforadopt- ing the explanations rather than merely accepting the information in the explana- tions? 84 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  157. [168]

    What elements of the tool should be cus- tomizable to specific user needs? Case difficulty Closed Questions

  158. [169]

    the mean difficulty feedback by the doctors (1: the task is easy, 2: the task is normal, 3: the task is difficult)

  159. [170]

    Participants reported difficulty

  160. [171]

    Participants reported the perceived degree of difficulty (or complexity) of the case Metric [63] Authors reported difficulty

  161. [177]

    Authors reported difficulty

  162. [178]

    Authors reported difficulty Cognitive load Closed Questions

  163. [179]

    The barplot with feature contribution is easy to interpret

  164. [180]

    The colour bar with the score is easy to interpret

  165. [181]

    The scatterplot with all patients is easy to interpret

  166. [182]

    Explanations are easily understandable

  167. [183]

    Explanations help reducing the learning time on the system

  168. [184]

    Which XAI solutions(s) were easier to in- terpret via visual maps?

  169. [185]

    The visualizations are clear and I don’t spend much time on their interpretation

  170. [187]

    NASA TLX [Effort and Frustration]

  171. [188]

    a4) User experienc 85 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  172. [189]

    Did the AI assistance make the task less complex, demanding, or exacting?

  173. [190]

    Was the user interface easy or frustrating to use?

  174. [191]

    I was insecure, discouraged, and stressed while using the system

  175. [192]

    The system helped me think through and complete the assessment tasks with less effort

  176. [193]

    Was the AI system overwhelming or not?

  177. [194]

    Using RadiologyAI, I easily found the in- formation I wanted to help me decide on a diagnosis

    Using RadiologyAI was very frustrating. Using RadiologyAI, I easily found the in- formation I wanted to help me decide on a diagnosis. (reversed) Using RadiologyAI took too much time. Using RadiologyAI was easy. (reversed) Using RadiologyAI re- quired too much effort. Using Ra...

  178. [195]

    I think that most people would learn to understand the explanations very quickly Open Questions

  179. [197]

    What could be changed to further reduce workload? What features could make the process less rushed? What features could minimize feelings of irritation/annoyance

  180. [198]

    What features could make the task less de- manding, complex, and exacting? What features could help decrease the mental ac- tivity required to carry out the task? What features could improve/minimize fatigue? Complete- ness Metric [125] The area under PC (AUPC(H)) 86 Table B.1...

  181. [199]

    Even when my initial course of action was different to what CORONET recom- mended, I still had full confidence in my original decision

  182. [200]

    When my initial decision was the same as CORONET had recommended, I felt reas- sured

  183. [207]

    Participants reported acceptance of the re- sult

  184. [209]

    XAI increases clinicians’ confidence in their own diagnoses

  185. [210]

    his or her confidence in the proposed diag- nosis

  186. [212]

    Did the AI assistance increase your level of confidence when grading splenic injuries?

  187. [213]

    How confident you are with your diagnosis

  188. [214]

    Do you feel confident to decide to use the tool?

  189. [215]

    Participants reported confidence on their decision 87 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  190. [216]

    I feel very confident about using the system

  191. [220]

    Participants reported confidence on their decision

  192. [221]

    After seeing the explanation of prediction, I am more inclined to change my decision if the system prediction conflicts with my initial belief

  193. [222]

    After seeing the explanation of prediction, I feel more confident to make my decision if the system prediction supports my initial belief Open Ques- tions

  194. [223]

    Do you feel confident with the priority sug- gestions given by the application? User Behaviour

  195. [224]

    Participants were asked to estimate the patient’s chances of developing an acute MI in the near future on a scale from 0 to 100% and their confidence in the estimate on a sliding scale

  196. [225]

    The confidence shift was measured as the difference of the reported participant’s con- fidence in the estimate before and after receiving the AI advice Contrastivity Closed Ques- tions

  197. [226]

    Participants reported contrastivity Controllabil- ity Closed Ques- tions

  198. [227]

    I could change the level of detail on demand Open Questions

  199. [228]

    Do you have a sense of control when you provide feedback to the application?

  200. [229]

    We also conducted a synthetic data experiment on the glioma grading task

    To what degree is an interactive method preferable over a non-interactive one? 88 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper) Correctness Metric [125] For the truthfulness assessment, we con- ducted cumulativ...

  201. [230]

    This is the pre- requisite for G4

    G3: Truthfulness Explanations should truthfully reflect the AI model decision process. This is the pre- requisite for G4

  202. [231]

    The area under PC (AUPC(H))

  203. [232]

    modality Shapley value [defined in study]

  204. [233]

    An explanation should truth- fully reflect the model decision pro- cess

    uthfulness. An explanation should truth- fully reflect the model decision pro- cess. Curiosity User Behaviour

  205. [234]

    Notably, when faced with content scenarios, participants asked an average of 1.5 more questions than in risk scenarios

  206. [235]

    There were an average of 3.9 queries in content sce- narios compared to 2.4 in risk scenarios

    Teams using GutGPT in con- tent scenar- ios submitted more queries to the chatbot than teams in risk scenarios. There were an average of 3.9 queries in content sce- narios compared to 2.4 in risk scenarios. This indicates that the chatbot feature was used more heavily when mak...

  207. [236]

    The system with explanations would help me completing the assessment in less time

  208. [237]

    The system without explanations would help me completing the assessment in less time

  209. [238]

    Which XAI solutions(s) had the better per- formance in terms of running-time?

  210. [239]

    89 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

    I am able to make decisions faster thanks to XAI view of the system. 89 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  211. [240]

    Did you feel that the AI assistance acceler- ated or slowed the workflow?

  212. [241]

    I received the explanations in a timely and efficient manner User Behaviour

  213. [242]

    Average completion rate (depending on the top completion limit for each task)

  214. [243]

    Mean task completion time (in seconds)

  215. [244]

    Completion time when presented with a diagnostic choice for individual and overall questions

  216. [245]

    Completion time across the four conditions

  217. [246]

    Similar to yes- rate, completion time does not mea- sure performance

    5) Completion Time: The average time (in seconds) it took participants to agree or disagree with the system. Similar to yes- rate, completion time does not mea- sure performance. However, it tells us how long it took participants to make decisions, which may be related to part...

  218. [247]

    The main objective of this is to check the accuracy and the time taken to make the decisions

  219. [248]

    Duration of Decision Making

  220. [249]

    Column Steps indicates the minimum num- ber of steps (in terms of links to click, overviews to open, or questions to pose) required by each explanatory tool to pro- vide the correct answer

  221. [250]

    Explanation power Closed Questions

    Time The time spent in making decisions. Explanation power Closed Questions

  222. [251]

    The extent the map highlights structures or findings that allow for the correct diagnosis to be reached more easily than if the map was not seen

  223. [252]

    The barplot with feature contribution con- vinces me to accept or reject the model’s recommendation 90 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  224. [253]

    The colour bar with the score convinces me to accept or reject the model’s recommen- dation

  225. [254]

    The scatterplot with all patients convinces me to accept or reject the model’s recom- mendation

  226. [255]

    The group of sentences provides truthful statements and avoids providing informa- tion not supported by evidence (maxim of quality) Open Ques- tions

  227. [256]

    Are you convinced with the interpretations that the system provides? User Behaviour

  228. [257]

    Yes-rate does not measure performance

    (4) Yes Rate: the percentage of times par- ticipants agreed with the system. Yes-rate does not measure performance. Rather, it pro- vides insights into participants’ ten- dencies to agree with the system across interface conditions

  229. [258]

    It measures the percentage shift in judg- ment after advice, and it quantifies how much the participants follow the advice they receive

  230. [259]

    Change in mental model (CMM) Difference in perceived feature importance before and after viewing model explanation

  231. [260]

    we analyzed the number of times when AI explanations assisted participants to change and make ‘right’ or ‘wrong’ deci- sions 91 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  232. [261]

    Carnap’s main criteria of explication ade- quacy [11] are similar- ity, exactness and fruitfulness.5 Similarity means that the explica- tum should be detailed about the explicandum, in the sense that at least many of the intended uses of the expli- candum, brought out in the c...

  233. [262]

    Adherence The adherence of decision- making to recommendations

  234. [263]

    Di, j as the behavioral distrust measure of annotator i for batch j, quantified as the average absolute discrepancy between the AI decision and the anno- tator’s decision (Chu et al., 2020)

  235. [264]

    Information correctness Closed Questions

    Finally, the behavioral measure of trust bi, k of annotator i at sample k is defined as the inverse of the absolute dis- crepancy between the human annotator i and ML prediction for sample k. Information correctness Closed Questions

  236. [265]

    Participants reported plausibity

  237. [266]

    Participants reported medical relevance

  238. [267]

    Images in the video look like a chest X-ray

  239. [268]

    Does it look reasonable at a glance?

  240. [269]

    Are the highlighted sentences topically re- lated to the query by using either TF-IDF, BM25 or BioBERT?

  241. [270]

    Are the top sentences (with the best model between TF-IDF, BM25, or BioBERT) cor- rectly supported by scientific evidence (sci- entific journal articles)? 92 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  242. [271]

    Local outlier factor measures the closeness of the counterfactuals to the training data distri- bution of the desired class [30]

    The group of sentences provides truthful statements and avoids providing informa- tion not supported by evidence (maxim of quality) Metric [112] The range of validity is between 0 and 1, and a higher score will be desired. Local outlier factor measures the closeness of the cou...

  243. [272]

    Validity is defined as the fraction of the counterfactuals that are valid based on the desired target class [29]

  244. [273]

    Foreign Object Preservation (FOP) score [defined in study]

  245. [274]

    Frechet Inception Distance (FID) [defined in study]

  246. [275]

    overlap in ontological explanations between the clinicians and the AI

  247. [276]

    Degree of Truth [defined in study] Information expectedness Closed Questions

  248. [277]

    The extent the map highlights the struc- tures that present a fracture or those in which fractures are excluded according to the machine’s suggestion

  249. [278]

    Participants reported whether the fact was new to them or not

  250. [279]

    Images in the video look like the chest X- ray from a given subject

  251. [280]

    [System - Condition X] generates new in- sights on patient’s performance

  252. [281]

    [Tool - Condition X] generates new insights on patient’s performance 93 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  253. [282]

    How closely the highlighted areas of the heatmap match with your clinical judg- ment?

  254. [283]

    How closely does the highlighted area of the color map match with your clinical judg- ment?

  255. [284]

    Participants reported with the shown infor- mation

  256. [285]

    The system provided new insights on pa- tient’s performance for assessment

  257. [286]

    The AI predictions align with my medical opinion

  258. [287]

    The explanation of the prediction is clear and reasonable Metric [125] Feature portion (FP), Modality-specific fea- ture importance (MSFI) [defined in study]

  259. [288]

    MSFI [defined in study]

  260. [289]

    We evaluated our XAI’s alignment with the clinicians’ ontological explanations and annotated regions of in- terest (ROI) by assessing their overlap

    We used these explanations to determine the extent to which the XAI system and clinicians detected similar expla- nations on the same lesions. We evaluated our XAI’s alignment with the clinicians’ ontological explanations and annotated regions of in- terest (ROI) by assessing ...

  261. [290]

    Does a reference to a standard scale facili- tates critical assessment of a recommenda- tion?

  262. [291]

    Did the explanations correspond well to your perception of the important waveform patterns? Intention to use Closed Questions

  263. [293]

    Participants reported whether they want to continue using the system

  264. [294]

    I intend to use the tool to have a second opinion on the patient risks I feel like I will use it in the future 94 Table B.11 – continued from previous page Property Measure- ment Type Pa- per Measurement (quote from the paper)

  265. [295]

    I would use [System - Condition X] to un- derstand and assess patient’s performance

  266. [296]

    I would use [Tool - Condition X] to under- stand and assess patient’s performance

  267. [297]

    I want to use this system for auto-decision- making in brain tumor diagnosis

  268. [298]

    Participants reported whether they would use the system

  269. [299]

    Would you be willing to use a productized tool like this in clinical practice?

  270. [300]

    Would you use a productized tool with these features in the future?

  271. [2020]

    URL: https://research.utwente.nl/en/publications/ interactive-explanations-of-internal-representations-of-neural-ne

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.