REVIEW 5 major objections 5 minor 86 references
Explainability and AI Confidence in Clinical Decision Support Systems: Effects on Trust, Diagnostic Performance, and Cognitive Load in Breast Cancer Care
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Low AI confidence makes clinicians pause; high confidence breeds overreliance.
desk verdict The high-confidence overreliance finding is likely an artifact of the outcome coding, but the low-confidence effects and the study design make this worth a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The experimental apparatus is a web-based CDSS built on a U-Net segmentation model and a CNN classifier trained on a public breast ultrasound dataset with 780 images and 81 percent accuracy. The mechanism carrying the argument is the staged explainability design: baseline (no AI), classification only, probability distribution, tumor localization, and enhanced localization with high and low confidence regions. The outcome measures are trust and agreement ratings per image, agreement-above-neutral treated as adopting the AI's label, and two NASA-TLX items (mental demand and stress) after each stage. Mixed-effects models estimate the effect of confidence level while controlling for participant-level variation.
What would settle it
Re-run the study with a design where participants must always enter their final diagnosis explicitly, without the agreement-rating shortcut. If the high-confidence performance decrement disappears or reverses when the final diagnosis is recorded directly, the paper's overreliance claim is falsified by its own outcome construction.
Extended reading notes
Core claim
The central claim is that AI confidence scores act as a trust-calibration signal with asymmetric effects: low confidence makes clinicians more cautious and less accepting of AI advice, whereas high confidence increases trust enough to produce measurable overreliance and a small drop in diagnostic performance. Using an interrupted time series design, the authors compared a no-support baseline with four explainability conditions and modeled trust, agreement, performance, and diagnosis duration as functions of confidence level. The effect sizes are modest — low confidence reduced trust by 0.16 points and agreement by 0.19 points on a 5-point scale and increased diagnosis duration, while high confidence reduced performance by a small but statistically significant margin, which the authors attribute to automation bias rather than to the AI's advice being wrong more often.
Load-bearing premise
The result that high confidence hurts performance rests on treating an agreement rating above neutral as the participant adopting the AI's diagnosis; if that mapping misrepresents what clinicians actually decided, the overreliance finding could be an artifact of how performance was coded rather than a real behavioral change.
Editorial extensions
If this is right
- Displaying low confidence will make clinicians more skeptical and slower, which may be desirable for error-prone AI but costly in time-critical settings.
- Displaying high confidence on a system with imperfect accuracy will nudge a small but real fraction of decisions toward incorrect AI labels, so high-confidence displays should be paired with safeguards.
- Tumor localization with probability estimates raises stress without improving mental demand, so explainability features are not neutral additions to the interface.
- Age, experience, job role, gender, and race are associated with different perceptions of AI usefulness and complexity, so a single explanation design is unlikely to fit all clinician groups.
Reading between the lines
- Editorial inference: the asymmetric confidence effect suggests a practical design rule — degrade or flag low-confidence outputs rather than suppress them, since low confidence already triggers useful caution.
- Editorial inference: the performance measurement issue could be resolved by a forced-choice final diagnosis per image; such a design would also let the authors separate agreement from adoption.
- Editorial inference: testing the same confidence thresholds with a more accurate, calibrated model would show whether overreliance scales with model accuracy or is a fixed human tendency.
- Editorial inference: the stress finding for localization suggests that visual complexity — not just confidence — drives cognitive load, and could be tested by isolating localization with and without probability numbers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a web-based interrupted time-series experiment with 28 U.S. healthcare professionals who diagnosed breast ultrasound images under an AI-based CDSS with increasing levels of explainability. The authors model how AI confidence scores (none, low, high) affect trust, agreement, diagnostic performance, and diagnosis duration, and they analyze how explainability affects mental demand and stress. They report that low confidence significantly decreased trust and agreement and increased diagnosis duration, while high confidence was associated with a small but significant decrease in performance, which they interpret as overreliance. They also report that the localization intervention increased stress and that demographic factors were associated with perceptions of AI.
Significance. The study addresses a practically important question in clinical decision support, and the low-confidence results—decreased trust and agreement and increased diagnosis duration—are internally consistent, within-participant findings that could inform designs for calibrated reliance. The experiment uses a clinically relevant task and a realistic AI model trained on breast ultrasound data, which strengthens ecological validity relative to abstract prediction tasks. However, the headline claim that high confidence caused overreliance and reduced diagnostic accuracy rests on a performance measure that imputes the AI label on any trial with above-neutral agreement, so the effect is not a clean behavioral outcome. The abstract also overstates the high-confidence trust result, and the discussion contradicts the mental-demand table. These issues are load-bearing for the paper's main conclusions.
major comments (5)
- [Section 3.4, Performance] The outcome labeled 'diagnostic performance' is not an observed final diagnosis on trials where agreement is above neutral; the participant's decision is 'considered aligned with the AI's suggestion' and is therefore scored as the AI label. Table 2's high-confidence performance coefficient (β = −0.015, p = 0.020) is consequently not a direct measure of decision accuracy: high confidence and agreement are modeled separately but are not independent, and the same agreement ratings are used to construct the performance outcome. Because no independent human diagnosis was elicited on those trials, the interpretation in Section 4.2 and the abstract that high confidence 'led to overreliance, reducing diagnostic accuracy' is not supported by the data as reported.
- [Abstract and Section 6] The abstract states that 'high confidence scores substantially increased trust' and Section 6 repeats that high confidence 'substantially improved trust,' but Table 2 reports β = 0.103, p = 0.169 for high-confidence trust, which Section 4.2 correctly describes as not statistically significant. The abstract and conclusion should be corrected to match the reported result. The conclusion's statement that 'AI confidence score also elevated stress levels' is also unsupported, since stress is modeled by intervention in Table 1 and never by confidence score.
- [Section 5.1 versus Table 1] The Discussion states that the fourth condition 'resulted in increased mental demand,' but Table 1 reports β = −0.107, p = 0.501 for the 4th Intervention, a nonsignificant decrease, and Section 4.1 describes a slight decrease. This is an internal contradiction that affects the interpretation of RQ1; the text should be aligned with the table.
- [Sections 3.3 and 4.2] The experiment uses a fixed intervention order for all participants, and the mixed-effects models in Table 2 include no time or order term. The low- and high-confidence effects are therefore confounded with practice, fatigue, and cumulative interface exposure, so the causal language in Section 5.2 is stronger than the design supports. Please add a sensitivity analysis with session or trial order as a covariate, or explicitly temper the causal claims.
- [Section 4.3, Table 3] The demographic ANOVAs appear to treat repeated post-intervention observations as independent units; for example, gender on complexity perception is reported as F = 92.97, p = 6.61e-16, which is implausible with 28 participants unless each observation is counted as an independent case. Please specify the unit of analysis and the number of observations per cell, and use participant-level or mixed-effects models before drawing the demographic conclusions in RQ3.
minor comments (5)
- [Section 3.4] The agreement and trust scales are described as 5-point Likert scales 'from 0 to 5,' which is six response options; please clarify the actual scale and the neutral point.
- [Sections 3.3 and 3.4] Section 3.3 says participants give their own diagnosis when agreement is 'below 3,' while Section 3.4 refers to agreement 'above neutral'; specify how the neutral response (e.g., agreement = 3) is handled.
- [Tables 1 and 2] The intervention numbering is unclear: Section 3.3 lists Baseline plus Interventions I–IV, but Tables 1–2 use '1st/2nd/3rd/4th Intervention' without a consistent mapping; label all conditions the same way.
- [Section 3.4, AI Confidence Score] The 90% threshold used to dichotomize low versus high AI confidence is not justified; report a sensitivity analysis or motivate the cutoff.
- [General] There are multiple typographical errors, including 'mental memand' in Section 4.3 and 'V ariable' in the Table 1 and Table 2 captions; please proofread the manuscript.
Circularity Check
No significant circularity: the study is empirical, the central behavioral effects are new measurements, and the only construct-level concern (performance coding from agreement) is a measurement-validity issue, not a circular reduction.
full rationale
This paper is a web-based user experiment with no derivation chain, no fitted model whose parameters are relabeled as predictions, and no uniqueness theorem or ansatz imported from prior work. The self-citations ([68] for the CDSS architecture, and [10,25,53] for background and prior related findings) are not load-bearing: the AI system is described concretely with its public training dataset and an 81% accuracy figure, and the trust, agreement, cognitive-load, and diagnosis-duration outcomes are newly collected behavioral data. The one construct that warrants scrutiny is the Section 3.4 performance definition, under which agreement above neutral is coded as adoption of the AI label, while only below-neutral trials elicit an independent clinician diagnosis. This makes the performance variable a composite of agreement and AI correctness rather than a pure measure of independent diagnostic skill, which is a genuine measurement-validity limitation and is not acknowledged in Section 5.4. However, it is not circular: the high-confidence performance coefficient (beta = -0.015, p = 0.020) is not equal by construction to the confidence-agreement effect (beta = 0.108, p = 0.121), and the direction of the performance effect depends on empirical AI and human accuracy rates, so the finding does not reduce to its own inputs. No step in the paper satisfies the defined circularity patterns, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- confidence threshold =
90%
- agreement cutoff =
above neutral, exact Likert value unspecified
assumptions (4)
- domain assumption Ground-truth labels in the public breast ultrasound dataset are correct.
- domain assumption The fixed order of interventions (baseline, classification, probability, localization, confidence localization) does not confound intervention effects.
- standard math Mixed-effects and ANOVA model assumptions hold, including normal errors, independent observations, and positive variance components.
- domain assumption The 28 participants are representative of the clinical population of interest.
Cite this review
Pith. "Pith review of Explainability and AI Confidence in Clinical Decision Support Systems: Effects on Trust, Diagnostic Performance, and Cognitive Load in Breast Cancer Care." pith.science (2026). https://pith.science/paper/3HAPWIAX
@misc{pith2026250116693,
author = {Pith},
title = {Pith review of: Explainability and AI Confidence in Clinical Decision Support Systems: Effects on Trust, Diagnostic Performance, and Cognitive Load in Breast Cancer Care},
year = {2026},
howpublished = {\url{https://pith.science/paper/3HAPWIAX}},
note = {Machine review of arXiv:2501.16693}
}
read the original abstract
Artificial Intelligence (AI) has demonstrated potential in healthcare, particularly in enhancing diagnostic accuracy and decision-making through Clinical Decision Support Systems (CDSSs). However, the successful implementation of these systems relies on user trust and reliance, which can be influenced by explainable AI. This study explores the impact of varying explainability levels on clinicians trust, cognitive load, and diagnostic performance in breast cancer detection. Utilizing an interrupted time series design, we conducted a web-based experiment involving 28 healthcare professionals. The results revealed that high confidence scores substantially increased trust but also led to overreliance, reducing diagnostic accuracy. In contrast, low confidence scores decreased trust and agreement while increasing diagnosis duration, reflecting more cautious behavior. Some explainability features influenced cognitive load by increasing stress levels. Additionally, demographic factors such as age, gender, and professional role shaped participants' perceptions and interactions with the system. This study provides valuable insights into how explainability impact clinicians' behavior and decision-making. The findings highlight the importance of designing AI-driven CDSSs that balance transparency, usability, and cognitive demands to foster trust and improve integration into clinical workflows.
Figures
Reference graph
Works this paper leans on
-
[1]
Interna- tional evaluation of an AI system for breast cancer screening,
S. M. McKinney, M. Sieniek, V. Godbole, J. God- win, N. Antropova, H. Ashrafian, T. Back, M. Chesus, G. S. Corrado, and A. Darzi, “Interna- tional evaluation of an AI system for breast cancer screening,” Nature, vol. 577, no. 7788, pp. 89–94,
-
[2]
Attitudes towards trusting artificial intelligence insights and factors to prevent the passive adherence of GPs: a pi- lot study,
M. Micocci, S. Borsci, V. Thakerar, S. Walne, Y. Manshadi, F. Edridge, D. Mullarkey, P. Buckle, and G. B. Hanna, “Attitudes towards trusting artificial intelligence insights and factors to prevent the passive adherence of GPs: a pi- lot study,” Journal of Clinical Medicine , vol. 10, no. 14, p. 3101, 2021. Publisher: MDPI
2021
-
[3]
Deep learn- ing for digital pathology image analysis: A com- prehensive tutorial with selected use cases,
A. Janowczyk and A. Madabhushi, “Deep learn- ing for digital pathology image analysis: A com- prehensive tutorial with selected use cases,” Jour- nal of pathology informatics , vol. 7, no. 1, p. 29,
-
[4]
Deep learning so- lutions for skin cancer detection and diagnosis,
H. Nahata and S. P. Singh, “Deep learning so- lutions for skin cancer detection and diagnosis,” Machine Learning with Health Care Perspective: Machine Learning and Healthcare , pp. 159–182,
-
[5]
Voxel-based dose prediction with multi-patient atlas selection for automated radiotherapy treatment planning,
C. McIntosh and T. G. Purdie, “Voxel-based dose prediction with multi-patient atlas selection for automated radiotherapy treatment planning,” Physics in Medicine & Biology , vol. 62, no. 2, p. 415, 2016. Publisher: IOP Publishing
2016
-
[6]
A. Choudhury, O. Asan, and J. E. Medow, “Effect of risk, expectancy, and trust on clini- cians’ intent to use an artificial intelligence sys- tem – Blood Utilization Calculator,” Applied Er- gonomics, vol. 101, p. 103708, May 2022
work page 2022
-
[7]
A. ˇCartolovni, A. Tomiˇ ci´ c, and E. L. Mosler, “Ethical, legal, and social considerations of AI- based medical decision-support tools: A scoping review,” International Journal of Medical Infor- matics, vol. 161, p. 104738, 2022. Publisher: El- sevier
work page 2022
-
[8]
Do as AI say: susceptibility in deployment of clinical decision- aids,
S. Gaube, H. Suresh, M. Raue, A. Merritt, S. J. Berkowitz, E. Lermer, J. F. Coughlin, J. V. Gut- tag, E. Colak, and M. Ghassemi, “Do as AI say: susceptibility in deployment of clinical decision- aids,” npj Digital Medicine , vol. 4, pp. 1–8, Feb
Show all 86 references
-
[9]
High-performance medicine: the convergence of human and artificial intelligence,
E. J. Topol, “High-performance medicine: the convergence of human and artificial intelligence,” Nature medicine, vol. 25, no. 1, pp. 44–56, 2019. Publisher: Nature Publishing Group US New York
2019
-
[10]
Arti- ficial Intelligence and Human Trust in Healthcare: Focus on Clinicians,
O. Asan, A. E. Bayrak, and A. Choudhury, “Arti- ficial Intelligence and Human Trust in Healthcare: Focus on Clinicians,” Journal of Medical Internet Research, vol. 22, p. e15154, June 2020. Company: Journal of Medical Internet Research Distributor: Journal of Medical Internet ...
2020
-
[11]
Humans and au- tomation: Use, misuse, disuse, abuse,
R. Parasuraman and V. Riley, “Humans and au- tomation: Use, misuse, disuse, abuse,” Human factors, vol. 39, no. 2, pp. 230–253, 1997. Pub- lisher: SAGE Publications Sage CA: Los Angeles, CA
1997
-
[12]
Ex- tending the technology acceptance model to as- sess automation,
M. Ghazizadeh, J. D. Lee, and L. N. Boyle, “Ex- tending the technology acceptance model to as- sess automation,” Cognition, Technology & Work, pp. 39–49, 2012. Publisher: Springer
2012
-
[13]
Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology,
F. D. Davis, “Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology,” MIS Quarterly , pp. 319–340, 1989. Publisher: Management Information Systems Re- search Center, University of Minnesota
1989
-
[14]
Human trust in artificial intelligence: Review of empiri- cal research,
E. Glikson and A. W. Woolley, “Human trust in artificial intelligence: Review of empiri- cal research,” Academy of Management Annals , pp. 627–660, 2020. Publisher: Briarcliff Manor, NY
2020
-
[15]
Factors in- fluencing trust in medical artificial intelligence for healthcare professionals: A narrative review,
V. Tucci, J. Saary, and T. E. Doyle, “Factors in- fluencing trust in medical artificial intelligence for healthcare professionals: A narrative review,” J. Med. Artif. Intell , vol. 5, no. 4, 2022
2022
-
[16]
The ef- fects of clinical information presentation on physi- cians’ and nurses’ decision-making in ICUs,
A. Miller, C. Scheinkestel, and C. Steele, “The ef- fects of clinical information presentation on physi- cians’ and nurses’ decision-making in ICUs,” Ap- plied ergonomics, vol. 40, no. 4, pp. 753–761, 2009. Publisher: Elsevier
2009
-
[17]
Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making,
Y. Zhang, Q. V. Liao, and R. K. Bellamy, “Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making,” in Proceedings of the 2020 conference on fair- ness, accountability, and transparency , pp. 295– 305, 2020
2020
-
[18]
Towards improving trust in context-aware systems by displaying system con- fidence,
S. Antifakos, N. Kern, B. Schiele, and A. Schwaninger, “Towards improving trust in context-aware systems by displaying system con- fidence,” pp. 9–14, 2005
2005
-
[19]
Ex- ploring Strategic National Research and Devel- opment Factors for Sustainable Adoption of Cel- lular Agriculture Technology,
M. Jalali Sepehr and N. Mashhadi Nejad, “Ex- ploring Strategic National Research and Devel- opment Factors for Sustainable Adoption of Cel- lular Agriculture Technology,” Narges, Explor- ing Strategic National Research and Development Factors for Sustainable Adoption of Cellul...
2024
-
[20]
Task- Technology Fit and Individual Performance,
D. L. Goodhue and R. L. Thompson, “Task- Technology Fit and Individual Performance,” MIS Quarterly , vol. 19, pp. 213–236, 1995. Pub- lisher: Management Information Systems Re- search Center, University of Minnesota
1995
-
[21]
User Acceptance of Information Technology: Toward a Unified View,
V. Venkatesh, M. G. Morris, G. B. Davis, and F. D. Davis, “User Acceptance of Information Technology: Toward a Unified View,” MIS Quar- terly, pp. 425–478, 2003. Publisher: Management Information Systems Research Center, University of Minnesota
2003
-
[22]
Measurement of trust in automation: A narrative review and reference guide,
S. C. Kohn, E. J. De Visser, E. Wiese, Y.-C. Lee, and T. H. Shaw, “Measurement of trust in automation: A narrative review and reference guide,” Frontiers in psychology, vol. 12, p. 604977,
-
[23]
Trust and TAM in Online Shopping: An Inte- grated Model,
D. Gefen, E. Karahanna, and D. W. Straub, “Trust and TAM in Online Shopping: An Inte- grated Model,” MIS Q. , pp. 51–90, 2003
2003
-
[24]
Inte- grating Transparency, Trust, and Acceptance: The Intelligent Systems Technology Acceptance Model (ISTAM),
E. S. Vorm and D. J. Y. Combs, “Inte- grating Transparency, Trust, and Acceptance: The Intelligent Systems Technology Acceptance Model (ISTAM),” International Journal of Hu- man–Computer Interaction , pp. 1828–1845, Dec
-
[25]
Publisher: Frontiers Media SA
-
[26]
Trust in automation: designing for appropriate reliance,
J. D. Lee and K. A. See, “Trust in automation: designing for appropriate reliance,” Human Fac- tors, vol. 46, no. 1, pp. 50–80, 2004
2004
-
[27]
An Integrative Model of Organiza- tional Trust,
R. Mayer, “An Integrative Model of Organiza- tional Trust,” Academy of Management Review , 1995
1995
-
[28]
Trust between humans and ma- chines, and the design of decision aids,
B. M. Muir, “Trust between humans and ma- chines, and the design of decision aids,” Inter- national journal of man-machine studies , vol. 27, no. 5-6, pp. 527–539, 1987. Publisher: Elsevier
1987
-
[29]
Trust, Workload, and Performance in Human–Artificial Intelligence Partnering: The Role of Artificial Intelligence Attributes in Solv- ing Classification Problems,
M. Lotfalian Saremi, I. Ziv, O. Asan, and A. E. Bayrak, “Trust, Workload, and Performance in Human–Artificial Intelligence Partnering: The Role of Artificial Intelligence Attributes in Solv- ing Classification Problems,” Journal of Mechan- ical Design, vol. 147, July 2024
2024
-
[30]
How the different explanation classes impact trust calibration: The case of clinical deci- sion support systems,
M. Naiseh, D. Al-Thani, N. Jiang, and R. Ali, “How the different explanation classes impact trust calibration: The case of clinical deci- sion support systems,” International Journal of Human-Computer Studies , vol. 169, p. 102941,
-
[31]
Explainable recommendation: when design meets trust calibration,
M. Naiseh, D. Al-Thani, N. Jiang, and R. Ali, “Explainable recommendation: when design meets trust calibration,” World Wide Web , pp. 1857–1884, 2021. Publisher: Springer
2021
-
[32]
Humans and Algorithms Detecting Fake News: Effects of Individual and Contextual Confi- dence on Trust in Algorithmic Advice,
C. Snijders, R. Conijn, E. de Fouw, and K. van Berlo, “Humans and Algorithms Detecting Fake News: Effects of Individual and Contextual Confi- dence on Trust in Algorithmic Advice,” Interna- tional Journal of Human–Computer Interaction , pp. 1483–1494, Apr. 2023. Publisher: Tay...
2023
-
[33]
How to evaluate trust in AI-assisted decision making? A survey of empirical methodologies,
O. Vereschak, G. Bailly, and B. Caramiaux, “How to evaluate trust in AI-assisted decision making? A survey of empirical methodologies,” Proceed- ings of the ACM on Human-Computer Interac- tion, vol. 5, no. CSCW2, pp. 1–39, 2021. Pub- lisher: ACM New York, NY, USA
2021
-
[34]
Influences on User Trust in Healthcare Artificial Intelligence: A Systematic Review,
E. Jermutus, D. Kneale, J. Thomas, and S. Michie, “Influences on User Trust in Healthcare Artificial Intelligence: A Systematic Review,” Wellcome Open Research, vol. 7, Feb. 2022. Pub- lisher: F1000 Research Ltd
2022
-
[35]
A Survey of Important Factors in Human - Artificial Intel- ligence Trust for Engineering System Design,
M. Lotfalian Saremi and A. E. Bayrak, “A Survey of Important Factors in Human - Artificial Intel- ligence Trust for Engineering System Design,” in IDETC-CIE2021, (Volume 6: 33rd International Conference on Design Theory and Methodology (DTM)), Aug. 2021. V006T06A056
2021
-
[36]
A Meta-Analysis of Factors Influ- encing the Development of Trust in Automation: Implications for Understanding Autonomy in Fu- ture Systems,
K. E. Schaefer, J. Y. C. Chen, J. L. Szalma, and P. A. Hancock, “A Meta-Analysis of Factors Influ- encing the Development of Trust in Automation: Implications for Understanding Autonomy in Fu- ture Systems,” Human Factors, pp. 377–400, May
-
[37]
Trust in automation: integrating empirical evidence on factors that in- fluence trust,
K. A. Hoff and M. Bashir, “Trust in automation: integrating empirical evidence on factors that in- fluence trust,” Human Factors, vol. 57, pp. 407– 434, May 2015
2015
-
[38]
The risk of algorithm trans- parency: How algorithm complexity drives the ef- fects on the use of advice,
C. A. Lehmann, C. B. Haubitz, A. F¨ ugener, and U. W. Thonemann, “The risk of algorithm trans- parency: How algorithm complexity drives the ef- fects on the use of advice,” Production and Oper- ations Management, vol. 31, pp. 3419–3434, Sept
-
[39]
The effects of explainability and caus- ability on perception, trust, and acceptance: Implications for explainable AI,
D. Shin, “The effects of explainability and caus- ability on perception, trust, and acceptance: Implications for explainable AI,” International Journal of Human-Computer Studies , vol. 146, p. 102551, 2021
2021
-
[40]
Affective Design Anal- ysis of Explainable Artificial Intelligence (XAI): A User-Centric Perspective,
E. Bernardo and R. Seva, “Affective Design Anal- ysis of Explainable Artificial Intelligence (XAI): A User-Centric Perspective,” Informatics, vol. 10, p. 32, Mar. 2023. Number: 1 Publisher: Multi- disciplinary Digital Publishing Institute
2023
-
[41]
Explainability does not miti- gate the negative impact of incorrect AI advice in a personnel selection task,
J. Cecil, E. Lermer, M. F. C. Hudecek, J. Sauer, and S. Gaube, “Explainability does not miti- gate the negative impact of incorrect AI advice in a personnel selection task,” Scientific Reports, vol. 14, p. 9736, Apr. 2024
2024
-
[42]
Explainable artificial intelligence improves human decision-making: re- sults from a mushroom picking experiment at a public art festival,
B. Leichtmann, A. Hinterreiter, C. Humer, M. Streit, and M. Mara, “Explainable artificial intelligence improves human decision-making: re- sults from a mushroom picking experiment at a public art festival,” International Journal of Hu- man–Computer Interaction, pp. 4787–4804, ...
2024
-
[43]
The rationality of ex- planation or human capacity? Understanding the impact of explainable artificial intelligence on human-AI trust and decision performance,
P. Wang and H. Ding, “The rationality of ex- planation or human capacity? Understanding the impact of explainable artificial intelligence on human-AI trust and decision performance,” Information Processing & Management , vol. 61, p. 103732, July 2024
2024
-
[44]
Publisher: SAGE Publications
-
[45]
Deep inside convolutional networks: Visualising image classification models and saliency maps,
K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside convolutional networks: Visualising image classification models and saliency maps,” arXiv preprint arXiv:1312.6034 , 2013
2013 arXiv
-
[46]
Grad-CAM: visual explanations from deep networks via gradient-based localization,
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedan- tam, D. Parikh, and D. Batra, “Grad-CAM: visual explanations from deep networks via gradient-based localization,” International jour- nal of computer vision , vol. 128, pp. 336–359,
-
[47]
Ax- iomatic attribution for deep networks,
M. Sundararajan, A. Taly, and Q. Yan, “Ax- iomatic attribution for deep networks,” pp. 3319– 3328, PMLR, 2017
2017
-
[48]
Causability and explainability of artificial intelligence in medicine,
A. Holzinger, G. Langs, H. Denk, K. Zatloukal, and H. M¨ uller, “Causability and explainability of artificial intelligence in medicine,” Wiley Interdis- ciplinary Reviews. Data Mining and Knowledge Discovery, vol. 9, no. 4, p. e1312, 2019
2019
-
[49]
Exam- ples are not enough, learn to criticize! criticism for interpretability,
B. Kim, R. Khanna, and O. O. Koyejo, “Exam- ples are not enough, learn to criticize! criticism for interpretability,” Advances in neural informa- tion processing systems, vol. 29, 2016
2016
-
[50]
” Why should i trust you?
M. T. Ribeiro, S. Singh, and C. Guestrin, “” Why should i trust you?” Explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining , pp. 1135–1144, 2016
2016
-
[51]
Greedy Function Approxima- tion: A Gradient Boosting Machine,
J. H. Friedman, “Greedy Function Approxima- tion: A Gradient Boosting Machine,” The Annals of Statistics , vol. 29, pp. 1189–1232, 2001. Pub- lisher: Institute of Mathematical Statistics
2001
-
[52]
Effects of Explainable Artificial Intelligence on trust and human behav- ior in a high-risk decision task,
B. Leichtmann, C. Humer, A. Hinterreiter, M. Streit, and M. Mara, “Effects of Explainable Artificial Intelligence on trust and human behav- ior in a high-risk decision task,” Computers in Human Behavior , vol. 139, p. 107539, 2023
2023
-
[53]
The Impact of AI Explanations on Clinicians Trust and Diagnostic Accuracy in Breast Cancer,
O. Rezaeian, O. Asan, and A. E. Bayrak, “The Impact of AI Explanations on Clinicians Trust and Diagnostic Accuracy in Breast Cancer,” arXiv preprint arXiv:2412.11298 , 2024
2024 arXiv
-
[54]
Counterfactual explanations without opening the black box: Automated decisions and the GDPR,
S. Wachter, B. Mittelstadt, and C. Russell, “Counterfactual explanations without opening the black box: Automated decisions and the GDPR,” Harv. JL & Tech. , vol. 31, p. 841, 2017. Publisher: HeinOnline
2017
-
[55]
Examining the effect of explanation on satisfaction and trust in AI di- agnostic systems,
L. Alam and S. Mueller, “Examining the effect of explanation on satisfaction and trust in AI di- agnostic systems,” BMC Medical Informatics and Decision Making, vol. 21, p. 178, June 2021
2021
-
[56]
A unified approach to interpreting model predictions,
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , pp. 4768–4777, Curran Associates Inc., 2017. event-place: Long Beach, California, USA
2017
-
[57]
Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making,
X. Wang and M. Yin, “Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making,” in Proceedings of the 26th International Conference on Intelligent User Interfaces, pp. 318–328, 2021
2021
-
[58]
The explainability paradox: Chal- lenges for xAI in digital pathology,
T. Evans, C. O. Retzlaff, C. Geißler, M. Kargl, M. Plass, H. M¨ uller, T.-R. Kiehl, N. Zerbe, and A. Holzinger, “The explainability paradox: Chal- lenges for xAI in digital pathology,” Future Gen- eration Computer Systems, vol. 133, pp. 281–296, Aug. 2022
2022
-
[59]
Supporting trust calibration and the effective use of deci- sion aids by presenting dynamic system confi- dence information,
J. M. McGuirl and N. B. Sarter, “Supporting trust calibration and the effective use of deci- sion aids by presenting dynamic system confi- dence information,” Human factors, vol. 48, no. 4, pp. 656–665, 2006. Publisher: SAGE Publications Sage CA: Los Angeles, CA
2006
-
[60]
Impact of Model Interpretability and Outcome Feedback on Trust in AI,
D. Ahn, A. Almaatouq, M. Gulabani, and K. Hosanagar, “Impact of Model Interpretability and Outcome Feedback on Trust in AI,” in Pro- ceedings of the CHI Conference on Human Fac- tors in Computing Systems , pp. 1–25, 2024
2024
-
[61]
Impact of cognitive workload and situation awareness on clinicians’ willingness to use an artificial intelligence system in clinical practice,
A. Choudhury and O. Asan, “Impact of cognitive workload and situation awareness on clinicians’ willingness to use an artificial intelligence system in clinical practice,” IISE Transactions on Health- care Systems Engineering, pp. 89–100, Apr. 2023. Publisher: Taylor & Francis
2023
-
[62]
The ef- fects of example-based explanations in a machine learning interface,
C. J. Cai, J. Jongejan, and J. Holbrook, “The ef- fects of example-based explanations in a machine learning interface,” in Proceedings of the 24th in- ternational conference on intelligent user inter- faces, pp. 258–262, 2019
2019
-
[63]
Using explain- able AI to unravel classroom dialogue analysis: Effects of explanations on teachers’ trust, technol- ogy acceptance and cognitive load,
D. Wang, C. Bian, and G. Chen, “Using explain- able AI to unravel classroom dialogue analysis: Effects of explanations on teachers’ trust, technol- ogy acceptance and cognitive load,” British Jour- nal of Educational Technology , 2024. Publisher: Wiley Online Library
2024
-
[64]
Impact of explainable ai on cog- nitive load: Insights from an empirical study,
L.-V. Herm, “Impact of explainable ai on cog- nitive load: Insights from an empirical study,” arXiv preprint arXiv:2304.08861 , 2023
2023 arXiv
-
[65]
Effects of multimodal explanations for autonomous driving on driving performance, cognitive load, expertise, confidence, and trust,
R. Kaufman, J. Costa, and E. Kimani, “Effects of multimodal explanations for autonomous driving on driving performance, cognitive load, expertise, confidence, and trust,” Scientific reports, vol. 14, no. 1, p. 13061, 2024. Publisher: Nature Publish- ing Group UK London
2024
-
[66]
When confidence meets accuracy: Exploring the effects of multiple performance indicators on trust in machine learn- ing models,
A. Rechkemmer and M. Yin, “When confidence meets accuracy: Exploring the effects of multiple performance indicators on trust in machine learn- ing models,” pp. 1–14, 2022
2022
-
[67]
In- terrupted time-series analysis and its application to behavioral data,
D. P. Hartmann, J. M. Gottman, R. R. Jones, W. Gardner, A. E. Kazdin, and R. S. Vaught, “In- terrupted time-series analysis and its application to behavioral data,” Journal of Applied Behavior Analysis, vol. 13, no. 4, pp. 543–559, 1980
1980
-
[68]
Cognitive Architecture and Instruc- tional Design,
J. Sweller, J. J. G. van Merrienboer, and F. G. W. C. Paas, “Cognitive Architecture and Instruc- tional Design,” Educational Psychology Review , vol. 10, pp. 251–296, Sept. 1998
1998
-
[69]
U- net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U- net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Ger- many, October 5-9, 2015, Proceedings, Part III 18...
2015
-
[70]
Convolutional neural networks for breast cancer detection in mammography: A survey,
L. Abdelrahman, M. Al Ghamdi, F. Collado- Mesa, and M. Abdel-Mottaleb, “Convolutional neural networks for breast cancer detection in mammography: A survey,” Computers in Biology and Medicine, vol. 131, p. 104248, Apr. 2021
2021
-
[71]
Ad- dressing Small and Imbalanced Medical Image Datasets Using Generative Models: A Compar- ative Study of DDPM and PGGANs with Ran- dom and Greedy K Sampling,
I. Khazrak, S. Takhirova, M. M. Rezaee, M. Yadollahi, R. C. Green II, and S. Niu, “Ad- dressing Small and Imbalanced Medical Image Datasets Using Generative Models: A Compar- ative Study of DDPM and PGGANs with Ran- dom and Greedy K Sampling,” arXiv preprint arXiv:2412.12532, 2024
2024
-
[72]
Explainable Artificial Intelli- gence (XAI): How the Visualization of AI Pre- dictions Affects User Cognitive Load and Confi- dence,
A. Hudon, T. Demazure, A. Karran, P.-M. L´ eger, and S. S´ en´ ecal, “Explainable Artificial Intelli- gence (XAI): How the Visualization of AI Pre- dictions Affects User Cognitive Load and Confi- dence,” in Information Systems and Neuroscience (F. D. Davis, R. Riedl, J. vom Br...
2021
-
[73]
Dataset of breast ultrasound images,
W. Al-Dhabyani, M. Gomaa, H. Khaled, and A. Fahmy, “Dataset of breast ultrasound images,” Data in Brief , vol. 28, p. 104863, 2020
2020
-
[74]
An architecture to support graduated levels of trust for cancer diagnosis with ai,
O. Rezaeian, A. E. Bayrak, and O. Asan, “An architecture to support graduated levels of trust for cancer diagnosis with ai,” in Interna- tional Conference on Human-Computer Interac- tion, pp. 344–351, Springer, 2024
2024
-
[75]
Proxy tasks and subjective measures can be misleading in evaluating explainable AI sys- tems,
Z. Bu¸ cinca, P. Lin, K. Z. Gajos, and E. L. Glass- man, “Proxy tasks and subjective measures can be misleading in evaluating explainable AI sys- tems,” in Proceedings of the 25th International Conference on Intelligent User Interfaces, IUI ’20, (New York, NY, USA), pp. 454–46...
2020
-
[76]
To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision- making,
Z. Bu¸ cinca, M. B. Malaya, and K. Z. Gajos, “To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision- making,” Proceedings of the ACM on Human- computer Interaction, vol. 5, pp. 1–21, 2021. Pub- lisher: ACM New York, NY, USA
2021
-
[77]
Designing for confidence: The impact of visualizing artificial intelligence decisions,
A. J. Karran, T. Demazure, A. Hudon, S. Senecal, and P.-M. L´ eger, “Designing for confidence: The impact of visualizing artificial intelligence decisions,” Frontiers in neuroscience , vol. 16, p. 883385, 2022. Publisher: Frontiers Media SA
2022
-
[78]
Deep convolu- tional neural network based medical image clas- sification for disease diagnosis,
S. S. Yadav and S. M. Jadhav, “Deep convolu- tional neural network based medical image clas- sification for disease diagnosis,” Journal of Big Data, vol. 6, p. 113, Dec. 2019
2019
-
[79]
Prediction of Human Behavior in Human–Robot Interaction Using Psychological Scales for Anx- iety and Negative Attitudes Toward Robots,
T. Nomura, T. Kanda, T. Suzuki, and K. Kato, “Prediction of Human Behavior in Human–Robot Interaction Using Psychological Scales for Anx- iety and Negative Attitudes Toward Robots,” IEEE Transactions on Robotics, vol. 24, pp. 442– 451, Apr. 2008
2008
-
[80]
Development of NASA-TLX (Task Load Index): Results of empirical and theoretical re- search,
S. Hart, “Development of NASA-TLX (Task Load Index): Results of empirical and theoretical re- search,” Human mental workload/Elsevier , 1988
1988
-
[84]
Influence of Gender and Age on the Attitudes of Children towards Humanoid Robots,
F.-W. Tung, “Influence of Gender and Age on the Attitudes of Children towards Humanoid Robots,” in Human-Computer Interaction. Users and Applications (J. A. Jacko, ed.), (Berlin, Hei- delberg), pp. 637–646, Springer Berlin Heidel- berg, 2011
2011
-
[86]
Smart Devices Require Skilled Users: Home Automation Performance Tests among the Dutch Population,
A. J. A. M. v. Deursen, P. S. de Boer, and T. J. L. van Rompay, “Smart Devices Require Skilled Users: Home Automation Performance Tests among the Dutch Population,” Interna- tional Journal of Human–Computer Interaction , pp. 1–12, 2024. Publisher: Taylor & Francis
2024
-
[2016]
Publisher: SAGE Publications Inc
-
[2020]
Publisher: Nature Publishing Group UK London
-
[2021]
Number: 1 Publisher: Nature Publishing Group
-
[2022]
Publisher: Taylor & Francis
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.