REVIEW 3 major objections 6 minor 1 cited by
The Impact of AI Explanations on Clinicians Trust and Diagnostic Accuracy in Breast Cancer
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A study of 28 clinicians finds that adding explanation layers to an AI breast-cancer diagnosis tool can reduce diagnostic accuracy, and the simplest interface is the best.
desk verdict New clinical dataset, but the fixed intervention order confounds the main finding; deserves revision rather than rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a four-rung explanation ladder built on top of the same AI recommender: (1) class prediction only, (2) class prediction with a probability distribution, (3) added tumor localization, and (4) added low/high confidence localization. Each clinician saw all four rungs in that fixed order, with a no-AI baseline first. A mixed-effects model with intervention as a fixed effect and participant as a random grouping is used to compare the later interventions to the first, attempting to isolate the effect of added explanation on trust, accuracy, agreement, decision time, understandability, and perceived accuracy.
What would settle it
Run the same experiment with the four explanation conditions presented in randomized or counterbalanced order across clinicians. If the diagnostic accuracy decline at the third and fourth explanation levels disappears or reverses under random ordering, the fixed sequence, not the explanation content, was the cause.
Extended reading notes
Core claim
The central claim is that explainability is not a monotonic good in AI-assisted breast cancer diagnosis. Using an interrupted time-series design in which 28 clinicians diagnosed breast tissue images under no support, then under classification-only, then classification plus probabilities, then plus tumor localization, then plus confidence levels, the authors found that the third and fourth explanation levels produced statistically significant drops in diagnostic accuracy relative to the first, with coefficients of -0.073 (p = 0.005) and -0.068 (p = 0.009). The fourth level also significantly reduced understandability, perceived accuracy, and agreement and lengthened decision time, while trust did not change significantly across interventions. Overall, all AI-assisted conditions outperformed the no-AI baseline, and the authors single out the classification-only interface as the ideal design.
Load-bearing premise
The comparisons assume that the fixed order in which all clinicians saw the explanation levels did not itself affect performance, so fatigue, practice, or learning are not responsible for the declines attributed to richer explanations.
Editorial extensions
If this is right
- Clinicians' diagnostic accuracy can decline when an AI support system adds probability and localization details on top of a plain classification recommendation.
- Trust, understandability, and perceived accuracy do not reliably improve as explanation richness increases, so explanation design should be treated as a trade-off rather than a monotonic benefit.
- The presence of an AI recommendation, even without explanations, improves diagnostic accuracy over relying on clinical judgment alone in this setting.
- Self-reported demographic differences, such as gender differences in AI familiarity, need not translate into differences in actual trust or diagnostic performance.
Reading between the lines
- Because all clinicians saw the explanation levels in the same order, the causal reading depends on an assumption the paper does not test; a counterbalanced replication would separate explanation level from fatigue or practice effects.
- The paper does not vary the accuracy of the underlying AI; with a more accurate system, richer explanations might not show the same decline, so the 'simplest is best' recommendation is bounded to this model and task.
- A practical extension would replace the fixed ladder with adaptive explanations that give more detail only on request, and measure whether that restores the trust gains without the accuracy loss.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an online experiment with 28 clinicians who performed breast cancer diagnosis tasks under five conditions: a no-AI baseline and four CDSS interventions with increasing levels of explainability (classification only, plus probability distribution, plus tumor localization, plus confidence levels). The authors analyze self-reported trust, understandability, perceived accuracy, and behavioral measures (diagnostic accuracy, agreement, decision time) using mixed-effects models, chi-square tests, and ANOVA. They report that higher explainability does not monotonically improve trust or accuracy, with Interventions III and IV showing significantly lower diagnostic accuracy than Intervention I, and they conclude that Intervention I (the simplest design) is ideal. Demographic analyses suggest gender and experience relate to self-reported AI familiarity but not to behavioral measures.
Significance. If the causal claims were supported, the study would make a useful contribution to the human-AI interaction and XAI literature, particularly for clinical decision support. The use of a realistic diagnostic task, multiple trust measures (self-report and behavioral), and a clinician sample are strengths. The paper also explicitly distinguishes self-reported from behavioral outcomes, which is valuable. However, the study is not pre-registered, does not provide code or data, and—as detailed below—its central causal inference is undermined by a design confound that cannot be repaired with the collected data. The significance of the contribution therefore depends on whether the authors can reframe the results as descriptive rather than causal.
major comments (3)
- [Section 3.1 / Section 4.2 / Section 5.1] The fixed intervention order is a fundamental confound. All participants receive the baseline and then Interventions I through IV in the same increasing-explainability sequence (Section 3.1). The mixed-effects model in Table 5 uses Intervention I as the reference category and includes no term for time or session order. Because intervention level and session order are perfectly collinear within every participant, the coefficients for Interventions III and IV (diagnostic accuracy -0.073, p=0.005; -0.068, p=0.009) absorb any fatigue, practice, boredom, or attention effects. The conclusion in Section 5.1 that 'the 1st intervention with the simplest design would be the ideal choice' is therefore not causally supported; the observed declines may simply reflect time-on-task. Since every participant followed the same order, the design offers no way to separate order from intervention, and adding a time covariate would make the intervention effects unidentifiable. This threatens the paper's central claim that increasing explanations can hurt performance.
- [Section 3.3.2] Diagnostic accuracy is not measured directly for trials where agreement is above neutral. The text states that if the participant's agreement level is higher than neutral, 'we considered their decision to align with the AI's suggestion'; only when agreement is below neutral are participants asked to provide their own decision. This creates a mechanical dependency: agreement and performance are partly the same variable. In the 4th intervention, agreement is significantly lower (Table 5, -0.149, p=0.048), so a larger fraction of trials enters the performance calculation as actual (and possibly less accurate) human decisions rather than imputed AI decisions. The observed accuracy decline in Interventions III and IV could therefore be an artifact of this measurement rule rather than a true change in decision quality. The authors should either collect explicit decisions on every trial or model performance conditional on the agreement threshold.
- [Section 4.2 (Performance paragraph)] The claim that accuracy when using AI is 'generally higher compared to the baseline intervention without AI, all improvements being statistically significant' is not supported by any statistic shown. Table 5 only contrasts Interventions II, III, and IV with Intervention I; the baseline condition is not included in the mixed-effects model reported. Without reporting the coefficients or tests for the baseline comparison, this claim is unverifiable and should be either documented or removed.
minor comments (6)
- [Section 4.2 (Trust paragraph)] There is an inconsistency between the text and Table 5: the text reports the 3rd Intervention trust coefficient as p = 0.187, while Table 5 gives p = 0.465. One of these is a typo and should be corrected.
- [Section 4.1 / Table 3] The chi-square tests are based on 28 participants with many categories (e.g., five age brackets, three experience brackets). Several expected cell counts are likely below 5, which violates the assumptions of the chi-square test. The authors should report expected counts or use an exact test.
- [Section 4.2 / Table 5] No correction for multiple comparisons is applied across the many mixed-effects tests, chi-square tests, and ANOVAs. Some p-values near 0.05 (e.g., perceived accuracy for the 4th intervention, p=0.048; agreement for the 4th intervention, p=0.048) may not survive a Benjamini-Hochberg correction. A pre-specified analysis plan or explicit multiplicity adjustment would strengthen the results.
- [Section 5.2] The limitations section does not acknowledge the fixed intervention order or the agreement-based performance imputation, both of which are major threats to causal interpretation. These should be explicitly discussed.
- [Abstract / Conclusion / Section 5.2] There are several language issues: 'such breast cancer' in the Abstract, 'regrad' in the Conclusion, and 'The clinicians were interaction' in Section 5.2. These need copyediting.
- [References] Some reference entries include 'Publisher:' metadata and formatting is inconsistent across entries (e.g., [1], [12], [21]). Ensure a consistent citation style.
Circularity Check
No circularity: the results are empirical measurements from a behavioral experiment, and the cited prior work is background context, not a load-bearing reduction.
full rationale
The paper's central claims are not derived from its inputs by construction. The main finding—that increasing explanation levels do not always improve trust or diagnostic accuracy, with significant accuracy reductions for the 3rd and 4th interventions—comes from a mixed-effects model fit to collected behavioral data (Section 4.2, Table 5), not from a definitional equivalence or a fitted parameter renamed as a prediction. The AI system is described as building on the authors' prior work [42], and its 81% accuracy is cited as background context for the CDSS; the clinician trust and accuracy outcome measures are assessed independently against ground truth during the experiment. The reference to a prior study [45] for gender effects is corroborative rather than load-bearing. The fixed intervention order described in Section 3.1 is a potential validity threat and confound, but confounding is not circularity: the statistical comparison between interventions is not equivalent to the hypothesis by construction. No self-definitional step, fitted-input-called-prediction, or self-citation chain reduces the central claim to its own inputs. Therefore the paper shows no significant circularity.
Assumptions & free parameters
free parameters (1)
- Neutral agreement threshold =
Not explicitly specified; reported as 'above neutral' on a 0-5 Likert scale
assumptions (4)
- domain assumption Self-reported Likert responses validly measure trust, familiarity, understandability, and perceived accuracy
- domain assumption The fixed order of interventions does not introduce carryover or order effects
- domain assumption The 81% accurate AI model on 780 ultrasound images is a suitable decision aid for the task
- domain assumption Performance imputation from agreement ratings is valid
Cite this review
Pith. "Pith review of The Impact of AI Explanations on Clinicians Trust and Diagnostic Accuracy in Breast Cancer." pith.science (2026). https://pith.science/paper/QX5GJWXJ
@misc{pith2026241211298,
author = {Pith},
title = {Pith review of: The Impact of AI Explanations on Clinicians Trust and Diagnostic Accuracy in Breast Cancer},
year = {2026},
howpublished = {\url{https://pith.science/paper/QX5GJWXJ}},
note = {Machine review of arXiv:2412.11298}
}
read the original abstract
Advances in machine learning have created new opportunities to develop artificial intelligence (AI)-based clinical decision support systems using past clinical data and improve diagnosis decisions in life-threatening illnesses such breast cancer. Providing explanations for AI recommendations is a possible way to address trust and usability issues in black-box AI systems. This paper presents the results of an experiment to assess the impact of varying levels of AI explanations on clinicians' trust and diagnosis accuracy in a breast cancer application and the impact of demographics on the findings. The study includes 28 clinicians with varying medical roles related to breast cancer diagnosis. The results show that increasing levels of explanations do not always improve trust or diagnosis performance. The results also show that while some of the self-reported measures such as AI familiarity depend on gender, age and experience, the behavioral assessments of trust and performance are independent of those variables.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Explainability and AI Confidence in Clinical Decision Support Systems: Effects on Trust, Diagnostic Performance, and Cognitive Load in Breast Cancer Care
Low AI confidence reduced clinician trust and agreement and increased decision time, while high confidence corresponded to a small drop in diagnostic accuracy in a 28-participant web experiment.
Reference graph
Works this paper leans on
-
[1]
Interna- tional evaluation of an AI system for breast cancer screening,
S. M. McKinney, M. Sieniek, V. Godbole, J. God- win, N. Antropova, H. Ashrafian, T. Back, M. Chesus, G. S. Corrado, and A. Darzi, “Interna- tional evaluation of an AI system for breast cancer screening,” Nature, vol. 577, no. 7788, pp. 89–94,
-
[2]
M. Micocci, S. Borsci, V. Thakerar, S. Walne, Y. Manshadi, F. Edridge, D. Mullarkey, P. Buckle, and G. B. Hanna, “Attitudes towards trusting artificial intelligence insights and factors to prevent the passive adherence of GPs: a pi- lot study,” Journal of Clinical Medicine , vol. 10, no. 14, p. 3101, 2021. Publisher: MDPI
work page 2021
-
[3]
A. Janowczyk and A. Madabhushi, “Deep learn- ing for digital pathology image analysis: A com- prehensive tutorial with selected use cases,” Jour- nal of pathology informatics , vol. 7, no. 1, p. 29,
-
[4]
Deep learning so- lutions for skin cancer detection and diagnosis,
H. Nahata and S. P. Singh, “Deep learning so- lutions for skin cancer detection and diagnosis,” Machine Learning with Health Care Perspective: Machine Learning and Healthcare , pp. 159–182,
-
[5]
C. McIntosh and T. G. Purdie, “Voxel-based dose prediction with multi-patient atlas selection for automated radiotherapy treatment planning,” Physics in Medicine & Biology , vol. 62, no. 2, p. 415, 2016. Publisher: IOP Publishing
work page 2016
-
[6]
Effect of risk, expectancy, and trust on clini- cians’ intent to use an artificial intelligence sys- tem – Blood Utilization Calculator,
A. Choudhury, O. Asan, and J. E. Medow, “Effect of risk, expectancy, and trust on clini- cians’ intent to use an artificial intelligence sys- tem – Blood Utilization Calculator,” Applied Er- gonomics, vol. 101, p. 103708, May 2022
2022
-
[7]
C. Woodcock, B. Mittelstadt, D. Busbridge, and G. Blank, “The impact of explanations on layper- son trust in Artificial Intelligence–driven symp- tom checker apps: Experimental study,” Jour- nal of Medical Internet Research , vol. 23, no. 11, p. e29386, 2021. Publisher: JMIR Publications Toronto, Canada
work page 2021
-
[8]
Ethical, legal, and social considerations of AI- based medical decision-support tools: A scoping review,
A. ˇCartolovni, A. Tomiˇ ci´ c, and E. L. Mosler, “Ethical, legal, and social considerations of AI- based medical decision-support tools: A scoping review,” International Journal of Medical Infor- matics, vol. 161, p. 104738, 2022. Publisher: El- sevier
2022
Show all 48 references
-
[9]
Do as AI say: susceptibility in deployment of clinical decision- aids,
S. Gaube, H. Suresh, M. Raue, A. Merritt, S. J. Berkowitz, E. Lermer, J. F. Coughlin, J. V. Gut- tag, E. Colak, and M. Ghassemi, “Do as AI say: susceptibility in deployment of clinical decision- aids,” npj Digital Medicine , vol. 4, pp. 1–8, Feb
-
[10]
High-performance medicine: the convergence of human and artificial intelligence,
E. J. Topol, “High-performance medicine: the convergence of human and artificial intelligence,” Nature medicine, vol. 25, no. 1, pp. 44–56, 2019. Publisher: Nature Publishing Group US New York
2019
-
[11]
Integrating Protein Sequence and Expression Level to Analysis Molecular Char- acterization of Breast Cancer Subtypes,
H. Sholehrasa, “Integrating Protein Sequence and Expression Level to Analysis Molecular Char- acterization of Breast Cancer Subtypes,” arXiv preprint arXiv:2410.01755, 2024
2024
-
[12]
Humans and au- tomation: Use, misuse, disuse, abuse,
R. Parasuraman and V. Riley, “Humans and au- tomation: Use, misuse, disuse, abuse,” Human factors, vol. 39, no. 2, pp. 230–253, 1997. Pub- lisher: SAGE Publications Sage CA: Los Angeles, CA
1997
-
[13]
Factors in- fluencing trust in medical artificial intelligence for healthcare professionals: A narrative review,
V. Tucci, J. Saary, and T. E. Doyle, “Factors in- fluencing trust in medical artificial intelligence for healthcare professionals: A narrative review,” J. Med. Artif. Intell , vol. 5, no. 4, 2022
2022
-
[14]
The ef- fects of clinical information presentation on physi- cians’ and nurses’ decision-making in ICUs,
A. Miller, C. Scheinkestel, and C. Steele, “The ef- fects of clinical information presentation on physi- cians’ and nurses’ decision-making in ICUs,” Ap- plied ergonomics, vol. 40, no. 4, pp. 753–761, 2009. Publisher: Elsevier
2009
-
[15]
Trust in automation: designing for appropriate reliance,
J. D. Lee and K. A. See, “Trust in automation: designing for appropriate reliance,” Human Fac- tors, vol. 46, no. 1, pp. 50–80, 2004
2004
-
[16]
Rela- tionship between automation trust and operator performance for the novice and expert in space- craft rendezvous and docking (R VD),
J. Niu, H. Geng, Y. Zhang, and X. Du, “Rela- tionship between automation trust and operator performance for the novice and expert in space- craft rendezvous and docking (R VD),” Applied Ergonomics, vol. 71, pp. 1–8, Sept. 2018
2018
-
[17]
Factors affecting trust in high-vulnerability human-robot interaction contexts: A struc- tural equation modelling approach,
W. Kim, N. Kim, J. B. Lyons, and C. S. Nam, “Factors affecting trust in high-vulnerability human-robot interaction contexts: A struc- tural equation modelling approach,” Applied Er- gonomics, vol. 85, p. 103056, May 2020
2020
-
[18]
Trust between humans and ma- chines, and the design of decision aids,
B. M. Muir, “Trust between humans and ma- chines, and the design of decision aids,” Inter- national journal of man-machine studies , vol. 27, no. 5-6, pp. 527–539, 1987. Publisher: Elsevier
1987
-
[19]
An Integrative Model of Organiza- tional Trust,
R. Mayer, “An Integrative Model of Organiza- tional Trust,” Academy of Management Review , 1995
1995
-
[20]
Arti- ficial Intelligence and Human Trust in Healthcare: Focus on Clinicians,
O. Asan, A. E. Bayrak, and A. Choudhury, “Arti- ficial Intelligence and Human Trust in Healthcare: Focus on Clinicians,” Journal of Medical Internet Research, vol. 22, p. e15154, June 2020
2020
-
[21]
Trust- ing Automation: Designing for Respon- sivity and Resilience,
E. K. Chiou and J. D. Lee, “Trust- ing Automation: Designing for Respon- sivity and Resilience,” Human Factors , vol. 65, no. 1, pp. 137–165, 2023. eprint: https://doi.org/10.1177/00187208211009995
2023 doi
-
[22]
Measurement of trust in automation: A narrative review and reference guide,
S. C. Kohn, E. J. De Visser, E. Wiese, Y.-C. Lee, and T. H. Shaw, “Measurement of trust in automation: A narrative review and reference guide,” Frontiers in psychology, vol. 12, p. 604977,
-
[23]
Trust in automation: integrating empirical evidence on factors that in- fluence trust,
K. A. Hoff and M. Bashir, “Trust in automation: integrating empirical evidence on factors that in- fluence trust,” Human Factors, vol. 57, pp. 407– 434, May 2015
2015
-
[24]
Influences on User Trust in Healthcare Artificial Intelligence: A Systematic Review,
E. Jermutus, D. Kneale, J. Thomas, and S. Michie, “Influences on User Trust in Healthcare Artificial Intelligence: A Systematic Review,” Wellcome Open Research, vol. 7, Feb. 2022. Pub- lisher: F1000 Research Ltd
2022
-
[25]
Publisher: Frontiers Media SA
-
[26]
Clinical integration of machine learning for curative-intent radiation treatment of patients with prostate cancer,
C. McIntosh, L. Conroy, M. C. Tjong, T. Craig, A. Bayley, C. Catton, M. Gospodarowicz, J. Helou, N. Isfahanian, V. Kong, T. Lam, S. Ra- man, P. Warde, P. Chung, A. Berlin, and T. G. Purdie, “Clinical integration of machine learning for curative-intent radiation treatment of pa...
2021
-
[27]
The effects of explainability and caus- ability on perception, trust, and acceptance: Implications for explainable AI,
D. Shin, “The effects of explainability and caus- ability on perception, trust, and acceptance: Implications for explainable AI,” International Journal of Human-Computer Studies , vol. 146, p. 102551, 2021
2021
-
[28]
A Survey of Important Factors in Human - Artificial Intel- ligence Trust for Engineering System Design,
M. Lotfalian Saremi and A. E. Bayrak, “A Survey of Important Factors in Human - Artificial Intel- ligence Trust for Engineering System Design,” in IDETC-CIE2021, (Volume 6: 33rd International Conference on Design Theory and Methodology (DTM)), Aug. 2021. V006T06A056
2021
-
[29]
Machine Learning Interpretability: A Survey on Methods and Metrics,
D. V. Carvalho, E. M. Pereira, and J. S. Cardoso, “Machine Learning Interpretability: A Survey on Methods and Metrics,” Electronics, vol. 8, no. 8, 2019
2019
-
[30]
A unified approach to interpreting model predictions,
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS’17, (Red Hook, NY, USA), pp. 4768–4777, Curran Asso- ciates Inc., 2017. event-place: Long B...
2017
-
[31]
Affective Design Anal- ysis of Explainable Artificial Intelligence (XAI): A User-Centric Perspective,
E. Bernardo and R. Seva, “Affective Design Anal- ysis of Explainable Artificial Intelligence (XAI): A User-Centric Perspective,” Informatics, vol. 10, p. 32, Mar. 2023. Number: 1 Publisher: Multi- disciplinary Digital Publishing Institute
2023
-
[32]
The rationality of ex- planation or human capacity? Understanding the impact of explainable artificial intelligence on human-AI trust and decision performance,
P. Wang and H. Ding, “The rationality of ex- planation or human capacity? Understanding the impact of explainable artificial intelligence on human-AI trust and decision performance,” Information Processing & Management , vol. 61, p. 103732, July 2024
2024
-
[33]
The ef- fects of example-based explanations in a machine learning interface,
C. J. Cai, J. Jongejan, and J. Holbrook, “The ef- fects of example-based explanations in a machine learning interface,” in Proceedings of the 24th in- ternational conference on intelligent user inter- faces, pp. 258–262, 2019
2019
-
[34]
” Why should i trust you?
M. T. Ribeiro, S. Singh, and C. Guestrin, “” Why should i trust you?” Explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining , pp. 1135–1144, 2016
2016
-
[35]
Effects of Explainable Artificial Intelligence on trust and human behav- ior in a high-risk decision task,
B. Leichtmann, C. Humer, A. Hinterreiter, M. Streit, and M. Mara, “Effects of Explainable Artificial Intelligence on trust and human behav- ior in a high-risk decision task,” Computers in Human Behavior , vol. 139, p. 107539, 2023
2023
-
[36]
Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making,
Y. Zhang, Q. V. Liao, and R. K. Bellamy, “Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making,” in Proceedings of the 2020 conference on fair- ness, accountability, and transparency , pp. 295– 305, 2020
2020
-
[37]
Explainability does not miti- gate the negative impact of incorrect AI advice in a personnel selection task,
J. Cecil, E. Lermer, M. F. C. Hudecek, J. Sauer, and S. Gaube, “Explainability does not miti- gate the negative impact of incorrect AI advice in a personnel selection task,” Scientific Reports, vol. 14, p. 9736, Apr. 2024
2024
-
[38]
Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making,
X. Wang and M. Yin, “Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making,” in Proceedings of the 26th International Conference on Intelligent User Interfaces, pp. 318–328, 2021
2021
-
[39]
Examining the effect of explanation on satisfaction and trust in AI di- agnostic systems,
L. Alam and S. Mueller, “Examining the effect of explanation on satisfaction and trust in AI di- agnostic systems,” BMC Medical Informatics and Decision Making, vol. 21, p. 178, June 2021
2021
-
[40]
Impact of Model Interpretability and Outcome Feedback on Trust in AI,
D. Ahn, A. Almaatouq, M. Gulabani, and K. Hosanagar, “Impact of Model Interpretability and Outcome Feedback on Trust in AI,” in Pro- ceedings of the CHI Conference on Human Fac- tors in Computing Systems , pp. 1–25, 2024
2024
-
[41]
In- terrupted time-series analysis and its application to behavioral data,
D. P. Hartmann, J. M. Gottman, R. R. Jones, W. Gardner, A. E. Kazdin, and R. S. Vaught, “In- terrupted time-series analysis and its application to behavioral data,” Journal of Applied Behavior Analysis, vol. 13, no. 4, pp. 543–559, 1980
1980
-
[42]
An architecture to support graduated levels of trust for cancer diagnosis with ai,
O. Rezaeian, A. E. Bayrak, and O. Asan, “An architecture to support graduated levels of trust for cancer diagnosis with ai,” in Interna- tional Conference on Human-Computer Interac- tion, pp. 344–351, Springer, 2024
2024
-
[43]
The explainability paradox: Chal- lenges for xAI in digital pathology,
T. Evans, C. O. Retzlaff, C. Geißler, M. Kargl, M. Plass, H. M¨ uller, T.-R. Kiehl, N. Zerbe, and A. Holzinger, “The explainability paradox: Chal- lenges for xAI in digital pathology,” Future Gen- eration Computer Systems, vol. 133, pp. 281–296, Aug. 2022
2022
-
[44]
Dataset of breast ultrasound images,
W. Al-Dhabyani, M. Gomaa, H. Khaled, and A. Fahmy, “Dataset of breast ultrasound images,” Data in Brief , vol. 28, p. 104863, 2020
2020
-
[45]
Trust, Workload, and Performance in Human–Artificial Intelligence Partnering: The Role of Artificial Intelligence Attributes in Solv- ing Classification Problems,
M. Lotfalian Saremi, I. Ziv, O. Asan, and A. E. Bayrak, “Trust, Workload, and Performance in Human–Artificial Intelligence Partnering: The Role of Artificial Intelligence Attributes in Solv- ing Classification Problems,” Journal of Mechan- ical Design, vol. 147, July 2024
2024
-
[46]
U- net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U- net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Ger- many, October 5-9, 2015, Proceedings, Part III 18...
2015
-
[2020]
Publisher: Nature Publishing Group UK London
-
[2021]
Number: 1 Publisher: Nature Publishing Group
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.