Pith. sign in

REVIEW 3 major objections 4 minor 140 references

Artificial Intelligence Should Genuinely Support Clinical Reasoning and Decision Making To Bridge the Translational Gap

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that medical AI's translational gap — the chasm between high benchmark performance and real clinical impact — is caused by technology-centric development, and should be closed by building AI that supports clinicians'…

desk verdict A useful, well-written Perspective with an overreaching causal diagnosis; the framing is worth engaging even though the central claim is asserted rather than demonstrated. read the letter →

arxiv 2506.05030 v1 pith:F2XDB6EV submitted 2025-06-05 cs.HC cs.AIcs.CYcs.LG

classification cs.HCcs.AIcs.CYcs.LG
keywords clinicalreasoningsupporttranslationalgapsociotechnicalAIaugmentedintelligencecognitiveforcingdecisionsystemspediatricsepsishuman-centred
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This Perspective argues that medical AI fails to reach patients not because predictive models are inaccurate, but because the field builds technology-centric systems that compete with, rather than support, how doctors actually reason and decide. The authors propose a sociotechnical reframing in which data-driven tools are designed to augment clinicians' cognitive and epistemic activities — reasoning under uncertainty, mental simulation, and belief updating — and to fit into the existing workflow, institutions, and responsibilities of care. They use paediatric sepsis as an illustrative case because it combines ambiguous definitions, high-stakes time pressure, and a heterogeneous population, exposing the mismatch between benchmark-driven modelling and bedside needs. If the paper is right, the path to clinical impact runs through human-centred design of reasoning support rather than through further accuracy gains on benchmark tasks.

What carries the argument

The central object is a sociotechnical conceptualisation of AI as a reasoning-support system rather than an autonomous decision maker, organised around three modes of integration (autonomy, assistance with a human-in-the-loop, and augmentation with a machine-in-the-loop) and selected per cognitive and epistemic activity according to a task's automation readiness level. The load-bearing mechanisms are the cognitive-science techniques proposed for supporting clinicians: cognitive forcing (interrupting heuristic reasoning to force consideration of disconfirming evidence and alternative hypotheses), AI-assisted mental simulation (projecting alternative future patient trajectories conditioned on different values of missing or unknown variables), and continuous belief updating. The argument also imports the distinction between structured, semi-structured, and unstructured decisions, and between stable and wicked environments, to justify why end-to-end automation is usually the wrong fit for clinical diagnosis.

What would settle it

A prospective, controlled deployment comparing a cognitive-support system (e.g., one delivering mental simulation and cognitive-forcing prompts for suspected sepsis) against a conventional predictive alert tool, measuring clinician uptake, decision quality, and patient outcomes: if the conventional tool matches the cognitive-support tool on adoption and outcomes, the claim that the reframing is necessary to bridge the gap would be undercut.

Watch

Extended reading notes

Core claim

The paper's central claim is normative: AI systems ought to seamlessly integrate into and augment established medical workflows and real-life reasoning and decision-making processes, rather than disrupt them. The authors posit that the prevailing technology-centric approach — optimising models for superhuman predictive performance on carefully chosen benchmarks — is the root cause of the translational gap, rendering such systems fundamentally incompatible with clinical practice. They argue that medical diagnostic reasoning is usually semi-structured, occurs in unstable (open) worlds with incomplete and uncertain information, and relies on cognitive functions that can be supported by AI, such as cognitive forcing to counteract anchoring and premature closure, prospective mental simulation of patient trajectories, and progressive belief updating, all within a systems ecology of institutions, protocols, and responsibilities. The intended consequence is that AI is judged by real-world impact and acceptability, as a reliable tool under human responsibility, rather than by benchmark scores.

Load-bearing premise

The claim hinges on the premise that the translational gap is caused by technology-centric development and that reframing AI as cognitive support will, by itself, materially improve adoption and outcomes; the paper offers reasoning and examples but no controlled evidence for this causal chain.

Editorial extensions

If this is right

  • If adopted, AI evaluation would shift from benchmark accuracy to measures of decision consistency, reduction of decision noise, and integration into clinical workflow.
  • AI tools designed for mental simulation would let clinicians explore patient trajectories under different assumptions, including explicitly handling missing data rather than imputing it away.
  • Ante-hoc interpretable models would be preferred over post-hoc explanations for high-stakes clinical support, because their behaviour is guaranteed to reflect the model's true operation.
  • The framework generalises beyond medicine to other high-stakes, semi-structured decision domains where replacing humans is undesirable.
  • For paediatric sepsis specifically, AI could be embedded in sepsis teams and huddles to aid detection, management, and treatment consistency, while reducing unnecessary antibiotic exposure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the causal claim is right, the research community's current reward structure — benchmark leaderboards and accuracy contests — is itself part of the translational problem, not merely a neutral evaluation tool.
  • A testable extension would be to build a cognitive-support sepsis tool that presents alternative trajectories and cognitive-forcing prompts, then compare clinician diagnostic accuracy and workflow adoption against a conventional predictive alert system in a prospective study.
  • The argument implies that regulatory and reimbursement pathways must reward cognitive support and workflow integration, since adoption is a property of the entire institutional ecosystem rather than of model design alone.
  • One can also read the paper as predicting that current human-in-the-loop systems that simply ask clinicians to accept or reject recommendations will underperform systems that actively scaffold the clinician's reasoning process.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This Perspective argues that the persistent translational gap in medical AI is caused by technology-centric development, which optimizes benchmark predictive performance while ignoring clinicians' real reasoning and decision-making processes. The authors propose a sociotechnical reframing in which AI systems augment, rather than automate, clinicians' cognitive and epistemic functions—such as reasoning under uncertainty, mental simulation, and belief updating—and they use paediatric sepsis as a running case study. The paper contributes a tripartite distinction among autonomy, assistance, and augmentation, and it outlines concrete mechanisms (e.g., cognitive forcing, counterfactual simulation) through which such support might operate.

Significance. If the central claim is correct, the paper usefully redirects medical AI research away from benchmark chasing and toward a human-centred, cognition-aware design agenda. It draws on a broad interdisciplinary literature and offers a clear conceptual vocabulary for discussing AI integration levels. The paper is well-structured and thoughtful, and it honestly acknowledges the lack of existing implementation frameworks. The main weakness is that its causal diagnosis—that technology-centric development is the principal cause of the translational gap—is asserted rather than demonstrated, and the paper's own success stories (benchmark-trained classifiers) undercut that diagnosis. These issues make the proposal a normative hypothesis rather than an evidence-based conclusion.

major comments (3)
  1. [Introduction; Medical Artificial Intelligence Adoption Challenges] The Introduction asserts that 'the prevailing technology-centric approaches underpin this challenge' and that this renders AI 'fundamentally incompatible with clinical practice' (Abstract and Introduction). This is a load-bearing causal claim, yet the paper offers no comparative or prospective evidence for it. The later section 'Medical Artificial Intelligence Adoption Challenges' itself lists a range of other candidate causes—regulatory approval, data access, implementation costs, alert fatigue, trust, reproducibility—without weighing them. If these factors dominate, the proposed shift to cognitive support may not close the translational gap. The authors should either soften the causal claim to a hypothesis or provide concrete evidence that technology-centric design is the primary bottleneck.
  2. [Medical Artificial Intelligence Adoption Challenges] The paper's three flagship successes—diabetic retinopathy detection (ref 5), skin cancer classification (ref 6), and lymph-node metastasis detection (ref 7)—are, by the paper's own definition, technology-centric: they are benchmark-trained deep learning classifiers. The paper acknowledges that these successes are 'not necessarily representative' and pertain mostly to visual domains, but it does not explain why these cases do not contradict the claim that technology-centric approaches are 'fundamentally incompatible' with clinical practice. This is a tension that should be addressed explicitly: if benchmark-trained classifiers succeeded in these domains, what distinguishes the domains where they fail, and how does cognitive support address those distinguishing features?
  3. [Human Decision Making and Artificial Intelligence; From Benchmark to Bedside] The proposed cognitive-support tools (mental simulation, cognitive forcing, belief updating, digital twins) are described at a conceptual level, but the paper admits in 'Human Decision Making and Artificial Intelligence' that 'we generally lack the corresponding (technical) frameworks, guidelines and protocols' for implementing such systems. The 'From Benchmark to Bedside' section cites only 'rudimentary research' and notes that this line of work 'by and large overlooks the broader systems ecology.' As a consequence, the central claim that these tools will materially improve adoption and outcomes is a conjecture rather than a result. The paper should frame this as an open research agenda and identify at least one concrete evaluation pathway (e.g., pilot studies with process and outcome measures) that could test the hypothesis.
minor comments (4)
  1. [Keywords] The keyword list uses interpuncts to separate terms; this is nonstandard. Please use commas or semicolons for clarity.
  2. [Systems Ecology and Artificial Intelligence] The phrase 'the race to the bottom' is introduced without definition. Since it is not a standard term in the medical AI literature, it should be explained or replaced with a more descriptive label.
  3. [Systems Ecology and Artificial Intelligence] The term 'ante-hoc interpretable' appears before it is formally introduced. Define it at first use, perhaps with a brief contrast to post-hoc explainability.
  4. [From Benchmark to Bedside] The concept of 'digital twins of individual clinicians' is intriguing but underdeveloped. The paper does not explain how such a twin would be created, validated, or used in practice. A reference or a one-sentence operational sketch would help.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: this normative Perspective has no derivation chain that reduces to its inputs; its self-citations are present but not load-bearing.

full rationale

This is a Perspective with no quantitative derivation, benchmark fitting, or formal theorem; the only 'chain' is a normative argument that AI should support clinicians' cognitive and epistemic functions rather than optimize benchmark accuracy. That argument is explicitly grounded in external cognitive science (Croskerry 2003, 2008; Kahneman & Klein 2009; Gigerenzer 2023; Crandall & Getchell-Reiter 1993), HCI and decision-support literature (Parasuraman et al. 2000; Ackerman 2000; van Baalen et al. 2021), and clinical and regulatory sources; it does not define the problem in terms of the recommended solution, nor fit a parameter and call it a prediction. The authors cite their own prior work (Keenan & Sokol 2023; Sokol & Vogt 2023; Sokol & Vogt 2025; Sokol et al. 2025), but in each case the self-cited item is one supporting reference among several external citations and is not the load-bearing premise. The causal diagnosis that technology-centric development underlies the translational gap is asserted rather than demonstrated, which is an evidence or correctness limitation, not a circularity. The paper's own success examples (diabetic retinopathy, skin cancer, lymph-node metastasis) are benchmark-trained classifiers, which weakens the empirical claim but again does not make it circular. No step reduces by construction to its input, so the correct circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters, formal model, or invented entities. Its central argument rests on background domain assumptions from cognitive science, HCI, and sociotechnical systems research, which are plausible but not established by the paper itself.

assumptions (4)
  • domain assumption The translational gap in medical AI is substantially caused by technology-centric development rather than by regulatory, economic, or structural factors.
    Stated in the Introduction ('the prevailing technology-centric approaches underpin this challenge'); asserted with citations but not empirically demonstrated in this paper.
  • domain assumption Supporting cognitive and epistemic functions, rather than automating tasks, will improve adoption and clinical outcomes.
    This is the core hypothesis of the Perspective, introduced in 'Systems Ecology and AI' and elaborated in 'From Benchmark to Bedside'; no prospective evidence is provided.
  • domain assumption Clinical diagnostic errors are largely due to structural factors and cognitive biases that AI tools can mitigate.
    Invoked in 'Systems Ecology and AI' with references to Croskerry and other cognitive science work; accepted background in that literature but not tested here.
  • domain assumption The stable-world principle distinguishes tasks amenable to automation, and clinical reasoning is semi-structured and therefore requires augmentation.
    Used in 'Human Decision Making and Artificial Intelligence' to argue that end-to-end AI modelling is inappropriate for many clinical tasks; a framing assumption rather than a formal theorem.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Artificial Intelligence Should Genuinely Support Clinical Reasoning and Decision Making To Bridge the Translational Gap." pith.science (2026). https://pith.science/paper/F2XDB6EV

@misc{pith2026250605030,
  author       = {Pith},
  title        = {Pith review of: Artificial Intelligence Should Genuinely Support Clinical Reasoning and Decision Making To Bridge the Translational Gap},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F2XDB6EV}},
  note         = {Machine review of arXiv:2506.05030}
}
read the original abstract

Artificial intelligence promises to revolutionise medicine, yet its impact remains limited because of the pervasive translational gap. We posit that the prevailing technology-centric approaches underpin this challenge, rendering such systems fundamentally incompatible with clinical practice, specifically diagnostic reasoning and decision making. Instead, we propose a novel sociotechnical conceptualisation of data-driven support tools designed to complement doctors' cognitive and epistemic activities. Crucially, it prioritises real-world impact over superhuman performance on inconsequential benchmarks.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

140 extracted references · 76 canonical work pages

  1. [1]

    Sarker, I. H. Machine learning: Algorithms, real-world applications and re- search directions.SN Computer Science2, 160 (2021)

  2. [2]

    C.Intelligence-based medicine: Artificial intelligence and human cogni- tion in clinical medicine and healthcare(Academic Press, Cambridge, MA, USA, 2020)

    Chang, A. C.Intelligence-based medicine: Artificial intelligence and human cogni- tion in clinical medicine and healthcare(Academic Press, Cambridge, MA, USA, 2020)

  3. [3]

    M., Lambert, C

    Bica, I., Alaa, A. M., Lambert, C. & van der Schaar, M. From real-world pa- tient data to individualized treatment effects using machine learning: Current and future methods to address underlying challenges.Clinical Pharmacology & Therapeutics109, 87–100 (2021)

  4. [4]

    Johnson, M.et al.The potential and pitfalls of artificial intelligence in clinical pharmacology.CPT: Pharmacometrics & Systems Pharmacology12, 279–284 (2023)

  5. [5]

    Gulshan, V.et al.Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs.JAMA316, 2402–2410 (2016)

  6. [6]

    Esteva, A.et al.Dermatologist-level classification of skin cancer with deep neural networks.Nature542, 115–118 (2017)

  7. [7]

    Golden, J. A. Deep learning algorithms for detection of lymph node metastases from breast cancer: Helping artificial intelligence be seen.JAMA318, 2184– 2186 (2017)

  8. [8]

    Patient care information systems and health care work: A sociotech- nical approach.International Journal of Medical Informatics55, 87–101 (1999)

    Berg, M. Patient care information systems and health care work: A sociotech- nical approach.International Journal of Medical Informatics55, 87–101 (1999)

Show all 140 references
  1. [9]

    Wiens, J.et al.Do no harm: A roadmap for responsible machine learning for health care.Nature Medicine25, 1337–1340 (2019)

  2. [10]

    Wardi, G.et al.Bringing the promise of artificial intelligence to critical care: What the experience with sepsis analytics can teach us.Critical Care Medicine 51, 985–991 (2023)

  3. [11]

    All models are wrong and yours are useless: Making clinical prediction models impactful for patients.npj Precision Oncology8, 54 (2024)

    Markowetz, F. All models are wrong and yours are useless: Making clinical prediction models impactful for patients.npj Precision Oncology8, 54 (2024)

  4. [12]

    The theoretical framework of cognitive informatics.International Journal of Cognitive Informatics and Natural Intelligence (IJCINI)1, 1–27 (2007)

    Wang, Y. The theoretical framework of cognitive informatics.International Journal of Cognitive Informatics and Natural Intelligence (IJCINI)1, 1–27 (2007)

  5. [13]

    A., Badawi, O., Gordon, A

    Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C. & Faisal, A. A. The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care.Nature Medicine24, 1716–1720 (2018)

  6. [14]

    Topol, E. J. High-performance medicine: The convergence of human and arti- ficial intelligence.Nature Medicine25, 44–56 (2019)

  7. [15]

    & Rodriguez, M

    Tsirtsis, S. & Rodriguez, M. Finding counterfactually optimal action sequences in continuous state spaces.Advances in Neural Information Processing Systems 36, 3220–3247 (2023)

  8. [16]

    & Al-Fuqaha, A

    Qayyum, A., Qadir, J., Bilal, M. & Al-Fuqaha, A. Secure and robust machine learning for healthcare: A survey.IEEE Reviews in Biomedical Engineering14, 156–180 (2020)

  9. [17]

    & Galstyan, A

    Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K. & Galstyan, A. A survey on bias and fairness in machine learning.ACM Computing Surveys (CSUR)54, 1–35 (2021)

  10. [18]

    Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.Nature Machine Intelligence1, 206–215 (2019)

    Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.Nature Machine Intelligence1, 206–215 (2019)

  11. [19]

    Why AI failed to live up to its potential during the pandemic

    Chakravorti, B. Why AI failed to live up to its potential during the pandemic. Harvard Business Review(2022)

  12. [20]

    & Nsoesie, E

    Ghassemi, M. & Nsoesie, E. O. In medicine, how do we machine learn anything real?Patterns3, 100392 (2022)

  13. [21]

    L., Ercole, A., Zhao, J

    Volovici, V., Syn, N. L., Ercole, A., Zhao, J. J. & Liu, N. Steps to avoid overuse and misuse of machine learning in clinical research.Nature Medicine28, 1996–1999 (2022)

  14. [22]

    C.et al.Implemented machine learning tools to inform decision- making for patient care in hospital settings: A scoping review.BMJ Open13, e065845 (2023)

    Tricco, A. C.et al.Implemented machine learning tools to inform decision- making for patient care in hospital settings: A scoping review.BMJ Open13, e065845 (2023)

  15. [23]

    A., Levin, J., Kahn, J

    Sivaraman, V., Bukowski, L. A., Levin, J., Kahn, J. M. & Perer, A. Ignore, trust, or negotiate: Understanding clinician acceptance of AI-based treatment recom- mendations in health care. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 1–18 (ACM, 2023)

  16. [24]

    T., Hoffman, R

    Mueller, S. T., Hoffman, R. R., Clancey, W., Emrey, A. & Klein, G. Explanation in human–AI systems: A literature meta-review, synopsis of key ideas and pub- lications, and bibliography for explainable AI. Tech. Rep. AD1073994, Florida Institute for Human and Machine Cognition (2019)

  17. [25]

    Akata, Z.et al.A research agenda for hybrid intelligence: Augmenting human intellect with collaborative, adaptive, responsible, and explainable artificial in- telligence.Computer53, 18–28 (2020)

  18. [26]

    Croskerry, P.The cognitive autopsy: A root cause analysis of medical decision making(Oxford University Press, Oxford, UK, 2020)

  19. [27]

    & Verhoef, P

    van Baalen, S., Boon, M. & Verhoef, P. From clinical decision support to clinical reasoning support systems.Journal of Evaluation in Clinical Practice27, 520– 9 Artificial Intelligence Should Genuinely Support Clinical Reasoning and Decision Making To Bridge the Translational ...

  20. [28]

    new” and “old

    Simkute, A., Surana, A., Luger, E., Evans, M. & Jones, R. XAI for learning: Narrowing down the digital divide between “new” and “old” experts. InAd- junct Proceedings of the 2022 Nordic Human–Computer Interaction Conference, 1–6 (2022)

  21. [29]

    A conversation with Ken Holstein: Fostering human–AI comple- mentarity.Ubiquity2023, 1–6 (2023)

    Anjum, B. A conversation with Ken Holstein: Fostering human–AI comple- mentarity.Ubiquity2023, 1–6 (2023)

  22. [30]

    P., Lyell, D., Widyantoro, B., Berkovsky, S

    Susanto, A. P., Lyell, D., Widyantoro, B., Berkovsky, S. & Magrabi, F. Effects of machine learning-based clinical decision support systems on decision-making, care delivery, and patient outcomes: A scoping review.Journal of the American Medical Informatics Association30, 2050–...

  23. [31]

    Wosny, M., Strasser, L. M. & Hastings, J. Experience of health care professionals using digital tools in the hospital: Qualitative systematic review.JMIR Human Factors10, e50357 (2023)

  24. [32]

    V., Vorvoreanu, M., Subramonyam, H

    Liao, Q. V., Vorvoreanu, M., Subramonyam, H. & Wilcox, L. UX matters: The critical role of UX in responsible AI.Interactions31, 22–27 (2024)

  25. [33]

    The Edmond & Lily Safra Center for Ethics, Harvard University

    Siddarth, D.et al.How AI fails us.Justice, Health, and Democracy Impact Ini- tiative & Carr Center for Human Rights Policy(2021). The Edmond & Lily Safra Center for Ethics, Harvard University

  26. [34]

    Munn, L.Automation is a myth(Stanford University Press, Redwood City, CA, USA, 2022)

  27. [35]

    Sheehan, B.et al.Informing the design of clinical decision support services for evaluation of children with minor blunt head trauma in the emergency depart- ment: A sociotechnical analysis.Journal of Biomedical Informatics46, 905–913 (2013)

  28. [36]

    & Mohrman, S

    Winby, S. & Mohrman, S. A. Digital sociotechnical system design.The Journal of Applied Behavioral Science54, 399–423 (2018)

  29. [37]

    Pasmore, W., Winby, S., Mohrman, S. A. & Vanasse, R. Reflections: Sociotech- nical systems design and organization change.Journal of Change Management 19, 67–85 (2019)

  30. [38]

    Seeber, I.et al.Machines as teammates: A research agenda on AI in team col- laboration.Information & Management57, 103174 (2020)

  31. [39]

    Shneiderman, B.Human-centered AI(Oxford University Press, Oxford, UK, 2022)

  32. [40]

    Explainable AI is dead, long live explainable AI! Hypothesis-driven decision support using evaluative AI

    Miller, T. Explainable AI is dead, long live explainable AI! Hypothesis-driven decision support using evaluative AI. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, 333–342 (2023)

  33. [41]

    & Pfeiffer, S

    Herrmann, T. & Pfeiffer, S. Keeping the organization in the loop: A socio- technical extension of human-centered artificial intelligence.AI & Society38, 1523–1542 (2023)

  34. [42]

    & Termine, A

    Ferrario, A., Facchini, A. & Termine, A. Experts or authorities? The strange case of the presumed epistemic superiority of artificial intelligence systems.Minds and Machines34, 30 (2024)

  35. [43]

    & Sokol, K

    Keenan, B. & Sokol, K. Mind the gap! Bridging explainable artificial intelligence and human understanding with Luhmann’s functional theory of communica- tion.arXiv preprint arXiv:2302.03460(2023)

  36. [44]

    Zhou, L.et al.From artificial intelligence (AI) to intelligence augmentation (IA): Design principles, potential risks, and emerging issues.AIS Transactions on Human–Computer Interaction15, 111–135 (2023)

  37. [45]

    Patel, V. L. & Groen, G. J. Knowledge based solution strategies in medical rea- soning.Cognitive Science10, 91–116 (1986)

  38. [46]

    & Barosi, G

    Ramoni, M., Stefanelli, M., Magnani, L. & Barosi, G. An epistemological frame- work for medical knowledge-based systems.IEEE Transactions on Systems, Man, and Cybernetics22, 1361–1375 (1992)

  39. [47]

    Kuhn, G. J. Diagnostic errors.Academic Emergency Medicine9, 740–750 (2002)

  40. [48]

    Cognitive forcing strategies in clinical decisionmaking.Annals of Emergency Medicine41, 110–120 (2003)

    Croskerry, P. Cognitive forcing strategies in clinical decisionmaking.Annals of Emergency Medicine41, 110–120 (2003)

  41. [49]

    Klein, J. G. Five pitfalls in decisions about diagnosis and prescribing.BMJ330, 781–783 (2005)

  42. [50]

    Groopman, J. E. & Prichard, M.How doctors think(Houghton Mifflin, Boston, MA, USA, 2007)

  43. [51]

    & Norman, G

    Croskerry, P. & Norman, G. Overconfidence in clinical decision making.The American Journal of Medicine121, S24–S29 (2008)

  44. [52]

    Tetlock, P. E. & Gardner, D.Superforecasting: The art and science of prediction (Random House, New York, NY, USA, 2016)

  45. [53]

    Singer, M.et al.The third international consensus definitions for sepsis and septic shock (sepsis-3).JAMA315, 801–810 (2016)

  46. [54]

    Fleischmann-Struzek, C.et al.The global burden of paediatric and neonatal sepsis: A systematic review.The Lancet Respiratory Medicine6, 223–230 (2018)

  47. [55]

    Morin, L.et al.The current and future state of pediatric sepsis definitions: An international survey.Pediatrics149, e2021052565 (2022)

  48. [56]

    J.et al.International consensus criteria for pediatric sepsis and septic shock.JAMA331, 665–674 (2024)

    Schlapbach, L. J.et al.International consensus criteria for pediatric sepsis and septic shock.JAMA331, 665–674 (2024)

  49. [57]

    & International Consensus Conference on Pediatric Sepsis

    Goldstein, B., Giroir, B., Randolph, A. & International Consensus Conference on Pediatric Sepsis. International pediatric sepsis consensus conference: Def- initions for sepsis and organ dysfunction in pediatrics.Pediatric Critical Care Medicine6, 2–8 (2005)

  50. [58]

    Tennant, R.et al.A scoping review on pediatric sepsis prediction technologies in healthcare.npj Digital Medicine7, 1–16 (2024)

  51. [59]

    Woods-Hill, C. Z.et al.Diagnostic stewardship for blood cultures in the pe- diatric intensive care unit: Lessons in implementation from the BrighT STAR Collaborative.Antimicrobial Stewardship & Healthcare Epidemiology4, e148 (2024)

  52. [60]

    Chiotos, K.et al.Antibiotic indications and appropriateness in the pediatric intensive care unit: A 10-center point prevalence study.Clinical Infectious Dis- eases76, e1021–e1030 (2023)

  53. [61]

    F., Buonocore, G., Maier, R

    Klingenberg, C., Kornelisse, R. F., Buonocore, G., Maier, R. F. & Stocker, M. Culture-negative early-onset neonatal sepsis – At the crossroad between effi- cient sepsis care and antimicrobial stewardship.Frontiers in Pediatrics6, 285 (2018)

  54. [62]

    & Coletti, C

    Lee, R., Al Rifaie, R., Subedi, K. & Coletti, C. Comparative analysis of bacteremic and non-bacteremic sepsis: A retrospective study.Cureus16, e76418 (2024)

  55. [63]

    Schlapbach, L. J. & Kissoon, N. Defining pediatric sepsis.JAMA Pediatrics172, 313–314 (2018)

  56. [64]

    Elish, M. C. The stakes of uncertainty: Developing and integrating machine learning in clinical care. InEthnographic Praxis in Industry Conference Proceed- ings, vol. 2018, 364–380 (Wiley Online Library, 2018)

  57. [65]

    R., Palaniyar, N

    Banerjee, S., Mohammed, A., Wong, H. R., Palaniyar, N. & Kamaleswaran, R. Machine learning identifies complicated sepsis course and subsequent mortality based on 20 genes in peripheral blood immune cells at 24 h post-ICU admission. Frontiers in Immunology12, 592303 (2021)

  58. [66]

    Griffin, M. P. & Moorman, J. R. Toward the early diagnosis of neonatal sepsis and sepsis-like illness using novel heart rate analysis.Pediatrics107, 97–104 (2001)

  59. [67]

    J.et al.Prediction of pediatric sepsis mortality within 1 h of intensive care admission.Intensive Care Medicine43, 1085–1096 (2017)

    Schlapbach, L. J.et al.Prediction of pediatric sepsis mortality within 1 h of intensive care admission.Intensive Care Medicine43, 1085–1096 (2017)

  60. [68]

    Joshi, R.et al.Predicting neonatal sepsis using features of heart rate variability, respiratory characteristics, and ECG-derived estimates of infant motion.IEEE Journal of Biomedical and Health Informatics24, 681–692 (2019)

  61. [69]

    Kanjilal, S.et al.A decision algorithm to promote outpatient antimicrobial stewardship for uncomplicated urinary tract infection.Science Translational Medicine12, eaay5067 (2020)

  62. [70]

    W.et al.Development of a machine learning model using elec- tronic health record data to identify antibiotic use among hospitalized patients

    Moehring, R. W.et al.Development of a machine learning model using elec- tronic health record data to identify antibiotic use among hospitalized patients. JAMA Network Open4, e213460 (2021)

  63. [71]

    Adams, R.et al.Prospective, multi-site study of patient outcomes after im- plementation of the TREWS machine learning-based early warning system for sepsis.Nature Medicine28, 1455–1460 (2022)

  64. [72]

    & Jurman, G

    Chicco, D. & Jurman, G. Survival prediction of patients with sepsis from age, sex, and septic episode number alone.Scientific Reports10, 17156 (2020)

  65. [73]

    M., Caterino, J

    Liu, R., Hunold, K. M., Caterino, J. M. & Zhang, P. Estimating treatment effects for time-to-treatment antibiotic stewardship in sepsis.Nature Machine Intelli- gence5, 421–431 (2023)

  66. [74]

    K.et al.Diagnosis trajectories of prior multi-morbidity predict sepsis mortality.Scientific Reports6, 36624 (2016)

    Beck, M. K.et al.Diagnosis trajectories of prior multi-morbidity predict sepsis mortality.Scientific Reports6, 36624 (2016)

  67. [75]

    & Krauthammer, M

    Allam, A., Feuerriegel, S., Rebhan, M. & Krauthammer, M. Analyzing patient trajectories with artificial intelligence.Journal of Medical Internet Research23, e29812 (2021)

  68. [76]

    Data-driven clinical decision processes: It’s time.Journal of Translational Medicine17, 1–2 (2019)

    Capobianco, E. Data-driven clinical decision processes: It’s time.Journal of Translational Medicine17, 1–2 (2019)

  69. [77]

    & Jenkins, J

    Spatharou, A., Hieronimus, S. & Jenkins, J. Transforming healthcare with AI: The impact on the workforce and organizations.McKinsey & Company10, 1– 131 (2020)

  70. [78]

    InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–16 (ACM, 2021)

    Bansal, G.et al.Does the whole exceed its parts? The effect of AI explanations on complementary team performance. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–16 (ACM, 2021)

  71. [79]

    Could machine learning fuel a reproducibility crisis in science?Na- ture608, 250–251 (2022)

    Gibney, E. Could machine learning fuel a reproducibility crisis in science?Na- ture608, 250–251 (2022)

  72. [80]

    The reproducibility issues that haunt health-care AI.Nature613, 402– 403 (2023)

    Sohn, E. The reproducibility issues that haunt health-care AI.Nature613, 402– 403 (2023)

  73. [81]

    & Klein, G

    Kahneman, D. & Klein, G. Conditions for intuitive expertise: A failure to dis- agree.American Psychologist64, 515–526 (2009)

  74. [82]

    O’Neil, C.Weapons of math destruction: How big data increases inequality and threatens democracy(Crown, New York, NY, USA, 2016)

  75. [83]

    & Kirchner, L

    Angwin, J., Larson, J., Mattu, S. & Kirchner, L. Machine bias. InEthics of Data and Analytics, 254–264 (Auerbach Publications, New York, NY, USA, 2022)

  76. [84]

    Dastin, J. Amazon scraps secret AI recruiting tool that showed bias against 10 Artificial Intelligence Should Genuinely Support Clinical Reasoning and Decision Making To Bridge the Translational Gap women. InEthics of Data and Analytics, 296–299 (Auerbach Publications, New Yor...

  77. [85]

    Robo-debt illegality: The seven veils of failed guarantees of the rule of law?Alternative Law Journal44, 4–10 (2019)

    Carney, T. Robo-debt illegality: The seven veils of failed guarantees of the rule of law?Alternative Law Journal44, 4–10 (2019)

  78. [86]

    Geiger, G.et al.Suspicion machines: Unprecedented experiment on welfare surveillance algorithm reveals discrimination.Lighthouse Reports(2023)

  79. [87]

    & Narayanan, A.Fairness and machine learning: Limita- tions and opportunities(MIT Press, Cambridge, MA, USA, 2023)

    Barocas, S., Hardt, M. & Narayanan, A.Fairness and machine learning: Limita- tions and opportunities(MIT Press, Cambridge, MA, USA, 2023)

  80. [88]

    & Gebru, T

    Buolamwini, J. & Gebru, T. Gender shades: Intersectional accuracy disparities in commercial gender classification. InProceedings of the 2018 ACM Conference on Fairness, Accountability and Transparency, 77–91 (PMLR, 2018)

  81. [89]

    & Sunstein, C

    Kahneman, D., Sibony, O. & Sunstein, C. R.Noise: A flaw in human judgment (Hachette, London, UK, 2021)

  82. [90]

    Byrne, R. M. Good explanations in explainable artificial intelligence (XAI): Ev- idence from human explanatory reasoning. InIJCAI, 6536–6544 (2023)

  83. [91]

    & Hamarneh, G

    Jin, W., Li, X. & Hamarneh, G. Why is plausibility surprisingly problematic as an XAI criterion?arXiv preprint arXiv:2303.17707(2025)

  84. [92]

    L.et al.Cognitive interventions to reduce diagnostic error: A nar- rative review.BMJ Quality & Safety21, 535–557 (2012)

    Graber, M. L.et al.Cognitive interventions to reduce diagnostic error: A nar- rative review.BMJ Quality & Safety21, 535–557 (2012)

  85. [93]

    & Kopp, S

    Battefeld, D. & Kopp, S. Formalizing cognitive biases in medical diagnostic reasoning.Proceedings of the 8th Workshop on Formal and Cognitive Reasoning 3242, 102–118 (2022)

  86. [94]

    Ackerman, M. S. The intellectual challenge of CSCW: The gap between social requirements and technical feasibility.Human–Computer Interaction15, 179– 203 (2000)

  87. [95]

    & Riedl, M

    Ehsan, U., Saha, K., De Choudhury, M. & Riedl, M. O. Charting the sociotechni- cal gap in explainable AI: A framework to address the gap in XAI.Proceedings of the ACM on Human–Computer Interaction7, 1–32 (2023)

  88. [96]

    Patel, V. L. & Cohen, T. A. Clinical cognition and AI: From emulation to sym- biosis. InIntelligent Systems in Medicine and Health: The Role of AI, 109–133 (Springer, Cham, Switzerland, 2022)

  89. [97]

    The Turing trap: The promise & peril of human-like artificial intelligence

    Brynjolfsson, E. The Turing trap: The promise & peril of human-like artificial intelligence. InAugmented Education in the Global Age, 103–116 (Routledge, New York, NY, USA, 2023)

  90. [98]

    Ironies of automation.Automatica19, 775–779 (1983)

    Bainbridge, L. Ironies of automation.Automatica19, 775–779 (1983)

  91. [99]

    & Naumann, F

    Hacker, P., Krestel, R., Grundmann, S. & Naumann, F. Explainable AI under contract and tort law: Legal incentives and technical challenges.Artificial In- telligence and Law28, 415–439 (2020)

  92. [100]

    & Boon, M

    van Baalen, S. & Boon, M. An epistemological shift: From evidence-based medicine to epistemological responsibility.Journal of Evaluation in Clinical Practice21, 433–439 (2015)

  93. [101]

    Tomsett, R.et al.Rapid trust calibration through interpretable and uncertainty- aware AI.Patterns1(2020)

  94. [102]

    In AI we trust: Ethics, artificial intelligence, and reliability.Science and Engineering Ethics26, 2749–2767 (2020)

    Ryan, M. In AI we trust: Ethics, artificial intelligence, and reliability.Science and Engineering Ethics26, 2749–2767 (2020)

  95. [103]

    Esposito, E.Artificial communication: How algorithms produce social intelligence (MIT Press, Cambridge, MA, USA, 2022)

  96. [104]

    Rudin, C.et al.Interpretable machine learning: Fundamental principles and 10 grand challenges.Statistic Surveys16, 1–85 (2022)

  97. [105]

    Human-centered artificial intelligence: Reliable, safe & trust- worthy.International Journal of Human–Computer Interaction36, 495–504 (2020)

    Shneiderman, B. Human-centered artificial intelligence: Reliable, safe & trust- worthy.International Journal of Human–Computer Interaction36, 495–504 (2020)

  98. [106]

    & Vogt, J

    Sokol, K. & Vogt, J. E. (Un)reasonable allure of ante-hoc interpretability for high-stakes domains: Transparency is necessary but insufficient for compre- hensibility. In3rd ICML Workshop on Interpretable Machine Learning in Health- care(2023)

  99. [107]

    Parasuraman, R., Sheridan, T. B. & Wickens, C. D. A model for types and levels of human interaction with automation.IEEE Transactions on Systems, Man, and Cybernetics – Part A: Systems and Humans30, 286–297 (2000)

  100. [108]

    Sheridan, T. B. & Verplank, W. L. Human and computer control of undersea teleoperators. Tech. Rep. ADA057655, Massachusetts Institute of Technology Cambridge Man–Machine Systems Laboratory (1978)

  101. [109]

    Intelligent decision support systems.Multicriteria Decision Aid and Artificial Intelligence: Links, Theory and Applications25–44 (2013)

    Phillips-Wren, G. Intelligent decision support systems.Multicriteria Decision Aid and Artificial Intelligence: Links, Theory and Applications25–44 (2013)

  102. [110]

    Lawrence, N. D. Data readiness levels.arXiv preprint arXiv:1705.02245(2017)

  103. [111]

    Levels of automation (2022)

    United States Department of Transportation: National Highway Traffic Safety Administration (NHTSA). Levels of automation (2022). URLhttps://www. nhtsa.gov/document/levels-automation

  104. [112]

    Adop- tion model for analytics maturity (AMAM) (2021)

    Healthcare Information and Management Systems Society (HIMSS). Adop- tion model for analytics maturity (AMAM) (2021). URLhttps://www. himss.org/what-we-do-solutions/digital-health-transformation/ maturity-models/adoption-model-analytics-maturity-amam

  105. [113]

    Martínez-Plumed, F.et al.CRISP-DM twenty years later: From data mining processes to data science trajectories.IEEE Transactions on Knowledge and Data Engineering(2019)

  106. [114]

    Kahneman, D.Thinking, fast and slow(Macmillan, London, UK, 2011)

  107. [115]

    M.Educating intuition(University of Chicago Press, Chicago, IL, USA, 2001)

    Hogarth, R. M.Educating intuition(University of Chicago Press, Chicago, IL, USA, 2001)

  108. [116]

    & Gamblian, V

    Crandall, B. & Gamblian, V. Guide to early sepsis assessment in the NICU. Instruction Manual Prepared for the Ohio Department of Development Under the Ohio SBIR Bridge Grant Program by Klein Associates Inc(1991)

  109. [117]

    & Getchell-Reiter, K

    Crandall, B. & Getchell-Reiter, K. Critical decision method: A technique for eliciting concrete assessment indicators from the intuition of NICU nurses.Ad- vances in Nursing Science16, 42–51 (1993)

  110. [118]

    V., Simsek, O., Buckmann, M

    Katsikopoulos, K. V., Simsek, O., Buckmann, M. & Gigerenzer, G.Classifica- tion in the wild: The science and art of transparent decision making(MIT Press, Cambridge, MA, USA, 2021)

  111. [119]

    Psychological AI: Designing algorithms informed by human psychology.Perspectives on Psychological Science19, 839–848 (2023)

    Gigerenzer, G. Psychological AI: Designing algorithms informed by human psychology.Perspectives on Psychological Science19, 839–848 (2023)

  112. [120]

    & Thompson, C

    Dowding, D. & Thompson, C. Evidence-based decisions: The role of decision analysis. InEssential Decision Making and Clinical Judgement for Nurses, 173– 195 (Churchill Livingstone, London, UK, 2009)

  113. [121]

    Al-Azzawi, R., Halvorsen, P. A. & Risør, T. Context and general practitioner decision-making – A scoping review of contextual influence on antibiotic pre- scribing.BMC Family Practice22, 1–10 (2021)

  114. [122]

    Stocker, M.et al.Less is more: Antibiotics at the beginning of life.Nature Communications14, 2423 (2023)

  115. [123]

    E., Scott, H

    Martin, B., DeWitt, P. E., Scott, H. F., Parker, S. & Bennett, T. D. Machine learn- ing approach to predicting absence of serious bacterial infection at PICU ad- mission.Hospital Pediatrics12, 590–603 (2022)

  116. [124]

    It’s safer to

    Cabral, C., Lucas, P. J., Ingram, J., Hay, A. D. & Horwood, J. “It’s safer to... ” Par- ent consulting and clinician antibiotic prescribing decisions for children with respiratory tract infections: An analysis across four qualitative studies.Social Science & Medicine136, 156–1...

  117. [125]

    S.et al.Clinical reasoning behind antibiotic use in PICUs: A quali- tative study.Pediatric Critical Care Medicine23, e126–e135 (2022)

    Fontela, P. S.et al.Clinical reasoning behind antibiotic use in PICUs: A quali- tative study.Pediatric Critical Care Medicine23, e126–e135 (2022)

  118. [126]

    & Bjerrum, L

    Llor, C. & Bjerrum, L. Antimicrobial resistance: Risk associated with antibiotic overuse and initiatives to reduce the problem.Therapeutic Advances in Drug Safety5, 229–241 (2014)

  119. [127]

    & O’Donoghue, T

    Frederick, S., Loewenstein, G. & O’Donoghue, T. Time discounting and time preference: A critical review.Journal of Economic Literature40, 351–401 (2002)

  120. [128]

    & Hey- dari, H

    Mahboub-Ahari, A., Pourreza, A., Akbari Sari, A., Rahimi Foroushani, A. & Hey- dari, H. Stated time preferences for health: A systematic review and meta anal- ysis of private and social discount rates.Journal of Research in Health Sciences 14, 181–186 (2014)

  121. [129]

    Cook, M. B. & Smallman, H. S. Human factors of the confirmation bias in intel- ligence analysis: Decision support from graphical evidence landscapes.Human Factors50, 745–754 (2008)

  122. [130]

    It is a moving process

    Corti, L.et al.“It is a moving process”: Understanding the evolution of explain- ability needs of clinicians in pulmonary medicine. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 1–21 (ACM, 2024)

  123. [131]

    & Lim, B

    Wang, D., Yang, Q., Abdul, A. & Lim, B. Y. Designing theory-driven user-centric explainable AI. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1–15 (ACM, 2019)

  124. [132]

    Buçinca, Z., Malaya, M. B. & Gajos, K. Z. To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making.Pro- ceedings of the ACM on Human–Computer Interaction5, 1–21 (2021)

  125. [133]

    Bertrand, A., Belloum, R., Eagan, J. R. & Maxwell, W. How cognitive biases affect XAI-assisted decision-making: A systematic review. InProceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, 78–91 (2022)

  126. [134]

    & Vogt, J

    Sokol, K. & Vogt, J. E. NeXAI: Naturalistic explainable artificial intelligence for data-driven reasoning and decision support.arXiv preprint(2025)

  127. [135]

    & Rescher, N

    Helmer, O. & Rescher, N. On the epistemology of the inexact sciences.Manage- ment Science6, 25–52 (1959)

  128. [136]

    & Xuan, Y

    Sokol, K., Small, E. & Xuan, Y. Navigating explanatory multiverse through counterfactual path geometry.Machine Learning114, 1–33 (2025)

  129. [137]

    F.A history of medical informatics in the United States, 1950 to 1990 (American Medical Informatics Association, Indianapolis, IN, USA, 1995)

    Collen, M. F.A history of medical informatics in the United States, 1950 to 1990 (American Medical Informatics Association, Indianapolis, IN, USA, 1995)

  130. [138]

    E., Goedhart, R

    Röber, T. E., Goedhart, R. & Birbil, Ş. İ. Clinicians’ voice: Fundamental consid- erations for XAI in healthcare.arXiv preprint arXiv:2411.04855(2024)

  131. [139]

    L.et al.Improved hospital mortality rates after the implementation of emergency department sepsis teams.The American Journal of Emergency Medicine51, 218–222 (2022)

    Simon, E. L.et al.Improved hospital mortality rates after the implementation of emergency department sepsis teams.The American Journal of Emergency Medicine51, 218–222 (2022)

  132. [140]

    E., Barry, H., Scanlan, J

    Currie, K. E., Barry, H., Scanlan, J. M. & Harvey, E. M. Impact of a multidisci- plinary sepsis huddle in the emergency department.The American Journal of Emergency Medicine64, 150–154 (2023). 11

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.