{"id":"9936deac-1e49-451a-a38d-cd618d03247e","arxiv_id":"2506.05030","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A Perspective arguing that clinical AI should be reframed as cognitive and epistemic support for clinicians, not as autonomous predictors.","lead":"This Perspective argues that medical AI should support doctors' reasoning and decision-making instead of replacing them. It proposes a sociotechnical framework, illustrated with paediatric sepsis, for building tools that complement clinical cognition and workflows.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's normative vision is clear, but its load-bearing causal diagnosis—that technology-centric development causes the translational gap and cognitive-support reframing will close it—is asserted, not demonstrated, and its own success examples pull against it.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing causal premise: the translational gap is caused by technology-centric development, and reframing AI as cognitive support will close it. I agree with that assessment and with the UNVERDICTED verdict, because the paper is a Perspective whose central claim is normative but whose supporting causal chain is empirical and unsupported. My additional observation is internal: the three celebrated success stories cited in the paper (refs 5–7) are themselves technology-centric, benchmark-driven classifiers, yet they have translated into practice. That creates a tension within the paper's own evidence and further weakens the claim that technology-centric design is the main cause of the translational gap. The concern is not that the authors are wrong to propose cognitive support; it is that the paper does not establish that this reframing is necessary or sufficient for impact. The proposed concrete test—coding existing cases by paradigm versus task environment, plus a prospective comparison of the same model deployed in predictive-alert and mental-simulation modes—would distinguish a paradigm-driven explanation from a task-environment-driven one and thereby settle whether the concern lands. Since the paper makes no empirical claim to verify, leaving the verdict UNVERDICTED is appropriate.","tokens_in":19617,"tokens_out":3730,"duration_ms":50464,"concrete_test":"Compile the paper's own cases—refs 5–7 (translated predictive classifiers) and the sepsis literature (non-translated)—and code each by design paradigm (benchmark-optimized output vs. cognitive support) and task environment (stable-world/structured vs. semi-structured/wicked; visual vs. non-visual; self-contained vs. workflow-embedded). If translated systems are technology-centric but occupy stable, self-contained visual tasks while failures are workflow-embedded regardless of paradigm, the causal diagnosis fails. A prospective pilot of the same sepsis model as (i) predictive alert and (ii) mental-simulation support, measuring adherence and decision quality, would directly test whether the proposed reframing changes adoption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a prescription: AI should support clinicians' cognitive and epistemic functions rather than optimize benchmark accuracy. That prescription only follows if two empirical premises are true: (1) the translational gap is substantially caused by technology-centric system design, and (2) redesigning systems around cognitive support will materially raise adoption and improve outcomes. The paper asserts premise (1) in the Introduction ('the prevailing technology-centric approaches underpin this challenge') and premise (2) throughout, but offers no comparative or prospective evidence. The 'Medical Artificial Intelligence Adoption Challenges' section itself lists other candidate causes—regulatory approval, data access, implementation cost, alert fatigue, trust, reproducibility—without weighting them, so the causal attribution is unestablished. The paper's own three flagship successes (diabetic retinopathy, skin cancer, lymph-node metastasis; refs 5–7) are benchmark-trained predictive classifiers, i.e., technology-centric by the paper's criterion, which cuts against the claim that such an approach is fundamentally incompatible with clinical practice. This does not invalidate the normative proposal, but it leaves the bridge from 'current AI underperforms in the clinic' to 'cognitive-support AI will perform better' unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Perspective argues that the persistent translational gap in medical AI is caused by technology-centric development, which optimizes benchmark predictive performance while ignoring clinicians' real reasoning and decision-making processes. The authors propose a sociotechnical reframing in which AI systems augment, rather than automate, clinicians' cognitive and epistemic functions—such as reasoning under uncertainty, mental simulation, and belief updating—and they use paediatric sepsis as a running case study. The paper contributes a tripartite distinction among autonomy, assistance, and augmentation, and it outlines concrete mechanisms (e.g., cognitive forcing, counterfactual simulation) through which such support might operate.","tokens_in":19754,"tokens_out":3491,"duration_ms":42984,"significance":"If the central claim is correct, the paper usefully redirects medical AI research away from benchmark chasing and toward a human-centred, cognition-aware design agenda. It draws on a broad interdisciplinary literature and offers a clear conceptual vocabulary for discussing AI integration levels. The paper is well-structured and thoughtful, and it honestly acknowledges the lack of existing implementation frameworks. The main weakness is that its causal diagnosis—that technology-centric development is the principal cause of the translational gap—is asserted rather than demonstrated, and the paper's own success stories (benchmark-trained classifiers) undercut that diagnosis. These issues make the proposal a normative hypothesis rather than an evidence-based conclusion.","major_comments":[{"comment":"The Introduction asserts that 'the prevailing technology-centric approaches underpin this challenge' and that this renders AI 'fundamentally incompatible with clinical practice' (Abstract and Introduction). This is a load-bearing causal claim, yet the paper offers no comparative or prospective evidence for it. The later section 'Medical Artificial Intelligence Adoption Challenges' itself lists a range of other candidate causes—regulatory approval, data access, implementation costs, alert fatigue, trust, reproducibility—without weighing them. If these factors dominate, the proposed shift to cognitive support may not close the translational gap. The authors should either soften the causal claim to a hypothesis or provide concrete evidence that technology-centric design is the primary bottleneck.","section":"Introduction; Medical Artificial Intelligence Adoption Challenges"},{"comment":"The paper's three flagship successes—diabetic retinopathy detection (ref 5), skin cancer classification (ref 6), and lymph-node metastasis detection (ref 7)—are, by the paper's own definition, technology-centric: they are benchmark-trained deep learning classifiers. The paper acknowledges that these successes are 'not necessarily representative' and pertain mostly to visual domains, but it does not explain why these cases do not contradict the claim that technology-centric approaches are 'fundamentally incompatible' with clinical practice. This is a tension that should be addressed explicitly: if benchmark-trained classifiers succeeded in these domains, what distinguishes the domains where they fail, and how does cognitive support address those distinguishing features?","section":"Medical Artificial Intelligence Adoption Challenges"},{"comment":"The proposed cognitive-support tools (mental simulation, cognitive forcing, belief updating, digital twins) are described at a conceptual level, but the paper admits in 'Human Decision Making and Artificial Intelligence' that 'we generally lack the corresponding (technical) frameworks, guidelines and protocols' for implementing such systems. The 'From Benchmark to Bedside' section cites only 'rudimentary research' and notes that this line of work 'by and large overlooks the broader systems ecology.' As a consequence, the central claim that these tools will materially improve adoption and outcomes is a conjecture rather than a result. The paper should frame this as an open research agenda and identify at least one concrete evaluation pathway (e.g., pilot studies with process and outcome measures) that could test the hypothesis.","section":"Human Decision Making and Artificial Intelligence; From Benchmark to Bedside"}],"minor_comments":[{"comment":"The keyword list uses interpuncts to separate terms; this is nonstandard. Please use commas or semicolons for clarity.","section":"Keywords"},{"comment":"The phrase 'the race to the bottom' is introduced without definition. Since it is not a standard term in the medical AI literature, it should be explained or replaced with a more descriptive label.","section":"Systems Ecology and Artificial Intelligence"},{"comment":"The term 'ante-hoc interpretable' appears before it is formally introduced. Define it at first use, perhaps with a brief contrast to post-hoc explainability.","section":"Systems Ecology and Artificial Intelligence"},{"comment":"The concept of 'digital twins of individual clinicians' is intriguing but underdeveloped. The paper does not explain how such a twin would be created, validated, or used in practice. A reference or a one-sentence operational sketch would help.","section":"From Benchmark to Bedside"}],"recommendation":"major_revision","confidential_remarks":"The paper is a Perspective, so I did not require original empirical data. However, the causal claim and the proposal's feasibility are central to its message, and both need strengthening or reframing. The paper is already published in npj Digital Medicine; if this is a resubmission to another venue, that prior publication should be disclosed to the editor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth reading. It makes a clear case that clinical AI should be built to support doctors' reasoning and decision making, not just to score better on benchmarks. The authors bring together naturalistic decision making, cognitive forcing, hybrid intelligence, and systems ecology into a coherent vision, and their three-way distinction among autonomy, assistance, and augmentation is a genuinely handy way to talk about integration. The sepsis case study grounds the discussion in a real, high-stakes problem where the mismatch between classification-style AI and clinical reasoning is obvious. It is also honest about being a Perspective, not a study.\n\nThe soft spot is the load-bearing causal claim. The paper says technology-centric approaches \"underpin\" the translational gap and render AI \"fundamentally incompatible\" with clinical practice, but that is a hypothesis, not an established fact. The authors themselves list other plausible causes—regulation, reimbursement, data access, implementation cost, alert fatigue—without weighting any of them. More tellingly, their own flagship success stories (diabetic retinopathy, skin cancer, lymph node detection) are exactly the kind of benchmark-trained predictive classifiers they call incompatible. Those examples do not disprove the argument, but the tension is never addressed. If such systems made it into practice, then technology-centric development is not categorically doomed; it just works in narrow, well-scoped visual tasks. The paper would be stronger if it acknowledged that and framed its proposal as a way to broaden success beyond those niches.\n\nThe proposed alternative, AI-assisted mental simulation, is still a sketch. That is fine for a Perspective, but it should be labeled as an agenda, not a solution. The authors do that in the conclusion, where they call the Perspective \"far from exhaustive in itself.\" The few self-citations are to relevant prior work on interpretability and explanation, so no concern there.\n\nOverall, this is a solid position piece for anyone working on clinical decision support. It will likely be cited as a reference point for the augmentation agenda. I would bring it to a reading group for the discussion alone. It deserves serious peer review if it were being submitted now, and it deserves a careful read as published—but treat the causal diagnosis as an open question, not a settled fact.","headline":"A useful, well-written Perspective with an overreaching causal diagnosis; the framing is worth engaging even though the central claim is asserted rather than demonstrated.","tokens_in":20324,"tokens_out":2026,"would_cite":true,"duration_ms":27473,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that medical AI's translational gap — the chasm between high benchmark performance and real clinical impact — is caused by technology-centric development, and should be closed by building AI that supports clinicians'…","keywords":["clinical reasoning support","translational gap","sociotechnical AI","augmented intelligence","cognitive forcing","decision support systems","pediatric sepsis","human-centred AI"],"falsifier":"A prospective, controlled deployment comparing a cognitive-support system (e.g., one delivering mental simulation and cognitive-forcing prompts for suspected sepsis) against a conventional predictive alert tool, measuring clinician uptake, decision quality, and patient outcomes: if the conventional tool matches the cognitive-support tool on adoption and outcomes, the claim that the reframing is necessary to bridge the gap would be undercut.","tokens_in":19345,"feed_emoji":"🩺","tokens_out":4799,"duration_ms":52701,"temperature":0.7,"pith_summary":"This Perspective argues that medical AI fails to reach patients not because predictive models are inaccurate, but because the field builds technology-centric systems that compete with, rather than support, how doctors actually reason and decide. The authors propose a sociotechnical reframing in which data-driven tools are designed to augment clinicians' cognitive and epistemic activities — reasoning under uncertainty, mental simulation, and belief updating — and to fit into the existing workflow, institutions, and responsibilities of care. They use paediatric sepsis as an illustrative case because it combines ambiguous definitions, high-stakes time pressure, and a heterogeneous population, exposing the mismatch between benchmark-driven modelling and bedside needs. If the paper is right, the path to clinical impact runs through human-centred design of reasoning support rather than through further accuracy gains on benchmark tasks.","feed_headline":"AI should support clinical reasoning, not chase benchmarks","feed_subtitle":"The translational gap closes when AI augments doctors' judgment instead of replacing it.","key_machinery":"The central object is a sociotechnical conceptualisation of AI as a reasoning-support system rather than an autonomous decision maker, organised around three modes of integration (autonomy, assistance with a human-in-the-loop, and augmentation with a machine-in-the-loop) and selected per cognitive and epistemic activity according to a task's automation readiness level. The load-bearing mechanisms are the cognitive-science techniques proposed for supporting clinicians: cognitive forcing (interrupting heuristic reasoning to force consideration of disconfirming evidence and alternative hypotheses), AI-assisted mental simulation (projecting alternative future patient trajectories conditioned on different values of missing or unknown variables), and continuous belief updating. The argument also imports the distinction between structured, semi-structured, and unstructured decisions, and between stable and wicked environments, to justify why end-to-end automation is usually the wrong fit for clinical diagnosis.","core_discovery":"The paper's central claim is normative: AI systems ought to seamlessly integrate into and augment established medical workflows and real-life reasoning and decision-making processes, rather than disrupt them. The authors posit that the prevailing technology-centric approach — optimising models for superhuman predictive performance on carefully chosen benchmarks — is the root cause of the translational gap, rendering such systems fundamentally incompatible with clinical practice. They argue that medical diagnostic reasoning is usually semi-structured, occurs in unstable (open) worlds with incomplete and uncertain information, and relies on cognitive functions that can be supported by AI, such as cognitive forcing to counteract anchoring and premature closure, prospective mental simulation of patient trajectories, and progressive belief updating, all within a systems ecology of institutions, protocols, and responsibilities. The intended consequence is that AI is judged by real-world impact and acceptability, as a reliable tool under human responsibility, rather than by benchmark scores.","pith_inferences":["If the causal claim is right, the research community's current reward structure — benchmark leaderboards and accuracy contests — is itself part of the translational problem, not merely a neutral evaluation tool.","A testable extension would be to build a cognitive-support sepsis tool that presents alternative trajectories and cognitive-forcing prompts, then compare clinician diagnostic accuracy and workflow adoption against a conventional predictive alert system in a prospective study.","The argument implies that regulatory and reimbursement pathways must reward cognitive support and workflow integration, since adoption is a property of the entire institutional ecosystem rather than of model design alone.","One can also read the paper as predicting that current human-in-the-loop systems that simply ask clinicians to accept or reject recommendations will underperform systems that actively scaffold the clinician's reasoning process."],"forward_implications":["If adopted, AI evaluation would shift from benchmark accuracy to measures of decision consistency, reduction of decision noise, and integration into clinical workflow.","AI tools designed for mental simulation would let clinicians explore patient trajectories under different assumptions, including explicitly handling missing data rather than imputing it away.","Ante-hoc interpretable models would be preferred over post-hoc explanations for high-stakes clinical support, because their behaviour is guaranteed to reflect the model's true operation.","The framework generalises beyond medicine to other high-stakes, semi-structured decision domains where replacing humans is undesirable.","For paediatric sepsis specifically, AI could be embedded in sepsis teams and huddles to aid detection, management, and treatment consistency, while reducing unnecessary antibiotic exposure."],"supporting_citations":[{"why":"Supplies the concept of clinical reasoning support systems, distinct from conventional decision support, on which the paper's central reframing is built.","marker":"[27]"},{"why":"Provides the cognitive forcing strategies that the paper adopts as a core mechanism for AI to improve clinical reasoning.","marker":"[48]"},{"why":"Defines conditions for intuitive expertise and distinguishes reliable from wicked environments, grounding the paper's argument about when automation is inappropriate.","marker":"[81]"},{"why":"Supports the claim that simple transparent models or heuristics can outperform complex data-driven systems in open-world decisions, justifying the preference for interpretable AI.","marker":"[119]"},{"why":"Supplies the sociotechnical approach to patient care information systems that underpins the paper's insistence on workflow integration.","marker":"[8]"},{"why":"Provides the case for ante-hoc interpretable models over black-box explainability in high-stakes decisions, which the paper endorses.","marker":"[18]"},{"why":"Documents the recent state of paediatric sepsis prediction technologies, which the paper uses as its representative case study of the translational gap.","marker":"[58]"},{"why":"Supports the view that AI is social and relational rather than autonomous, a premise of the paper's augmentation framework.","marker":"[33]"}],"fun_headline_variants":["AI should support clinical reasoning, not chase benchmarks","The translational gap closes when AI augments judgment","Real impact, not superhuman scores, for medical AI","AI that works with doctors' minds, not against them","From benchmark chaser to clinical partner: AI's fix"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim hinges on the premise that the translational gap is caused by technology-centric development and that reframing AI as cognitive support will, by itself, materially improve adoption and outcomes; the paper offers reasoning and examples but no controlled evidence for this causal chain.","fun_headline_variants_meta":{"raw":{"variants":["AI should support clinical reasoning, not chase benchmarks","The translational gap closes when AI augments judgment","Real impact, not superhuman scores, for medical AI","AI that works with doctors' minds, not against them","From benchmark chaser to clinical partner: AI's fix"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000649,"raw_usage":{"total_tokens":2892,"prompt_tokens":775,"completion_tokens":2117,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":391,"completion_tokens_details":{"reasoning_tokens":2040}},"tokens_in":391,"tokens_out":2117,"duration_ms":18842,"temperature":1.0,"reasoning_tokens":2040,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:27:26.300945+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective, controlled deployment comparing a cognitive-support system (e.g., one delivering mental simulation and cognitive-forcing prompts for suspected sepsis) against a conventional predictive alert tool, measuring clinician uptake, decision quality, and patient outcomes: if the conventional tool matches the cognitive-support tool on adoption and outcomes, the claim that the reframing is necessary to bridge the gap would be undercut.","supporting_citations":[{"cited_title":"& Klein, G","cited_arxiv_id":null,"evidence_quote":"Defines conditions for intuitive expertise and distinguishes reliable from wicked environments, grounding the paper's argument about when automation is inappropriate."},{"cited_title":"Psychological AI: Designing algorithms informed by human psychology.Perspectives on Psychological Science19, 839–848 (2023)","cited_arxiv_id":null,"evidence_quote":"Supports the claim that simple transparent models or heuristics can outperform complex data-driven systems in open-world decisions, justifying the preference for interpretable AI."}],"review_version":1}