REVIEW 2 major objections 6 minor 29 references
No current large language model has demonstrated safe autonomous triage of undifferentiated patients.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Passing medical exams does not make LLMs safe for autonomous triage: they fail to seek missing red flags, and current benchmarks do not test that.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection A well-argued evidence-gap perspective with a useful new frame; the causal training-objective claim is the soft underbelly, but the safety case doesn't rest on it. the 2 major comments →
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that the safety gap for autonomous triage is not medical knowledge but an objective mismatch: an LLM optimized to predict the most probable next token and to be helpful has no native representation of the cost of missing a rare, dangerous diagnosis. Under incomplete patient histories, models fail at the behaviors safe triage requires—widening the differential, seeking the missing red flag, lowering the threshold for escalation, and deferring when information is insufficient. Because existing evaluations use complete, expertly curated cases and confidence-gated populations that exclude the ambiguous patient, they cannot detect this failure. The paper proposes that
What carries the argument
The central mechanism is 'reasoning from silence': making safe inferences from diagnostically informative missing information—recognizing that what a patient has not said may reflect incomplete elicitation rather than true absence of disease. The paper uses this concept to explain why next-token-prediction training fails triage: the model conditions on what is present, not on what is absent, so absence exerts no force on confidence. The other load-bearing piece is the proposed evaluation design: withhold information, score the right objective with asymmetric cost, and validate the simulator in three layers (face, distributional, predictive).
Load-bearing premise
The argument depends on the premise that the training objective—predicting the most probable next token and being helpfully agreeable—is what causes models to ignore missing red flags, so that no amount of prompting or wrapping can fully fix the deficit.
What would settle it
A prospective randomized trial in a real emergency or primary-care setting where an LLM triages undifferentiated, incompletely disclosing patients, with the model required to ask for missing details and escalate high-harm cases; if the model's under-triage rate of must-not-miss diagnoses is comparable to or better than clinicians' on the same encounters, the central 'evidence does not exist' claim would be overturned for that system. A cheaper check: measure whether a model's differential width widens and its confidence falls as a fixed history is systematically reduced from complete to 20% co
If this is right
- If correct, exam-passing and curated-benchmark accuracy cannot be used as evidence of readiness for autonomous triage, no matter how high the scores.
- Deploying LLM symptom checkers or autonomous primary-care systems without evidence from incomplete-information, harm-weighted evaluations risks confident, plausible answers that close the case before a must-not-miss diagnosis is excluded.
- Assistant-like behaviors—credulity, agreeableness, miscalibration—compound the core deficit, so evaluations must test resistance to minimized and adversarial patient histories, not just accuracy.
- The paper's proposed standard—withhold information, score asymmetric cost, validate simulators—offers a concrete precondition that any autonomous triage claim must meet before it could justify deployment.
Where Pith is reading between the lines
- If the objective-mismatch account is right, then scale alone will not close the gap; the field may need a different training signal that rewards information seeking and harm-weighted deferral, not just helpfulness.
- The same reasoning-from-silence logic likely applies beyond triage—to any high-stakes LLM decision under incomplete inputs, where the unstated detail is the one that matters.
- A testable extension: build a benchmark that systematically varies information completeness and measures whether a model's differential width and escalation threshold respond to missingness; the paper predicts they will not.
- The paper leaves open whether structured elicitation policies layered on top of current models can compensate; that is testable and is not settled by the paper's argument.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This Perspective argues that no current evidence demonstrates that LLMs can safely perform autonomous triage of self-presenting, undifferentiated patients without a clinician in the loop. The paper distinguishes knowledge-oriented benchmarks (licensing exams, curated vignettes) from the real task of sequential decision-making under incomplete histories and asymmetric error costs, and argues that existing evaluations—including those of AMIE and autonomous primary-care systems—pre-filter patients, supply complete histories, or confidence-gate outputs, so they cannot detect failures of 'reasoning from silence.' It also identifies a cluster of assistant-like behaviors (credulity, agreeableness, miscalibration) that compound the deficit, and proposes three evaluation requirements: withhold information by design, score harm-weighted behaviors rather than top-k accuracy, and validate synthetic simulators through face, distributional, and predictive validity.
Significance. The evidence-gap claim is well supported: the paper marshals a randomized study showing a drop from 94.9% to <34.5% condition identification when real users provide histories [2], and two systematic reviews showing that only a small fraction of LLM evaluations use real-world prospective data [15,16]. If accepted, the paper provides a timely caution against premature deployment and a concrete, operationalizable evaluation agenda. Although it is a perspective rather than a new empirical study, its strengths include a falsifiable central claim, explicit positive proposals, and a clear breakdown of how specific design choices (pre-screening, confidence gating, complete histories) hide tail risks. The mechanistic explanation of the deficit is the least supported part, but it is not required for the central claim.
major comments (2)
- [Sections 1 and 2] The paper presents a causal claim as established fact: 'a model trained to continue the most probable text has no native representation of the danger it was never told about' (§2), and 'the deeper problem is what the substrate was never built to do' (§1). The cited evidence ([11], [12]) demonstrates that current models behave this way, but it does not show that the training objective itself prevents mitigation; §5 acknowledges that structured elicitation policies exist (ref. 24). Since the central evidence-gap claim does not depend on this mechanistic explanation, please frame it as a hypothesis ('we hypothesize that...') rather than as a settled property of LLMs.
- [Declarations] 'All authors declare no conflict of interest' appears inconsistent with the author affiliations listing Atman Labs, a commercial AI company (author block, lines 3-4). Employment, equity, or funding from Atman Labs is a financial interest highly relevant to a Perspective arguing for stricter evidence before autonomous triage deployment. This must be disclosed or explicitly rebutted before publication.
minor comments (6)
- [Section 1] The sentence reporting '94.9% ... but fewer than 34.5%' could clarify that both percentages refer to condition-identification rates, and that the first is on curated inputs while the second is on real-user dialogue; the current phrasing is slightly ambiguous.
- [Section 2] Reference [11] is a medRxiv preprint; the 1,000-consultation stress-test and the 97.5% accuracy figure rest on it. Please mark it as a preprint and temper 'a large stress-test confirms' to reflect the provisional status.
- [Section 5] 'This journal's editors have put it more plainly still' is ambiguous because the manuscript does not identify the target journal; specify that ref. [29] is a Nature Medicine editorial.
- [Section 5] 'Even the most advanced dynamic simulators still furnish the clinical facts' is unclear; specify that the simulator provides the clinical facts to the model rather than requiring the model to elicit them.
- [Declarations] The data availability statement claims a dataset 'used for this review' that will be shared; as this is a perspective with no original dataset, remove or clarify this statement.
- [Section 1] Consider defining 'autonomous triage' explicitly (degree of clinician backstop, setting, and population) at first use, since the paper discusses consumer symptom checkers, autonomous primary care systems, and ED agents, which differ in autonomy and risk.
Circularity Check
No significant circularity: the safety-evidence-gap claim rests on external studies and is not derived from self-referential definitions, fitted parameters, or load-bearing self-citations.
full rationale
This is a Perspective/argumentative review, not a formal derivation. The central claim is that evidence of safety for autonomous LLM triage does not yet exist, and that claim is supported by external empirical studies and systematic reviews (e.g., refs 2, 6–10, 15, 16), not by equations or fitted parameters. The mechanistic assertion about next-token prediction is explanatory, not a logical consequence of the evidence-gap claim, and the paper explicitly acknowledges that elicitation policies exist (refs 23–24) while arguing that they require proper evaluation. The only self-citation is ref [5] (Hillis & Payne), which is used for the peripheral point that human involvement should be deliberately specified along an autonomy spectrum; it does not support the central safety-evidence claim. No step in the paper reduces a prediction to an input by construction. Therefore no circularity is present.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption An LLM trained for next-token continuation and helpfulness does not, by default, optimize for safe action under asymmetric cost or actively seek missing information.
- domain assumption In autonomous triage, the decisive red flag is often not volunteered by the patient and must be actively elicited.
- domain assumption Existing evaluation studies are representative enough to conclude that curated benchmark performance does not transfer to incomplete-history triage.
- ad hoc to paper Simulator validity is established by face, distributional, and predictive validity.
Cite this review
Pith. "Pith review of Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support." pith.science (2026). https://pith.science/paper/E6R4RZBV
@misc{pith2026260728677,
author = {Pith},
title = {Pith review of: Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support},
year = {2026},
howpublished = {\url{https://pith.science/paper/E6R4RZBV}},
note = {Machine review of arXiv:2607.28677}
}
read the original abstract
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.
Figures
Reference graph
Works this paper leans on
-
[1]
Performance of a large language model on the reasoning tasks of a physician
Brodeur PG, Buckley TA, Kanjee Z, Goh E, Ling EB, Jain P, et al. Performance of a large language model on the reasoning tasks of a physician. Science. 2026;392(6797):524-7
2026
-
[2]
Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
Bean AM, Payne RE, Parsons G, Kirk HR, Ciro J, Mosquera-Gómez R, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat Med. 2026;32(2):609-15
2026
-
[3]
Measuring what Matters: Construct Validity in Large Language Model Benchmarks2025 November 3, 2025
Bean AM, Kearns RO, Romanou A, Hafner FS, Mayne H, Batzner J, et al. Measuring what Matters: Construct Validity in Large Language Model Benchmarks2025 November 3, 2025
2025
-
[4]
On the robustness of medical term representations in locally deployable language models
Auger SD, Graham NSN, Scott G. On the robustness of medical term representations in locally deployable language models. medRxiv. 2026
2026
-
[5]
Health AI needs meaningful human involvement: lessons from war
Hillis JM, Payne K. Health AI needs meaningful human involvement: lessons from war. Nat Med. 2024;30(12):3397-8
2024
-
[6]
From Concept to Clinic: Real World Evidence for Autonomous AI Deployment in Primary Care Telemedicine
Saenz A, Schumacher E, Naik D, Khosla N, Kannan A. From Concept to Clinic: Real World Evidence for Autonomous AI Deployment in Primary Care Telemedicine. medRxiv. 2026:2026.03.18.26348749
2026
-
[7]
Towards autonomous medical artificial intelligence agents
Ferber D, Hilgers L, Höper C, Kinny-Köster B, Eckardt JN, Egger-Heidrich K, et al. Towards autonomous medical artificial intelligence agents. Nature. 2026. 12
2026
-
[8]
Brodeur P, Koshy JM, Palepu A, Saab K, Homiar A, Ruparel R, et al. A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic. arXiv. 2026;2603.08448v3
arXiv 2026
-
[9]
Safety of a large language model-based clinical decision support system in African primary healthcare
Agweyu A, Mwaniki P, Musau W, Korom R, Isaaka L, Wanyama C, et al. Safety of a large language model-based clinical decision support system in African primary healthcare. Nat Health. 2026;1(6):607-18
2026
-
[10]
Korom R, Kiptinness S, Adan N, Said K, Ithuli C, Rotich O, et al. AI-based Clinical Decision Support for Primary Care: A Real-World Study2025 22 Jul 2025; (arXiv:2507.16947v1 [cs.CL]). Available from: https://arxiv.org/abs/2507.16947
Pith/arXiv arXiv 2025
-
[11]
Medical errors in large language models revealed using 1,000 synthetic clinical transcripts
Auger SD, Scott G. Medical errors in large language models revealed using 1,000 synthetic clinical transcripts. medRxiv. 2026:2026.03.23.26349082
2026
-
[12]
Large Language Models lack essential metacognition for reliable medical reasoning
Griot M, Hemptinne C, Vanderdonckt J, Yuksel D. Large Language Models lack essential metacognition for reliable medical reasoning. Nat Commun. 2025;16(1):642
2025
-
[13]
Towards conversational diagnostic artificial intelligence
Tu T, Schaekermann M, Palepu A, Saab K, Freyberg J, Tanno R, et al. Towards conversational diagnostic artificial intelligence. Nature. 2025;642(8067):442-50
2025
-
[14]
Towards Conversational AI for Disease Management
Liévin V, Palepu A, Weng WH, Saab K, Stutz D, Cheng Y, et al. Towards Conversational AI for Disease Management. Nature. 2026
2026
-
[15]
Testing and Evalua- tion of Health Care Applications of Large Language Models: A Systematic Review
Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and Evalua- tion of Health Care Applications of Large Language Models: A Systematic Review. Jama. 2025;333(4):319-28
2025
-
[16]
LLM-assisted systematic review of large language models in clinical medicine
Chen SF, Alyakin A, Seas A, Yang E, Choi JJ, Lee JV, et al. LLM-assisted systematic review of large language models in clinical medicine. Nat Med. 2026;32(3):1152-9
2026
-
[17]
Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support
Omar M, Sorin V, Collins JD, Reich D, Freeman R, Gavin N, et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun Med (Lond). 2025;5(1):330
2025
-
[18]
MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Zhou H, Zou X, Wu J, Wu S, Wu J, Segal BM, et al. MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context. bioRxiv. 2026:2026.05.25.727671
2026
-
[19]
Training language models to be warm can reduce accuracy and increase sycophancy
Ibrahim L, Hafner FS, Rocher L. Training language models to be warm can reduce accuracy and increase sycophancy. Nature. 2026;652(8112):1159-65
2026
-
[20]
Competing Biases underlie Overconfidence and Underconfidence in LLMs
Kumaran D, Fleming SM, Markeeva L, Heyward J, Banino A, Mathur M, et al. Competing Biases underlie Overconfidence and Underconfidence in LLMs. Nature Machine Intelligence. 2026;8(4):614-27
2026
-
[21]
Assessment of Large Language Models in Clinical Reasoning: A Novel Benchmarking Study
McCoy LG, Swamy R, Sagar N, Wang M, Bacchi S, Fong JMN, et al. Assessment of Large Language Models in Clinical Reasoning: A Novel Benchmarking Study. NEJM AI. 2025;2(10):AIdbp2500120
2025
-
[22]
First, do NOHARM: towards clinically safe large language models
Wu D, Haredasht FN, Maharaj SK, Jain P, Tran J, Gwiazdon M, et al. First, do NOHARM: towards clinically safe large language models. ArXiv. 2025
2025
-
[23]
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
Li SS, Balachandran V, Feng S, Ilgen JS, Pierson E, Koh PW, et al. MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning. 2024
2024
-
[24]
SymptomAI: Toward a Con- versational AI Agent for Everyday Symptom Assessment
Breda J, Yousif F, Hawkins B, Cotoi M, Liu M, Luo R, et al. SymptomAI: Toward a Con- versational AI Agent for Everyday Symptom Assessment. arXiv preprint arXiv:260504012. 2026
2026
-
[25]
A POMDP Formulation of Preference Elicitation Problems
Boutilier C. A POMDP Formulation of Preference Elicitation Problems. In: Dechter R, Kearns MJ, Sutton RS, editors. Proceedings of the Eighteenth National Conference on Artificial 13 Intelligence and Fourteenth Conference on Innovative Applications of Artificial Intelligence; July 28-August 1, 2002; Edmonton, Alberta, Canada: AAAI Press; 2002. p. 239-46
2002
-
[26]
A clinical environment simulator for dynamic AI evaluation
Luo L, Kim SE, Zhang X, Kernbach JM, Kenia R, Acosta JN, et al. A clinical environment simulator for dynamic AI evaluation. Nat Med. 2026;32(3):820-7
2026
-
[27]
Mayne H, Kearns RO, Yang Y, Bean AM, Delaney E, Russell C, et al. LLMs Don’t Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations2025 2025-09-11. Available from: https://arxiv.org/abs/2509.09396
Pith/arXiv arXiv 2025
-
[28]
AI, Health, and Health Care Today and Tomorrow: The JAMA Summit Report on Artificial Intelligence
Angus DC, Khera R, Lieu T, Liu V, Ahmad FS, Anderson B, et al. AI, Health, and Health Care Today and Tomorrow: The JAMA Summit Report on Artificial Intelligence. JAMA. 2025;334(18):1650-64
2025
-
[29]
Nature Medicine
Show us the evidence for the value of medical AI. Nature Medicine. 2026;32(4):1163-. 14
2026
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.