REVIEW 3 major objections 4 minor 33 references
A retrieval-augmented LLM improved physicians' answers, but a citation that seemed to support a wrong answer sharply weakened their resistance.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-05 00:10 UTC pith:I6IXKOTH
load-bearing objection A preregistered, well-run reader study with a genuinely new reliance metric, but the safety claim outruns the observational support ratings; worth refereeing seriously. the 3 major comments →
Large language models improve physician accuracy but lead to false reliance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the core discovery is that physicians' perceived citation support predicts both beneficial and harmful reliance on LLM advice. Using 'at least one citation judged supportive' as the cue, the authors show that physicians who were initially wrong and saw a supportive citation reached the correct final answer in 76.9% of cases versus 34.0% without; but physicians who were initially right and saw what they took to be a supportive citation behind an incorrect LLM answer kept their correct answer in only 34.8% of cases versus 92.0% without. The same grounding signal therefore has a double edge: it identifies more accurate CORA outputs and enables error correction, and it
What carries the argument
CORA (Citation-Oriented Retrieval Assistant), an agentic retrieval-augmented generation system that retrieves passages from clinical guidelines, a dermatology textbook, and case reports, evaluates evidence sufficiency, reformulates queries when needed, reranks, and generates an answer labeled with cited source identifiers. The reader study's analytical engine is the binary per-source support rating: each physician judges whether each displayed citation contains enough information to support the LLM's answer, and the paper classifies an answer as supported when at least one citation is so judged. Stratifying reliance behavior—adoption of correct advice and resistance to incorrect advice—by th
Load-bearing premise
The central asymmetry rests on taking physicians' 'supports' ratings as a measure of the evidence-answer relationship independent of whether physicians already agree with the answer; support was rated, not experimentally manipulated, so the dissociation could partly reflect confirmation rather than evidence appraisal.
What would settle it
A randomized experiment that pairs correct and incorrect LLM answers with citations that genuinely support, merely match the topic, contradict, or are absent, and measures revision decisions. If physicians' resistance to incorrect advice does not fall when a topically relevant but non-establishing citation is attached compared with no citation, the grounding-miscalibration claim fails. Alternatively, if 'supportive' ratings no longer predict reliance when raters are blind to the LLM's answer, the result would be confirmation bias rather than citation grounding.
If this is right
- Aggregate post-assistance accuracy is an incomplete endpoint: a system can raise mean accuracy while making its residual errors harder to catch, so evaluations should report reliance separately for cases where clinician and model disagree and where the model is wrong.
- The gap between apparent and actual support becomes a design target: systems should flag weak retrieval, indicate whether a source proves the answer rather than merely matching its topic, and require an explicit check before an answer is revised.
- Because the same evidence-cue logic likely applies beyond citations—to saliency overlays, retrieved sources, and confidence scores—grounding miscalibration is probably a general risk for medical AI decision support, not a quirk of dermatology or of retrieval augmentation.
- Physicians in the assisted condition did not exceed CORA's own accuracy, reproducing a pattern seen in LLM assistance without citations; reliance stratification offers an explanation for that ceiling.
- Retrieval-grounding is widely treated as the principal safeguard against LLM error in clinical settings, but this result indicates the safeguard can backfire: the cue meant to enable scrutiny can substitute for it.
Where Pith is reading between the lines
- The causal reading is not yet established: citation support was rated, not experimentally manipulated, so 'supportive' ratings may partly track agreement with CORA's answer. A randomized crossing of answer correctness and citation quality is the direct test.
- If the asymmetry is causal, a display of a topically relevant but non-establishing citation is potentially worse than no citation at all, because it converts an initially correct physician answer into an incorrect one.
- A practical monitoring metric follows: track how often physicians revise an initially correct answer to a wrong LLM recommendation, stratify by whether displayed evidence is judged supportive, and treat a rising rate as a safety signal.
- The mechanism could be probed by having physicians rate citation support while blind to the LLM's answer, or by presenting the same source text worded to support versus contradict the answer; the paper's design reveals the pattern but not the cognitive route.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents CORA, an agentic retrieval-augmented LLM for dermatology, and evaluates it in two stages: a benchmark comparison against five base LLMs on a compiled multiple-choice set (DermBenchQA) and on a contamination-resistant open-ended set derived from post-2025 case reports (DermCaseQA), followed by a preregistered within-subjects reader study with 46 physicians and 736 paired decisions. CORA was non-inferior to baselines on matched items and improved physician accuracy from 70.8% to 82.6%. The central safety claim is that perceived citation support predicts both beneficial reliance on correct advice (RAIR rising from 34.0% to 76.9%) and harmful deference to incorrect advice (RSR falling from 92.0% to 34.8%), a pattern the authors term 'grounding miscalibration.' The Discussion explicitly limits the conclusion to an association—'physicians' judgement that a citation supported the answer predicted whether they deferred to it'—but the Abstract and Results use stronger causal language.
Significance. If the central claim holds, the paper makes an important contribution by moving beyond aggregate accuracy to reliance-stratified evaluation of retrieval-augmented clinical LLMs. The study has notable strengths: preregistration on OSF, physician-clustered bootstrap confidence intervals, openly available data and code, a contamination-resistant evaluation set with careful date-gating, and transparent reporting of small strata and limitations. The proposed construct of 'grounding miscalibration' is potentially valuable for the design and regulation of clinical decision support. However, the headline safety conclusion rests on a non-manipulated, self-reported measure of citation support, so the causal interpretation is not yet secured. The manuscript itself concedes this in the Discussion, which should be reconciled with the abstract's framing.
major comments (3)
- [Abstract and Results ('Perceived citation support...') vs Methods (Reader study)] The abstract states that 'citations created an important asymmetry,' and the Results state that perceived support 'may also create a key safety risk.' But citation support was rated, not experimentally manipulated, and the rating was collected after the CORA answer was revealed. A physician who finds the answer plausible may rate its citations as supportive and also adopt it; 'support' may partly proxy for agreement or confirmation bias. The Discussion narrows the claim to 'predicted whether they deferred,' but the causal framing remains in the abstract and the main text. This is load-bearing for the grounding-miscalibration concept. Please either replace causal language with associational language throughout or add an analysis that addresses the endogeneity (e.g., restricting to ratings by physicians whose unaided answer disagreed with CORA, or using an objectively coded support measure
- [Fig. 3f and Results paragraph 'Perceived citation support...'] The harmful-deference stratum is small: RSR with support is 8/23 (34.8%; 95% CI 17.2, 52.6%) versus 23/25 (92.0%) without support, a total of 48 decisions and only 23 supported-incorrect cases. The 57.2 pp decrease is driven by 15 abandonment events. With denominators this small and responses clustered within 46 physicians, the cluster bootstrap cannot fully convey the fragility of the estimate. Report the number of physicians contributing to the 23 decisions, an exact confidence interval for the 8/23 proportion, and a sensitivity analysis such as leave-one-physician-out or a question-level mixed-effects model. The current presentation may overstate the precision of a potentially important safety signal.
- [Methods (Reader study) and Discussion (Limitations)] The design lacks a citation-free or non-retrieval condition, as acknowledged, but the specific Fig. 3f comparison is also confounded by question-level properties. Supported versus unsupported incorrect recommendations may differ in how obviously wrong the answer is, how well the cited text matches the topic, or the difficulty of the underlying question; any of these could produce the observed association without a causal role for perceived support. The manuscript should explicitly acknowledge this residual confounding and, if possible, stratify by an objective measure of citation support (e.g., an entailment-based NLP score) or by question difficulty. This is necessary to support the 'grounding miscalibration' mechanism rather than an alternative explanation.
minor comments (4)
- [Main text, first paragraph of Results] Typo: 'Ag entic RAG' should be 'Agentic RAG.'
- [Fig. 2 and Methods (Benchmark evaluation)] The non-inferiority margin of 1 pp and the Wald test are described, but the reported Holm–Bonferroni corrected P-values are given only as inequalities. Please report the exact corrected P-values or state them in a supplementary table for reproducibility.
- [Methods (Reader study, participants)] Participants were instructed not to consult other AI tools or external resources, but this could not be technically verified in a remote design. The manuscript notes the instruction but should also note explicitly that noncompliance is possible and would affect the assisted-condition estimates.
- [Extended Data Fig. 2] Inter-rater agreement per citation is described as low-moderate. Since question-level support is derived by majority vote from three physicians per question, please report the distribution of majority margins (2-1 vs 3-0) and whether the main results are robust to requiring unanimous support.
Circularity Check
No circularity: reliance metrics are conditional definitions and citation support is a measured variable, not a fitted or derived outcome.
full rationale
The paper's central quantities are empirical measurements, not reductions to inputs. RAIR and RSR are conditionally defined rates: 'Among decisions where physicians were initially incorrect and CORA was correct, relative AI reliance (RAIR) is defined as the proportion they revised to CORA's answer. Among decisions where physicians were initially correct but CORA was incorrect, relative self-reliance (RSR) is defined as the proportion who retained their own correct answer.' These definitions separate reliance from baseline accuracy by restricting denominators to cases where reliance changes the outcome; they do not make the reported RAIR/RSR values equal to overall accuracy or to each other by construction. Citation support is likewise an independently measured rating ('Does this cited text contain enough information that supports the LLM's answer?'), collected before the final decision but not defined in terms of it. The paper explicitly acknowledges the observational nature of the support measure and narrows its conclusion: 'Because incorrect recommendations were uncommon and citation support was rated rather than experimentally manipulated, the harmful-reliance estimates are imprecise and observational' and 'our findings support a narrower conclusion: physicians' judgement that a citation supported the answer predicted whether they deferred to it.' The appropriate-reliance framework is cited from Schemmer et al., an external source, and is not used to smuggle in an ansatz. There is no self-citation chain, no uniqueness theorem imported from the authors, and no fitted parameter relabeled as a prediction. The only potential concern—shared-method variance between support ratings and final decisions—is a measurement/validity limitation, not circular reasoning, and the paper flags it explicitly. The derivation of the main claims is therefore self-contained relative to its empirical inputs.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Physician citation-support ratings measure the evidential relationship between the cited text and the LLM answer, rather than merely agreement with the answer.
- domain assumption The date-gate correctly identifies the earliest public exposure of each DermCaseQA source, so the post-cutoff set is contamination-resistant.
- domain assumption Automated LLM judge (Claude Sonnet 4.6) accurately scores open-ended answers against reference diagnoses.
- domain assumption The retrieval-sufficiency criterion yields a subset representative enough for the benchmark claims.
Cite this review
Pith. "Pith review of Large language models improve physician accuracy but lead to false reliance." pith.science (2026). https://pith.science/paper/I6IXKOTH
@misc{pith2026260800817,
author = {Pith},
title = {Pith review of: Large language models improve physician accuracy but lead to false reliance},
year = {2026},
howpublished = {\url{https://pith.science/paper/I6IXKOTH}},
note = {Machine review of arXiv:2608.00817}
}
read the original abstract
Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.
Reference graph
Works this paper leans on
-
[1]
Singhal, K. et al. Large language models encode clinical knowledge. Nature 620 , 172–180 (2023)
work page 2023
-
[2]
Thirunavukarasu, A. J. et al. Large language models in medicine | Nature Medicine. Nat. Med. 29 , 1930–1940 (2023)
work page 1930
-
[3]
Shool, S. et al. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med. Inform. Decis. Mak. 25 , 117 (2025)
work page 2025
-
[4]
Asgari, E. et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. Npj Digit. Med. 8 , 274 (2025)
work page 2025
-
[5]
Chen, Q. et al. Benchmarking large language models for biomedical natural language processing applications and recommendations. Nat. Commun. 16 , 3280 (2025)
work page 2025
-
[6]
Ferber, D. et al. GPT-4 for Information Retrieval and Comparison of Medical Oncology Guidelines. NEJM AI 1 , AIcs2300235 (2024)
work page 2024
-
[7]
Liu, S., McCoy, A. B. & Wright, A. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines. J. Am. Med. Inform. Assoc. 32 , 605–615 (2025)
work page 2025
-
[8]
Kresevic, S. et al. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. Npj Digit. Med. 7 , 102 (2024)
work page 2024
-
[9]
Asai, A., Wu, Z., Wang, Y., Sil, A. & Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. Int. Conf. Learn. Represent. 2024 , 9112–9141 (2024)
work page 2024
-
[10]
Xia, Y., Zhou, J., Shi, Z., Chen, J. & Huang, H. Improving retrieval augmented language model with self-reasoning. in Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Symposium on Educational Advances in Artificial Intelligence vol. 39 ...
work page 2025
-
[11]
Wu, K. et al. An automated framework for assessing how well LLMs cite relevant medical references. Nat. Commun. 16 , 3615 (2025)
work page 2025
-
[12]
Goh, E. et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw. Open 7 , e2440969 (2024)
work page 2024
-
[13]
Goh, E. et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat. Med. 31 , 1233–1238 (2025)
work page 2025
-
[14]
Ong, J. C. L. et al. Large language model as clinical decision support system augments medication safety in 16 clinical specialties. Cell Rep. Med. 6 , (2025)
work page 2025
-
[15]
Lammert, J. et al. Expert-Guided Large Language Models for Clinical Decision Support in Precision Oncology. JCO Precis. Oncol. e2400478 (2024) doi:10.1200/PO-24-00478
-
[16]
Hetz, M. J. et al. Superhuman performance on urology board questions using an explainable language model enhanced with European Association of Urology guidelines. ESMO Real World Data Digit. Oncol. 6 , (2024)
work page 2024
-
[17]
Zhai, G. et al. AI for evidence-based treatment recommendation in oncology: a blinded evaluation of large language models and agentic workflows. Front. Artif. Intell. 8 , (2025)
work page 2025
-
[18]
Zakka, C. et al. Almanac — Retrieval-Augmented Language Models for Clinical Medicine. NEJM AI 1 , AIoa2300068 (2024)
work page 2024
-
[19]
Tayebi Arasteh, S. et al. RadioRAG: Online Retrieval–Augmented Generation for Radiology Question Answering. Radiol. Artif. Intell. 7 , e240476 (2025)
work page 2025
-
[20]
Schemmer, M., Kuehl, N., Benz, C., Bartos, A. & Satzger, G. Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. in Proceedings of the 28th International Conference on Intelligent User Interfaces 410–422 (Association for Computing Machinery, New York, NY, USA, 2023). doi:10.1145/3581641.3584066
arXiv 2023
-
[21]
Goddard, K., Roudsari, A. & Wyatt, J. C. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J. Am. Med. Inform. Assoc. 19 , 121–127 (2012)
work page 2012
-
[22]
Gaube, S. et al. Do as AI say: susceptibility in deployment of clinical decision-aids. Npj Digit. Med. 4 , 31 (2021)
work page 2021
-
[23]
Qazi, I. A. et al. Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial. NEJM AI 3 , AIoa2501001 (2026)
work page 2026
-
[24]
Buçinca, Z., Malaya, M. B. & Gajos, K. Z. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proc. ACM Hum.-Comput. Interact. 5 , 188:1-188:21 (2021)
work page 2021
-
[25]
Bansal, G. et al. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. in Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems 1–16 (Association for Computing Machinery, New York, NY, USA, 2021). doi:10.1145/3411764.3445717
arXiv 2021
-
[26]
Prinster, D. et al. Care to Explain? AI Explanation Types Differentially Impact Chest Radiograph Diagnostic Performance and Physician Trust in AI. Radiology 313 , e233261 (2024)
work page 2024
-
[27]
Ding, Y. et al. Citations and Trust in LLM Generated Responses. Proc. AAAI Conf. Artif. Intell. 39 , 23787–23795 (2025)
work page 2025
-
[28]
Griot, M., Hemptinne, C., Vanderdonckt, J. & Yuksel, D. Large Language Models lack essential metacognition for reliable medical reasoning. Nat. Commun. 16 , 642 (2025)
work page 2025
-
[29]
Bedi, S. et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA 333 , 319–328 (2025)
work page 2025
-
[30]
Carl, N. et al. Enhancing clinicians’ trust in large language models via transparent source attribution: A randomized controlled evaluation in uro-oncology. Eur. J. Cancer 233 , (2026)
work page 2026
-
[31]
Pal, A., Umapathi, L. K. & Sankarasubbu, M. MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering. in Proceedings of the Conference on Health, Inference, and Learning 248–260 (PMLR, 2022)
work page 2022
-
[32]
Jin, D. et al. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Appl. Sci. 11 , 6421 (2021)
work page 2021
-
[33]
Hendrycks, D. et al. Measuring Massive Multitask Language Understanding. in (2020). Extended Data Extended Data Figure 1: Accuracy by experience levels Per-physician accuracy gains with LLM assistance by experience level. Each line connects one physician’s unaided (left) and CORA-assisted (right). Mean gains (pp) are shown above each stratum. Improvement ...
work page 2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.