REVIEW 3 major objections 5 minor 44 references
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper introduces a 750-case EHR benchmark showing that correct verdicts can mask indefensible investigations, with defect-free accuracy 4.8-14.8 points lower than raw accuracy.
desk verdict A genuinely useful benchmark with a real label-construction soft spot; worth serious refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the process rubric: per-scenario MUST-DO and MUST-NOT criteria, weighted by clinician authors, graded by an LLM judge against the final report together with the logged tool-call trajectory. From this rubric the benchmark derives defect-free accuracy, which credits a run only when the verdict is correct and no MUST-NOT criterion is sprung. The complementary mechanism is the four-way verdict space, whose two Indeterminate classes - Lack of Data versus Medically Ambiguous - make calibrated abstention a first-class, scorable outcome rather than a post hoc confidence threshold. Policy grounding is enforced by a fixed corpus of governing documents with line-level provenance, so policy citations can be checked for resolution and support by rule or by judge. Together these let the benchmark separate outcome from process: a correct label, a sound investigation, and a defensible abstention are each measured separately.
What would settle it
Independently re-adjudicate all 750 cases with blinded clinicians, or a large stratified sample such as 200 cases, and compare their verdicts to the reference labels; if clinician agreement on the Indeterminate classes dropped well below the 91% seen on the 75-case calibration sample, or if clinicians assigned Yes or No to a substantial share of cases the ensemble called Indeterminate, the under-abstention finding would be an artifact of the reference labels rather than a property of the systems.
Extended reading notes
Core claim
The central claim is that raw verdict accuracy systematically overstates the quality of agentic clinical investigation over EHRs. Using a benchmark built from real patient records and clinician-validated adjudication criteria, the paper shows that when a verdict is correct but the investigation uses a prohibited shortcut - for example, asserting a condition is excluded without the test result that would settle it - the run is defective. Defect-free accuracy, which requires both a correct verdict and no MUST-NOT violation, falls 4.8-14.8 points below accuracy on the same runs, and this correction changes the system ordering. In addition, all sixteen evaluated systems over-commit: on cases where the reference verdict is Indeterminate, they return Yes or No at rates (21.3-55.3%) far above their over-abstention rates on definitive cases (7.3-16.5%). The benchmark is offered as a deployment-oriented standard for measuring whether a clinical agent can retrieve needed evidence, ground claims in it, apply the governing policy, and defer when the record cannot support a unique conclusion.
Load-bearing premise
The headline findings assume the 750 reference verdicts are correct, but only 75 cases plus 25 double-reviewed cases were checked by the Clinical Board; the remaining roughly 650 labels come from a three-model adjudication ensemble without direct clinician verification, so if those model-assisted labels are systematically wrong - especially on the two Indeterminate classes - the measured under-abstention and the defect-free gap would be artifacts of labeling rather than properties of the systems.
Editorial extensions
If this is right
- No evaluated system comes close to autonomous deployment: the best resolves only 76.1% of the 750 cases, so human-supervised review remains necessary for retrospective clinical audit.
- A correct verdict is not a reliable signal of a sound investigation: up to one in five correct verdicts for the most affected system violates a prohibited shortcut, so quality controls should audit the trace, not just the label.
- The under-abstention asymmetry is systematic: all sixteen systems over-commit on deferral cases more than they over-abstain on definitive ones, so interventions should push agents toward deferral when evidence is missing or ambiguous.
- Repeated attempts do not fix this: three-attempt reliability shows systems are mostly stably wrong rather than randomly wrong, so running a query again will not surface the error.
- Policy omission is the norm: every system cites genuinely governing documents but omits roughly half of the other governing documents, meaning verdicts can be correct for the wrong reasons.
Reading between the lines
- If the defect-free gap is driven by a small set of recurring shortcuts, then a targeted intervention - such as requiring confirmation of an order time against a radiology report or an ascitic-fluid result before ruling out peritonitis - could raise defect-free accuracy without changing raw accuracy; this is testable by adding such requirements to a harness.
- The two Indeterminate classes may index different failure mechanisms: over-committing on Lack of Data cases suggests retrieval or coverage failures, while over-committing on Medically Ambiguous cases suggests criterion-application failures; separating them could guide where to invest in an agent.
- Because the case mix is intentionally stratified, the accuracy numbers are not prevalence estimates; a deployment cohort with a different base rate of indeterminate cases would likely show different accuracy and different over-commitment rates.
- The same evaluation recipe - governed logged tools, explicit adjudication criteria, traceable evidence, and scored deferral - transfers to other audit settings such as financial or legal review, where the defensibility of a decision matters as much as its conclusion; a concrete port would require instantiating a policy corpus and process rubric for that domain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces CliniCARE-Bench, a benchmark for retrospective clinical audit over real longitudinal EHR data. It consists of 25 clinician-validated scenarios instantiated as 750 patient-specific cases from MIMIC-IV, with each case requiring one of four verdicts: Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous. Systems interact with a governed, logged tool environment for record retrieval, computation, and policy access, and are scored not only on verdict accuracy but also on defect-free accuracy, process adherence, evidence and policy grounding, abstention behavior, reliability, and resource use. Reference verdicts are produced by an ensemble of three frontier model harnesses and calibrated against a Clinical Board review on a stratified subset of 100 cases. The paper evaluates 16 agentic systems, reporting four-way accuracy between 65.3% and 76.1%, a defect-free accuracy that is 4.8–14.8 percentage points lower and reorders the leaderboard, and a universal over-commitment pattern in which every system issues definitive verdicts on reference-abstention cases at a higher rate than it abstains on reference-definitive cases.
Significance. If the benchmark and its reference labels are reliable, this is a valuable and timely contribution. It moves clinical agent evaluation from static knowledge tests and simulated encounters to replayable, process-aware audit over real longitudinal records, and it jointly scores investigation quality, grounding, policy use, and calibrated abstention. The paper is unusually thorough in its validation efforts: the Clinical Board calibration, the judge-reliability panels, the repeated-run analyses, the policy-corpus ablation, and the test-time-compute comparison are all reported with appropriate caveats. The headline dissociations, particularly the gap between raw accuracy and defect-free accuracy and the systematic under-abstention, are clinically important and, if correct, would justify the benchmark's deployment-oriented framing. The main risk to these claims is the validity of the reference verdicts for the majority of cases that were not directly clinician-reviewed, together with the self-referential role of GPT-5.5 as ensemble member, evaluated system, and judge; these issues are acknowledged in the paper but are not fully resolved.
major comments (3)
- [Sections 4.4, 4.5, and 6.2 (Tables 3 and Figure 6)] The reference verdicts for 650 of 750 cases are not directly clinician-reviewed; only 75 cases were used for verdict-level calibration, with an additional 25 double-reviewed. On the 2-vs-1 split-vote cases, clinician agreement with the reference was 18/24 (75%), and the paper states that every disagreement resolved to the dissenting model rather than an outside label. Because the over-commitment rate is defined over reference-abstention runs and defect-free accuracy requires exact label match, systematic label errors on the 183 un-reviewed split-vote cases, especially on the Indeterminate classes, would directly bias both headline findings. The paper acknowledges this limitation in Section 7 but does not quantify its impact. I recommend a sensitivity analysis that recomputes over-commitment and defect-free accuracy under plausible label flips for split-vote cases, or direct clinician review of all 183 split-vote cases; without this, the claim that every system under-abstains is not robustly established.
- [Section 4.5 (Clinical Review and Label Calibration)] The calibration review is explicitly not independent: clinicians were given the scenario specification and the three ensemble reports and did not interface with the patient record themselves. This means the ensemble's retrieval and extraction errors are never independently checked. This is particularly consequential for Indeterminate: Lack of Data verdicts, which depend on the absence of evidence: if the ensemble failed to retrieve a decisive record, the reference label would be systematically wrong. The paper states this design is intentional, but it limits the calibration to medical judgment over model-surfaced evidence. I recommend a supplementary audit in which clinicians directly query the raw MIMIC record for a sample of split-vote Lack-of-Data cases, in order to rule out extraction-driven label error.
- [Sections 5.5, 5.7, and 6.7 (process rubric and judge)] Defect-free accuracy, the second headline metric, depends on the LLM judge's grading of MUST-NOT process criteria. The production judge is GPT-5.5, which is also an evaluated system and a member of the reference-verdict ensemble. The paper's judge-reliability validation reports moderate agreement for the process rubric (Fleiss kappa approximately 0.65, with 16% of criteria contested) and finds no own-family favoritism only for policy support, not for process criteria. Because the defect gap reorders the leaderboard, a systematic grading bias in MUST-NOT criteria for one system family could change the central conclusion. Please report process-grading agreement broken down by evaluated system family, or run the process rubric with an independent judge from a different model family, to rule out this source of bias.
minor comments (5)
- [Section 4.4 and Table 2] The paper should state explicitly whether the GPT-5.5 model used in the reference-verdict ensemble is the same checkpoint as the evaluated GPT-5.5 system, and should provide version or date information for all ensemble members, so that the degree of overlap between the labeling process and the evaluated systems is fully transparent.
- [Figure 5 and Table 3] The caption of Figure 5 and the definition of the defect gap should clarify that both raw accuracy and defect-free accuracy in the figure are computed over report-present runs, matching the Gap column in Table 3, so that readers do not compare the figure's raw-accuracy values directly with the all-run Acc column.
- [Section 6.6 and Appendix E] The development-subset experiments use Opus 4.8 for paired ablations (Table 8) but Opus 5 for the sliced full-cohort runs reported in Appendix E.1 (Table 16); while the text explains this, a single sentence in Section 6.6 pointing to the two different sources would prevent readers from mistakenly treating Table 8 and Table 16 as directly comparable.
- [Section 5.5 (process score)] The process score formula floors negative totals at zero, so a run that violates one MUST-NOT criterion and a run that violates several are indistinguishable once the total is zero; the text should state this limitation explicitly, since the defect-free metric counts any MUST-NOT violation equally.
- [Section 7 and release plans] The paper repeatedly states that benchmark data and evaluation code will be released, but no repository or release timeline is provided; for a benchmark paper, a concrete availability statement with an anonymized URL would strengthen reproducibility claims.
Circularity Check
Reference verdicts and the production judge are generated in part by GPT-5.5, which is also an evaluated system, giving the leaderboard a self-referential component; the Clinical Board calibration and multi-judge panels keep the central abstention and defect-free findings from being fully forced.
-
self definitional
[Section 6.7 (roles); Section 4.4 (label construction)]
"GPT-5.5 serves three roles: one of three systems in the reference-verdict ensemble, an evaluated system through Codex, and the production judge for the semantically adjudicated metrics."
By the Section 4.4 construction, the reference verdict is the aggregate of three model outputs, one of which is GPT-5.5's own vote. For the GPT-5.5 system, the predicted verdict is the same output, so leaderboard accuracy and abstention metrics are scored against a label that contains that output as a constituent. The reduction is partial: another ensemble member must concur for the majority label, and the 11 three-distinct cases were escalated to clinicians, while the 183 split 2-vs-1 cases were not directly adjudicated. The central under-abstention result is not forced, because it is measured over all sixteen systems and reappears on the development subset; process and policy scores use multi-judge panels. The self-reference biases, but does not by itself determine, GPT-5.5's rank.
full rationale
The circularity burden is concentrated in the overlap among label generator, judge, and evaluated system. The paper is candid about this: Section 5.7 states that the production judge is GPT-5.5, Section 6.7 lists its three roles and says replication with an independent judge remains the direct test, and Section 4.5 explicitly notes that the Clinical Board calibration is 'not truly independent' because reviewers used model-surfaced evidence rather than the raw record. Against that, the central claims are not equations that reduce to their inputs. The under-abstention asymmetry is computed for all sixteen systems against the same reference labels and reproduces on the 148-case development subset; the defect-free gap is scored from Clinical Board-authored MUST-NOT criteria and validated by a five-judge panel with reported agreement; the grounding and policy metrics pair rule-based checks with LLM judgment, and the self-referential judge is tested for own-family favoritism, with no evidence found. The one construction-level issue is GPT-5.5's own score against an ensemble reference that contains its own vote, which is a genuine but partial self-reference. No parameter is fitted to force the headline results, so a moderate score of 4 is appropriate rather than a higher circularity verdict.
Assumptions & free parameters
assumptions (5)
- domain assumption MIMIC-IV v3.1 records contain the patient-level evidence needed to instantiate all 25 audit scenarios.
- domain assumption Reference verdicts produced by the three-model ensemble are accurate for cases not directly reviewed by clinicians.
- domain assumption LLM judges assign stable and unbiased grades for claim support, finding coverage, process criteria, and policy support.
- domain assumption The policy corpus includes the correct and complete governing standards for every scenario.
- domain assumption Stratified sampling of 30 cases per scenario is sufficient for the benchmark's aggregate comparative claims.
Cite this review
Pith. "Pith review of CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR." pith.science (2026). https://pith.science/paper/J5KIDMTY
@misc{pith2026260807796,
author = {Pith},
title = {Pith review of: CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR},
year = {2026},
howpublished = {\url{https://pith.science/paper/J5KIDMTY}},
note = {Machine review of arXiv:2608.07796}
}
read the original abstract
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Sara Mahdavi, Jason Wei, et al
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, et al. Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023. doi: 10.1038/s41586-023-06291-2
-
[2]
Ca- pabilities of gpt-4 on medical challenge problems.ArXiv, abs/2303.13375, 2023
Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. Ca- pabilities of gpt-4 on medical challenge problems.ArXiv, abs/2303.13375, 2023. URL https: //api.semanticscholar.org/CorpusID:257687695
arXiv 2023
-
[3]
Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim, Rahul K. Arora, Foivos Tsimpourlas, Pre- ston Bowman, Michael Sharman, Chio-Kin Tong, Kavin Karthik, Arnav Dugar, Akshay Jagadeesh, Khaled Saab, Johannes Heidecke, Ashley Alexander, Nate Gross, and Karan Singhal. Healthbench professional: Evaluating large language models on real clinician chats.ArXiv, ...
arXiv 2026
-
[4]
Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments
Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments. ArXiv, abs/2405.07960, 2024. URLhttps://api.semanticscholar.org/CorpusID:269757778
arXiv 2024
-
[5]
Luyang Luo, Sung Eun Kim, Xiaoman Zhang, Julius Kernbach, Roshan Kenia, Julián Nicolás Acosta, Larry A Nathanson, Adrian D Haimovich, Adam Rodman, Ethan Goh, Jonathan H. Chen, Nigam H. Shah, David A. Kim, James Zou, Faisal Mahmood, Jakob Nikolas Kather, Matthew P . Lungren, Vivek Natarajan, Eric J. Topol, and Pranav Rajpurkar. A clinical environment simul...
-
[6]
Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow, Juan M. Banda, Nikesh Kotecha, Timothy Keyes, Yifan Mai, Mert Oez, Hao Qiu, Shrey Jain, Leonardo Schettini, Mehr Kashyap, Jason Alan Fries, Akshay Swaminathan, Philip Chung, Fateme Nateghi, Asad Aali, Ash- win Nayak, Shivam Vedak, Sneha S. Jain, Birju Patel, Oluseyi Fayanju, Shreya J. S...
arXiv 2025
-
[7]
Fleming, Alejandro Lozano, William J
Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak, Birju S. Patel, Chia- Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay S. Chaudhari, Nima Ag...
-
[8]
EHRSHOT: An EHR benchmark for few-shot evaluation of foundation models
Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason Alan Fries, and Nigam Shah. EHRSHOT: An EHR benchmark for few-shot evaluation of foundation models. InThirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URLhttps://openreview. net/forum?id=CsXC6IcdwI
work page 2023
Show all 44 references
-
[9]
Goodell, Yeasul Kim, S
Philip Chung, Akshay Swaminathan, Alex J. Goodell, Yeasul Kim, S. Momsen Reincke, Lichy Han, Ben Deverett, Mohammad Amin Sadeghi, Abdel-Badih Ariss, Marc Ghanem, David Seong, Andrew A. Lee, Caitlin E. Coombes, Brad Bradshaw, Mahir A. Sufian, Hyo Jung Hong, Teresa P . Nguyen, M...
2026
-
[10]
FactEHR: A dataset for evaluating factuality in clinical notes using LLMs
Monica Munnangi, Akshay Swaminathan, Jason Alan Fries, Jenelle A Jindal, Sanjana Narayanan, Ivan Lopez, Lucia Tu, Philip Chung, Jesutofunmi Omiye, Mehr Kashyap, and Nigam Shah. FactEHR: A dataset for evaluating factuality in clinical notes using LLMs. In Monica Agrawal, Kaival...
2025
-
[11]
Hejie Cui, Alyssa Unell, Bowen Chen, Jason Alan Fries, Emily Alsentzer, Oluwasanmi Koyejo, and Nigam H. Shah. Timer: temporal instruction modeling and evaluation for longitudinal clinical records.NPJ Digital Medicine, 8, 2025. doi: 10.1038/s41746-025-01965-9. URL https: //api....
2025 doi
-
[12]
Clindet-bench: Beyond abstention, evaluating judgment determinability of llms in clinical decision-making, 2026
Yusuke Watanabe, Yohei Kobashi, Takeshi Kojima, Yusuke Iwasawa, Yasushi Okuno, and Yutaka Matsuo. Clindet-bench: Beyond abstention, evaluating judgment determinability of llms in clinical decision-making, 2026. URLhttps://arxiv.org/abs/2602.22771
2026
-
[13]
Knowing when to abstain: Medical llms under clinical uncertainty, 2026
Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, and Zonghai Yao. Knowing when to abstain: Medical llms under clinical uncertainty, 2026. URL https://arxiv.org/abs/2601.12471
2026
-
[14]
MIMIC-IV.PhysioNet, October 2024
Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Brian Gow, Benjamin Moody, Steven Horng, Leo Anthony Celi, and Roger Mark. MIMIC-IV.PhysioNet, October 2024. doi: 10.13026/kpb9-mt58. URLhttps://doi.org/10.13026/kpb9-mt58. Version 3.1
2024 doi
-
[15]
KDIGO clinical practice guideline for acute kidney injury.Kidney International Supplements, 2012
Kidney Disease: Improving Global Outcomes (KDIGO) Acute Kidney Injury Work Group. KDIGO clinical practice guideline for acute kidney injury.Kidney International Supplements, 2012
2012
-
[16]
Severe Sepsis and Septic Shock: Management Bundle Measure, 2020
Center for Medicare & Medicaid Services. Severe Sepsis and Septic Shock: Management Bundle Measure, 2020. URL https://www.cms.gov/priorities/innovation/media/document/ bpci-advanced-alt-fs-my4-sepsis
2020
-
[17]
Carson, Simon J
Jeffrey L. Carson, Simon J. Stanworth, Gordon Guyatt, Stacey Valentine, Jane Dennis, Sara Bakhtary, Claudia S. Cohn, Allan Dubon, Brenda J. Grossman, Gaurav K. Gupta, Aaron S. Hess, Jessica L. 38 Scale AI Research Jacobson, Lewis J. Kaplan, Yulia Lin, Ryan A. Metcalf, Colin H....
2023
-
[18]
Heidenreich, Biykem Bozkurt, David Aguilar, Larry A
Paul A. Heidenreich, Biykem Bozkurt, David Aguilar, Larry A. Allen, Joni J. Byun, Monica M. Colvin, Anita Deswal, Mark H. Drazner, Shannon M. Dunlay, Linda R. Evers, James C. Fang, Savitri E. Fedson, Gregg C. Fonarow, Salim S. Hayek, Adrian F. Hernandez, Prateeti Khazanie, Mic...
2022
-
[19]
Evaluation and mitigation of the limitations of large language models in clinical decision-making.Nature medicine, 30(9):2613–2622, 2024
Paul Hager, Friederike Jungmann, Robbie Holland, Kunal Bhagat, Inga Hubrecht, Manuel Knauer, Jakob Vielhauer, Marcus Makowski, Rickmer Braren, Georgios Kaissis, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making.Nature medi...
2024
-
[20]
Fidelity of medical reasoning in large language models.JAMA Network Open, 8(8):e2526021, 2025
Suhana Bedi, Yixing Jiang, Philip Chung, Sanmi Koyejo, and Nigam Shah. Fidelity of medical reasoning in large language models.JAMA Network Open, 8(8):e2526021, 2025
2025
-
[21]
Yu Gu, Jingjing Fu, Xiaodong Liu, Jeya Maria Jose Valanarasu, Noel C. F. Codella, Reuben Tan, Qianchu Liu, Ying Jin, Sheng Zhang, Jinyu Wang, Rui Wang, Lei Song, Guanghui Qin, Naoto Usuyama, Cliff Wong, Hao Cheng, HoHin Lee, Praneeth Sanapathi, Sarah Hilado, Tristan Naumann, J...
2026
-
[22]
Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, and Karan Singhal. Healthbench: Evaluating large language models towards improved human...
2025 arXiv
-
[23]
Rethinking clinical trials for medical ai with dynamic deployments of adaptive systems.npj Digital Medicine, 8(1):252, 2025
Jacob T Rosenthal, Ashley Beecy, and Mert R Sabuncu. Rethinking clinical trials for medical ai with dynamic deployments of adaptive systems.npj Digital Medicine, 8(1):252, 2025
2025
-
[24]
Towards autonomous medical artificial intelligence agents.Nature, 2026
Dyke Ferber, Lars Hilgers, Christiane Höper, Benedict Kinny-Köster, Jan-Niklas Eckardt, et al. Towards autonomous medical artificial intelligence agents.Nature, 2026. doi: 10.1038/ s41586-026-10675-5. URL https://www.nature.com/articles/s41586-026-10675-5 . Advance online publication
2026
-
[25]
Ben Tamo, Xukai Zhao, Jinzhuo Wang, and May Dongmei Wang
Yuxing Lu, Yushuhong Lin, Wenqi Shi, J. Ben Tamo, Xukai Zhao, Jinzhuo Wang, and May Dongmei Wang. Clinenv: An interactive multi-stage long horizon ehr environment for agents, 2026. URL https://arxiv.org/abs/2606.02568
2026 arXiv
-
[26]
Ehrsql: A practical text-to-sql benchmark for electronic health records.Advances in Neural Information Processing Systems, 35:15589–15601, 2022
Gyubok Lee, Hyeonji Hwang, Seongsu Bae, Yeonsu Kwon, Woncheol Shin, Seongjun Yang, Minjoon Seo, Jong-Yeup Kim, and Edward Choi. Ehrsql: A practical text-to-sql benchmark for electronic health records.Advances in Neural Information Processing Systems, 35:15589–15601, 2022. 39 S...
2022
-
[27]
Ho, Carl Yang, and May Dongmei Wang
Wenqi Shi, Ran Xu, Yuchen Zhuang, Yue Yu, Jieyu Zhang, Hang Wu, Yuanda Zhu, Joyce C. Ho, Carl Yang, and May Dongmei Wang. EHRAgent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. InProceedings of the 2024 Conference on ...
2024 doi
-
[28]
Medagentbench: A virtual ehr environment to benchmark medical llm agents.NEJM AI, page AIdbp2500144, 2025
Yixing Jiang, Kameron C Black, Gloria Geng, Danny Park, James Zou, Andrew Y Ng, and Jonathan H Chen. Medagentbench: A virtual ehr environment to benchmark medical llm agents.NEJM AI, page AIdbp2500144, 2025
2025
-
[29]
Pollard, Alistair Johnson, Edward Choi, Yugang Jia, and Jong Ha Lee
Gyubok Lee, Elea Bach, Eric Yang, Tom J. Pollard, Alistair Johnson, Edward Choi, Yugang Jia, and Jong Ha Lee. Fhir-agentbench: Benchmarking llm agents for realistic interoperable ehr question answering.ArXiv, abs/2509.19319, 2025. URL https://api.semanticscholar.org/CorpusID: ...
2025
-
[30]
Ehr-complex: Benchmarking medical agents for complex clinical reasoning, 2026
Yitong Qiao, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu, and Kui Ren. Ehr-complex: Benchmarking medical agents for complex clinical reasoning, 2026. URL https://arxiv.org/abs/ 2606.23301
2026 arXiv
-
[31]
Mohiuddin, Austin J
Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler, Kavita Renduchintala, Ashwin Nayak, Pras- antha L. Vemu, Shivam C. Vedak, Kameron C. Black, John L. Havlik, Isaac Ogunmola, Stephen P . Ma, Roopa Dhatt, and Jonathan H. Chen. Physicianbench: Evaluating llm agents in real-wor...
2026 arXiv
-
[32]
Longmedbench: Benchmarking medical agents for long-horizon clinical decision-making, 2026
Yanzhen Chen, Zihan Xu, Xiaocheng Zhang, Zhiting Fan, Weiqi Zhai, Hongxia Xu, and Zuozhu Liu. Longmedbench: Benchmarking medical agents for long-horizon clinical decision-making, 2026. URLhttps://arxiv.org/abs/2607.09322
2026 arXiv
-
[33]
A dataset for addressing patient’s information needs related to clinical course of hospitalization.Scientific Data, 13:523, 2026
Sarvesh Soni and Dina Demner-Fushman. A dataset for addressing patient’s information needs related to clinical course of hospitalization.Scientific Data, 13:523, 2026. doi: 10.1038/ s41597-026-06639-z. URLhttps://doi.org/10.1038/s41597-026-06639-z
2026 doi
-
[34]
Clicare: Grounding large language models in clinical guidelines for decision support over longitudinal cancer electronic health records, 2026
Dongchen Li, Jitao Liang, Wei Li, Xiaoyu Wang, Longbing Cao, and Kun Yu. Clicare: Grounding large language models in clinical guidelines for decision support over longitudinal cancer electronic health records, 2026. URLhttps://arxiv.org/abs/2507.22533
2026
-
[35]
Overview of the EHRSQL 2024 shared task on reliable text-to-SQL modeling on electronic health records
Gyubok Lee, Sunjun Kweon, Seongsu Bae, and Edward Choi. Overview of the EHRSQL 2024 shared task on reliable text-to-SQL modeling on electronic health records. InProceedings of the 6th Clinical Natural Language Processing Workshop, pages 644–654. Association for Computational L...
2024 doi
-
[36]
Agents’ last exam.arXiv preprint arXiv:2606.05405, 2026
Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, et al. Agents’ last exam.arXiv preprint arXiv:2606.05405, 2026
2026
-
[37]
Suhana Bedi, Iddah Mlauzi, Daniel Shin, Sanmi Koyejo, and Nigam H. Shah. The optimization paradox in clinical AI multi-agent systems. InAgentic & GenAI Evaluation Workshop, KDD 2025 (Evaluation and Trustworthiness of Agentic and Generative AI Models), 2025. URL https://openrev...
2025
-
[38]
A framework for formalizing LLM agent security, 2026
Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Neil Gong, Chenguang Wang, and Dawn Song. A framework for formalizing LLM agent security, 2026. URL https://arxiv.org/abs/2603. 19469
2026
-
[39]
Steinberg, Michael Wornow, Taeil Kim, Haroun Zakaria Ahmed, Peter V
Suhana Bedi, Ryan Welch, Ethan H. Steinberg, Michael Wornow, Taeil Kim, Haroun Zakaria Ahmed, Peter V . Sterling, Bravim K. Purohit, Qurat-Ul-Ain Akram, Angelic Acosta, Esther Nubla, Priti Sharma, Mike Pfeffer, Oluwasanmi Koyejo, and Nigam H. Shah. Healthadminbench: Evaluating...
2026 arXiv
-
[40]
MIMIC-IV-Note: Deidentified free-text clinical notes.PhysioNet, January 2023
Alistair Johnson, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. MIMIC-IV-Note: Deidentified free-text clinical notes.PhysioNet, January 2023. doi: 10.13026/1n74-ne17. URL https://doi.org/10.13026/1n74-ne17. Version 2.2
2023 doi
-
[41]
MIMIC-IV-ED.PhysioNet, January 2023
Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Leo Anthony Celi, Roger Mark, and Steven Horng. MIMIC-IV-ED.PhysioNet, January 2023. doi: 10.13026/5ntk-km72. URL https://doi.org/10. 13026/5ntk-km72. Version 2.2
2023 doi
-
[42]
first qualifying
Gordon F. Tomaselli, Kenneth W. Mahaffey, Adam Cuker, Paul P . Dobesh, John U. Doherty, John W. Eikelboom, Roberta Florido, Ty J. Gluckman, William J. Hucker, Roxana Mehran, Steven R. Messé, Alexander C. Perino, Fatima Rodriguez, Ravindra Sarode, Deborah M. Siegal, and Barbara...
2020
-
[1245]
URLhttps://aclanthology.org/2024.emnlp-main.1245/
2024
-
[2024]
URLhttps://doi.org/10.1609/aaai.v38i20.30205
doi: 10.1609/AAAI.V38I20.30205. URLhttps://doi.org/10.1609/aaai.v38i20.30205
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.