Pith. sign in

REVIEW 3 major objections 6 minor 299 references

Medical agents become clinically trustworthy by scaling the tools and gyms they can reach, and by improving through interaction, not by growing model size alone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 06:28 UTC pith:JEMXH7JT

load-bearing objection Useful deployment-first survey with a clear spine; the environment-scaling priority is a roadmap claim, not a measured law. the 3 major comments →

arxiv 2607.11175 v1 pith:JEMXH7JT submitted 2026-07-13 cs.AI

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

classification cs.AI
keywords medical imaging agentsvision-language modelsself-evolutionclinical scalingautonomy taxonomyagentic benchmarkstrustworthy clinical AIsurvey
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This survey argues that medical imaging agents will only be ready for real hospitals when we stop treating them as smarter single-shot predictors and start treating them as sequential decision systems that act under incomplete clinical evidence. The authors organize the field around a single "scaling spine" with three axes: how agents are wired together, how deeply their perceive-reason-act loop runs, and how rich the reachable environment of tools, data, and simulators is. They claim that clinical environment scaling—connecting agents to PACS, EHR, FHIR systems and interactive clinical gyms—is the most practical lever available today, and that lasting progress comes when agents improve by practicing in those environments rather than only by adding parameters. Across radiology, pathology, ophthalmology, and hospital workflows, the paper maps a path from assisted tools toward self-evolving clinical systems, while cataloguing the failure modes that autonomy multiplies.

Core claim

Real-world readiness of a medical agent is amplified along three complementary axes—framework topology, cognitive-loop depth, and reachable environment size—and clinical environment scaling (tools, data, and interactive gyms inside PACS, EHR, and FHIR ecosystems) is the most immediately actionable yet least explored lever, with clinical self-evolution through environment interaction as the destination rather than parameter growth alone.

What carries the argument

The scaling spine: readiness is treated as a function of topology richness N_topo, cognitive-loop depth D_loop, and environment size |E|, so that framework wiring, harness-driven iteration, and tool/data access are orthogonal levers that together lift a passive model toward a self-evolving clinical agent.

Load-bearing premise

That enriching the clinical tool and data environment often buys more real-world readiness per unit effort than further fine-tuning or more complex agent teams, and that the three scaling axes stay largely independent under real hospital constraints.

What would settle it

Hold backbone model and clinical task fixed, systematically vary only the richness of available tools, archives, and interactive gyms across matched deployments, and check whether measured readiness gains track environment size more than topology or loop depth; if they do not, the prioritization of environment scaling fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Hospitals and builders should invest first in standardized tool interfaces, clinical gyms, and interoperable data layers rather than only larger models.
  • Evaluation must move from static accuracy to multi-turn, tool-using, end-to-end workflow benchmarks with process-level and safety metrics.
  • Self-evolving agents that improve from environment interaction become a primary research target rather than a side topic.
  • Risk analysis for hallucination, cascade failures, and fairness must be framed as properties of the full orchestration loop, not of a single model.
  • The assisted–cooperative–fully autonomous taxonomy becomes a practical language for clinical adoption and regulatory design.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If environment scaling is truly the cheapest lever, hospital IT integration and protocol standardization may matter more for agent readiness than another round of parameter growth.
  • The partial-observability framing suggests medical agents will borrow more from sequential decision-making and interactive simulators than from classic single-pass medical imaging pipelines.
  • Contamination-resistant, multi-level agentic benchmarks could become the de facto gate for clinical AI claims the way static leaderboards once gated perception models.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This survey organizes medical imaging agents from a deployment-first stance rather than a capability-first taxonomy. It formalizes agents as six-module sequential decision systems under partial observability, proposes a three-level autonomy taxonomy (assisted, cooperative, fully autonomous), and threads the literature along a conceptual “scaling spine” of framework topology N_topo, cognitive-loop depth D_loop, and environment size |E| (Eq. 1; §2.3). Clinical environment scaling—tools, data, and gyms in PACS/EHR/FHIR ecosystems—is argued to be the most immediately actionable lever, with clinical self-evolution via environment interaction as the aspirational destination. The manuscript consolidates 300+ references (search window Jan 2022–Jun 2026; disclosed multi-reviewer screening), catalogs benchmarks and systems (Tables 1–3), reviews specialty applications and deployment/regulatory constraints, and maps risks along the agent pipeline.

Significance. If the organizing claims hold as a research roadmap, the paper offers a useful alternative to concurrent capability-centric medical-agent surveys: it puts contamination-resistant benchmarks, interactive clinical gyms, and interoperability first, and it elevates environment scaling and self-evolution as explicit research priorities rather than appendices. Strengths that should be credited include a disclosed retrieval protocol (1,174→942→300+), extensive benchmark and system catalogs (Tables 1–3), an explicit autonomy taxonomy (Fig. 1), a companion GitHub, and consistently cautious language that Level-3 autonomy and self-evolution remain aspirational. The synthesis of 2025–2026 agentic RL, gyms, and protocol work (MCP/FHIR) is timely for the community. The distinctive thesis—that |E| is often the highest-leverage axis for agents already in tool-rich clinical stacks—is significant as a prioritization claim, but its force depends on how carefully it is framed relative to the evidence the survey actually provides.

major comments (3)
  1. §2.3 and Eq. (1) present readiness as factoring into three “largely independent” / “orthogonal” axes and argue that enriching |E| often buys more readiness per unit effort than further finetuning or topology growth for agents in PACS/EHR/FHIR ecosystems (also Fig. 2; §6.7). This prioritization is the paper’s distinctive load-bearing thesis, yet the manuscript never holds backbone and task fixed while varying only environment richness, nor does it systematically code the systems in Tables 1–3 by (N_topo, D_loop, |E|). In §§4–7, tool enrichment, loop depth, and multi-agent wiring routinely co-vary. Please either (i) reframe the priority of environment scaling as a hypothesis/roadmap claim with explicit falsification criteria, or (ii) add a short comparative synthesis of a subset of cataloged systems that attempts to disentangle the three axes (even qualitatively). Without that, the spine r
  2. §3.2–3.4 and Table 1 organize benchmarks by autonomy/realism levels and correctly stress Level-2/3 gaps, but the survey’s claim that environment scaling is “underexplored” sits uneasily next to the rapid 2025–2026 arrival of MedAgentGym, Healthcare AI GYM, ClinEnv, MedAgentBench, ABRA, etc., which the paper itself catalogs. Clarify what “underexplored” means operationally (e.g., lack of controlled ablations of tool/data richness; lack of multi-site gyms; lack of reward design for imaging-specific verifiers) so the roadmap does not understate the very literature it surveys while still motivating the proposed lever.
  3. §2.1 formalizes A=(P,R,L,M,T,F) under partial observability but stops short of a usable decision-process statement (state/belief, action space, observation model, reward/objective). Given that later sections invoke agentic multi-turn RL and verifiable rewards (§6.4–6.6; Table 2) as the training counterpart, a brief explicit mapping from the six modules to a POMDP/belief-MDP (or an honest statement that the formalization is taxonomic rather than operational) would make the “sequential decision making under partial observability” claim load-bearing rather than decorative, and would tighten the link between §2 and the RL/self-evolution material.
minor comments (6)
  1. Fig. 2 and Eq. (1): label Eq. (1) explicitly as a conceptual readiness map, not a scaling law, to avoid readers treating N_topo, D_loop, |E| as measurable, calibrated quantities.
  2. Table 1 / Table 3: several 2026 entries list venues as arXiv or workshop; a short note on inclusion criteria for preprints versus peer-reviewed work would help readers weight the catalog.
  3. §1 “Relation to concurrent surveys”: the claim that none couples a disclosed retrieval protocol with an environment-centered axis is strong; a short comparison table (scope, axes, retrieval protocol, self-evolution coverage) would make the differentiation checkable.
  4. §5.7–5.8 and §6.6 partially overlap on self-evolution mechanisms (skills, memory, code rewrite). A single cross-reference box stating what is inference-time substrate vs training-time evolution would reduce redundancy.
  5. Minor consistency: “self-evolving” / “self-evolution” hyphenation and capitalization vary across abstract, keywords, and §5.8/§6.6; normalize.
  6. §9.2 cites a 76.6% vs 51.3% hallucination result from Kim et al.; state the exact metric and setting in-text so the counterintuitive claim is interpretable without opening the reference.

Circularity Check

0 steps flagged

No significant circularity: conceptual survey spine, not a fitted or self-definitional derivation.

full rationale

This is a literature survey that organizes medical-agent work along a conceptual “scaling spine” (framework topology N_topo, loop depth D_loop, environment size |E|; Eq. 1, §2.3, Fig. 2) and argues that clinical environment scaling is the most immediately actionable lever, with self-evolution as an aspirational frontier. Eq. 1 is explicitly a compact readiness map (“we write compactly as”), not a calibrated law or a prediction obtained by fitting free parameters to data. The prioritization of |E| is supported by external environment-scaling literature [62, 97], tool-interface standards (MCP/FHIR), and catalogs of benchmarks, gyms, and systems (Tables 1–3; §§3–7), not by a uniqueness theorem or load-bearing self-citation chain that forbids alternatives. Author self-citations (e.g., MedCausalX, MedEyes) appear as surveyed applications, which is normal and non-load-bearing for the spine thesis. There is no self-definitional loop (X defined as Y then “derived”), no fitted input renamed as prediction, no ansatz smuggled in as a forced result, and no renaming of a known empirical law presented as a first-principles derivation. Weaknesses of the thesis (unmeasured orthogonality of axes; co-variation of tools, loop depth, and topology in the surveyed systems) are evidence and correctness issues, not circular reductions. Score 0 with empty steps is the honest finding.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 3 invented entities

As a survey, the paper’s load-bearing content is definitional and organizational rather than fitted. The main commitments are the POMDP-style agent formalization, the three-level autonomy taxonomy, the three-axis readiness map, and the priority claim for environment scaling and self-evolution. No numerical free parameters are fit to clinical outcomes. Invented entities are taxonomies and axes used to structure the literature, not physical objects.

axioms (5)
  • domain assumption Medical agency is sequential decision-making under partial observability over images, reports, and tool outputs rather than single-pass prediction.
    Section 2.1 formalizes A=(P,R,L,M,T,F) and treats partial observability as definitional for clinical agents.
  • ad hoc to paper Readiness factors approximately into orthogonal framework, capability-loop, and environment axes (Eq. 1).
    Section 2.3 introduces the scaling spine as the survey’s global map; orthogonality is asserted rather than derived from data.
  • ad hoc to paper Enriching tools/data/interfaces often yields more readiness per unit effort than further parameter scaling for agents already in tool-rich clinical ecosystems.
    Stated in Introduction and Sections 2.3 and 6.7; supported by citation to environment-scaling literature, not a controlled measurement in this paper.
  • domain assumption A three-level assisted/cooperative/fully-autonomous taxonomy is an appropriate transfer from driving and surgical robotics to medical agents.
    Section 2.2 cites SAE-style and surgical autonomy levels as background for Fig. 1.
  • domain assumption Self-evolution via environment interaction is a primary long-term path to trustworthy clinical agents beyond frozen foundation models.
    Abstract, Sections 5.8 and 6.6; treated as aspirational frontier informed by general-domain agent gyms.
invented entities (3)
  • Scaling spine (N_topo, D_loop, |E| readiness map) no independent evidence
    purpose: Unify heterogeneous medical-agent literature under three complementary scaling axes culminating in self-evolution.
    Eq. 1 and Fig. 2 are paper-specific organizing constructs; independent evidence is only qualitative literature synthesis.
  • Three-level medical-agent autonomy taxonomy (assisted/cooperative/fully autonomous) no independent evidence
    purpose: Provide a shared clinical autonomy ladder for comparing systems and deployment risk.
    Adapted from driving/robotics taxonomies but specialized here for medical agents (Fig. 1); useful taxonomy, not an empirically validated clinical standard.
  • Clinical environment scaling as primary actionable lever no independent evidence
    purpose: Prioritize tools, data, FHIR/PACS interfaces, and clinical gyms over parameter growth.
    Central agenda claim of the survey; falsifiable only via future controlled environment-scaling studies not present here.

pith-pipeline@v1.1.0-grok45 · 54121 in / 3394 out tokens · 37032 ms · 2026-07-14T06:28:43.670432+00:00 · methodology

0 comments
read the original abstract

The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from the capability first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision making systems under partial observability, together with a three level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self improving agents, agent gyms, and test time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this work provides a roadmap toward trustworthy, self improving medical imaging systems for real clinical practice.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

299 extracted references · 105 linked inside Pith

  1. [1]

    Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023

  2. [2]

    Longhealth: A question answering benchmark with long clinical documents.Journal of Healthcare Informatics Research, 9(3):280–296, 2025

    Lisa Adams, Felix Busch, Tianyu Han, Jean-Baptiste Excoffier, Matthieu Ortala, Alexander Löser, Hugo JWL Aerts, Jakob Nikolas Kather, Daniel Truhn, and Keno Bressem. Longhealth: A question answering benchmark with long clinical documents.Journal of Healthcare Informatics Research, 9(3):280–296, 2025

  3. [3]

    Synthagent: A multi-agent llm framework for realistic patient simulation–a case study in obesity with mental health comorbidities.arXiv preprint arXiv:2602.08254, 2026

    Arman Aghaee, Sepehr Asgarian, and Jouhyun Jeon. Synthagent: A multi-agent llm framework for realistic patient simulation–a case study in obesity with mental health comorbidities.arXiv preprint arXiv:2602.08254, 2026

  4. [4]

    Oaagent: Multimodal llm agent for predicting knee osteoarthritis progression

    Pegah Ahadian, Mingrui Yang, Eva Powlison, Xiaojuan Li, Wei Xu, and Qiang Guan. Oaagent: Multimodal llm agent for predicting knee osteoarthritis progression. InProceedings of the ACM/IEEE International Conference on Connected Health: Applications, Systems and Engineering Technologies, pages 144–148, 2025

  5. [5]

    Self-evolving multi-agent simulations for realistic clinical interactions.arXiv preprint arXiv:2503.22678, 2025

    Mohammad Almansoori, Komal Kumar, and Hisham Cholakkal. Self-evolving multi-agent simulations for realistic clinical interactions.arXiv preprint arXiv:2503.22678, 2025

  6. [6]

    Effective context engineering for AI agents

    Anthropic. Effective context engineering for AI agents. Anthropic Engineering Blog, 2025. https://www. anthropic.com/engineering/effective-context-engineering-for-ai-agents, accessed 2026-06-14

  7. [7]

    Effective harnesses for long-running agents

    Anthropic. Effective harnesses for long-running agents. Anthropic Engineering Blog, 2025. https://www. anthropic.com/engineering/effective-harnesses-for-long-running-agents, accessed 2026-06-14

  8. [8]

    Healthbench: Evaluating large language models towards improved human health.arXiv preprint arXiv:2505.08775, 2025

    Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, et al. Healthbench: Evaluating large language models towards improved human health.arXiv preprint arXiv:2505.08775, 2025

  9. [9]

    Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

  10. [10]

    Agentic ai in healthcare: A comprehensive survey of foundations, taxonomy, and applications.Authorea Preprints, 2025

    Shruti Banerjie, Yuxin Zhu, Isaac Freeman, Julyssa Villa Machado, Abdulaziz Ahmed, Abeed Sarker, and Mohammed Al-Garadi. Agentic ai in healthcare: A comprehensive survey of foundations, taxonomy, and applications.Authorea Preprints, 2025

  11. [11]

    Magda: Multi-agent guideline-driven diagnostic assistance

    David Bani-Harouni, Nassir Navab, and Matthias Keicher. Magda: Multi-agent guideline-driven diagnostic assistance. InInternational workshop on foundation models for general medical AI, pages 163–172. Springer, 2024

  12. [12]

    Maira-2: Grounded radiology report generation.arXiv preprint arXiv:2406.04449, 2024

    Shruthi Bannur, Kenza Bouzid, Daniel C Castro, Anton Schwaighofer, Anja Thieme, Sam Bond-Taylor, Maxim- ilian Ilse, Fernando Pérez-García, Valentina Salvatelli, Harshita Sharma, et al. Maira-2: Grounded radiology report generation.arXiv preprint arXiv:2406.04449, 2024

  13. [13]

    Radgpt: Constructing 3d image-text tumor datasets

    Pedro RAS Bassi, Mehmet Can Yavuz, Ibrahim Ethem Hamamci, Sezgin Er, Xiaoxi Chen, Wenxuan Li, Bjoern Menze, Sergio Decherchi, Andrea Cavalli, Kang Wang, et al. Radgpt: Constructing 3d image-text tumor datasets. InProceedings of the IEEE/CVF international conference on computer vision, pages 23720–23730, 2025

  14. [14]

    Medhelm: Holistic evaluation of large language models for medical tasks.arXiv preprint arXiv:2505.23802, 2025

    Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow, Juan M Banda, Nikesh Kotecha, Timothy Keyes, Yifan Mai, Mert Oez, et al. Medhelm: Holistic evaluation of large language models for medical tasks.arXiv preprint arXiv:2505.23802, 2025

  15. [15]

    Healthadminbench: Evaluating computer-use agents on healthcare administration tasks.arXiv preprint arXiv:2604.09937, 2026

    Suhana Bedi, Ryan Welch, Ethan Steinberg, Michael Wornow, Taeil Matthew Kim, Haroun Ahmed, Peter Sterling, Bravim Purohit, Qurat Akram, Angelic Acosta, et al. Healthadminbench: Evaluating computer-use agents on healthcare administration tasks.arXiv preprint arXiv:2604.09937, 2026

  16. [16]

    Graph of thoughts: Solving elaborate problems with large language models

    Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. Graph of thoughts: Solving elaborate problems with large language models. InProceedings of the AAAI conference on artificial intelligence, volume 38, pages 17682–17690, 2024

  17. [17]

    Agentic systems in radiology: Design, applications, evaluation, and challenges.arXiv preprint arXiv:2510.09404, 2025

    Christian Bluethgen, Dave Van Veen, Daniel Truhn, Jakob Nikolas Kather, Michael Moor, Malgorzata Polacin, Akshay Chaudhari, Thomas Frauenfelder, Curtis P Langlotz, Michael Krauthammer, et al. Agentic systems in radiology: Design, applications, evaluation, and challenges.arXiv preprint arXiv:2510.09404, 2025. 32

  18. [18]

    Making the most of text semantics to improve biomedical vision–language processing

    Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. Making the most of text semantics to improve biomedical vision–language processing. InEuropean conference on computer vision, pages 1–21. Springer, 2022

  19. [19]

    Visual alignment of medical vision-language models for grounded radiology report generation

    Sarosij Bose, Ravi K Rajendran, Biplob Debnath, Konstantinos Karydis, Amit K Roy-Chowdhury, and Srimat Chakradhar. Visual alignment of medical vision-language models for grounded radiology report generation. arXiv preprint arXiv:2512.16201, 2025

  20. [20]

    Adversarial robustness analysis of vision-language models in medical image segmentation.arXiv preprint arXiv:2505.02971, 2025

    Anjila Budathoki and Manish Dhakal. Adversarial robustness analysis of vision-language models in medical image segmentation.arXiv preprint arXiv:2505.02971, 2025

  21. [21]

    Reinforcement learning with verifiable yet noisy rewards under imperfect verifiers.arXiv preprint arXiv:2510.00915, 2025

    Xin-Qiang Cai, Wei Wang, Feng Liu, Tongliang Liu, Gang Niu, and Masashi Sugiyama. Reinforcement learning with verifiable yet noisy rewards under imperfect verifiers.arXiv preprint arXiv:2510.00915, 2025

  22. [22]

    Atpo: Adaptive tree policy optimization for multi-turn medical dialogue.arXiv preprint arXiv:2603.02216, 2026

    Ruike Cao, Shaojie Bai, Fugen Yao, Liang Dong, Jian Xu, and Li Xiao. Atpo: Adaptive tree policy optimization for multi-turn medical dialogue.arXiv preprint arXiv:2603.02216, 2026

  23. [23]

    Predetermined change control plans: Guiding principles for advancing safe, effective, and high-quality ai-ml technologies.JMIR AI, 4: e76854, 2025

    Eduardo Carvalho, Miguel Mascarenhas, Francisca Pinheiro, Ricardo Correia, Sandra Balseiro, Guilherme Barbosa, Ana Guerra, Dulce Oliveira, Rita Moura, André Martins dos Santos, et al. Predetermined change control plans: Guiding principles for advancing safe, effective, and high-quality ai-ml technologies.JMIR AI, 4: e76854, 2025

  24. [24]

    Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation.arXiv preprint arXiv:2403.08002, 2024

    Juan Manuel Zambrano Chaves, Shih-Cheng Huang, Yanbo Xu, Hanwen Xu, Naoto Usuyama, Sheng Zhang, Fei Wang, Yujia Xie, Mahmoud Khademi, Ziyi Yang, et al. Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation.arXiv preprint arXiv:2403.08002, 2024

  25. [25]

    Minimax-m1: Scaling test-time compute efficiently with lightning attention.arXiv preprint arXiv:2506.13585, 2025

    Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, et al. Minimax-m1: Scaling test-time compute efficiently with lightning attention.arXiv preprint arXiv:2506.13585, 2025

  26. [26]

    Clinicalbench: Can llms beat traditional ml models in clinical prediction?arXiv preprint arXiv:2411.06469, 2024

    Canyu Chen, Jian Yu, Shan Chen, Che Liu, Zhongwei Wan, Danielle Bitterman, Fei Wang, and Kai Shu. Clinicalbench: Can llms beat traditional ml models in clinical prediction?arXiv preprint arXiv:2411.06469, 2024

  27. [27]

    Chi-bench: Can ai agents automate end-to-end, long-horizon, policy-rich healthcare workflows?arXiv preprint arXiv:2605.16679, 2026

    Haolin Chen, Deon Metelski, Leon Qi, Tao Xia, Joonyul Lee, Steve Brown, Kevin Riley, Frank Wang, TY Liu, Hank Capps MD, et al. Chi-bench: Can ai agents automate end-to-end, long-horizon, policy-rich healthcare workflows?arXiv preprint arXiv:2605.16679, 2026

  28. [28]

    Ethical machine learning in healthcare.Annual review of biomedical data science, 4(1):123–144, 2021

    Irene Y Chen, Emma Pierson, Sherri Rose, Shalmali Joshi, Kadija Ferryman, and Marzyeh Ghassemi. Ethical machine learning in healthcare.Annual review of biomedical data science, 4(1):123–144, 2021

  29. [29]

    Pathagent: Toward interpretable analysis of whole-slide pathology images via large language model-based agentic reasoning.arXiv preprint arXiv:2511.17052, 2025

    Jingyun Chen, Linghan Cai, Zhikang Wang, Yi Huang, Songhan Jiang, Shenjin Huang, Hongpeng Wang, and Yongbing Zhang. Pathagent: Toward interpretable analysis of whole-slide pathology images via large language model-based agentic reasoning.arXiv preprint arXiv:2511.17052, 2025

  30. [30]

    Mdteamgpt: A self-evolving llm- based multi-agent framework for multi-disciplinary team medical consultation.arXiv preprint arXiv:2503.13856, 2025

    Kai Chen, Xinfeng Li, Tianpei Yang, Hewei Wang, Wei Dong, and Yang Gao. Mdteamgpt: A self-evolving llm- based multi-agent framework for multi-disciplinary team medical consultation.arXiv preprint arXiv:2503.13856, 2025

  31. [31]

    Medsentry: Understanding and mitigating safety risks in medical llm multi-agent systems.arXiv preprint arXiv:2505.20824, 2025

    Kai Chen, Taihang Zhen, Hewei Wang, Kailai Liu, Xinfeng Li, Jing Huo, Tianpei Yang, Jinfeng Xu, Wei Dong, and Yang Gao. Medsentry: Understanding and mitigating safety risks in medical llm multi-agent systems.arXiv preprint arXiv:2505.20824, 2025

  32. [32]

    Towards a general-purpose foundation model for computational pathology.Nature medicine, 30(3):850–862, 2024

    Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathology.Nature medicine, 30(3):850–862, 2024

  33. [33]

    Medbrowsecomp: Benchmarking medical deep research and computer use

    Shan Chen, Pedro Moreira, Yuxin Xiao, Sam Schmidgall, Jeremy Warner, Hugo Aerts, Thomas Hartvigsen, Jack Gallifant, and Danielle S Bitterman. Medbrowsecomp: Benchmarking medical deep research and computer use. arXiv preprint arXiv:2505.14963, 2025

  34. [34]

    Radfabric: Agentic ai system with reasoning capability for radiology.arXiv preprint arXiv:2506.14142, 2025

    Wenting Chen, Yi Dong, Zhaojun Ding, Yucheng Shi, Yifan Zhou, Fang Zeng, Yijun Luo, Tianyu Lin, Yihang Su, Yichen Wu, et al. Radfabric: Agentic ai system with reasoning capability for radiology.arXiv preprint arXiv:2506.14142, 2025. 33

  35. [35]

    Automated rubrics for reliable evaluation of medical dialogue systems.arXiv preprint arXiv:2601.15161, 2026

    Yinzhu Chen, Abdine Maiga, Hossein A Rahmani, and Emine Yilmaz. Automated rubrics for reliable evaluation of medical dialogue systems.arXiv preprint arXiv:2601.15161, 2026

  36. [36]

    Meissa: Multi-modal medical agentic intelligence.arXiv preprint arXiv:2603.09018, 2026

    Yixiong Chen, Xinyi Bai, Yue Pan, Zongwei Zhou, and Alan Yuille. Meissa: Multi-modal medical agentic intelligence.arXiv preprint arXiv:2603.09018, 2026

  37. [37]

    Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24185–24198, 2024

  38. [38]

    Chexagent: Towards a foundation model for chest x-ray interpretation

    Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck, Magdalini Paschali, Louis Blankemeier, Dave Van Veen, Jeya Maria Jose Valanarasu, Alaa Youssef, Joseph Paul Cohen, Eduardo Pontes Reis, et al. Chexagent: Towards a foundation model for chest x-ray interpretation. InAAAI 2024 Spring Symposium on Clinical Foundation Models, 2024

  39. [39]

    Engelhard, Somesh Jha, Anivarya Kumar, and David Page

    Jihye Choi, Nils Palumbo, Prasad Chalasani, Matthew M. Engelhard, Somesh Jha, Anivarya Kumar, and David Page. Malade: Orchestration of llm-powered agents with retrieval augmented generation for pharmacovigilance,

  40. [40]

    URLhttps://arxiv.org/abs/2408.01869

  41. [41]

    Medai: Evaluating txagent’s therapeutic agentic reasoning in the neurips cure-bench competition.arXiv preprint arXiv:2512.11682, 2025

    Tim Cofala, Christian Kalfar, Jingge Xiao, Johanna Schrader, Michelle Tang, and Wolfgang Nejdl. Medai: Evaluating txagent’s therapeutic agentic reasoning in the neurips cure-bench competition.arXiv preprint arXiv:2512.11682, 2025

  42. [42]

    SAE international, 2021

    On-Road Automated Driving (ORAD) Committee.Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles. SAE international, 2021

  43. [43]

    Same verdict, different reasons: Llm-as-a-judge and clinician disagreement on medical chatbot completeness

    Alexandra DeLucia, Heyuan Huang, Sonal Joshi, Mahsa Yarmohammadi, Ahmed Hassoon, and Mark Dredze. Same verdict, different reasons: Llm-as-a-judge and clinician disagreement on medical chatbot completeness. arXiv preprint arXiv:2604.16383, 2026

  44. [44]

    Medka: A knowledge graph-augmented approach to improve factuality in medical large language models.Journal of Biomedical Informatics, page 104871, 2025

    Yiyan Deng, Shen Zhao, Yongming Miao, Junjie Zhu, and Jin Li. Medka: A knowledge graph-augmented approach to improve factuality in medical large language models.Journal of Biomedical Informatics, page 104871, 2025

  45. [45]

    Radvlm: a multitask conversational vision-language model for radiology.arXiv preprint arXiv:2502.03333, 2025

    Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruipérez-Campillo, Moritz Vandenhirtz, Sonia Laguna, Alain Ryser, Koji Fujimoto, Mizuho Nishio, Thomas M Sutter, Julia E Vogt, et al. Radvlm: a multitask conversational vision-language model for radiology.arXiv preprint arXiv:2502.03333, 2025

  46. [46]

    Agentic ai in radiology: emerging potential and unresolved challenges.British Journal of Radiology, 98(1174):1582–1584, 2025

    Nicholas Dietrich. Agentic ai in radiology: emerging potential and unresolved challenges.British Journal of Radiology, 98(1174):1582–1584, 2025

  47. [47]

    Empowering multi-turn tool-integrated agentic reasoning with group turn policy optimization.arXiv preprint arXiv:2511.14846, 2025

    Yifeng Ding, Hung Le, Songyang Han, Kangrui Ruan, Zhenghui Jin, Varun Kumar, Zijian Wang, and Anoop Deoras. Empowering multi-turn tool-integrated agentic reasoning with group turn policy optimization.arXiv preprint arXiv:2511.14846, 2025

  48. [48]

    Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025

    Guanting Dong, Licheng Bao, Zhongyuan Wang, Kangzhi Zhao, Xiaoxi Li, Jiajie Jin, Jinghan Yang, Hangyu Mao, Fuzheng Zhang, Kun Gai, et al. Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025

  49. [49]

    Agent-world: Scaling real-world environment synthesis for evolving general agent intelligence.arXiv preprint arXiv:2604.18292, 2026

    Guanting Dong, Junting Lu, Junjie Huang, Wanjun Zhong, Longxiang Liu, Shijue Huang, Zhenyu Li, Yang Zhao, Xiaoshuai Song, Xiaoxi Li, et al. Agent-world: Scaling real-world environment synthesis for evolving general agent intelligence.arXiv preprint arXiv:2604.18292, 2026

  50. [50]

    Ophin-500k: Curating web-scale visual instructions for scaling ophthalmic multimodal large language models.arXiv preprint arXiv:2605.27916, 2026

    Xuanzhao Dong, Wenhui Zhu, Xiwen Chen, Hao Wang, Xin Li, Yujian Xiong, Jiajun Cheng, Jingjing Wang, Xiaobing Yu, Haiyu Wu, et al. Ophin-500k: Curating web-scale visual instructions for scaling ophthalmic multimodal large language models.arXiv preprint arXiv:2605.27916, 2026

  51. [51]

    Care: Towards clinical accountability in multi-modal medical reasoning with an evidence-grounded agentic framework.arXiv preprint arXiv:2603.01607, 2026

    Yuexi Du, Jinglu Wang, Shujie Liu, Nicha C Dvornek, and Yan Lu. Care: Towards clinical accountability in multi-modal medical reasoning with an evidence-grounded agentic framework.arXiv preprint arXiv:2603.01607, 2026

  52. [52]

    Llms can simulate standardized patients via agent coevolution

    Zhuoyun Du, LujieZheng LujieZheng, Renjun Hu, Yuyang Xu, Xiawei Li, Ying Sun, Wei Chen, Jian Wu, Haolei Cai, and Haochao Ying. Llms can simulate standardized patients via agent coevolution. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17278–17306, 2025. 34

  53. [53]

    Abul Ehtesham, Aditi Singh, Gaurav Kumar Gupta, and Saket Kumar. A survey of agent interoperability protocols: Model context protocol (mcp), agent communication protocol (acp), agent-to-agent protocol (a2a), and agent network protocol (anp).arXiv preprint arXiv:2505.02279, 2025

  54. [54]

    Enhancing clinical decision support and ehr insights through llms and the model context protocol: An open-source mcp-fhir framework

    Abul Ehtesham, Aditi Singh, and Saket Kumar. Enhancing clinical decision support and ehr insights through llms and the model context protocol: An open-source mcp-fhir framework. In2025 IEEE World AI IoT Congress (AIIoT), pages 0205–0211. IEEE, 2025

  55. [55]

    Red-teaming medical ai: Systematic adversarial evaluation of llm safety guardrails in clinical contexts.medRxiv, pages 2026–02, 2026

    Tashfeen Ekram. Red-teaming medical ai: Systematic adversarial evaluation of llm safety guardrails in clinical contexts.medRxiv, pages 2026–02, 2026

  56. [56]

    Medical ai consensus: A multi-agent framework for radiology report generation and evaluation.arXiv preprint arXiv:2509.17353, 2025

    Ahmed T Elboardy, Ghada Khoriba, and Essam A Rashed. Medical ai consensus: A multi-agent framework for radiology report generation and evaluation.arXiv preprint arXiv:2509.17353, 2025

  57. [57]

    Agentic large language models for training-free neuro-radiological image analysis.arXiv preprint arXiv:2604.16729, 2026

    Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, and Jan C Peeken. Agentic large language models for training-free neuro-radiological image analysis.arXiv preprint arXiv:2604.16729, 2026

  58. [58]

    Medrax: Medical reasoning agent for chest x-ray.arXiv preprint arXiv:2502.02673, 2025

    Adibvafa Fallahpour, Jun Ma, Alif Munim, Hongwei Lyu, and Bo Wang. Medrax: Medical reasoning agent for chest x-ray.arXiv preprint arXiv:2502.02673, 2025

  59. [59]

    Evolving medical imaging agents via experience-driven self-skill discovery.arXiv preprint arXiv:2603.05860, 2026

    Lin Fan, Pengyu Dai, Zhipeng Deng, Haolin Wang, Xun Gong, Yefeng Zheng, and Yafei Ou. Evolving medical imaging agents via experience-driven self-skill discovery.arXiv preprint arXiv:2603.05860, 2026

  60. [60]

    Medodyssey: A medical domain benchmark for long context evaluation up to 200k tokens

    Yongqi Fan, Hongli Sun, Kui Xue, Xiaofan Zhang, Shaoting Zhang, and Tong Ruan. Medodyssey: A medical domain benchmark for long context evaluation up to 200k tokens. InFindings of the Association for Computational Linguistics: NAACL 2025, pages 32–56, 2025

  61. [61]

    Chestx-reasoner: Advancing radiology foundation models with reasoning through step-by-step verification.arXiv preprint arXiv:2504.20930, 2025

    Ziqing Fan, Cheng Liang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Chestx-reasoner: Advancing radiology foundation models with reasoning through step-by-step verification.arXiv preprint arXiv:2504.20930, 2025

  62. [62]

    A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems.arXiv preprint arXiv:2508.07407, 2025

    Jinyuan Fang, Yanwen Peng, Xi Zhang, Yingxu Wang, Xinhao Yi, Guibin Zhang, Yi Xu, Bin Wu, Siwei Liu, Zihao Li, et al. A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems.arXiv preprint arXiv:2508.07407, 2025

  63. [63]

    Towards general agentic intelligence via environment scaling.arXiv preprint arXiv:2509.13311, 2025

    Runnan Fang, Shihao Cai, Baixuan Li, Jialong Wu, Guangyu Li, Wenbiao Yin, Xinyu Wang, Xiaobin Wang, Liangcai Su, Zhen Zhang, et al. Towards general agentic intelligence via environment scaling.arXiv preprint arXiv:2509.13311, 2025

  64. [64]

    M 3 builder: A multi-agent system for automated machine learning in medical imaging

    Jinghao Feng, Qiaoyu Zheng, Chaoyi Wu, Ziheng Zhao, Ya Zhang, Yanfeng Wang, and Weidi Xie. M 3 builder: A multi-agent system for automated machine learning in medical imaging. InInternational Workshop on Agentic AI for Medicine, pages 115–124. Springer, 2025

  65. [65]

    Levels of autonomy for ai agents.arXiv preprint arXiv:2506.12469, 2025

    Kevin J Feng, David W McDonald, and Amy X Zhang. Levels of autonomy for ai agents.arXiv preprint arXiv:2506.12469, 2025

  66. [66]

    Doctoragent-rl: A multi-agent collaborative reinforcement learning system for multi-turn clinical dialogue

    Yichun Feng, Jiawei Wang, Lu Zhou, Zhen Lei, and Yixue Li. Doctoragent-rl: A multi-agent collaborative reinforcement learning system for multi-turn clinical dialogue. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 16952–16956. IEEE, 2026

  67. [67]

    Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology.Nature cancer, 6(8): 1337–1349, 2025

    Dyke Ferber, Omar SM El Nahhas, Georg Wölflein, Isabella C Wiest, Jan Clusmann, Marie-Elisabeth Leßmann, Sebastian Foersch, Jacqueline Lammert, Maximilian Tschochohei, Dirk Jäger, et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology.Nature cancer, 6(8): 1337–1349, 2025

  68. [68]

    Are: Scaling up agent environments and evaluations.arXiv preprint arXiv:2509.17158, 2025

    Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Laurençon, Maxime Lecanu, et al. Are: Scaling up agent environments and evaluations.arXiv preprint arXiv:2509.17158, 2025

  69. [69]

    A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.arXiv preprint arXiv:2507.21046, 2025

    Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, et al. A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.arXiv preprint arXiv:2507.21046, 2025

  70. [70]

    Pal: Program-aided language models

    Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: Program-aided language models. InInternational conference on machine learning, pages 10764–10799. PMLR, 2023. 35

  71. [71]

    Txagent: An ai agent for therapeutic reasoning across a universe of tools, 2025

    Shanghua Gao, Richard Zhu, Zhenglun Kong, Ayush Noori, Xiaorui Su, Curtis Ginder, Theodoros Tsiligkaridis, and Marinka Zitnik. Txagent: An ai agent for therapeutic reasoning across a universe of tools, 2025. URL https://arxiv.org/abs/2503.10970

  72. [72]

    Enhancing medical visual grounding via knowledge-guided spatial prompts.arXiv preprint arXiv:2604.01915, 2026

    Yifan Gao, Tao Zhou, Yi Zhou, Ke Zou, Yizhe Zhang, and Huazhu Fu. Enhancing medical visual grounding via knowledge-guided spatial prompts.arXiv preprint arXiv:2604.01915, 2026

  73. [73]

    Clinicalagents: Multi-agent orchestration for clinical decision making with dual-memory.arXiv preprint arXiv:2603.26182, 2026

    Zhuohan Ge, Haoyang Li, Yubo Wang, Nicole Hu, Chen Jason Zhang, and Qing Li. Clinicalagents: Multi-agent orchestration for clinical decision making with dual-memory.arXiv preprint arXiv:2603.26182, 2026

  74. [74]

    Improving the safety and trustworthiness of medical ai via multi-agent evaluation loops.arXiv preprint arXiv:2601.13268, 2026

    Zainab Ghafoor, Md Shafiqul Islam, Koushik Howlader, Md Rasel Khondokar, Tanusree Bhattacharjee, Sayantan Chakraborty, Adrito Roy, Ushashi Bhattacharjee, and Tirtho Roy. Improving the safety and trustworthiness of medical ai via multi-agent evaluation loops.arXiv preprint arXiv:2601.13268, 2026

  75. [75]

    Pathfinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology

    Fatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom O Ikezogwo, Beibin Li, Tejoram Vivekanan- dan, Joann G Elmore, Ranjay Krishna, and Linda Shapiro. Pathfinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 234...

  76. [76]

    Meddialogrubrics: A comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models.arXiv preprint arXiv:2601.03023, 2026

    Lecheng Gong, Weimin Fang, Ting Yang, Dongjie Tao, Chunxiao Guo, Peng Wei, Bo Xie, Jinqun Guan, Zixiao Chen, Fang Shi, et al. Meddialogrubrics: A comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models.arXiv preprint arXiv:2601.03023, 2026

  77. [77]

    The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  78. [78]

    Clinical-r1: Empowering large language models for faithful and comprehensive reasoning with clinical objective relative policy optimization.arXiv preprint arXiv:2512.00601, 2025

    Boyang Gu, Hongjian Zhou, Bradley Max Segal, Jinge Wu, Zeyu Cao, Hantao Zhong, Lei Clifton, Fenglin Liu, and David A Clifton. Clinical-r1: Empowering large language models for faithful and comprehensive reasoning with clinical objective relative policy optimization.arXiv preprint arXiv:2512.00601, 2025

  79. [79]

    Medagentaudit: Diagnosing and quantifying collaborative failure modes in medical multi-agent systems.arXiv preprint arXiv:2510.10185, 2025

    Lei Gu, Yinghao Zhu, Haoran Sang, Zixiang Wang, Dehao Sui, Wen Tang, Ewen Harrison, Junyi Gao, Lequan Yu, and Liantao Ma. Medagentaudit: Diagnosing and quantifying collaborative failure modes in medical multi-agent systems.arXiv preprint arXiv:2510.10185, 2025

  80. [80]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025

Showing first 80 references.