REVIEW 3 major objections 6 minor 299 references
Medical agents become clinically trustworthy by scaling the tools and gyms they can reach, and by improving through interaction, not by growing model size alone.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 06:28 UTC pith:JEMXH7JT
load-bearing objection Useful deployment-first survey with a clear spine; the environment-scaling priority is a roadmap claim, not a measured law. the 3 major comments →
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Real-world readiness of a medical agent is amplified along three complementary axes—framework topology, cognitive-loop depth, and reachable environment size—and clinical environment scaling (tools, data, and interactive gyms inside PACS, EHR, and FHIR ecosystems) is the most immediately actionable yet least explored lever, with clinical self-evolution through environment interaction as the destination rather than parameter growth alone.
What carries the argument
The scaling spine: readiness is treated as a function of topology richness N_topo, cognitive-loop depth D_loop, and environment size |E|, so that framework wiring, harness-driven iteration, and tool/data access are orthogonal levers that together lift a passive model toward a self-evolving clinical agent.
Load-bearing premise
That enriching the clinical tool and data environment often buys more real-world readiness per unit effort than further fine-tuning or more complex agent teams, and that the three scaling axes stay largely independent under real hospital constraints.
What would settle it
Hold backbone model and clinical task fixed, systematically vary only the richness of available tools, archives, and interactive gyms across matched deployments, and check whether measured readiness gains track environment size more than topology or loop depth; if they do not, the prioritization of environment scaling fails.
If this is right
- Hospitals and builders should invest first in standardized tool interfaces, clinical gyms, and interoperable data layers rather than only larger models.
- Evaluation must move from static accuracy to multi-turn, tool-using, end-to-end workflow benchmarks with process-level and safety metrics.
- Self-evolving agents that improve from environment interaction become a primary research target rather than a side topic.
- Risk analysis for hallucination, cascade failures, and fairness must be framed as properties of the full orchestration loop, not of a single model.
- The assisted–cooperative–fully autonomous taxonomy becomes a practical language for clinical adoption and regulatory design.
Where Pith is reading between the lines
- If environment scaling is truly the cheapest lever, hospital IT integration and protocol standardization may matter more for agent readiness than another round of parameter growth.
- The partial-observability framing suggests medical agents will borrow more from sequential decision-making and interactive simulators than from classic single-pass medical imaging pipelines.
- Contamination-resistant, multi-level agentic benchmarks could become the de facto gate for clinical AI claims the way static leaderboards once gated perception models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey organizes medical imaging agents from a deployment-first stance rather than a capability-first taxonomy. It formalizes agents as six-module sequential decision systems under partial observability, proposes a three-level autonomy taxonomy (assisted, cooperative, fully autonomous), and threads the literature along a conceptual “scaling spine” of framework topology N_topo, cognitive-loop depth D_loop, and environment size |E| (Eq. 1; §2.3). Clinical environment scaling—tools, data, and gyms in PACS/EHR/FHIR ecosystems—is argued to be the most immediately actionable lever, with clinical self-evolution via environment interaction as the aspirational destination. The manuscript consolidates 300+ references (search window Jan 2022–Jun 2026; disclosed multi-reviewer screening), catalogs benchmarks and systems (Tables 1–3), reviews specialty applications and deployment/regulatory constraints, and maps risks along the agent pipeline.
Significance. If the organizing claims hold as a research roadmap, the paper offers a useful alternative to concurrent capability-centric medical-agent surveys: it puts contamination-resistant benchmarks, interactive clinical gyms, and interoperability first, and it elevates environment scaling and self-evolution as explicit research priorities rather than appendices. Strengths that should be credited include a disclosed retrieval protocol (1,174→942→300+), extensive benchmark and system catalogs (Tables 1–3), an explicit autonomy taxonomy (Fig. 1), a companion GitHub, and consistently cautious language that Level-3 autonomy and self-evolution remain aspirational. The synthesis of 2025–2026 agentic RL, gyms, and protocol work (MCP/FHIR) is timely for the community. The distinctive thesis—that |E| is often the highest-leverage axis for agents already in tool-rich clinical stacks—is significant as a prioritization claim, but its force depends on how carefully it is framed relative to the evidence the survey actually provides.
major comments (3)
- §2.3 and Eq. (1) present readiness as factoring into three “largely independent” / “orthogonal” axes and argue that enriching |E| often buys more readiness per unit effort than further finetuning or topology growth for agents in PACS/EHR/FHIR ecosystems (also Fig. 2; §6.7). This prioritization is the paper’s distinctive load-bearing thesis, yet the manuscript never holds backbone and task fixed while varying only environment richness, nor does it systematically code the systems in Tables 1–3 by (N_topo, D_loop, |E|). In §§4–7, tool enrichment, loop depth, and multi-agent wiring routinely co-vary. Please either (i) reframe the priority of environment scaling as a hypothesis/roadmap claim with explicit falsification criteria, or (ii) add a short comparative synthesis of a subset of cataloged systems that attempts to disentangle the three axes (even qualitatively). Without that, the spine r
- §3.2–3.4 and Table 1 organize benchmarks by autonomy/realism levels and correctly stress Level-2/3 gaps, but the survey’s claim that environment scaling is “underexplored” sits uneasily next to the rapid 2025–2026 arrival of MedAgentGym, Healthcare AI GYM, ClinEnv, MedAgentBench, ABRA, etc., which the paper itself catalogs. Clarify what “underexplored” means operationally (e.g., lack of controlled ablations of tool/data richness; lack of multi-site gyms; lack of reward design for imaging-specific verifiers) so the roadmap does not understate the very literature it surveys while still motivating the proposed lever.
- §2.1 formalizes A=(P,R,L,M,T,F) under partial observability but stops short of a usable decision-process statement (state/belief, action space, observation model, reward/objective). Given that later sections invoke agentic multi-turn RL and verifiable rewards (§6.4–6.6; Table 2) as the training counterpart, a brief explicit mapping from the six modules to a POMDP/belief-MDP (or an honest statement that the formalization is taxonomic rather than operational) would make the “sequential decision making under partial observability” claim load-bearing rather than decorative, and would tighten the link between §2 and the RL/self-evolution material.
minor comments (6)
- Fig. 2 and Eq. (1): label Eq. (1) explicitly as a conceptual readiness map, not a scaling law, to avoid readers treating N_topo, D_loop, |E| as measurable, calibrated quantities.
- Table 1 / Table 3: several 2026 entries list venues as arXiv or workshop; a short note on inclusion criteria for preprints versus peer-reviewed work would help readers weight the catalog.
- §1 “Relation to concurrent surveys”: the claim that none couples a disclosed retrieval protocol with an environment-centered axis is strong; a short comparison table (scope, axes, retrieval protocol, self-evolution coverage) would make the differentiation checkable.
- §5.7–5.8 and §6.6 partially overlap on self-evolution mechanisms (skills, memory, code rewrite). A single cross-reference box stating what is inference-time substrate vs training-time evolution would reduce redundancy.
- Minor consistency: “self-evolving” / “self-evolution” hyphenation and capitalization vary across abstract, keywords, and §5.8/§6.6; normalize.
- §9.2 cites a 76.6% vs 51.3% hallucination result from Kim et al.; state the exact metric and setting in-text so the counterintuitive claim is interpretable without opening the reference.
Circularity Check
No significant circularity: conceptual survey spine, not a fitted or self-definitional derivation.
full rationale
This is a literature survey that organizes medical-agent work along a conceptual “scaling spine” (framework topology N_topo, loop depth D_loop, environment size |E|; Eq. 1, §2.3, Fig. 2) and argues that clinical environment scaling is the most immediately actionable lever, with self-evolution as an aspirational frontier. Eq. 1 is explicitly a compact readiness map (“we write compactly as”), not a calibrated law or a prediction obtained by fitting free parameters to data. The prioritization of |E| is supported by external environment-scaling literature [62, 97], tool-interface standards (MCP/FHIR), and catalogs of benchmarks, gyms, and systems (Tables 1–3; §§3–7), not by a uniqueness theorem or load-bearing self-citation chain that forbids alternatives. Author self-citations (e.g., MedCausalX, MedEyes) appear as surveyed applications, which is normal and non-load-bearing for the spine thesis. There is no self-definitional loop (X defined as Y then “derived”), no fitted input renamed as prediction, no ansatz smuggled in as a forced result, and no renaming of a known empirical law presented as a first-principles derivation. Weaknesses of the thesis (unmeasured orthogonality of axes; co-variation of tools, loop depth, and topology in the surveyed systems) are evidence and correctness issues, not circular reductions. Score 0 with empty steps is the honest finding.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Medical agency is sequential decision-making under partial observability over images, reports, and tool outputs rather than single-pass prediction.
- ad hoc to paper Readiness factors approximately into orthogonal framework, capability-loop, and environment axes (Eq. 1).
- ad hoc to paper Enriching tools/data/interfaces often yields more readiness per unit effort than further parameter scaling for agents already in tool-rich clinical ecosystems.
- domain assumption A three-level assisted/cooperative/fully-autonomous taxonomy is an appropriate transfer from driving and surgical robotics to medical agents.
- domain assumption Self-evolution via environment interaction is a primary long-term path to trustworthy clinical agents beyond frozen foundation models.
invented entities (3)
-
Scaling spine (N_topo, D_loop, |E| readiness map)
no independent evidence
-
Three-level medical-agent autonomy taxonomy (assisted/cooperative/fully autonomous)
no independent evidence
-
Clinical environment scaling as primary actionable lever
no independent evidence
read the original abstract
The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments. This work departs from the capability first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision making systems under partial observability, together with a three level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self improving agents, agent gyms, and test time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this work provides a roadmap toward trustworthy, self improving medical imaging systems for real clinical practice.
Reference graph
Works this paper leans on
-
[1]
Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023
Pith/arXiv arXiv 2023
-
[2]
Longhealth: A question answering benchmark with long clinical documents.Journal of Healthcare Informatics Research, 9(3):280–296, 2025
Lisa Adams, Felix Busch, Tianyu Han, Jean-Baptiste Excoffier, Matthieu Ortala, Alexander Löser, Hugo JWL Aerts, Jakob Nikolas Kather, Daniel Truhn, and Keno Bressem. Longhealth: A question answering benchmark with long clinical documents.Journal of Healthcare Informatics Research, 9(3):280–296, 2025
2025
-
[3]
Arman Aghaee, Sepehr Asgarian, and Jouhyun Jeon. Synthagent: A multi-agent llm framework for realistic patient simulation–a case study in obesity with mental health comorbidities.arXiv preprint arXiv:2602.08254, 2026
arXiv 2026
-
[4]
Oaagent: Multimodal llm agent for predicting knee osteoarthritis progression
Pegah Ahadian, Mingrui Yang, Eva Powlison, Xiaojuan Li, Wei Xu, and Qiang Guan. Oaagent: Multimodal llm agent for predicting knee osteoarthritis progression. InProceedings of the ACM/IEEE International Conference on Connected Health: Applications, Systems and Engineering Technologies, pages 144–148, 2025
2025
-
[5]
Mohammad Almansoori, Komal Kumar, and Hisham Cholakkal. Self-evolving multi-agent simulations for realistic clinical interactions.arXiv preprint arXiv:2503.22678, 2025
arXiv 2025
-
[6]
Effective context engineering for AI agents
Anthropic. Effective context engineering for AI agents. Anthropic Engineering Blog, 2025. https://www. anthropic.com/engineering/effective-context-engineering-for-ai-agents, accessed 2026-06-14
2025
-
[7]
Effective harnesses for long-running agents
Anthropic. Effective harnesses for long-running agents. Anthropic Engineering Blog, 2025. https://www. anthropic.com/engineering/effective-harnesses-for-long-running-agents, accessed 2026-06-14
2025
-
[8]
Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, et al. Healthbench: Evaluating large language models towards improved human health.arXiv preprint arXiv:2505.08775, 2025
Pith/arXiv arXiv 2025
-
[9]
Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Pith/arXiv arXiv 2025
-
[10]
Agentic ai in healthcare: A comprehensive survey of foundations, taxonomy, and applications.Authorea Preprints, 2025
Shruti Banerjie, Yuxin Zhu, Isaac Freeman, Julyssa Villa Machado, Abdulaziz Ahmed, Abeed Sarker, and Mohammed Al-Garadi. Agentic ai in healthcare: A comprehensive survey of foundations, taxonomy, and applications.Authorea Preprints, 2025
2025
-
[11]
Magda: Multi-agent guideline-driven diagnostic assistance
David Bani-Harouni, Nassir Navab, and Matthias Keicher. Magda: Multi-agent guideline-driven diagnostic assistance. InInternational workshop on foundation models for general medical AI, pages 163–172. Springer, 2024
2024
-
[12]
Maira-2: Grounded radiology report generation.arXiv preprint arXiv:2406.04449, 2024
Shruthi Bannur, Kenza Bouzid, Daniel C Castro, Anton Schwaighofer, Anja Thieme, Sam Bond-Taylor, Maxim- ilian Ilse, Fernando Pérez-García, Valentina Salvatelli, Harshita Sharma, et al. Maira-2: Grounded radiology report generation.arXiv preprint arXiv:2406.04449, 2024
Pith/arXiv arXiv 2024
-
[13]
Radgpt: Constructing 3d image-text tumor datasets
Pedro RAS Bassi, Mehmet Can Yavuz, Ibrahim Ethem Hamamci, Sezgin Er, Xiaoxi Chen, Wenxuan Li, Bjoern Menze, Sergio Decherchi, Andrea Cavalli, Kang Wang, et al. Radgpt: Constructing 3d image-text tumor datasets. InProceedings of the IEEE/CVF international conference on computer vision, pages 23720–23730, 2025
2025
-
[14]
Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow, Juan M Banda, Nikesh Kotecha, Timothy Keyes, Yifan Mai, Mert Oez, et al. Medhelm: Holistic evaluation of large language models for medical tasks.arXiv preprint arXiv:2505.23802, 2025
Pith/arXiv arXiv 2025
-
[15]
Suhana Bedi, Ryan Welch, Ethan Steinberg, Michael Wornow, Taeil Matthew Kim, Haroun Ahmed, Peter Sterling, Bravim Purohit, Qurat Akram, Angelic Acosta, et al. Healthadminbench: Evaluating computer-use agents on healthcare administration tasks.arXiv preprint arXiv:2604.09937, 2026
Pith/arXiv arXiv 2026
-
[16]
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. Graph of thoughts: Solving elaborate problems with large language models. InProceedings of the AAAI conference on artificial intelligence, volume 38, pages 17682–17690, 2024
2024
-
[17]
Christian Bluethgen, Dave Van Veen, Daniel Truhn, Jakob Nikolas Kather, Michael Moor, Malgorzata Polacin, Akshay Chaudhari, Thomas Frauenfelder, Curtis P Langlotz, Michael Krauthammer, et al. Agentic systems in radiology: Design, applications, evaluation, and challenges.arXiv preprint arXiv:2510.09404, 2025. 32
arXiv 2025
-
[18]
Making the most of text semantics to improve biomedical vision–language processing
Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. Making the most of text semantics to improve biomedical vision–language processing. InEuropean conference on computer vision, pages 1–21. Springer, 2022
2022
-
[19]
Visual alignment of medical vision-language models for grounded radiology report generation
Sarosij Bose, Ravi K Rajendran, Biplob Debnath, Konstantinos Karydis, Amit K Roy-Chowdhury, and Srimat Chakradhar. Visual alignment of medical vision-language models for grounded radiology report generation. arXiv preprint arXiv:2512.16201, 2025
arXiv 2025
-
[20]
Anjila Budathoki and Manish Dhakal. Adversarial robustness analysis of vision-language models in medical image segmentation.arXiv preprint arXiv:2505.02971, 2025
Pith/arXiv arXiv 2025
-
[21]
Xin-Qiang Cai, Wei Wang, Feng Liu, Tongliang Liu, Gang Niu, and Masashi Sugiyama. Reinforcement learning with verifiable yet noisy rewards under imperfect verifiers.arXiv preprint arXiv:2510.00915, 2025
Pith/arXiv arXiv 2025
-
[22]
Ruike Cao, Shaojie Bai, Fugen Yao, Liang Dong, Jian Xu, and Li Xiao. Atpo: Adaptive tree policy optimization for multi-turn medical dialogue.arXiv preprint arXiv:2603.02216, 2026
arXiv 2026
-
[23]
Predetermined change control plans: Guiding principles for advancing safe, effective, and high-quality ai-ml technologies.JMIR AI, 4: e76854, 2025
Eduardo Carvalho, Miguel Mascarenhas, Francisca Pinheiro, Ricardo Correia, Sandra Balseiro, Guilherme Barbosa, Ana Guerra, Dulce Oliveira, Rita Moura, André Martins dos Santos, et al. Predetermined change control plans: Guiding principles for advancing safe, effective, and high-quality ai-ml technologies.JMIR AI, 4: e76854, 2025
2025
-
[24]
Juan Manuel Zambrano Chaves, Shih-Cheng Huang, Yanbo Xu, Hanwen Xu, Naoto Usuyama, Sheng Zhang, Fei Wang, Yujia Xie, Mahmoud Khademi, Ziyi Yang, et al. Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation.arXiv preprint arXiv:2403.08002, 2024
Pith/arXiv arXiv 2024
-
[25]
Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, et al. Minimax-m1: Scaling test-time compute efficiently with lightning attention.arXiv preprint arXiv:2506.13585, 2025
Pith/arXiv arXiv 2025
-
[26]
Canyu Chen, Jian Yu, Shan Chen, Che Liu, Zhongwei Wan, Danielle Bitterman, Fei Wang, and Kai Shu. Clinicalbench: Can llms beat traditional ml models in clinical prediction?arXiv preprint arXiv:2411.06469, 2024
Pith/arXiv arXiv 2024
-
[27]
Haolin Chen, Deon Metelski, Leon Qi, Tao Xia, Joonyul Lee, Steve Brown, Kevin Riley, Frank Wang, TY Liu, Hank Capps MD, et al. Chi-bench: Can ai agents automate end-to-end, long-horizon, policy-rich healthcare workflows?arXiv preprint arXiv:2605.16679, 2026
Pith/arXiv arXiv 2026
-
[28]
Ethical machine learning in healthcare.Annual review of biomedical data science, 4(1):123–144, 2021
Irene Y Chen, Emma Pierson, Sherri Rose, Shalmali Joshi, Kadija Ferryman, and Marzyeh Ghassemi. Ethical machine learning in healthcare.Annual review of biomedical data science, 4(1):123–144, 2021
2021
-
[29]
Jingyun Chen, Linghan Cai, Zhikang Wang, Yi Huang, Songhan Jiang, Shenjin Huang, Hongpeng Wang, and Yongbing Zhang. Pathagent: Toward interpretable analysis of whole-slide pathology images via large language model-based agentic reasoning.arXiv preprint arXiv:2511.17052, 2025
arXiv 2025
-
[30]
Kai Chen, Xinfeng Li, Tianpei Yang, Hewei Wang, Wei Dong, and Yang Gao. Mdteamgpt: A self-evolving llm- based multi-agent framework for multi-disciplinary team medical consultation.arXiv preprint arXiv:2503.13856, 2025
Pith/arXiv arXiv 2025
-
[31]
Kai Chen, Taihang Zhen, Hewei Wang, Kailai Liu, Xinfeng Li, Jing Huo, Tianpei Yang, Jinfeng Xu, Wei Dong, and Yang Gao. Medsentry: Understanding and mitigating safety risks in medical llm multi-agent systems.arXiv preprint arXiv:2505.20824, 2025
Pith/arXiv arXiv 2025
-
[32]
Towards a general-purpose foundation model for computational pathology.Nature medicine, 30(3):850–862, 2024
Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathology.Nature medicine, 30(3):850–862, 2024
2024
-
[33]
Medbrowsecomp: Benchmarking medical deep research and computer use
Shan Chen, Pedro Moreira, Yuxin Xiao, Sam Schmidgall, Jeremy Warner, Hugo Aerts, Thomas Hartvigsen, Jack Gallifant, and Danielle S Bitterman. Medbrowsecomp: Benchmarking medical deep research and computer use. arXiv preprint arXiv:2505.14963, 2025
Pith/arXiv arXiv 2025
-
[34]
Wenting Chen, Yi Dong, Zhaojun Ding, Yucheng Shi, Yifan Zhou, Fang Zeng, Yijun Luo, Tianyu Lin, Yihang Su, Yichen Wu, et al. Radfabric: Agentic ai system with reasoning capability for radiology.arXiv preprint arXiv:2506.14142, 2025. 33
Pith/arXiv arXiv 2025
-
[35]
Yinzhu Chen, Abdine Maiga, Hossein A Rahmani, and Emine Yilmaz. Automated rubrics for reliable evaluation of medical dialogue systems.arXiv preprint arXiv:2601.15161, 2026
Pith/arXiv arXiv 2026
-
[36]
Meissa: Multi-modal medical agentic intelligence.arXiv preprint arXiv:2603.09018, 2026
Yixiong Chen, Xinyi Bai, Yue Pan, Zongwei Zhou, and Alan Yuille. Meissa: Multi-modal medical agentic intelligence.arXiv preprint arXiv:2603.09018, 2026
arXiv 2026
-
[37]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24185–24198, 2024
2024
-
[38]
Chexagent: Towards a foundation model for chest x-ray interpretation
Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck, Magdalini Paschali, Louis Blankemeier, Dave Van Veen, Jeya Maria Jose Valanarasu, Alaa Youssef, Joseph Paul Cohen, Eduardo Pontes Reis, et al. Chexagent: Towards a foundation model for chest x-ray interpretation. InAAAI 2024 Spring Symposium on Clinical Foundation Models, 2024
2024
-
[39]
Engelhard, Somesh Jha, Anivarya Kumar, and David Page
Jihye Choi, Nils Palumbo, Prasad Chalasani, Matthew M. Engelhard, Somesh Jha, Anivarya Kumar, and David Page. Malade: Orchestration of llm-powered agents with retrieval augmented generation for pharmacovigilance,
-
[40]
URLhttps://arxiv.org/abs/2408.01869
-
[41]
Tim Cofala, Christian Kalfar, Jingge Xiao, Johanna Schrader, Michelle Tang, and Wolfgang Nejdl. Medai: Evaluating txagent’s therapeutic agentic reasoning in the neurips cure-bench competition.arXiv preprint arXiv:2512.11682, 2025
arXiv 2025
-
[42]
SAE international, 2021
On-Road Automated Driving (ORAD) Committee.Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles. SAE international, 2021
2021
-
[43]
Alexandra DeLucia, Heyuan Huang, Sonal Joshi, Mahsa Yarmohammadi, Ahmed Hassoon, and Mark Dredze. Same verdict, different reasons: Llm-as-a-judge and clinician disagreement on medical chatbot completeness. arXiv preprint arXiv:2604.16383, 2026
Pith/arXiv arXiv 2026
-
[44]
Medka: A knowledge graph-augmented approach to improve factuality in medical large language models.Journal of Biomedical Informatics, page 104871, 2025
Yiyan Deng, Shen Zhao, Yongming Miao, Junjie Zhu, and Jin Li. Medka: A knowledge graph-augmented approach to improve factuality in medical large language models.Journal of Biomedical Informatics, page 104871, 2025
2025
-
[45]
Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruipérez-Campillo, Moritz Vandenhirtz, Sonia Laguna, Alain Ryser, Koji Fujimoto, Mizuho Nishio, Thomas M Sutter, Julia E Vogt, et al. Radvlm: a multitask conversational vision-language model for radiology.arXiv preprint arXiv:2502.03333, 2025
arXiv 2025
-
[46]
Agentic ai in radiology: emerging potential and unresolved challenges.British Journal of Radiology, 98(1174):1582–1584, 2025
Nicholas Dietrich. Agentic ai in radiology: emerging potential and unresolved challenges.British Journal of Radiology, 98(1174):1582–1584, 2025
2025
-
[47]
Yifeng Ding, Hung Le, Songyang Han, Kangrui Ruan, Zhenghui Jin, Varun Kumar, Zijian Wang, and Anoop Deoras. Empowering multi-turn tool-integrated agentic reasoning with group turn policy optimization.arXiv preprint arXiv:2511.14846, 2025
Pith/arXiv arXiv 2025
-
[48]
Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025
Guanting Dong, Licheng Bao, Zhongyuan Wang, Kangzhi Zhao, Xiaoxi Li, Jiajie Jin, Jinghan Yang, Hangyu Mao, Fuzheng Zhang, Kun Gai, et al. Agentic entropy-balanced policy optimization.arXiv preprint arXiv:2510.14545, 2025
arXiv 2025
-
[49]
Guanting Dong, Junting Lu, Junjie Huang, Wanjun Zhong, Longxiang Liu, Shijue Huang, Zhenyu Li, Yang Zhao, Xiaoshuai Song, Xiaoxi Li, et al. Agent-world: Scaling real-world environment synthesis for evolving general agent intelligence.arXiv preprint arXiv:2604.18292, 2026
Pith/arXiv arXiv 2026
-
[50]
Xuanzhao Dong, Wenhui Zhu, Xiwen Chen, Hao Wang, Xin Li, Yujian Xiong, Jiajun Cheng, Jingjing Wang, Xiaobing Yu, Haiyu Wu, et al. Ophin-500k: Curating web-scale visual instructions for scaling ophthalmic multimodal large language models.arXiv preprint arXiv:2605.27916, 2026
Pith/arXiv arXiv 2026
-
[51]
Yuexi Du, Jinglu Wang, Shujie Liu, Nicha C Dvornek, and Yan Lu. Care: Towards clinical accountability in multi-modal medical reasoning with an evidence-grounded agentic framework.arXiv preprint arXiv:2603.01607, 2026
arXiv 2026
-
[52]
Llms can simulate standardized patients via agent coevolution
Zhuoyun Du, LujieZheng LujieZheng, Renjun Hu, Yuyang Xu, Xiawei Li, Ying Sun, Wei Chen, Jian Wu, Haolei Cai, and Haochao Ying. Llms can simulate standardized patients via agent coevolution. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17278–17306, 2025. 34
2025
-
[53]
Abul Ehtesham, Aditi Singh, Gaurav Kumar Gupta, and Saket Kumar. A survey of agent interoperability protocols: Model context protocol (mcp), agent communication protocol (acp), agent-to-agent protocol (a2a), and agent network protocol (anp).arXiv preprint arXiv:2505.02279, 2025
Pith/arXiv arXiv 2025
-
[54]
Enhancing clinical decision support and ehr insights through llms and the model context protocol: An open-source mcp-fhir framework
Abul Ehtesham, Aditi Singh, and Saket Kumar. Enhancing clinical decision support and ehr insights through llms and the model context protocol: An open-source mcp-fhir framework. In2025 IEEE World AI IoT Congress (AIIoT), pages 0205–0211. IEEE, 2025
2025
-
[55]
Red-teaming medical ai: Systematic adversarial evaluation of llm safety guardrails in clinical contexts.medRxiv, pages 2026–02, 2026
Tashfeen Ekram. Red-teaming medical ai: Systematic adversarial evaluation of llm safety guardrails in clinical contexts.medRxiv, pages 2026–02, 2026
2026
-
[56]
Ahmed T Elboardy, Ghada Khoriba, and Essam A Rashed. Medical ai consensus: A multi-agent framework for radiology report generation and evaluation.arXiv preprint arXiv:2509.17353, 2025
arXiv 2025
-
[57]
Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, and Jan C Peeken. Agentic large language models for training-free neuro-radiological image analysis.arXiv preprint arXiv:2604.16729, 2026
Pith/arXiv arXiv 2026
-
[58]
Medrax: Medical reasoning agent for chest x-ray.arXiv preprint arXiv:2502.02673, 2025
Adibvafa Fallahpour, Jun Ma, Alif Munim, Hongwei Lyu, and Bo Wang. Medrax: Medical reasoning agent for chest x-ray.arXiv preprint arXiv:2502.02673, 2025
Pith/arXiv arXiv 2025
-
[59]
Lin Fan, Pengyu Dai, Zhipeng Deng, Haolin Wang, Xun Gong, Yefeng Zheng, and Yafei Ou. Evolving medical imaging agents via experience-driven self-skill discovery.arXiv preprint arXiv:2603.05860, 2026
arXiv 2026
-
[60]
Medodyssey: A medical domain benchmark for long context evaluation up to 200k tokens
Yongqi Fan, Hongli Sun, Kui Xue, Xiaofan Zhang, Shaoting Zhang, and Tong Ruan. Medodyssey: A medical domain benchmark for long context evaluation up to 200k tokens. InFindings of the Association for Computational Linguistics: NAACL 2025, pages 32–56, 2025
2025
-
[61]
Ziqing Fan, Cheng Liang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Chestx-reasoner: Advancing radiology foundation models with reasoning through step-by-step verification.arXiv preprint arXiv:2504.20930, 2025
Pith/arXiv arXiv 2025
-
[62]
Jinyuan Fang, Yanwen Peng, Xi Zhang, Yingxu Wang, Xinhao Yi, Guibin Zhang, Yi Xu, Bin Wu, Siwei Liu, Zihao Li, et al. A comprehensive survey of self-evolving ai agents: A new paradigm bridging foundation models and lifelong agentic systems.arXiv preprint arXiv:2508.07407, 2025
Pith/arXiv arXiv 2025
-
[63]
Towards general agentic intelligence via environment scaling.arXiv preprint arXiv:2509.13311, 2025
Runnan Fang, Shihao Cai, Baixuan Li, Jialong Wu, Guangyu Li, Wenbiao Yin, Xinyu Wang, Xiaobin Wang, Liangcai Su, Zhen Zhang, et al. Towards general agentic intelligence via environment scaling.arXiv preprint arXiv:2509.13311, 2025
arXiv 2025
-
[64]
M 3 builder: A multi-agent system for automated machine learning in medical imaging
Jinghao Feng, Qiaoyu Zheng, Chaoyi Wu, Ziheng Zhao, Ya Zhang, Yanfeng Wang, and Weidi Xie. M 3 builder: A multi-agent system for automated machine learning in medical imaging. InInternational Workshop on Agentic AI for Medicine, pages 115–124. Springer, 2025
2025
-
[65]
Levels of autonomy for ai agents.arXiv preprint arXiv:2506.12469, 2025
Kevin J Feng, David W McDonald, and Amy X Zhang. Levels of autonomy for ai agents.arXiv preprint arXiv:2506.12469, 2025
Pith/arXiv arXiv 2025
-
[66]
Doctoragent-rl: A multi-agent collaborative reinforcement learning system for multi-turn clinical dialogue
Yichun Feng, Jiawei Wang, Lu Zhou, Zhen Lei, and Yixue Li. Doctoragent-rl: A multi-agent collaborative reinforcement learning system for multi-turn clinical dialogue. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 16952–16956. IEEE, 2026
2026
-
[67]
Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology.Nature cancer, 6(8): 1337–1349, 2025
Dyke Ferber, Omar SM El Nahhas, Georg Wölflein, Isabella C Wiest, Jan Clusmann, Marie-Elisabeth Leßmann, Sebastian Foersch, Jacqueline Lammert, Maximilian Tschochohei, Dirk Jäger, et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology.Nature cancer, 6(8): 1337–1349, 2025
2025
-
[68]
Are: Scaling up agent environments and evaluations.arXiv preprint arXiv:2509.17158, 2025
Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Laurençon, Maxime Lecanu, et al. Are: Scaling up agent environments and evaluations.arXiv preprint arXiv:2509.17158, 2025
arXiv 2025
-
[69]
Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, et al. A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.arXiv preprint arXiv:2507.21046, 2025
Pith/arXiv arXiv 2025
-
[70]
Pal: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: Program-aided language models. InInternational conference on machine learning, pages 10764–10799. PMLR, 2023. 35
2023
-
[71]
Txagent: An ai agent for therapeutic reasoning across a universe of tools, 2025
Shanghua Gao, Richard Zhu, Zhenglun Kong, Ayush Noori, Xiaorui Su, Curtis Ginder, Theodoros Tsiligkaridis, and Marinka Zitnik. Txagent: An ai agent for therapeutic reasoning across a universe of tools, 2025. URL https://arxiv.org/abs/2503.10970
Pith/arXiv arXiv 2025
-
[72]
Yifan Gao, Tao Zhou, Yi Zhou, Ke Zou, Yizhe Zhang, and Huazhu Fu. Enhancing medical visual grounding via knowledge-guided spatial prompts.arXiv preprint arXiv:2604.01915, 2026
arXiv 2026
-
[73]
Zhuohan Ge, Haoyang Li, Yubo Wang, Nicole Hu, Chen Jason Zhang, and Qing Li. Clinicalagents: Multi-agent orchestration for clinical decision making with dual-memory.arXiv preprint arXiv:2603.26182, 2026
Pith/arXiv arXiv 2026
-
[74]
Zainab Ghafoor, Md Shafiqul Islam, Koushik Howlader, Md Rasel Khondokar, Tanusree Bhattacharjee, Sayantan Chakraborty, Adrito Roy, Ushashi Bhattacharjee, and Tirtho Roy. Improving the safety and trustworthiness of medical ai via multi-agent evaluation loops.arXiv preprint arXiv:2601.13268, 2026
arXiv 2026
-
[75]
Pathfinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology
Fatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom O Ikezogwo, Beibin Li, Tejoram Vivekanan- dan, Joann G Elmore, Ranjay Krishna, and Linda Shapiro. Pathfinder: A multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 234...
2025
-
[76]
Lecheng Gong, Weimin Fang, Ting Yang, Dongjie Tao, Chunxiao Guo, Peng Wei, Bo Xie, Jinqun Guan, Zixiao Chen, Fang Shi, et al. Meddialogrubrics: A comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models.arXiv preprint arXiv:2601.03023, 2026
arXiv 2026
-
[77]
The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Pith/arXiv arXiv 2024
-
[78]
Boyang Gu, Hongjian Zhou, Bradley Max Segal, Jinge Wu, Zeyu Cao, Hantao Zhong, Lei Clifton, Fenglin Liu, and David A Clifton. Clinical-r1: Empowering large language models for faithful and comprehensive reasoning with clinical objective relative policy optimization.arXiv preprint arXiv:2512.00601, 2025
arXiv 2025
-
[79]
Lei Gu, Yinghao Zhu, Haoran Sang, Zixiang Wang, Dehao Sui, Wen Tang, Ewen Harrison, Junyi Gao, Lequan Yu, and Liantao Ma. Medagentaudit: Diagnosing and quantifying collaborative failure modes in medical multi-agent systems.arXiv preprint arXiv:2510.10185, 2025
Pith/arXiv arXiv 2025
-
[80]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.