Pith. sign in

REVIEW 4 major objections 6 minor 24 references

A scoping review of 557 studies concludes that medical agentic AI is an emerging technical paradigm rather than a mature clinical solution, because most current evidence comes from benchmarks, simulations, retrospective data, and small expe

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 02:15 UTC pith:TBCR7BAD

load-bearing objection A useful synthesis whose central distributional claim is asserted, not shown—worth refereeing, but only after the evidence map becomes auditable. the 4 major comments →

arxiv 2607.25489 v1 pith:TBCR7BAD submitted 2026-07-28 cs.CV

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

classification cs.CV
keywords agentic AImedical AIlarge language modelsmulti-agent systemsclinical validationevidence mappingscoping reviewclinical translation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Medical agentic AI systems — which plan, use tools, retrieve knowledge, remember context, and coordinate specialized agents around a clinical goal — are often discussed as if they were ready for the clinic. This review tries to establish what the evidence actually supports. It screened 1,649 exportable records and provisionally mapped 557 studies, classifying them by architecture, application, and how close each evaluation comes to real clinical use. The central finding is that most studies sit at low validation maturity: they rely on public benchmarks, simulated environments, retrospective datasets, or small-scale expert judgment, while process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are measured inconsistently. The sympathetic reading is that the field shows genuine technical promise, especially in medical imaging, but should be treated as an emerging research paradigm rather than a deployable clinical solution.

Core claim

The paper's central claim is that current medical agentic AI is better understood as a system-level orchestration paradigm than as a mature clinical technology. On its own terms, the review demonstrates this by mapping 557 studies across four architectural families — single-agent tool use, multi-agent collaboration, knowledge-augmented agents, and multimodal medical agents — and ranking each study's validation setting on a five-level maturity ladder from static benchmarks to prospective workflows. The decisive observation is that the evidence concentrates in the lower rungs: public datasets, simulated environments, retrospective records, and small expert evaluations dominate, while prospecti

What carries the argument

The analytical engine is the evidence map itself. Eligibility was defined functionally — a system counted as agentic only if it showed goal-directed multistep execution plus at least one explicit agentic mechanism such as planning, tool use, environment interaction, memory, feedback-based refinement, or multi-agent collaboration — rather than by whether it used the word 'agent.' Each included study was then classified along two axes: an agentic complexity score (0–8, counting planning, tool use, retrieval, memory, reflection, multi-agent collaboration, multimodal processing, and workflow integration) and a five-level clinical validation maturity scale (static benchmark, agent benchmark, retr

Load-bearing premise

The review's map is only as reliable as its assumption that the 557 included studies were correctly identified and classified — which depends on the 445 Europe PMC records that could not be exported or screened not biasing the evidence pool, and on screening and extraction performed largely by a single author being accurate.

What would settle it

Re-run the selection with the 445 unscreened Europe PMC records included and check whether the distribution of validation levels shifts; if a meaningful share of those studies had prospective or external clinical validation, the conclusion that the evidence base is dominated by low-maturity evaluations would weaken. A second check: if a multi-institutional prospective study of an agentic medical system demonstrated clear workflow benefit and safety, the claim that agentic AI is not a mature clinical solution would need to be updated.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Public-benchmark accuracy should no longer be treated as evidence of clinical readiness; evaluation must also report tool-selection reliability, code-execution success, retrieval relevance, error recovery, abstention and escalation behavior, and human revision burden.
  • Because the four architecture families are not mutually exclusive and many systems combine them, the reliability of an agentic system depends on the whole execution chain, not on the underlying foundation model alone.
  • In medical imaging, future systems must demonstrate visual grounding and cross-modal consistency — generated text must be traceable to identifiable image findings — and be validated across institutions and acquisition protocols.
  • Clinical translation should proceed through staged evidence generation: external retrospective validation, prospective silent testing, controlled workflow studies, and post-deployment surveillance, with reporting guided by study-design-appropriate standards.
  • Current medical agentic systems should be designed as supervised workflows with bounded action spaces, evidence provenance records, uncertainty-based abstention, and mandatory clinician review, not as autonomous decision-makers.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The mapping implies a sharper regulatory hypothesis than the authors state: if benchmark-level evidence cannot support clinical claims, then an agentic system's 'indication' should be defined by its workflow and oversight boundaries, not by the underlying model's benchmark score.
  • The finding that multi-agent benchmarks do not show consistent superiority over strong single-model baselines suggests that the field may currently over-invest in multi-agent collaboration; a testable extension is to compare single-agent and multi-agent versions of the same task with matched compute and measurement of error propagation.
  • A practical extension of the five-level maturity scale would be to require a minimal reporting checklist — tool invocations, retrieved sources, abstention decisions, clinician overrides, failure cases — before a study can be assigned to a validation level; this would make future evidence maps more reproducible.
  • As simulated EHR environments and virtual hospitals become more realistic, they could serve as an inexpensive screening stage that selects only the most promising agentic systems for prospective clinical studies, reducing the cost and risk of clinical translation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This scoping review with systematic evidence mapping characterizes the emerging field of agentic AI in medicine. After searching five electronic sources, the authors screened 1,649 exportable records and provisionally included 557 unique studies. The paper proposes a taxonomy of four architectural families (single-agent tool-use, multi-agent collaboration, knowledge-augmented, multimodal medical), defines a five-level clinical-validation maturity scale, and maps representative systems by agentic complexity and validation level. The central claim is that the evidence base is dominated by public benchmarks, simulated settings, retrospective data, and small-scale expert evaluation, and therefore medical agentic AI should be regarded as an emerging technical paradigm rather than a mature clinical solution. The discussion and conclusions call for prospective, workflow-integrated validation, structured evidence traceability, and clearer reporting standards.

Significance. If the distributional claim is correct, the paper provides a useful and timely synthesis that could help align research priorities with clinical translation needs. The review is explicitly cautious: the agentic complexity score (0–8) is repeatedly disclaimed as descriptive rather than validated, and the validation-maturity categories are presented as an analytical framing rather than a measured scale. The inclusion of formal concepts such as the POMDP formulation (Eq. 1), expected calibration error (Eq. 2), and decision curve analysis (Eq. 3) strengthens the conceptual grounding. The paper also gives credit to prior reviews and distinguishes enabling components (e.g., vision–language models, segmentation tools) from complete agentic workflows. However, the central empirical premise—the distribution of evidence across validation levels—is never quantified from the 557-study set, and the review's reproducibility is limited by unscreened Europe PMC records and single-author screening. These issues are acknowledged in the text but are load-bearing for the main translational message.

major comments (4)
  1. [§3.1, §3.2, Fig. 4] The central claim that the evidence base is 'dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation' is never supported by counts of the 557 included studies across the five validation levels defined in §2.6 or the six evaluation dimensions of §3.7. Figure 4 plots only a handful of representative systems and explicitly disclaims ranking; no table or supplementary file lists the included studies with their assigned classifications. As a result, the distributional premise that drives the conclusion 'emerging technical paradigm rather than a mature clinical solution' is not auditable from the paper alone. Please add a study-level evidence map or at least aggregate counts by validation level and evaluation dimension.
  2. [§2.4, §4.4] The unscreened Europe PMC records (n=445) are acknowledged as a limitation, but their potential impact on the 557-study set and on the distributional claim is not assessed. If these records disproportionately contain retrospective or benchmark studies, the 'dominated by' conclusion might be robust; if they disproportionately contain more mature evaluations, the conclusion could weaken. Please report the source-specific record counts, and either provide a sensitivity analysis based on title-level screening of Europe PMC records or justify why overlap with PubMed makes the unscreened set negligible.
  3. [Author Contributions, §2.6] Study screening and data extraction were performed by a single author. This creates a risk of systematic misclassification of validation level and agentic complexity, which directly affects the evidence map. Even a small random sample independently screened by a second reviewer, with inter-rater agreement reported, would substantially improve confidence in the classification. The current description does not include any reliability check.
  4. [§2.3, §4.4] Eligibility and classification rely on each primary study's self-reported workflow. The paper notes that under-reporting may influence assigned categories, but it does not state how missing information was handled (e.g., whether authors were contacted, or whether a 'not reported' category was used). This matters because a study that omits details of memory or feedback mechanisms could be misclassified as less agentic, biasing the complexity–validation map. Please specify the coding rules for missing or ambiguous information and report how many studies were affected.
minor comments (6)
  1. [Fig. 4] CXR-Agent is included in the figure as a contextual example but was explicitly not part of the formal evidence-mapping set (§3.6.1). To avoid confusion, the figure caption should state this more prominently, or the point should be visually distinguished from the included studies.
  2. [Table 1] Table 1 refers to a 'five-stage validation setting' while the text (§2.6) describes 'five ordered levels.' Use consistent terminology to prevent ambiguity.
  3. [Eq. (2)] In the ECE definition, the notation acc(B_m) and conf(B_m) is defined, but the formula would benefit from a brief note that the bins are usually equal-width confidence intervals. This is implied by the reference to binning, but it would aid readers not familiar with the literature.
  4. [Eq. (3)] Decision curve analysis: the equation uses TP(pt) and FP(pt) as counts, which is clear, but the sentence 'The term pt/(1−pt) represents the relative harm...' is a bit terse. Consider adding one sentence that the net benefit is plotted against threshold probability and compared with 'treat all' and 'treat none' strategies, as stated later.
  5. [Data availability] The data availability statement says 'No new datasets were generated or analyzed.' However, the systematic evidence-mapping set (557 studies with assigned classifications) is a new dataset that should be made available as a supplementary table to support reproducibility and future updates.
  6. [§2.2] The search syntax is described conceptually but the exact queries for each source are not reported. For a reproducible scoping review, provide the full search strings for at least one representative database, or include them as an appendix.

Circularity Check

0 steps flagged

No circular derivation: the review's claims are analytical summaries of an external evidence base, not predictions fitted from its own taxonomy.

full rationale

This is a scoping review with narrative synthesis and systematic categorical mapping, not a derivation or prediction chain. The central claim ('The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation') is a summary characterization of the 557 included studies. Nothing in the paper fits a parameter to data and then predicts that same data; the taxonomy (agentic capability, architecture, validation maturity) is explicitly described as descriptive and non-validated ('The score was used solely for comparative visualization and evidence mapping and was not treated as a validated measure'; §2.6, §4.4). The evidence map (Figure 4) plots representative systems and disclaims ranking. The paper's definitions of Agentic AI and its capabilities are drawn from external surveys and prior work (refs 22-24), not from the included studies' outcomes, so there is no self-definitional prediction. Self-citations are not present among the load-bearing justifications: every cited prior result (e.g., ReAct, AutoGen, MedAgents, EHRAgent, MALADE) is external to the authors (Tong et al.) and the paper does not import any 'uniqueness theorem' from the authors' prior work. The strongest critique of the paper is an evidentiary/empirical one—the marginal distribution of the 557 studies across validation levels is never tabled, and Europe PMC records were unscreened (Methods §2.4; Results §3.1; Limitations §4.4). That is a completeness and falsifiability concern, not a circularity concern. The recommendation to prioritize prospective validation is a normative inference from the (claimed but unquantified) distribution, and it does not reduce to a fitted input by the paper's own equations. Equally, the short quantitative machinery (ECE Eq. 2, DCA Eq. 3, POMDP Eq. 1) is standard external formalism used for context, with the paper explicitly disclaiming that the POMDP is 'an analytical abstraction ... rather than an implementation requirement.' The review's internal limitations are openly disclosed and do not conceal a self-referential derivation. No circular steps meet the evidentiary bar of quote-and-reduction, so the honest finding is no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

The review's central conclusions rest on the reliability of the screening and classification process and on the assumption that the unexamined Europe PMC records would not change the landscape. The paper defines its own taxonomy and complexity score while disclaiming them as validated; these are analytical conventions, not falsifiable predictions.

axioms (4)
  • domain assumption Eligibility can be reliably determined from the implemented workflow and evaluation described in each report, rather than from use of the terms 'agent' or 'Agentic AI'.
    Used throughout Methods §2.3–2.4. If primary studies under-describe their agentic mechanisms, the 557-study set and its classifications could shift.
  • ad hoc to paper The five-level clinical-validation maturity ordering (static benchmark → prospective workflow) is a meaningful ordinal scale.
    Introduced in §2.6. The paper itself notes it is descriptive and not validated, but the central 'low clinical maturity' conclusion depends on this ordering.
  • ad hoc to paper The unscreened Europe PMC records (n=445) would not materially alter the main conclusions about evaluation maturity.
    Methods §2.4 and Limitations §4.4. The review proceeds to conclusions despite roughly 20% of identified records never being screened, deduplicated, or assessed.
  • ad hoc to paper The agentic complexity score (0–8) provides a valid basis for the evidence map in Figure 4.
    §2.6. An unweighted sum of eight binary capability indicators; the paper disclaims it as a validated measure, yet it is used to support the complexity-versus-validation gap narrative.
invented entities (1)
  • Agentic complexity score (0–8) no independent evidence
    purpose: Ordinal visualization of system complexity for Figure 4 and for the claim that functional sophistication outpaces clinical validation.
    Unweighted sum of eight reported capabilities (planning, tool use, retrieval, memory, reflection, multi-agent collaboration, multimodal processing, workflow integration). No validation or calibration against external measures; the paper itself states it should not be interpreted as a validated measure.

pith-pipeline@v1.3.0-alltime-deepseek · 23840 in / 11093 out tokens · 113624 ms · 2026-08-01T02:15:26.337868+00:00 · methodology

0 comments
read the original abstract

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

Figures

Figures reproduced from arXiv: 2607.25489 by Cong Wang, Congyu Liao, Haifan Gong, Jing Qin, Wanshu Fan, Xiaofeng Liu, Yang Liu, Zheng Tong, Zhongbin Han.

Figure 1
Figure 1. Figure 1: Study identification and selection process. Searches of PubMed, IEEE Xplore, Europe PMC, arXiv, and medRxiv identified 2,192 records, and backward and forward citation searching yielded five additional reports. Of the 1,747 records available from the four exportable sources, 98 duplicate or substantially overlapping reports were removed, leaving 1,649 records for title-and-abstract screening. A total of 66… view at source ↗
Figure 2
Figure 2. Figure 2: Historical evolution of medical AI toward Agentic AI. Schematic overview of the progression from expert systems and conventional machine learning to deep learning, medical foundation models, and emerging medical Agentic AI. The upper trajectory and horizontal bars indicate approximate periods of emergence, continued use, and overlap among paradigms, rather than quantitative gains in model performance or cl… view at source ↗
Figure 3
Figure 3. Figure 3: Conceptual framework of medical Agentic AI. Multimodal clinical inputs, including medical images, clinical text, electronic health records, physiological signals, and structured data, provide the information environment for agentic workflows. Core capabilities include planning, reasoning, memory, retrieval, external tool use, feedback-based refinement, multi-agent collaboration, and simulation. These capab… view at source ↗
Figure 4
Figure 4. Figure 4: Evidence map of representative medical Agentic AI systems. Systems are positioned according to a descriptive agentic complexity score based on the reported presence of planning, tool use, retrieval, memory, reflection, multi-agent collaboration, multimodal processing, and workflow integration, and according to validation maturity ranging from static benchmarks to prospective workflow evaluation. Colors den… view at source ↗
Figure 5
Figure 5. Figure 5: Evaluation and clinical translation framework for medical Agentic AI. The upper continuum represents increasing validation maturity, from static benchmark evaluation and simulation to retrospective clinical data, external validation, and prospective real-world workflow assessment. The central panels define six complementary dimensions: task performance, process reliability, evidence traceability, safety, c… view at source ↗
Figure 6
Figure 6. Figure 6: Current challenges and research priorities for medical Agentic AI. Major barriers span multimodal data integration, planning and reasoning, tool and evidence reliability, uncertainty and safety, clinical validation, and real-world deployment. The evidence base remains dominated by public benchmarks, simulated environments, retrospective datasets, and limited expert evaluation. Key priorities include intero… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 2 canonical work pages

  1. [1]

    Topol, E. J. High-performance medicine: The convergence of human and artificial intelligence.Nat. Medicine25, 44–56, DOI: 10.1038/s41591-018-0300-7 (2019). 2.Rajpurkar, P., Chen, E., Banerjee, O. & Topol, E. J. AI in health and medicine.Nat. Medicine28, 31–38, DOI: 10.1038/s41591-021-01614-0 (2022). 3.Moor, M.et al.Foundation models for generalist medical...

  2. [5]

    Medicine31, 943–950, DOI: 10.1038/s41591-024-03423-7 (2025)

    Singhal, K.et al.Toward expert-level medical question answering with large language models.Nat. Medicine31, 943–950, DOI: 10.1038/s41591-024-03423-7 (2025). 6.Zhou, J.et al.Large language models in biomedicine and healthcare.npj Artif. Intell.1, 44, DOI: 10.1038/s44387-025-00047-1 (2025). 7.AlSaad, R.et al.Multimodal large language models in health care: ...

  3. [11]

    Medicine 3, 141, DOI: 10.1038/s43856-023-00370-1 (2023)

    Clusmann, J.et al.The future landscape of large language models in medicine.Commun. Medicine 3, 141, DOI: 10.1038/s43856-023-00370-1 (2023). 12.Yao, S.et al.ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations(2023). 13.Schick, T.et al.Toolformer: Language models can teach themselves to use too...

  4. [15]

    & O’Malley, L

    Arksey, H. & O’Malley, L. Scoping studies: Towards a methodological framework.Int. J. Soc. Res. Methodol.8, 19–32 (2005). 16.Tricco, A. C., Lillie, E., Zarin, W.et al.PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation.Annals Intern. Medicine169, 467–473 (2018). 17.Petersen, K., Feldt, R., Mujtaba, S. & Mattsson, M. Systematic map...

  5. [18]

    Medicine25, 16–18, DOI: 10.1038/s41591-018-0310-5 (2019)

    Gottesman, O.et al.Guidelines for reinforcement learning in healthcare.Nat. Medicine25, 16–18, DOI: 10.1038/s41591-018-0310-5 (2019). 19.Abramoff, M. D., Lavin, P. T., Birch, M., Shah, N. & Folk, J. C. Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices.npj Digit. Medicine1, 39, DOI: 10.1...

  6. [22]

    & Jennings, N

    Wooldridge, M. & Jennings, N. R. Intelligent agents: Theory and practice.The Knowl. Eng. Rev.10, 115–152 (1995)

  7. [23]

    Wang, L.et al.A survey on large language model based autonomous agents.Front. Comput. Sci.18, 186345 (2024). 24.Vatsal, S., Dubey, H. & Singh, A. Agentic AI in healthcare and medicine: A seven-dimensional taxonomy for empirical evaluation of LLM-based agents.IEEE AccessDOI: 10.1109/ACCESS.2026.3651218 (2026)

  8. [25]

    InAdvances in Neural Information Processing Systems, vol

    Yao, S.et al.Tree of thoughts: Deliberate problem solving with large language models. InAdvances in Neural Information Processing Systems, vol. 36, 11809–11822 (Curran Associates, Inc., 2023). 26.Shi, W.et al.EHRAgent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. InProceedings of the 2024 Confere...

  9. [27]

    Li, B.et al.MMedAgent: Learning to use medical tools with multi-modal agent. InFindings of the Association for Computational Linguistics: EMNLP 2024, 8745–8760, DOI: 10.18653/v1/2024.findings-emnlp.510 (Association for Computational Linguistics, Miami, Florida, USA, 2024)

  10. [28]

    InAdvances in Neural Information Processing Systems, vol

    Lewis, P.et al.Retrieval-augmented generation for knowledge-intensive NLP tasks. InAdvances in Neural Information Processing Systems, vol. 33, 9459–9474 (2020). 29.Izacard, G. & Grave, E. Leveraging passage retrieval with generative models for open domain question answering. InProceedings of the 16th Conference of the European Chapter of the Association f...

  11. [33]

    Tang, X.et al.MedAgents: Large language models as collaborators for zero-shot medical reasoning. InFindings of the Association for Computational Linguistics: ACL 2024, 599–621, DOI: 10.18653/v1/2024.findings-acl.33 (Association for Computational Linguistics, Bangkok, Thailand, 2024). 34.Kim, Y .et al.MDAgents: An adaptive collaboration of LLMs for medical...

  12. [35]

    InAdvances in Neural Information Processing Systems, vol

    Zhu, Y .et al.MedAgentBoard: Benchmarking multi-agent collaboration with conventional methods for diverse medical tasks. InAdvances in Neural Information Processing Systems, vol. 38 (2025). 36.Fallahpour, A., Ma, J., Munim, A., Lyu, H. & Wang, B. MedRAX: Medical reasoning agent for chest X-ray. InProceedings of the 42nd International Conference on Machine...

  13. [42]

    InProceedings of the IEEE/CVF International Conference on Computer Vision, 4015–4026 (2023)

    Kirillov, A.et al.Segment anything. InProceedings of the IEEE/CVF International Conference on Computer Vision, 4015–4026 (2023). 43.Ma, J.et al.Segment anything in medical images.Nat. Commun.15, 654, DOI: 10.1038/s41467-024-44824-z (2024). 44.Tiu, E.et al.Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learnin...

  14. [60]

    Bustos, A., Pertusa, A., Salinas, J. M. & de la Iglesia-Vayá, M. PadChest: A large chest X-ray image dataset with multi-label annotated reports.Med. Image Analysis66, 101797, DOI: 10.1016/j.media.2020.101797 (2020). 61.Wang, X.et al.ChestX-ray8: Hospital-scale chest X-ray database and benchmarks on weakly-supervised classification and localization of comm...

  15. [63]

    Clark, K.et al.The cancer imaging archive (TCIA): Maintaining and operating a public information repository.J. Digit. Imaging26, 1045–1057, DOI: 10.1007/s10278-013-9622-7 (2013). 64.Lau, J. J., Gayen, S., Ben Abacha, A. & Demner-Fushman, D. A dataset of clinically generated visual questions and answers about radiology images.Sci. Data5, 180251, DOI: 10.10...

  16. [69]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20730–20740 (2022)

    Tang, Y .et al.Self-supervised pre-training of swin transformers for 3D medical image analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20730–20740 (2022)

  17. [70]

    Intell.1, 9, DOI: 10.1038/s44387-025-00011-z (2025)

    Zhou, S.et al.Large language models for disease diagnosis: A scoping review.npj Artif. Intell.1, 9, DOI: 10.1038/s44387-025-00011-z (2025). 71.Liévin, V ., Hother, C. E., Motzfeldt, A. G. & Winther, O. Can large language models reason about medical questions?Patterns5, 100943, DOI: 10.1016/j.patter.2024.100943 (2024). 72.Mohammadi, M., Li, Y ., Lo, J. & Y...

  18. [74]

    Dice, L. R. Measures of the amount of ecologic association between species.Ecology26, 297–302 (1945). 75.Taha, A. A. & Hanbury, A. Metrics for evaluating 3D medical image segmentation: Analysis, selection, and tool.BMC Med. Imaging15, 29, DOI: 10.1186/s12880-015-0068-x (2015)

  19. [76]

    & Salcudean, S

    Karimi, D. & Salcudean, S. E. Reducing the hausdorff distance in medical image segmentation with convolutional neural networks.IEEE Transactions on Med. Imaging39, 499–513, DOI: 10.1109/TMI.2019.2930068 (2020). 77.Papineni, K., Roukos, S., Ward, T. & Zhu, W. J. BLEU: A method for automatic evaluation of machine translation. InProceedings of the 40th Annua...

  20. [91]

    Li, J.et al.Agent Hospital: A simulacrum of hospital with evolvable medical agents.arXiv preprint arXiv:2405.02957(2024)

  21. [92]

    InProceedings of the 31st International Conference on Computational Linguistics, 10183–10213 (2025)

    Fan, Z.et al.AI hospital: Benchmarking large language models in a multi-agent medical interaction simulator. InProceedings of the 31st International Conference on Computational Linguistics, 10183–10213 (2025). 93.Wiens, J.et al.Do no harm: A roadmap for responsible machine learning for health care.Nat. Medicine25, 1337–1340, DOI: 10.1038/s41591-019-0548-6...

  22. [95]

    W., Denniston, A

    Cruz Rivera, S., Liu, X., Chan, A. W., Denniston, A. K. & Calvert, M. J. Guidelines for clinical trial protocols for interventions involving artificial intelligence: The SPIRIT-AI extension.Nat. Medicine 26, 1351–1363, DOI: 10.1038/s41591-020-1037-7 (2020). 96.Vasey, B.et al.Reporting guideline for the early-stage clinical evaluation of decision support s...

  23. [103]

    104.Kelly, C

    Nagendran, M.et al.Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies.BMJ368, m689, DOI: 10.1136/bmj.m689 (2020). 104.Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence.BMC Medicine1...

  24. [106]

    Roberts, M.et al.Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans.Nat. Mach. Intell.3, 199–217, DOI: 10.1038/s42256-021-00307-0 (2021). 107.Varoquaux, G. & Cheplygina, V . Machine learning for medical imaging: Methodological failures and recommendations for the fut...