Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This scoping review of 144 papers claims generative AI in medicine is shifting from text-only large language models to multimodal systems that combine imaging, text, and structured data in a single model.

desk verdict Useful scoping review of medical generative AI, but the 'shift to multimodal' claim rests on a hand-curated subset and an off-by-one PRISMA count; worth refereeing after a revision. read the letter →

arxiv 2502.09242 v1 pith:2RQJO6RY submitted 2025-02-13 cs.AI

classification cs.AI
keywords generativeAIinmedicinemultimodallargelanguagemodelsscopingreviewclinicalevaluationmetricsradiologyreportgenerationmedicaldatasetsgeneralistcontrastivelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This scoping review claims that generative AI in medicine is moving from text-only large language models to multimodal systems that combine images, text, and structured data in one model. The authors argue this shift is already visible in diagnostic support, medical report generation, drug discovery, and conversational AI, and that evaluation practice has not kept up. To ground this claim they systematically screened 4,384 records and included 144 papers published up to the end of 2024. The review matters because it identifies where clinical deployment is likely to succeed and where it will stall: data heterogeneity, interpretability, ethics, and real-world validation remain open problems.

What carries the argument

The central objects are the two multimodal architectures the review uses to organize the field: contrastive language-image pretraining (CLIP) models, which align different modalities in a shared embedding space, and multimodal large language models (MLLMs), which encode images or other non-text data into the language model's embedding space. These carry the argument by defining what multimodal means and by structuring the taxonomy of methods, datasets, and applications. A second mechanism is the pairing of evaluation metrics: lexical metrics such as BLEU and ROUGE versus clinically grounded metrics such as entity-based scoring, grounded factual checks, and error-notation frameworks, which the review uses to show that evaluation lags behind model development.

What would settle it

Count the share of included papers that integrate at least two modalities by publication year; if the share does not rise across 2022 to 2024, the claimed shift is an artifact of curation rather than a property of the literature.

Watch

Extended reading notes

Core claim

The paper's central finding is that the center of gravity of medical generative AI has shifted from unimodal large language models to multimodal models, organized around two architectural families: contrastive models that align images with text in a shared embedding space, and multimodal large language models that project non-text features into the language model's embedding space. This shift is documented across 144 included papers and appears in a growing family of generalist models that handle classification, segmentation, report generation, and visual question answering within one architecture. The authors also find that standard lexical evaluation metrics such as BLEU and ROUGE miss clinical correctness, and that newer clinically grounded metrics, while promising, are not yet standardized across sites and specialties.

Load-bearing premise

The review's trend claims depend on the representativeness of the curated set of 144 papers, since 83 entered through a manual search described as capturing recent and high-impact work and the authors deliberately balanced inclusion across prevalent fields.

Editorial extensions

If this is right

  • Clinical AI development will increasingly center on models that ingest images, text, and structured data together rather than text-only assistants.
  • Evaluation of medical generative models should combine lexical metrics with clinically grounded metrics that check factual correctness and relevance.
  • Dataset construction should move beyond radiology-centric resources toward other specialties and modalities to avoid generalization failures.
  • Generalist models that unify multiple tasks and modalities are becoming the default architectural direction for medical AI.
  • Synthetic image generation will be used to augment scarce data and simulate rare conditions, provided its clinical utility is validated on downstream tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: If the field's center of gravity is indeed shifting to multimodal models, text-only clinical language models may become components of larger systems rather than standalone products.
  • Editorial inference: The review's selection method makes the shift claim sensitive to curation, so a formal bibliometric analysis of all published medical AI abstracts would test whether the multimodal trend is a property of the literature or an artifact of inclusion choices.
  • Editorial inference: Clinical adoption may hinge less on model capability than on the availability of evaluation benchmarks that predict real-world diagnostic utility, a gap the review identifies but does not quantify.
  • Editorial inference: The heavy representation of radiology suggests the next bottleneck is data: without analogous multimodal datasets in pathology, genomics, and primary care, the shift to multimodal AI will proceed unevenly across medicine.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This scoping review (arXiv:2502.09242) applies PRISMA-ScR guidance to survey generative AI in medicine, covering text-only LLMs, multimodal CLIP and MLLM architectures, associated datasets, and evaluation metrics. The authors report screening 4,384 database records and 84 additional records from manual searches, ultimately including 144 papers. The review synthesizes these papers in six tables organized by application, dataset, and evaluation approach, and the abstract and discussion make a central claim of a field-level 'shift from unimodal to multimodal approaches' in medical AI. The paper also describes persistent challenges such as data heterogeneity, interpretability, and evaluation gaps.

Significance. If the reported trend claim were reliable, the review would offer a valuable, up-to-date map of a rapidly moving field, with useful taxonomies of models, datasets, and clinical evaluation metrics. Its strengths include the reproducible database queries in Supplementary Table S.2, the explicit PRISMA-ScR checklist, and a structured categorization of 144 primary papers into coherent categories (Tables 1-6). The qualitative discussion of evaluation metrics (Section 6) is informative and reasonably current, and the identification of radiology-heavy dataset bias is a credible observation. However, the headline 'shift' claim depends on a literature-selection process that includes an unspecified manual search and a deliberate balancing step, which the reported methodology does not adequately support as a population-level trend. The review's descriptive synthesis remains useful regardless, but the central claim needs substantial clarification and additional transparency about how the included set was curated.

major comments (4)
  1. [Section 3 and Figure 2] The reported screening numbers are internally inconsistent. The text says 2,656 articles were excluded during title/abstract screening, while Figure 2 reports 2,657 excluded records. The text also states that 83 papers came from manual searches and that the total is 144, but 60 database papers plus 83 manual papers gives 143; Figure 2 shows 84 records from 'other methods' (1 website + 83 citations), which would give 144. These discrepancies affect the credibility of the PRISMA-ScR flow and must be reconciled before the review can be considered reproducible.
  2. [Sections 2.3-2.4 and Abstract] The central claim of a 'shift from unimodal to multimodal approaches' is not supported by the described selection design. The database search was constructed as two targeted subsearches, one for text-only LLMs and one for multimodal models, so the retrieved set is a union of two intentionally defined categories rather than a sample from which prevalence can be inferred. Moreover, 83 of the 144 included papers (58%) came from a manual search described only as capturing 'recent and high-impact publications' (Section 2.3), and Section 2.4 states that inclusion was deliberately balanced 'to ensure proportional inclusion from prevalent fields.' Under this design, the observed preponderance of multimodal papers could be an artifact of the inclusion rules. To keep the claim, the authors should either restrict all trend statements to the curated set with an explicit caveat that the set is not a random sample, or provide a sensitivity analysis based only on database-derived papers and a prespecified manual search protocol.
  3. [Section 2.4] The balancing criterion 'proportional inclusion from prevalent fields, such as X-ray report generation' is too vague to be a reproducible selection rule. The authors should specify how 'prevalent' was operationalized, who made the balancing decisions, what data were used to determine prevalence, and how many candidate papers were included or excluded because of this rule. Without this detail, a reader cannot determine whether the resulting 144-paper corpus is representative of the literature or of the authors' interests.
  4. [Sections 5 and 7] The 'shift' claim is presented as a temporal trend, but no time-based analysis is shown. The review spans 2020-2024, yet no table or figure reports the distribution of unimodal versus multimodal papers by publication year, nor is there an analysis of whether the share of multimodal papers changed over time within the included set. If the authors intend to claim a shift, they need to provide direct evidence of changing proportions over time, or explicitly reframe the claim as a comparison between two categories present in the included set.
minor comments (5)
  1. [Figure 3 caption] The caption contains a typo: 'ebbedding space' should be 'embedding space.'
  2. [Section 2.1] The eligibility criteria state that 'manually selected preprints with high relevance and potential impact' were included, but the criteria do not define 'relevance' or 'impact.' Consider providing an operational definition.
  3. [Section 6, Table 6] The column header 'Application' is so broad that entries like 'Report evaluation' and 'Conversation evaluation' are not fully descriptive; consider subdividing by evaluation target (e.g., text generation, image generation, calibration).
  4. [Section 5.4, Table 5] Some dataset sizes are given as raw counts without a clear unit (e.g., '224K triplets' vs. '25M pairs'); standardizing the unit label (e.g., 'image-text pairs,' 'patients') would improve clarity.
  5. [References] Several references are to arXiv preprints without a note of peer-reviewed status, which conflicts with the stated inclusion criterion of 'peer-reviewed conference and journal publications, alongside manually selected preprints.' Clarify which included papers are preprints and how they were evaluated for 'relevance and potential impact.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review synthesizes external literature, and its trend claims, while methodologically vulnerable, are not equivalent to the search design by construction.

full rationale

This is a scoping review, not a derivation with fitted parameters or equations, so the core circularity patterns do not apply. The central claim of a 'shift from unimodal to multimodal approaches' is a narrative synthesis of 144 selected papers, and the selection process did not force that conclusion: the two database subsearches were designed to capture both text-only LLMs and multimodal models, and the manual search aimed at 'recent and high-impact publications' with 'proportional inclusion from prevalent fields.' Those choices may create selection bias, and the PRISMA flow has a minor numerical inconsistency, but bias and reporting inconsistency are not circularity in the sense of a result being equivalent to its inputs by construction. The paper's self-citations (e.g., [2], [10], [16], [34], [73], [91], [104], [133], [142]) appear as illustrative examples in tables and narrative descriptions; none is load-bearing for the review's overall conclusions, and none imports a uniqueness theorem or forbids alternatives. The methodological limitations are acknowledged in Section 7, including 'overrepresentation of radiology-focused datasets and models,' which further indicates the authors did not present the selection as a derivation. Therefore no circular step is identifiable, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities: the paper introduces no model, method, or quantity. The load-bearing premises are methodological: representative literature selection and effective search. These are not backed by external evidence beyond the paper's own curation.

assumptions (3)
  • domain assumption The included set of 144 papers, including manually selected preprints, is representative of the research landscape in generative AI for medicine.
    Section 2.4 states that inclusion aimed for proportional representation across prevalent fields, and Section 3 reports that 83 of 144 papers came from non-database manual searches; the trend claims depend on this representativeness.
  • domain assumption The search strings with NOT 'ChatGPT' and NOT 'Review' filters identify the intended literature without systematic bias.
    Supplementary Table S.2 defines the queries; the review's conclusions about the field presuppose these queries effectively sample the literature.
  • domain assumption The evaluation metrics discussed are representative of current practices in medical generative AI.
    Section 6 generalizes about evaluation practices based on the included subset, which is heavily radiology-oriented, a point the authors acknowledge in the Discussion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine." pith.science (2026). https://pith.science/paper/2RQJO6RY

@misc{pith2026250209242,
  author       = {Pith},
  title        = {Pith review of: From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2RQJO6RY}},
  note         = {Machine review of arXiv:2502.09242}
}
read the original abstract

Generative artificial intelligence (AI) models, such as diffusion models and OpenAI's ChatGPT, are transforming medicine by enhancing diagnostic accuracy and automating clinical workflows. The field has advanced rapidly, evolving from text-only large language models for tasks such as clinical documentation and decision support to multimodal AI systems capable of integrating diverse data modalities, including imaging, text, and structured data, within a single model. The diverse landscape of these technologies, along with rising interest, highlights the need for a comprehensive review of their applications and potential. This scoping review explores the evolution of multimodal AI, highlighting its methods, applications, datasets, and evaluation in clinical settings. Adhering to PRISMA-ScR guidelines, we systematically queried PubMed, IEEE Xplore, and Web of Science, prioritizing recent studies published up to the end of 2024. After rigorous screening, 144 papers were included, revealing key trends and challenges in this dynamic field. Our findings underscore a shift from unimodal to multimodal approaches, driving innovations in diagnostic support, medical report generation, drug discovery, and conversational AI. However, critical challenges remain, including the integration of heterogeneous data types, improving model interpretability, addressing ethical concerns, and validating AI systems in real-world clinical settings. This review summarizes the current state of the art, identifies critical gaps, and provides insights to guide the development of scalable, trustworthy, and clinically impactful multimodal AI solutions in healthcare.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI4Research: A Survey of Artificial Intelligence for Scientific Research

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A survey that organizes AI-for-research work into five tasks, comprehension, survey, discovery, writing, and peer review, and compiles associated tools and benchmarks.

Reference graph

Works this paper leans on

172 extracted references · 41 canonical work pages · cited by 1 Pith paper

  1. [1]

    Nature Medicine 28(9), 1773–1784 (2022)

    Acosta, J.N., Falcone, G.J., Rajpurkar, P., Topol, E.J.: Multimodal biomedical ai. Nature Medicine 28(9), 1773–1784 (2022)

  2. [2]

    Nature Communications 15(1), 1603 (2024)

    Tayebi Arasteh, S., Han, T., Lotfinia, M., Kuhl, C., Kather, J.N., Truhn, D., Nebelung, S.: Large language models streamline automated machine learning for clinical studies. Nature Communications 15(1), 1603 (2024)

  3. [3]

    IEEE Transactions on Multimedia 25, 845–855 (2021)

    Cai, X., Liu, S., Han, J., Yang, L., Liu, Z., Liu, T.: Chestxraybert: A pretrained language model for chest radiology report summarization. IEEE Transactions on Multimedia 25, 845–855 (2021)

  4. [4]

    Cureus 15(6) (2023)

    Li, Y., Li, Z., Zhang, K., Dan, R., Jiang, S., Zhang, Y.: Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus 15(6) (2023)

  5. [5]

    nature 596(7873), 583–589 (2021)

    Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R.,ˇZ ´ ıdek, A., Potapenko, A.,et al.: Highly accurate protein structure prediction with alphafold. nature 596(7873), 583–589 (2021)

  6. [6]

    The Impact of AI Assistance on Radiology Reporting: A Pilot Study Using Simulated AI Draft Reports

    Acosta, J.N., Dogra, S., Adithan, S., Wu, K., Moritz, M., Kwak, S., Rajpurkar, P.: The Impact of AI Assistance on Radiology Reporting: A Pilot Study Using Simulated AI Draft Reports (2024). https://arxiv.org/abs/2412.12042

  7. [7]

    Nature medicine 30(4), 1134–1142 (2024)

    Van Veen, D., Van Uden, C., Blankemeier, L., Delbrouck, J.-B., Aali, A., Bluethgen, C., Pareek, A., Polacin, M., Reis, E.P., Seehofnerov´ a, A., et al.: Adapted large language models can outperform medical experts in clinical text summarization. Nature medicine 30(4), 1134–1142 (2024)

  8. [8]

    Scientific data 6(1), 317 (2019)

    Johnson, A.E., Pollard, T.J., Berkowitz, S.J., Greenbaum, N.R., Lungren, M.P., Deng, C.-y., Mark, R.G., Horng, S.: Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data 6(1), 317 (2019)

Show all 172 references
  1. [9]

    arXiv preprint arXiv:2311.10798 (2023)

    Huang, S.-C., Huo, Z., Steinberg, E., Chiang, C.-C., Lungren, M.P., Langlotz, C.P., Yeung, S., Shah, N.H., Fries, J.A.: Inspect: a multimodal dataset for pulmonary embolism diagnosis and prognosis. arXiv preprint arXiv:2311.10798 (2023)

  2. [10]

    Radiology 313(2), 233441 (2024)

    Tayebi Arasteh, S., Siepmann, R., Huppertz, M., Lotfinia, M., Puladi, B., Kuhl, C., Truhn, D., Nebelung, S.: The treasure trove hidden in plain sight: The utility of gpt-4 in chest radiograph evaluation. Radiology 313(2), 233441 (2024)

  3. [11]

    Radiology 309(1), 230806 (2023)

    Khader, F., M¨ uller-Franzes, G., Wang, T., Han, T., Tayebi Arasteh, S., Haar- burger, C., Stegmaier, J., Bressem, K., Kuhl, C., Nebelung, S.,et al.: Multimodal deep learning for integrating chest radiographs and clinical parameters: a case 21 for transformers. Radiology 309(1...

  4. [12]

    Nucleic acids research 49(D1), 916–923 (2021)

    Frankish, A., Diekhans, M., Jungreis, I., Lagarde, J., Loveland, J.E., Mudge, J.M., Sisu, C., Wright, J.C., Armstrong, J., Barnes, I., et al.: Gencode 2021. Nucleic acids research 49(D1), 916–923 (2021)

  5. [13]

    Nature Medicine, 1–13 (2024)

    Zhang, K., Zhou, R., Adhikarla, E., Yan, Z., Liu, Y., Yu, J., Liu, Z., Chen, X., Davison, B.D., Ren, H., et al.: A generalist vision–language foundation model for diverse biomedical tasks. Nature Medicine, 1–13 (2024)

  6. [14]

    arXiv preprint arXiv:2403.17834 (2024)

    Hamamci, I.E., Er, S., Almas, F., Simsek, A.G., Esirgun, S.N., Dogan, I., Das- delen, M.F., Wittmann, B., Simsar, E., Simsar, M., et al.: A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-shot detection of abnormalities. arXiv prepri...

  7. [15]

    Advances in Neural Information Processing Systems 36 (2024)

    Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., Gao, J.: Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36 (2024)

  8. [16]

    In: International Con- ference on Medical Image Computing and Computer-Assisted Intervention, pp

    ¨Ozsoy, E., Pellegrini, C., Keicher, M., Navab, N.: Oracle: Large vision-language models for knowledge-guided holistic or domain modeling. In: International Con- ference on Medical Image Computing and Computer-Assisted Intervention, pp. 455–465 (2024). Springer

  9. [17]

    arXiv preprint arXiv:2306.13549 (2023)

    Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., Chen, E.: A survey on multimodal large language models. arXiv preprint arXiv:2306.13549 (2023)

  10. [18]

    arXiv preprint arXiv:2408.01319 (2024)

    Wang, J., Jiang, H., Liu, Y., Ma, C., Zhang, X., Pan, Y., Liu, M., Gu, P., Xia, S., Li, W., et al.: A comprehensive review of multimodal large language models: Per- formance and challenges across different tasks. arXiv preprint arXiv:2408.01319 (2024)

  11. [19]

    npj Digital Medicine 5(1), 171 (2022)

    Kline, A., Wang, H., Li, Y., Dennis, S., Hutch, M., Xu, Z., Wang, F., Cheng, F., Luo, Y.: Multimodal machine learning in precision health: A scoping review. npj Digital Medicine 5(1), 171 (2022)

  12. [20]

    arXiv preprint arXiv:2404.03264 (2024)

    He, Y., Huang, F., Jiang, X., Nie, Y., Wang, M., Wang, J., Chen, H.: Foundation model for advancing healthcare: Challenges, opportunities, and future directions. arXiv preprint arXiv:2404.03264 (2024)

  13. [21]

    Annals of internal medicine 169(7), 467–473 (2018) 22

    Tricco, A.C., Lillie, E., Zarin, W., O’Brien, K.K., Colquhoun, H., Levac, D., Moher, D., Peters, M.D., Horsley, T., Weeks, L.,et al.: Prisma extension for scop- ing reviews (prisma-scr): checklist and explanation. Annals of internal medicine 169(7), 467–473 (2018) 22

  14. [22]

    bmj 372 (2021)

    Page, M.J., McKenzie, J.E., Bossuyt, P.M., Boutron, I., Hoffmann, T.C., Mul- row, C.D., Shamseer, L., Tetzlaff, J.M., Akl, E.A., Brennan, S.E., et al.: The prisma 2020 statement: an updated guideline for reporting systematic reviews. bmj 372 (2021)

  15. [23]

    Systematic reviews 5, 1–10 (2016)

    Ouzzani, M., Hammady, H., Fedorowicz, Z., Elmagarmid, A.: Rayyan—a web and mobile app for systematic reviews. Systematic reviews 5, 1–10 (2016)

  16. [24]

    Nature 619(7969), 357–362 (2023)

    Jiang, L.Y., Liu, X.C., Nejatian, N.P., Nasir-Moin, M., Wang, D., Abidin, A., Eaton, K., Riina, H.A., Laufer, I., Punjabi, P., et al.: Health system-scale lan- guage models are all-purpose prediction engines. Nature 619(7969), 357–362 (2023)

  17. [25]

    Advances in Neural Information Processing Systems (2017)

    Vaswani, A.: Attention is all you need. Advances in Neural Information Processing Systems (2017)

  18. [26]

    Bioinformatics 36(4), 1234–1240 (2020)

    Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C.H., Kang, J.: Biobert: a pre- trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4), 1234–1240 (2020)

  19. [27]

    arXiv preprint arXiv:2402.10373 (2024)

    Labrak, Y., Bazoge, A., Morin, E., Gourraud, P.-A., Rouvier, M., Dufour, R.: Biomistral: A collection of open-source pretrained large language models for medical domains. arXiv preprint arXiv:2402.10373 (2024)

  20. [28]

    npj Digital Medicine 7(1), 41 (2024)

    Wang, L., Chen, X., Deng, X., Wen, H., You, M., Liu, W., Li, Q., Li, J.: Prompt engineering in consistency and reliability with the evidence-based guideline for llms. npj Digital Medicine 7(1), 41 (2024)

  21. [29]

    Lee, H., Phatale, S., Mansoor, H., Lu, K.R., Mesnard, T., Ferret, J., Bishop, C., Hall, E., Carbune, V., Rastogi, A.: Rlaif: Scaling reinforcement learning from human feedback with ai feedback (2023)

  22. [30]

    arXiv preprint arXiv:2305.15075 (2023)

    Zhang, H., Chen, J., Jiang, F., Yu, F., Chen, Z., Li, J., Chen, G., Wu, X., Zhang, Z., Xiao, Q., et al.: Huatuogpt, towards taming language model to be a doctor. arXiv preprint arXiv:2305.15075 (2023)

  23. [31]

    arXiv preprint arXiv:2412.18925 (2024)

    Chen, J., Cai, Z., Ji, K., Wang, X., Liu, W., Wang, R., Hou, J., Wang, B.: Huatuogpt-o1, towards medical complex reasoning with llms. arXiv preprint arXiv:2412.18925 (2024)

  24. [32]

    Advances in Neural Information Processing Systems 33, 9459–9474 (2020)

    Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., K¨ uttler, H., Lewis, M., Yih, W.-t., Rockt¨ aschel, T., et al.: Retrieval-augmented gen- eration for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems 33, 9459–9474 (2020)

  25. [33]

    NEJM AI 1(2), 2300068 (2024)

    Zakka, C., Shad, R., Chaurasia, A., Dalal, A.R., Kim, J.L., Moor, M., Fong, R., Phillips, C., Alexander, K., Ashley, E., et al.: Almanac—retrieval-augmented 23 language models for clinical medicine. NEJM AI 1(2), 2300068 (2024)

  26. [34]

    https://arxiv.org/abs/2407.15621

    Arasteh, S.T., Lotfinia, M., Bressem, K., Siepmann, R., Adams, L., Ferber, D., Kuhl, C., Kather, J.N., Nebelung, S., Truhn, D.: RadioRAG: Factual Large Language Models for Enhanced Diagnostics in Radiology Using Online Retrieval Augmented Generation (2024). https://arxiv.org/a...

  27. [35]

    NPJ Digital Medicine 7(1), 100 (2024)

    Gilbert, S., Kather, J.N., Hogan, A.: Augmented non-hallucinating large lan- guage models as medical information curators. NPJ Digital Medicine 7(1), 100 (2024)

  28. [36]

    In: 2021 International Joint Conference on Neural Networks (IJCNN), pp

    Naseem, U., Khushi, M., Reddy, V., Rajendran, S., Razzak, I., Kim, J.: Bioal- bert: A simple and effective pre-trained language model for biomedical named entity recognition. In: 2021 International Joint Conference on Neural Networks (IJCNN), pp. 1–7 (2021). IEEE

  29. [37]

    Briefings in bioinformatics 23(6), 409 (2022)

    Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., Liu, T.-Y.: Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics 23(6), 409 (2022)

  30. [38]

    npj Digital Medicine 7(1), 16 (2024)

    Wang, H., Gao, C., Dantona, C., Hull, B., Sun, J.: Drg-llama: tuning llama model to predict diagnosis-related group for hospitalized patients. npj Digital Medicine 7(1), 16 (2024)

  31. [39]

    NPJ digital medicine 5(1), 194 (2022)

    Yang, X., Chen, A., PourNejatian, N., Shin, H.C., Smith, K.E., Parisien, C., Compas, C., Martin, C., Costa, A.B., Flores, M.G., et al.: A large language model for electronic health records. NPJ digital medicine 5(1), 194 (2022)

  32. [40]

    In: Proceedings of the ACM Conference on Health, Inference, and Learning, pp

    Johnson, A.E., Bulgarelli, L., Pollard, T.J.: Deidentification of free-text medical records using pre-trained bidirectional transformers. In: Proceedings of the ACM Conference on Health, Inference, and Learning, pp. 214–221 (2020)

  33. [41]

    NPJ Digital Medicine 7(1), 102 (2024)

    Kresevic, S., Giuffr` e, M., Ajcevic, M., Accardo, A., Croc` e, L.S., Shung, D.L.: Optimization of hepatological clinical guidelines interpretation by large lan- guage models: a retrieval augmented generation-based framework. NPJ Digital Medicine 7(1), 102 (2024)

  34. [42]

    AMIA Summits on Translational Science Proceedings 2021, 420 (2021)

    Mahendran, D., McInnes, B.T.: Extracting adverse drug events from clinical notes. AMIA Summits on Translational Science Proceedings 2021, 420 (2021)

  35. [43]

    Medical Image Analysis 99, 103383 (2025)

    Lanfredi, R.B., Mukherjee, P., Summers, R.M.: Enhancing chest x-ray datasets with privacy-preserving large language models and multi-type annotations: a data-driven approach for improved classification. Medical Image Analysis 99, 103383 (2025)

  36. [44]

    IEEE Transactions on Industrial 24 Informatics 18(8), 5600–5608 (2021)

    Liu, N., Hu, Q., Xu, H., Xu, X., Chen, M.: Med-bert: A pretraining framework for medical records named entity recognition. IEEE Transactions on Industrial 24 Informatics 18(8), 5600–5608 (2021)

  37. [45]

    arXiv preprint arXiv:2304.08247 (2023)

    Han, T., Adams, L.C., Papaioannou, J.-M., Grundmann, P., Oberhauser, T., L¨ oser, A., Truhn, D., Bressem, K.K.: Medalpaca–an open-source collec- tion of medical conversational ai models and training data. arXiv preprint arXiv:2304.08247 (2023)

  38. [46]

    arXiv preprint arXiv:2311.16079 (2023)

    Chen, Z., Cano, A.H., Romanou, A., Bonnet, A., Matoba, K., Salvi, F., Pagliar- dini, M., Fan, S., K¨ opf, A., Mohtashami, A., et al.: Meditron-70b: Scaling medical pretraining for large language models. arXiv preprint arXiv:2311.16079 (2023)

  39. [47]

    Nature 620(7972), 172–180 (2023)

    Singhal, K., Azizi, S., Tu, T., Mahdavi, S.S., Wei, J., Chung, H.W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al.: Large language models encode clinical knowledge. Nature 620(7972), 172–180 (2023)

  40. [48]

    Nature Communications 15(1), 8384 (2024)

    Qiu, P., Wu, C., Zhang, X., Lin, W., Wang, H., Zhang, Y., Wang, Y., Xie, W.: Towards building multilingual language model for medicine. Nature Communications 15(1), 8384 (2024)

  41. [49]

    Communications medicine 1(1), 11 (2021)

    Mu, Y., Tizhoosh, H.R., Tayebi, R.M., Ross, C., Sur, M., Leber, B., Camp- bell, C.J.: A bert model generates diagnostically relevant semantic embeddings from pathology synopses with active learning. Communications medicine 1(1), 11 (2021)

  42. [50]

    Journal of the American Medical Informatics Association, 045 (2024)

    Wu, C., Lin, W., Zhang, X., Zhang, Y., Xie, W., Wang, Y.: Pmc-llama: toward building open-source language models for medicine. Journal of the American Medical Informatics Association, 045 (2024)

  43. [51]

    medRxiv (2024)

    Jia, S., Bit, S., Searls, E., Claus, L.A., Fan, P., Jasodanand, V.H., Lauber, M.V., Veerapaneni, D., Wang, W.M., Au, R., et al.: Medpodgpt: A multilingual audio- augmented large language model for medical research and education. medRxiv (2024)

  44. [52]

    Radiology: Artificial Intelligence 4(4), 210258 (2022)

    Yan, A., McAuley, J., Lu, X., Du, J., Chang, E.Y., Gentili, A., Hsu, C.-N.: Radbert: adapting transformer-based language models to radiology. Radiology: Artificial Intelligence 4(4), 210258 (2022)

  45. [53]

    Radiology: Artificial Intelligence 6(2), 230205 (2024)

    Schmidt, R.A., Seah, J.C., Cao, K., Lim, L., Lim, W., Yeung, J.: Generative large language models for detection of speech recognition errors in radiology reports. Radiology: Artificial Intelligence 6(2), 230205 (2024)

  46. [54]

    In: MAbs, vol

    Prihoda, D., Maamary, J., Waight, A., Juan, V., Fayadat-Dilman, L., Svozil, D., Bitton, D.A.: Biophi: A platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning. In: MAbs, vol. 14, p. 2020203 (2022). Taylor & Francis

  47. [55]

    7: 25 using protein language models, regulatory cnns and other nucleotide-level scores to improve genome-wide variant predictions

    Schubach, M., Maass, T., Nazaretyan, L., R¨ oner, S., Kircher, M.: Cadd v1. 7: 25 using protein language models, regulatory cnns and other nucleotide-level scores to improve genome-wide variant predictions. Nucleic acids research 52(D1), 1143–1154 (2024)

  48. [56]

    Bioinformatics 37(15), 2112–2120 (2021)

    Ji, Y., Zhou, Z., Liu, H., Davuluri, R.V.: Dnabert: pre-trained bidirectional encoder representations from transformers model for dna-language in genome. Bioinformatics 37(15), 2112–2120 (2021)

  49. [57]

    Nature 618(7965), 616–624 (2023)

    Theodoris, C.V., Xiao, L., Chopra, A., Chaffin, M.D., Al Sayed, Z.R., Hill, M.C., Mantineo, H., Brydon, E.M., Zeng, Z., Liu, X.S., et al.: Transfer learning enables predictions in network biology. Nature 618(7965), 616–624 (2023)

  50. [58]

    Nature Biotechnology 42(2), 275–283 (2024)

    Hie, B.L., Shanker, V.R., Xu, D., Bruun, T.U., Weidenbacher, P.A., Tang, S., Wu, W., Pak, J.E., Kim, P.S.: Efficient evolution of human antibodies from general protein language models. Nature Biotechnology 42(2), 275–283 (2024)

  51. [59]

    In: International Conference on Machine Learning, pp

    Rao, R.M., Liu, J., Verkuil, R., Meier, J., Canny, J., Abbeel, P., Sercu, T., Rives, A.: Msa transformer. In: International Conference on Machine Learning, pp. 8844–8856 (2021). PMLR

  52. [60]

    Nature Biotechnology 41(8), 1099–1106 (2023)

    Madani, A., Krause, B., Greene, E.R., Subramanian, S., Mohr, B.P., Holton, J.M., Olmos, J.L., Xiong, C., Sun, Z.Z., Socher, R., et al.: Large language models generate functional protein sequences across diverse families. Nature Biotechnology 41(8), 1099–1106 (2023)

  53. [61]

    Nature communications 13(1), 4348 (2022)

    Ferruz, N., Schmidt, S., H¨ ocker, B.: Protgpt2 is a deep unsupervised language model for protein design. Nature communications 13(1), 4348 (2022)

  54. [62]

    IEEE transactions on pattern analysis and machine intelligence 44(10), 7112–7127 (2021)

    Elnaggar, A., Heinzinger, M., Dallago, C., Rehawi, G., Wang, Y., Jones, L., Gibbs, T., Feher, T., Angerer, C., Steinegger, M., et al.: Prottrans: Toward understanding the language of life through self-supervised learning. IEEE transactions on pattern analysis and machine intel...

  55. [63]

    Nature Machine Intelligence 4(10), 852–866 (2022)

    Yang, F., Wang, W., Wang, F., Fang, Y., Tang, D., Huang, J., Lu, H., Yao, J.: scbert as a large-scale pretrained deep language model for cell type annotation of single-cell rna-seq data. Nature Machine Intelligence 4(10), 852–866 (2022)

  56. [64]

    Computers in Biology and Medicine 179, 108926 (2024)

    Rathore, A.S., Choudhury, S., Arora, A., Tijare, P., Raghava, G.P.: Toxinpred 3.0: An improved method for predicting the toxicity of peptides. Computers in Biology and Medicine 179, 108926 (2024)

  57. [65]

    European Radiology 33(6), 4228–4236 (2023)

    Nowak, S., Biesner, D., Layer, Y., Theis, M., Schneider, H., Block, W., Wulff, B., Attenberger, U., Sifa, R., Sprinkart, A.: Transformer-based structuring of free- text radiology report databases. European Radiology 33(6), 4228–4236 (2023)

  58. [66]

    Scientific data 10(1), 1 (2023)

    Johnson, A.E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., 26 Pollard, T.J., Hao, S., Moody, B., Gow, B., et al.: Mimic-iv, a freely accessible electronic health record dataset. Scientific data 10(1), 1 (2023)

  59. [67]

    Scientific data 5(1), 1–13 (2018)

    Pollard, T.J., Johnson, A.E., Raffa, J.D., Celi, L.A., Mark, R.G., Badawi, O.: The eicu collaborative research database, a freely available multi-center database for critical care research. Scientific data 5(1), 1–13 (2018)

  60. [68]

    In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp

    Zeng, G., Yang, W., Ju, Z., Yang, Y., Wang, S., Zhang, R., Zhou, M., Zeng, J., Dong, X., Zhang, R., et al.: Meddialog: Large-scale medical dialogue datasets. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9241–9250 (2020)

  61. [69]

    Nucleic acids research 51(D1), 523–531 (2023)

    Uniprot: the universal protein knowledgebase in 2023. Nucleic acids research 51(D1), 523–531 (2023)

  62. [70]

    Nucleic acids research 41(D1), 36–42 (2012)

    Benson, D.A., Cavanaugh, M., Clark, K., Karsch-Mizrachi, I., Lipman, D.J., Ostell, J., Sayers, E.W.: Genbank. Nucleic acids research 41(D1), 36–42 (2012)

  63. [71]

    Journal of proteome research 14(6), 2707–2713 (2015)

    Edwards, N.J., Oberti, M., Thangudu, R.R., Cai, S., McGarvey, P.B., Jacob, S., Madhavan, S., Ketchum, K.A.: The cptac data portal: a resource for cancer proteomics research. Journal of proteome research 14(6), 2707–2713 (2015)

  64. [72]

    arXiv preprint arXiv:2406.04449 (2024)

    Bannur, S., Bouzid, K., Castro, D.C., Schwaighofer, A., Bond-Taylor, S., Ilse, M., P´ erez-Garc ´ ıa, F., Salvatelli, V., Sharma, H., Meissen, F., et al.: Maira-2: Grounded radiology report generation. arXiv preprint arXiv:2406.04449 (2024)

  65. [73]

    arXiv preprint arXiv:2311.18681 (2023)

    Pellegrini, C., ¨Ozsoy, E., Busam, B., Navab, N., Keicher, M.: Radialog: A large vision-language model for radiology report generation and conversational assistance. arXiv preprint arXiv:2311.18681 (2023)

  66. [74]

    Nature Biomedical Engineering 6(12), 1399–1406 (2022)

    Tiu, E., Talius, E., Patel, P., Langlotz, C.P., Ng, A.Y., Rajpurkar, P.: Expert- level detection of pathologies from unannotated chest x-ray images via self- supervised learning. Nature Biomedical Engineering 6(12), 1399–1406 (2022)

  67. [75]

    In: Machine Learning for Health, pp

    Endo, M., Krishnan, R., Krishna, V., Ng, A.Y., Rajpurkar, P.: Retrieval-based chest x-ray report generation using a pre-trained contrastive language-image model. In: Machine Learning for Health, pp. 209–219 (2021). PMLR

  68. [76]

    In: International Conference on Machine Learning, pp

    Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual mod- els from natural language supervision. In: International Conference on Machine Learning, pp. 8748–8763 (2021). PMLR

  69. [77]

    arXiv preprint arXiv:2406.06512 (2024) 27

    Blankemeier, L., Cohen, J.P., Kumar, A., Van Veen, D., Gardezi, S.J.S., Paschali, M., Chen, Z., Delbrouck, J.-B., Reis, E., Truyts, C., et al.: Merlin: A vision language foundation model for 3d computed tomography. arXiv preprint arXiv:2406.06512 (2024) 27

  70. [78]

    Advances in neural information processing systems 36 (2024)

    Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural information processing systems 36 (2024)

  71. [79]

    NEJM AI 1(3), 2300138 (2024)

    Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I.,et al.: Towards generalist biomedical ai. NEJM AI 1(3), 2300138 (2024)

  72. [80]

    arXiv preprint arXiv:2405.07988 (2024)

    Zhou, H.-Y., Adithan, S., Acosta, J.N., Topol, E.J., Rajpurkar, P.: A gen- eralist learner for multifaceted medical image interpretation. arXiv preprint arXiv:2405.07988 (2024)

  73. [81]

    arXiv preprint arXiv:2308.02463 (2023)

    Wu, C., Zhang, X., Zhang, Y., Wang, Y., Xie, W.: Towards generalist foundation model for radiology. arXiv preprint arXiv:2308.02463 (2023)

  74. [82]

    NEJM AI, 2400640 (2024)

    Zhang, S., Xu, Y., Usuyama, N., Xu, H., Bagga, J., Tinn, R., Preston, S., Rao, R., Wei, M., Valluri, N., et al.: A multimodal biomedical foundation model trained from fifteen million image–text pairs. NEJM AI, 2400640 (2024)

  75. [83]

    In: European Conference on Computer Vision, pp

    Boecking, B., Usuyama, N., Bannur, S., Castro, D.C., Schwaighofer, A., Hyland, S., Wetscherek, M., Naumann, T., Nori, A., Alvarez-Valle, J., et al.: Making the most of text semantics to improve biomedical vision–language processing. In: European Conference on Computer Vision, ...

  76. [84]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Bannur, S., Hyland, S., Liu, Q., Perez-Garcia, F., Ilse, M., Castro, D.C., Boecking, B., Sharma, H., Bouzid, K., Thieme, A., et al.: Learning to exploit temporal structure for biomedical vision-language processing. In: Proceedings of the IEEE/CVF Conference on Computer Vision ...

  77. [85]

    In: Machine Learning for Healthcare Conference, pp

    Zhang, Y., Jiang, H., Miura, Y., Manning, C.D., Langlotz, C.P.: Contrastive learning of medical visual representations from paired images and text. In: Machine Learning for Healthcare Conference, pp. 2–25 (2022). PMLR

  78. [86]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Javed, S., Mahmood, A., Ganapathi, I.I., Dharejo, F.A., Werghi, N., Ben- namoun, M.: Cplip: Zero-shot learning for histopathology with comprehensive vision-language alignment. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11450–11459 (2024)

  79. [87]

    arXiv preprint arXiv:2405.03162 (2024)

    Yang, L., Xu, S., Sellergren, A., Kohlberger, T., Zhou, Y., Ktena, I., Kiraly, A., Ahmed, F., Hormozdiari, F., Jaroensri, T., et al.: Advancing multimodal medical capabilities of gemini. arXiv preprint arXiv:2405.03162 (2024)

  80. [88]

    In: ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

    Liu, C., Wan, Z., Cheng, S., Zhang, M., Arcucci, R.: Etp: Learning transferable ecg representations via ecg-text pre-training. In: ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 8230–8234 (2024). IEEE 28

  81. [89]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Luo, Y., Shi, M., Khan, M.O., Afzal, M.M., Huang, H., Yuan, S., Tian, Y., Song, L., Kouhana, A., Elze, T., et al.: Fairclip: Harnessing fairness in vision-language learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12289–12301 (2024)

  82. [90]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Li, H., Chen, Y., Chen, Y., Yu, R., Yang, W., Wang, L., Ding, B., Han, Y.: Generalizable whole slide image classification with fine-grained visual-semantic interaction. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11398–11407 (2024)

  83. [91]

    In: Medical Imaging with Deep Learning, pp

    Keicher, M., Zaripova, K., Czempiel, T., Mach, K., Khakzar, A., Navab, N.: Flexr: few-shot classification with language embeddings for structured reporting of chest x-rays. In: Medical Imaging with Deep Learning, pp. 1493–1508 (2024). PMLR

  84. [92]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

    Huang, S.-C., Shen, L., Lungren, M.P., Yeung, S.: Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3942–3951 (2021)

  85. [93]

    Nature Communications 14(1), 4542 (2023)

    Zhang, X., Wu, C., Zhang, Y., Xie, W., Wang, Y.: Knowledge-enhanced visual- language pre-training on chest radiology images. Nature Communications 14(1), 4542 (2023)

  86. [94]

    Nature Communications 15(1), 7620 (2024)

    Huang, W., Li, C., Zhou, H.-Y., Yang, H., Liu, J., Liang, Y., Zheng, H., Zhang, S., Wang, S.: Enhancing representation in radiography-reports founda- tion model: A granular alignment algorithm using masked contrastive learning. Nature Communications 15(1), 7620 (2024)

  87. [95]

    IEEE Transactions on Medical Imaging (2024)

    Wang, P., Zhang, H., Yuan, Y.: Mcpl: Multi-modal collaborative prompt learn- ing for medical vision-language model. IEEE Transactions on Medical Imaging (2024)

  88. [96]

    arXiv preprint arXiv:2410.06542 (2024)

    Codella, N.C., Jin, Y., Jain, S., Gu, Y., Lee, H.H., Abacha, A.B., Santamaria- Pang, A., Guyman, W., Sangani, N., Zhang, S., et al.: Medimageinsight: An open-source embedding model for general domain medical imaging. arXiv preprint arXiv:2410.06542 (2024)

  89. [97]

    NPJ Digital Medicine 6(1), 226 (2023)

    Liu, F., Zhu, T., Wu, X., Yang, B., You, C., Wang, C., Lu, L., Liu, Z., Zheng, Y., Sun, X., et al.: A medical multimodal large language model for future pandemics. NPJ Digital Medicine 6(1), 226 (2023)

  90. [98]

    IEEE Journal of Biomedical and Health Informatics 26(12), 6070–6080 (2022) 29

    Moon, J.H., Lee, H., Shin, W., Kim, Y.-H., Choi, E.: Multi-modal understanding and generation for medical images and text via vision-language pre-training. IEEE Journal of Biomedical and Health Informatics 26(12), 6070–6080 (2022) 29

  91. [99]

    Nature Machine Intelligence 5(12), 1447–1457 (2023)

    Liu, S., Nie, W., Wang, C., Lu, J., Qiao, Z., Liu, L., Tang, J., Xiao, C., Anand- kumar, A.: Multi-modal molecule structure–text model for text-based retrieval and editing. Nature Machine Intelligence 5(12), 1447–1457 (2023)

  92. [100]

    Bioinformatics 40(Supplement 1), 357–368 (2024)

    Tang, X., Tran, A., Tan, J., Gerstein, M.B.: Mollm: a unified language model for integrating biomedical text with 2d and 3d molecular representations. Bioinformatics 40(Supplement 1), 357–368 (2024)

  93. [101]

    Nature medicine 29(9), 2307–2316 (2023)

    Huang, Z., Bianchi, F., Yuksekgonul, M., Montine, T.J., Zou, J.: A visual– language foundation model for pathology image analysis using medical twitter. Nature medicine 29(9), 2307–2316 (2023)

  94. [102]

    Nature, 1–8 (2024)

    Xu, H., Usuyama, N., Bagga, J., Zhang, S., Rao, R., Naumann, T., Wong, C., Gero, Z., Gonz´ alez, J., Gu, Y., et al.: A whole-slide foundation model for digital pathology from real-world data. Nature, 1–8 (2024)

  95. [103]

    arXiv preprint arXiv:2412.10372 (2024)

    Khattak, M.U., Kunhimon, S., Naseer, M., Khan, S., Khan, F.S.: Unimed-clip: Towards a unified image-text pretraining paradigm for diverse medical imaging modalities. arXiv preprint arXiv:2412.10372 (2024)

  96. [104]

    In: Inter- national Conference on Medical Image Computing and Computer-Assisted Intervention, pp

    Pellegrini, C., Keicher, M., ¨Ozsoy, E., Jiraskova, P., Braren, R., Navab, N.: Xplainer: From x-ray observations to explainable zero-shot diagnosis. In: Inter- national Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 420–429 (2023). Springer

  97. [105]

    Biomedical Engineering Letters 14(2), 209–220 (2024)

    Ran, A., Liu, H.: Joint spatio-temporal features constrained self-supervised elec- trocardiogram representation learning. Biomedical Engineering Letters 14(2), 209–220 (2024)

  98. [106]

    Biomedical Engineering Letters 13(2), 197–207 (2023)

    Kang, Y., Yang, G., Eom, H., Han, S., Baek, S., Noh, S., Shin, Y., Park, C.: Gan- based patient information hiding for an ecg authentication system. Biomedical Engineering Letters 13(2), 197–207 (2023)

  99. [107]

    Medical Image Analysis 82, 102630 (2022)

    Alsharid, M., Cai, Y., Sharma, H., Drukker, L., Papageorghiou, A.T., Noble, J.A.: Gaze-assisted automatic captioning of fetal ultrasound videos using three- way multi-modal deep neural networks. Medical Image Analysis 82, 102630 (2022)

  100. [108]

    arXiv preprint arXiv:2407.16684 (2024)

    Lei, J., Zhang, X., Wu, C., Dai, L., Zhang, Y., Zhang, Y., Wang, Y., Xie, W., Li, Y.: Autorg-brain: Grounded report generation for brain mri. arXiv preprint arXiv:2407.16684 (2024)

  101. [109]

    arXiv preprint arXiv:2406.13173 (2024)

    Cui, H., Mao, L., Liang, X., Zhang, J., Ren, H., Li, Q., Li, X., Yang, C.: Biomed- ical visual instruction tuning with clinician preference alignment. arXiv preprint arXiv:2406.13173 (2024)

  102. [110]

    arXiv preprint arXiv:2302.07257 (2023)

    Wang, S., Zhao, Z., Ouyang, X., Wang, Q., Shen, D.: Chatcad: Interactive 30 computer-aided diagnosis on medical image using large language models. arXiv preprint arXiv:2302.07257 (2023)

  103. [111]

    arXiv preprint arXiv:2401.12208 (2024)

    Chen, Z., Varma, M., Delbrouck, J.-B., Paschali, M., Blankemeier, L., Van Veen, D., Valanarasu, J.M.J., Youssef, A., Cohen, J.P., Reis, E.P., et al.: Chexa- gent: Towards a foundation model for chest x-ray interpretation. arXiv preprint arXiv:2401.12208 (2024)

  104. [112]

    In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp

    Gu, T., Liu, D., Li, Z., Cai, W.: Complex organ mask guided radiology report generation. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 7995–8004 (2024)

  105. [113]

    npj Digital Medicine 7(1), 111 (2024)

    Chen, X., Zhang, W., Xu, P., Zhao, Z., Zheng, Y., Shi, D., He, M.: Ffa-gpt: an automated pipeline for fundus fluorescein angiography interpretation and question-answer. npj Digital Medicine 7(1), 111 (2024)

  106. [114]

    arXiv preprint arXiv:2305.16037 (2023)

    Hamamci, I.E., Er, S., Sekuboyina, A., Simsar, E., Tezcan, A., Simsek, A.G., Esirgun, S.N., Almas, F., Dogan, I., Dasdelen, M.F., et al.: Generatect: Text- conditional generation of 3d chest ct volumes. arXiv preprint arXiv:2305.16037 (2023)

  107. [115]

    arXiv preprint arXiv:2303.00091 (2023)

    Huh, J., Park, S., Lee, J.E., Ye, J.C.: Improving medical speech-to-text accu- racy with vision-language pre-training model. arXiv preprint arXiv:2303.00091 (2023)

  108. [116]

    IEEE transactions on medical imaging (2023)

    Li, Z., Li, Y., Li, Q., Wang, P., Guo, D., Lu, L., Jin, D., Zhang, Y., Hong, Q.: Lvit: language meets vision transformer in medical image segmentation. IEEE transactions on medical imaging (2023)

  109. [117]

    arXiv preprint arXiv:2404.00578 (2024)

    Bai, F., Du, Y., Huang, T., Meng, M.Q.-H., Zhao, B.: M3d: Advancing 3d medical image analysis with multi-modal large language models. arXiv preprint arXiv:2404.00578 (2024)

  110. [118]

    arXiv preprint arXiv:2411.11362 (2024)

    Sharma, H., Salvatelli, V., Srivastav, S., Bouzid, K., Bannur, S., Castro, D.C., Ilse, M., Bond-Taylor, S., Ranjit, M.P., Falck, F., et al.: Maira-seg: Enhancing radiology report generation with segmentation-aware multimodal large language models. arXiv preprint arXiv:2411.113...

  111. [119]

    In: Machine Learning for Health (ML4H), pp

    Moor, M., Huang, Q., Wu, S., Yasunaga, M., Dalmia, Y., Leskovec, J., Zakka, C., Reis, E.P., Rajpurkar, P.: Med-flamingo: a multimodal medical few-shot learner. In: Machine Learning for Health (ML4H), pp. 353–367 (2023). PMLR

  112. [120]

    In: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pp

    Khare, Y., Bagal, V., Mathew, M., Devi, A., Priyakumar, U.D., Jawahar, C.: Mmbert: Multimodal bert pretraining for improved medical vqa. In: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pp. 1033–1036 (2021). IEEE 31

  113. [121]

    arXiv preprint arXiv:2411.11943 (2024)

    Cao, X., Liang, K., Liao, K.-D., Gao, T., Ye, W., Chen, J., Ding, Z., Cao, J., Rehg, J.M., Sun, J.: Medical video generation for disease progression simulation. arXiv preprint arXiv:2411.11943 (2024)

  114. [122]

    Nature 634(8033), 466–473 (2024)

    Lu, M.Y., Chen, B., Williamson, D.F., Chen, R.J., Zhao, M., Chow, A.K., Ike- mura, K., Kim, A., Pouli, D., Patel, A.,et al.: A multimodal generative ai copilot for human pathology. Nature 634(8033), 466–473 (2024)

  115. [123]

    In: Pro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp

    Yellapragada, S., Graikos, A., Prasanna, P., Kurc, T., Saltz, J., Samaras, D.: Pathldm: Text conditioned latent diffusion model for histopathology. In: Pro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 5182–5191 (2024)

  116. [124]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Seyfioglu, M.S., Ikezogwo, W.O., Ghezloo, F., Krishna, R., Shapiro, L.: Quilt- llava: Visual instruction tuning by extracting localized narratives from open- source histopathology videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp...

  117. [125]

    Meta-Radiology 1(3), 100033 (2023)

    Wang, Z., Liu, L., Wang, L., Zhou, L.: R2gengpt: Radiology report generation with frozen llms. Meta-Radiology 1(3), 100033 (2023)

  118. [126]

    arXiv preprint arXiv:2410.00441 (2024)

    Luo, L., Vairavamurthy, J., Zhang, X., Kumar, A., Ter-Oganesyan, R.R., Schroff, S.T., Shilo, D., Hossain, R., Moritz, M., Rajpurkar, P.: Rexplain: Translating radiology into patient-friendly video reports. arXiv preprint arXiv:2410.00441 (2024)

  119. [127]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Tanida, T., M¨ uller, P., Kaissis, G., Rueckert, D.: Interactive and explainable region-guided radiology report generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7433–7442 (2023)

  120. [128]

    Nature Biomedical Engineering, 1–13 (2024)

    Bluethgen, C., Chambon, P., Delbrouck, J.-B., Sluijs, R., Po lacin, M., Zam- brano Chaves, J.M., Abraham, T.M., Purohit, S., Langlotz, C.P., Chaudhari, A.S.: A vision–language foundation model for the generation of realistic chest x-ray images. Nature Biomedical Engineering, 1...

  121. [129]

    Nature Communications 15(1), 5649 (2024)

    Zhou, J., He, X., Sun, L., Xu, J., Chen, X., Chu, Y., Zhou, L., Liao, X., Zhang, B., Afvari, S., et al.: Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4. Nature Communications 15(1), 5649 (2024)

  122. [130]

    Information Fusion 113, 102602 (2025)

    Bai, L., Wang, G., Islam, M., Seenivasan, L., Wang, A., Ren, H.: Surgical- vqla++: Adversarial contrastive learning for calibrated robust visual question- localized answering in robotic surgery. Information Fusion 113, 102602 (2025)

  123. [131]

    In: Proceedings of the IEEE/CVF International Conference on 32 Computer Vision, pp

    Liu, J., Zhang, Y., Chen, J.-N., Xiao, J., Lu, Y., A Landman, B., Yuan, Y., Yuille, A., Tang, Y., Zhou, Z.: Clip-driven universal model for organ segmentation and tumor detection. In: Proceedings of the IEEE/CVF International Conference on 32 Computer Vision, pp. 21152–21164 (2023)

  124. [132]

    In: The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track

    Wang, Y., Dai, Y., Jones, C., Sair, H.I., Shen, J., Loizou, N., Hsu, W.-C., Imami, M.R., Jiao, Z., Zhang, P.J., et al.: Enhancing vision-language models for medical imaging: bridging the 3d gap with innovative slice selection. In: The Thirty-eight Conference on Neural Informat...

  125. [133]

    Scientific Reports 13(1), 7303 (2023)

    Khader, F., M¨ uller-Franzes, G., Tayebi Arasteh, S., Han, T., Haarburger, C., Schulze-Hagen, M., Schad, P., Engelhardt, S., Baeßler, B., Foersch, S., et al.: Denoising diffusion probabilistic models for 3d medical image generation. Scientific Reports 13(1), 7303 (2023)

  126. [134]

    In: Proceedings of the AAAI Conference on Artificial Intelligence, vol

    Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., et al.: Chexpert: A large chest radio- graph dataset with uncertainty labels and expert comparison. In: Proceedings of the AAAI Conference on Artificial ...

  127. [135]

    arXiv preprint arXiv:2408.02900 (2024)

    Xie, Y., Zhou, C., Gao, L., Wu, J., Li, X., Zhou, H.-Y., Liu, S., Xing, L., Zou, J., Xie, C., et al.: Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine. arXiv preprint arXiv:2408.02900 (2024)

  128. [136]

    Scientific Data 10(1), 158 (2023)

    Gupta, D., Attal, K., Demner-Fushman, D.: A dataset for medical instructional video classification and question answering. Scientific Data 10(1), 158 (2023)

  129. [137]

    In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Hu, Y., Li, T., Lu, Q., Shao, W., He, J., Qiao, Y., Luo, P.: Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22170–22183 (2024)

  130. [138]

    Medical image analysis 66, 101797 (2020)

    Bustos, A., Pertusa, A., Salinas, J.-M., De La Iglesia-Vaya, M.: Padchest: A large chest x-ray image dataset with multi-label annotated reports. Medical image analysis 66, 101797 (2020)

  131. [139]

    arXiv preprint arXiv:2003.10286 (2020)

    He, X., Zhang, Y., Mou, L., Xing, E., Xie, P.: Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:2003.10286 (2020)

  132. [140]

    arXiv preprint arXiv:2406.19280 (2024)

    Chen, J., Ouyang, R., Gao, A., Chen, S., Chen, G.H., Wang, X., Zhang, R., Cai, Z., Ji, K., Yu, G., et al.: Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale. arXiv preprint arXiv:2406.19280 (2024)

  133. [141]

    Advances in neural information processing systems36 (2024)

    Ikezogwo, W., Seyfioglu, S., Ghezloo, F., Geva, D., Sheikh Mohammed, F., Anand, P.K., Krishna, R., Shapiro, L.: Quilt-1m: One million image-text pairs for histopathology. Advances in neural information processing systems36 (2024)

  134. [142]

    In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp

    Pellegrini, C., Keicher, M., ¨Ozsoy, E., Navab, N.: Rad-restruct: A novel vqa 33 benchmark and method for structured radiology reporting. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 409–419 (2023). Springer

  135. [143]

    In: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pp

    Liu, B., Zhan, L.-M., Xu, L., Ma, L., Yang, Y., Wu, X.-M.: Slake: A semantically- labeled knowledge-enhanced dataset for medical visual question answering. In: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pp. 1650–1654 (2021). IEEE

  136. [144]

    Scientific data 5(1), 1–10 (2018)

    Lau, J.J., Gayen, S., Ben Abacha, A., Demner-Fushman, D.: A dataset of clini- cally generated visual questions and answers about radiology images. Scientific data 5(1), 1–10 (2018)

  137. [145]

    Accessed: 2024-11-04 (2024)

    codabench: MICCAI24 AMOS-MM: ABDOMINAL MULTIMODAL ANALY- SIS CHALLENGE. Accessed: 2024-11-04 (2024). https://www.codabench.org/ competitions/3137/

  138. [146]

    Advances in Neural Information Processing Systems 35, 36722–36732 (2022)

    Ji, Y., Bai, H., Ge, C., Yang, J., Zhu, Y., Zhang, R., Li, Z., Zhanng, L., Ma, W., Wan, X., et al.: Amos: A large-scale abdominal multi-organ benchmark for ver- satile medical image segmentation. Advances in Neural Information Processing Systems 35, 36722–36732 (2022)

  139. [147]

    In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp

    Chen, Y., Liu, C., Liu, X., Arcucci, R., Xiong, Z.: Bimcv-r: A landmark dataset for 3d ct text-image retrieval. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 124–134 (2024). Springer

  140. [148]

    arXiv preprint arXiv:2404.16754 (2024)

    Zhang, X., Wu, C., Zhao, Z., Lei, J., Zhang, Y., Wang, Y., Xie, W.: Radgenome- chest ct: A grounded vision-language dataset for chest ct analysis. arXiv preprint arXiv:2404.16754 (2024)

  141. [149]

    British journal of cancer 119(4), 508–516 (2018)

    Saha, A., Harowicz, M.R., Grimm, L.J., Kim, C.E., Ghate, S.V., Walsh, R., Mazurowski, M.A.: A machine learning approach to radiogenomics of breast cancer: a study of 922 subjects and 529 dce-mri features. British journal of cancer 119(4), 508–516 (2018)

  142. [150]

    Scientific data 7(1), 1–15 (2020)

    Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F.I., Samek, W., Schaeffter, T.: Ptb-xl, a large publicly available electrocardiography dataset. Scientific data 7(1), 1–15 (2020)

  143. [151]

    arXiv preprint arXiv:2302.04611 (2023)

    Liu, S., Li, Y., Li, Z., Gitter, A., Zhu, Y., Lu, J., Xu, Z., Nie, W., Ramanathan, A., Xiao, C., et al.: A text-guided protein design framework. arXiv preprint arXiv:2302.04611 (2023)

  144. [152]

    arXiv preprint arXiv:2411.15122 (2024) 34

    Zhang, X., Zhou, H.-Y., Yang, X., Banerjee, O., Acosta, J.N., Miller, J., Huang, O., Rajpurkar, P.: Rexrank: A public leaderboard for ai-powered radiology report generation. arXiv preprint arXiv:2411.15122 (2024) 34

  145. [153]

    Nature Communications 15(1), 10104 (2024)

    Ferber, D., W¨ olflein, G., Wiest, I.C., Ligero, M., Sainath, S., Ghaffari Laleh, N., El Nahhas, O.S., M¨ uller-Franzes, G., J¨ ager, D., Truhn, D.,et al.: In-context learning enables multimodal large language models to classify cancer pathology images. Nature Communications 1...

  146. [154]

    arXiv preprint arXiv:2405.03595 (2024)

    Ostmeier, S., Xu, J., Chen, Z., Varma, M., Blankemeier, L., Bluethgen, C., Michalson, A.E., Moseley, M., Langlotz, C., Chaudhari, A.S., et al.: Green: Generative radiology report evaluation and error notation. arXiv preprint arXiv:2405.03595 (2024)

  147. [155]

    medRxiv, 2024–06 (2024)

    Zhao, W., Wu, C., Zhang, X., Zhang, Y., Wang, Y., Xie, W.: Ratescore: A metric for radiology report generation. medRxiv, 2024–06 (2024)

  148. [156]

    Patterns 4(9) (2023)

    Yu, F., Endo, M., Krishnan, R., Pan, I., Tsai, A., Reis, E.P., Fonseca, E.K.U.N., Lee, H.M.H., Abad, Z.S.H., Ng, A.Y., et al.: Evaluating progress in automatic chest x-ray radiology report generation. Patterns 4(9) (2023)

  149. [157]

    In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp

    Papineni, K., Roukos, S., Ward, T., Zhu, W.-J.: Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311–318 (2002)

  150. [158]

    In: Text Summarization Branches Out, pp

    Lin, C.-Y.: Rouge: A package for automatic evaluation of summaries. In: Text Summarization Branches Out, pp. 74–81 (2004)

  151. [159]

    In: Proceedings of the Acl Work- shop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation And/or Summarization, pp

    Banerjee, S., Lavie, A.: Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In: Proceedings of the Acl Work- shop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation And/or Summarization, pp. 65–72 (2005)

  152. [160]

    arXiv preprint arXiv:2004.09167 (2020)

    Smit, A., Jain, S., Rajpurkar, P., Pareek, A., Ng, A.Y., Lungren, M.P.: Chexbert: combining automatic labelers and expert annotations for accurate radiology report labeling using bert. arXiv preprint arXiv:2004.09167 (2020)

  153. [161]

    Nature Medicine (2025) https://doi.org/10.1038/ s41591-024-03328-5

    Johri, S., Jeong, J., Tran, B.A., Schlessinger, D.I., Wongvibulsin, S., Barnes, L.A., Zhou, H.-Y., Cai, Z.R., Van Allen, E.M., Kim, D., Daneshjou, R., Rajpurkar, P.: An evaluation framework for clinical use of large language models in patient interaction tasks. Nature Medicine...

  154. [162]

    arXiv preprint arXiv:2405.20613 (2024)

    Huang, A., Banerjee, O., Wu, K., Reis, E.P., Rajpurkar, P.: Fineradscore: A radiology report line-by-line evaluation technique generating corrections with severity scores. arXiv preprint arXiv:2405.20613 (2024)

  155. [163]

    NPJ Digital Medicine 7(1), 124 (2024) 35

    Ong Ly, C., Unnikrishnan, B., Tadic, T., Patel, T., Duhamel, J., Kandel, S., Moayedi, Y., Brudno, M., Hope, A., Ross, H.,et al.: Shortcut learning in medical ai hinders generalization: method for estimating ai model generalization without external data. NPJ Digital Medicine 7(...

  156. [164]

    Advances in neural information processing systems 30 (2017)

    Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017)

  157. [165]

    In: Biocomputing 2025: Proceedings of the Pacific Symposium, pp

    Banerjee, O., Saenz, A., Wu, K., Clements, W., Zia, A., Buensalido, D., Kav- noudias, H., Abi-Ghanem, A.S., Ghawi, N.E., Luna, C., et al.: Rexamine-global: A framework for uncovering inconsistencies in radiology report generation met- rics. In: Biocomputing 2025: Proceedings o...

  158. [166]

    Journal of Medical Internet Research 25, 48009 (2023)

    Wang, C., Liu, S., Yang, H., Guo, J., Wu, Y., Liu, J.: Ethical considerations of using chatgpt in health care. Journal of Medical Internet Research 25, 48009 (2023)

  159. [167]

    The Lancet Digital Health 5(6), 333–335 (2023)

    Li, H., Moon, J.T., Purkayastha, S., Celi, L.A., Trivedi, H., Gichoya, J.W.: Ethics of large language models in medicine and medical research. The Lancet Digital Health 5(6), 333–335 (2023)

  160. [168]

    arXiv preprint arXiv:2411.18730 (2024)

    Paschali, M., Chen, Z., Blankemeier, L., Varma, M., Youssef, A., Bluethgen, C., Langlotz, C., Gatidis, S., Chaudhari, A.: Foundation models in radiology: What, how, when, why and why not. arXiv preprint arXiv:2411.18730 (2024)

  161. [169]

    arXiv preprint arXiv:2412.01233 (2024)

    Bluethgen, C., Van Veen, D., Zakka, C., Link, K., Fanous, A., Daneshjou, R., Frauenfelder, T., Langlotz, C., Gatidis, S., Chaudhari, A.: Best practices for large language models in radiology. arXiv preprint arXiv:2412.01233 (2024)

  162. [170]

    Nature medicine 30(9), 2613–2622 (2024)

    Hager, P., Jungmann, F., Holland, R., Bhagat, K., Hubrecht, I., Knauer, M., Vielhauer, J., Makowski, M., Braren, R., Kaissis, G., et al.: Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature medicine 30(9), 2613–2622 (2024)

  163. [171]

    arXiv preprint arXiv:2305.20050 (2023)

    Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., Cobbe, K.: Let’s verify step by step. arXiv preprint arXiv:2305.20050 (2023)

  164. [172]

    Chua, M., Kim, D., Choi, J., Lee, N.G., Deshpande, V., Schwab, J., Lev, M.H., Gonzalez, R.G., Gee, M.S., Do, S.: Tackling prediction uncertainty in machine learning for healthcare. Nature Biomedical Engineering 7(6), 711–718 (2023) 36 A Supplementary information A.1 PRISMA-ScR...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.