Pith. sign in

REVIEW 7 cited by

Multimodal Large Language Models for Medicine: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.21051 v1 pith:MSYKUDKS submitted 2025-04-29 cs.LG cs.CLcs.MM

Multimodal Large Language Models for Medicine: A Comprehensive Survey

classification cs.LG cs.CLcs.MM
keywords mllmsmedicalhealthcarecapabilitiescomprehensivedatadomaindomains
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, MLLMs have gained substantial attention from different domains. Researchers have begun to explore the potential of MLLMs in the medical and healthcare domain. In this paper, we first introduce the background and fundamental concepts related to LLMs and MLLMs, while emphasizing the working principles of MLLMs. Subsequently, we summarize three main directions of application within healthcare: medical reporting, medical diagnosis, and medical treatment. Our findings are based on a comprehensive review of 330 recent papers in this area. We illustrate the remarkable capabilities of MLLMs in these domains by providing specific examples. For data, we present six mainstream modes of data along with their corresponding evaluation benchmarks. At the end of the survey, we discuss the challenges faced by MLLMs in the medical and healthcare domain and propose feasible methods to mitigate or overcome these issues.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

    cs.CV 2026-04 unverdicted novelty 7.0

    Medical MLLMs degrade on image classification due to four failure modes in visual representation quality, connector projection fidelity, LLM comprehension, and semantic mapping alignment, quantified by feature probing...

  2. MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

    cs.LG 2026-05 unverdicted novelty 6.0

    MultiSeismo is a new multimodal seismic dataset with 16K events and SeisModal is a domain-adapted model that outperforms general multimodal models on seismic reasoning tasks.

  3. Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs

    cs.CL 2026-04 unverdicted novelty 5.0

    A Medical Entity Tree organizes medical knowledge to engineer higher-quality training data that boosts general MLLMs on medical benchmarks.

  4. Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models

    cs.AI 2025-11 conditional novelty 5.0

    Ablating just four neurons in LLaVA-1.5-7b's language-model down-projection layer triggers complete output collapse, with critical neurons concentrated in the language backbone.

  5. Generating Findings for Jaw Cysts in Dental Panoramic Radiographs Using a GPT-Based VLM: A Preliminary Study on Building a Two-Stage Self-Correction Loop with Structured Output (SLSO) Framework

    cs.CV 2025-10 unverdicted novelty 5.0

    The SLSO framework uses iterative structured output generation, consistency checks, and regeneration to improve GPT-VLM accuracy on jaw cyst findings in panoramic radiographs compared to standard chain-of-thought prompting.

  6. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

    cs.AI 2026-07 conditional novelty 4.5

    Medical agents should be scaled mainly by richer clinical environments and self-evolution loops, not parameter growth alone, under a three-level autonomy taxonomy.

  7. Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

    cs.CV 2026-04 unverdicted novelty 4.0

    Context alignment in medical VLMs raises AUC from 0.918 to 0.925, cuts hallucinated keywords from 1.14 to 0.25, shortens explanations to 15.3 words, and maintains calibrated uncertainty without raising model confidence.