Pith. sign in

REVIEW 41 cited by

HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19280 v4 pith:FFO3UGVT submitted 2024-06-27 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords medicaldatamultimodalmllmspubmedvisioncapabilitiesdatasetgpt-4v
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid development of multimodal large language models (MLLMs), such as GPT-4V, has led to significant advancements. However, these models still face challenges in medical multimodal capabilities due to limitations in the quantity and quality of medical vision-text data, stemming from data privacy concerns and high annotation costs. While pioneering approaches utilize PubMed's large-scale, de-identified medical image-text pairs to address these limitations, they still fall short due to inherent data noise. To tackle this, we refined medical image-text pairs from PubMed and employed MLLMs (GPT-4V) in an 'unblinded' capacity to denoise and reformat the data, resulting in the creation of the PubMedVision dataset with 1.3 million medical VQA samples. Our validation demonstrates that: (1) PubMedVision can significantly enhance the medical multimodal capabilities of current MLLMs, showing significant improvement in benchmarks including the MMMU Health & Medicine track; (2) manual checks by medical experts and empirical results validate the superior data quality of our dataset compared to other data construction methods. Using PubMedVision, we train a 34B medical MLLM HuatuoGPT-Vision, which shows superior performance in medical multimodal scenarios among open-source MLLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 41 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new five-level medical imaging benchmark, DrVD-Bench, shows that vision-language models lose accuracy sharply as reasoning complexity grows and often diagnose without grounding in lesion evidence.

  2. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A cascaded multi-encoder medical MLLM with native 3D fusion and RoI-grounded report metrics claims SOTA on most 2D/3D medical benchmarks and highest radiologist report rankings.

  3. OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

    cs.CL 2026-04 accept novelty 6.5 of 10

    OralAgent, a ReAct-style dental agent with 22 vision tools and a 134.8M-token textbook RAG corpus, reaches SOTA on MMOral-Uni, MMOral-OPG, and OralQA-ZH.

  4. MIRA: Medical Image Reflection for Agentic Diagnosis

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A two-stage pipeline, MCTS-based supervised fine-tuning plus GRPO reinforcement learning with a validation-gated reflection memory, makes an 8B medical vision-language model both more accurate and more selective in tool use.

  5. Evaluating and Understanding Model Editing for Medical Vision Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    M3Bench is a clinically grounded benchmark showing that gradient-based VLM editors generalize but break locality, while memory-based editors preserve locality but fail on composition and temporal tasks, with failures ...

  6. IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A dual-branch clinical data engine (Topic Finding Tree plus scene-driven dialogues) yields IRIS-120K and a 4B VLM that outperforms up to 34B medical VLMs on OSD VQA tasks.

  7. Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain

    cs.AI 2026-03 conditional novelty 6.0 of 10

    A fixed, sampling-free score — per-token log-probability variance times (1 + average |image-vs-text probability shift|) — detects medical-VQA hallucinations better than semantic-entropy baselines in 13 of 16 settings.

  8. Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

    cs.CV 2026-03 conditional novelty 6.0 of 10

    SpatialMed provides the first CT-based benchmark of 3D spatial reasoning for medical MLLMs, on which 14 models perform near chance, particularly for distance and volume estimation.

  9. CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space

    cs.LG 2025-12 conditional novelty 6.0 of 10

    A latent-space world model that forecasts treatment-conditioned tumor trajectories and selects therapies by iteratively minimizing its own predicted risk score.

  10. Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Tracking positive shifts in visual attention over information-rich query words yields a saliency map that, when used to boost visual and query attention during decoding, reduces object hallucination on CHAIR, POPE, an...

  11. Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

    cs.CV 2025-09 reject novelty 6.0 of 10

    MMOral is a large new dental X-ray instruction dataset and benchmark, but the proposed model's 24.73% improvement is from fine-tuning and then testing on the same data pool.

  12. Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 3D encoder pretrained with GPT-4V slice captions and partial optimal transport alignment beats vision-only SSL baselines on several medical tasks, but a key evaluation dataset may overlap with pretraining.

  13. CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A preference optimization strategy using confidence-based hard example mining, similarity retrieval, and synthetic counterfactual rationales improves chest X-ray VQA accuracy by 8.93% relative over supervised fine-tuning.

  14. Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MeWM combines a GPT-style policy, a diffusion tumor dynamics model, and a survival analysis heuristic to simulate post-treatment tumor appearance and select TACE treatment plans, improving physician F1-score by 13 points.

  15. MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MedBookVQA is a new 5,000-question, textbook-derived multimodal benchmark for testing medical AI systems, with labels for imaging modality, body anatomy, and clinical specialty.

  16. Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 8-stage chest X-ray VQA benchmark and a context-aware model trained on it.

  17. Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A3Tune aligns the visual attention of medical LVLMs to prompt-relevant regions via SAM and BioMedCLIP weak labels plus a Mixture-of-Experts over LoRA, improving VQA and report generation accuracy.

  18. Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Multimodal LLMs trained on medical images decomposed into modality, anatomy, and task can generalize to unseen combinations of those elements, and this compositional generalization partially explains multi-task traini...

  19. On Domain-Adaptive Post-Training for Multimodal Large Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A generate-then-filter, open-source-only synthesis pipeline plus single-stage post-training consistently improves MLLM performance across biomedicine, food, and remote sensing.

  20. CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A training-free decoding method that splits images into complementary SAM-based parts and adaptively contrasts token distributions to reduce hallucinations in vision-language models.

  21. MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

    cs.CV 2026-08 reject novelty 5.0 of 10

    A medical vision-language model that represents masks as discrete tokens and is trained on a large four-stream corpus reports strong unified results, though its benchmark comparisons are not yet fair or fully held out.

  22. Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    PMC-InterCPT builds a context-grounded biomedical interleaved corpus from PMC literature and shows it improves multimodal performance on Qwen3.5-4B-Base after CPT and SFT while using fewer tokens.

  23. Scalable and Private Federated Learning Using Distributed Differential Privacy and Secure Aggregation

    cs.CR 2026-04 unverdicted novelty 5.0 of 10

    DDP-SA combines client-side Laplace noise perturbation with full-threshold additive secret sharing to let federated learning servers reconstruct only aggregated noisy gradients without exposing individual client updates.

  24. M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

    eess.IV 2026-01 conditional novelty 5.0 of 10

    A medical-image benchmark that scores the step-by-step reasoning chains of multimodal LLMs shows current models explain poorly and chain-of-thought prompting frequently reduces diagnostic accuracy.

  25. Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning

    cs.CL 2025-08 conditional novelty 5.0 of 10

    RoMed and CCL: a 144k-question perturbation benchmark for medical VQA and a consistency-plus-contrastive training method that improves LLaVA-Med's accuracy and reduces answer variation.

  26. MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    MedMKEB is a medical multimodal knowledge-editing benchmark with four task types, on which existing editing methods underperform according to the authors.

  27. CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    CX-Mind combines curriculum reinforcement learning and rule-based process rewards to train a chest X-ray vision-language model that produces interleaved think-answer reasoning and reports state-of-the-art results acro...

  28. Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Selecting the intermediate layer where image-conditioned and text-only predictions diverge most, and adding that layer's contrastive visual signal back to the final logits, reduces object hallucinations in four large ...

  29. HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    HSENet improves 3D CT vision-language understanding by combining global and local 3D encoders with a centroid-based spatial token compressor, posting state-of-the-art results on CT-RATE and RadGenome-ChestCT.

  30. RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints

    cs.CV 2025-06 reject novelty 5.0 of 10

    RARL fine-tunes Qwen2-VL-2B on 716 medical samples with GRPO, LoRA, and a vaguely defined reasoning reward, claiming gains of 7.78% over SFT on reasoning and up to 27% on unseen VQA benchmarks.

  31. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

    cs.CV 2025-02 reject novelty 5.0 of 10

    HealthGPT unifies medical image comprehension and generation in a single autoregressive model using heterogeneous low-rank adaptation, reporting strong benchmark results.

  32. EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

    cs.CV 2025-01 reject novelty 5.0 of 10

    EndoChat is a grounded multimodal LLM for endoscopic surgery, trained on the new Surg-396K dataset and reported to outperform prior MLLMs, though its evaluation is confounded by training-data overlap.

  33. GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

    cs.CV 2024-11 reject novelty 5.0 of 10

    GMAI-VL-5.5M is a new 5.5M-sample medical image-text dataset built from 219 datasets via GPT-4o annotation-guided generation, and GMAI-VL is a three-stage LLaVA-style model reporting SOTA numbers, though the evaluatio...

  34. ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics

    cs.LG 2026-04 unverdicted novelty 4.0 of 10

    Standard Conditional Flow Matching loss is a misleading early plateau; physics-informed metrics keep improving, so ScatterPrism and multi-metric diagnostics are needed for kinematic fidelity.

  35. How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

    cs.CV 2025-07 reject novelty 4.0 of 10

    A ten-model, seven-benchmark medical VLM evaluation whose headline reasoning-vs-understanding finding is contradicted by its own tables.

  36. Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models

    cs.CV 2025-07 reject novelty 4.0 of 10

    Expert-CFG combines entropy-based uncertainty selection with classifier-free guidance over expert-highlighted text to refine MedVLM outputs, reporting gains on VQA-RAD, SLAKE, and PathVQA.

  37. MIRA: A Novel Framework for Fusing Modalities in Medical RAG

    cs.CV 2025-07 reject novelty 4.0 of 10

    A medical multimodal RAG pipeline with rethink-and-rearrange and online search; the claimed SOTA is contradicted by the paper's own PMC-VQA numbers.

  38. Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.

  39. Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A curriculum-based GRPO schedule that first trains on close-ended medical VQA and then on open-ended VQA improves benchmark scores over joint training and vanilla RL, though the open-ended metric is the training objec...

  40. Spatial navigation in preclinical Alzheimer's disease: A review

    q-bio.NC 2026-03 unverdicted novelty 3.0 of 10

    Spatial navigation performance, particularly path integration and wayfinding, correlates with AD biomarkers such as p-tau in cognitively unimpaired at-risk individuals and may enable earlier detection than episodic me...

  41. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.

Pith tools