Pith. sign in

REVIEW 10 cited by

MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.02900 v3 pith:QZQKWTIQ submitted 2024-08-06 cs.CV

classification cs.CV
keywords multimodalannotationsmedtrinity-25mmodelsmultigranularcomprehensivedatasetlarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotations for more than 65 diseases. These multigranular annotations encompass both global information, such as modality and organ detection, and local information like ROI analysis, lesion texture, and region-wise correlations. Unlike the existing multimodal datasets, which are limited by the availability of image-text pairs, we have developed the first automated pipeline that scales up multimodal data by generating multigranular visual and textual annotations in the form of image-ROI-description triplets without the need for any paired text descriptions. Specifically, data from over 30 different sources have been collected, preprocessed, and grounded using domain-specific expert models to identify ROIs related to abnormal regions. We then build a comprehensive knowledge base and prompt multimodal large language models to perform retrieval-augmented generation with the identified ROIs as guidance, resulting in multigranular textual descriptions. Compared to existing datasets, MedTrinity-25M provides the most enriched annotations, supporting a comprehensive range of multimodal tasks such as captioning and report generation, as well as vision-centric tasks like classification and segmentation. We propose LLaVA-Tri by pretraining LLaVA on MedTrinity-25M, achieving state-of-the-art performance on VQA-RAD, SLAKE, and PathVQA, surpassing representative SOTA multimodal large language models. Furthermore, MedTrinity-25M can also be utilized to support large-scale pre-training of multimodal medical AI models, contributing to the development of future foundation models in the medical domain. We will make our dataset available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 8.0 of 10

    Introduces the first large-scale multimodal benchmark MedLayXPlain-122K showing medical VLMs suffer significant lay-register degradation while general VLMs lack clinical precision.

  2. CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

    cs.CV 2026-05 conditional novelty 7.0 of 10

    Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.

  3. Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Introduces VietPET-RoI dataset with fine-grained RoI annotations for Vietnamese 3D PET/CT and HiRRA graph framework that improves report generation by modeling region dependencies, claiming large gains over prior models.

  4. Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework

    cs.CV 2026-04 conditional novelty 7.0 of 10

    Introduces the first large-scale 3D PET/CT dataset with fine-grained RoI annotations for Vietnamese and a graph-enhanced HiRRA framework that achieves SOTA report generation by modeling RoI dependencies.

  5. Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    ViToS uses dual-stream RL with cross-feedback optimization to prune medical image tokens to 77% length while reporting 108.27% and 104.16% relative performance on two 7B VLMs across seven benchmarks.

  6. MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Current medical multimodal models, including GPT-4o and Claude 3.5 Sonnet, fail simple perceptual tasks on medical images that human experts solve almost perfectly.

  7. Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    FundusExpert, an 8B ophthalmic MLLM trained on region-grounded cognitive-chain instructions, reports state-of-the-art QA and report-generation results, with a fitted data-scaling exponent of 0.068.

  8. TextSLIP: Text Self-Supervised CLIP for Medical Report Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Adding ESimCSE text contrastive learning to CLIP improves medical report generation BLEU scores on brain MRI over standard CLIP by about 1-3 points.

  9. Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models

    cs.CV 2026-07 reject novelty 5.0 of 10

    A synthetic chain-of-thought dataset generated from CT reports lets a 2D-pretrained medical MLLM improve on 3D CT spatial-reasoning benchmarks.

  10. MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence

    cs.CV 2026-03 reject novelty 5.0 of 10

    MEDIC-AD adds anomaly-aware and difference tokens to a medical VLM, claiming SOTA lesion detection, temporal tracking, and visual grounding; the zero-shot claim is undermined by likely train/test overlap.

Pith tools