Pith. sign in

REVIEW 3 cited by

MultiMed: Massively Multimodal and Multitask Medical Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.12682 v1 pith:XW4GKH6U submitted 2024-08-22 cs.LG cs.AIcs.CLcs.CVcs.MM

classification cs.LGcs.AIcs.CLcs.CVcs.MM
keywords medicalmultimedacrossmodalitiestasksbiomedicaldatamultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Biomedical data is inherently multimodal, consisting of electronic health records, medical imaging, digital pathology, genome sequencing, wearable sensors, and more. The application of artificial intelligence tools to these multifaceted sensing technologies has the potential to revolutionize the prognosis, diagnosis, and management of human health and disease. However, current approaches to biomedical AI typically only train and evaluate with one or a small set of medical modalities and tasks. This limitation hampers the development of comprehensive tools that can leverage the rich interconnected information across many heterogeneous biomedical sensors. To address this challenge, we present MultiMed, a benchmark designed to evaluate and enable large-scale learning across a wide spectrum of medical modalities and tasks. MultiMed consists of 2.56 million samples across ten medical modalities such as medical reports, pathology, genomics, and protein data, and is structured into eleven challenging tasks, including disease prognosis, protein structure prediction, and medical question answering. Using MultiMed, we conduct comprehensive experiments benchmarking state-of-the-art unimodal, multimodal, and multitask models. Our analysis highlights the advantages of training large-scale medical models across many related modalities and tasks. Moreover, MultiMed enables studies of generalization across related medical concepts, robustness to real-world noisy data and distribution shifts, and novel modality combinations to improve prediction performance. MultiMed will be publicly available and regularly updated and welcomes inputs from the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Multimodal LLMs trained on medical images decomposed into modality, anatomy, and task can generalize to unseen combinations of those elements, and this compositional generalization partially explains multi-task traini...

  2. Leveraging the Structure of Medical Data for Improved Representation Learning

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Using paired frontal and lateral chest X-rays as self-supervision signals improves downstream pathology classification over a supervised baseline on MIMIC-CXR.

  3. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools