Pith. sign in

REVIEW 3 cited by

A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11217 v2 pith:4HK24QXB submitted 2024-02-17 cs.CL cs.CV

classification cs.CLcs.CV
keywords medicalmed-mllmsmodelsspecialtiesanalysisasclepiusbenchmarkclinical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the complexity of real-world diagnostics across diverse specialties. To address this gap, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gastroenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs' capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Current medical multimodal models, including GPT-4o and Claude 3.5 Sonnet, fail simple perceptual tasks on medical images that human experts solve almost perfectly.

  2. SpatialFly: Implicit 3D Prior-Guided Visual Reparameterization for Continuous UAV Vision-and-Language Navigation

    cs.CV 2026-03 unverdicted novelty 4.0 of 10

    SpatialFly reparameterizes 2D visual tokens with implicit geometric priors and reports lower navigation error and higher success than prior UAV VLN systems on unseen splits.

  3. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools