Pith. sign in

REVIEW 8 cited by

MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.02718 v1 pith:H6FQWMT2 submitted 2024-08-05 cs.CV

classification cs.CV
keywords multi-imagemmiulvlmsmodelsunderstandingevaluationmultimodaltasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks. MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular LVLMs, including both open-source and proprietary models, reveals significant challenges in multi-image comprehension, particularly in tasks involving spatial understanding. Even the most advanced models, such as GPT-4o, achieve only 55.7% accuracy on MMIU. Through multi-faceted analytical experiments, we identify key performance gaps and limitations, providing valuable insights for future model and data improvements. We aim for MMIU to advance the frontier of LVLM research and development, moving us toward achieving sophisticated multimodal multi-image user interactions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

    cs.CV 2025-06 conditional novelty 7.0 of 10

    MMRB is the first benchmark combining multi-image inputs with chain-of-thought reasoning annotations, and its evaluation shows open-source MLLMs trail commercial models while multi-image reward models are unstable.

  2. HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical benchmark for multimodal models on human-centric visual understanding finds frontier models average under 60% and miss question-uncued visual evidence, with test-time scaling helping only marginally.

  3. PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PeRL applies reinforcement learning to a vision-language model with image-order permutation and difficulty-based data filtering, improving multi-image reasoning while keeping single-image performance.

  4. Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Context-to-Cue Direct Preference Optimization (CcDPO) reduces multi-image hallucinations in 7B multimodal LLMs by training on perturbed full-sequence captions and region-focused visual prompts, improving average multi...

  5. Medical Large Vision Language Models with Multi-Image Visual Ability

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Fine-tuning medical vision-language models on the Med-MIM multi-image instruction dataset improves their scores on the authors' multi-image benchmarks, but the held-in benchmark is drawn from the same data used for training.

  6. CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Multimodal context-learning benchmark CLBench-V separates grounding, information application, and knowledge acquisition; the best evaluated model scores 0.2847.

  7. MANBench: Is Your Multimodal Model Smarter than Human?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new bilingual 1,314-question benchmark finds the best multimodal model scores about 60%, below the average human score of 62%, beating humans only on knowledge and basic image-text tasks.

  8. Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models

    cs.CL 2025-08 unverdicted novelty 3.0 of 10

    The submission cannot be reviewed as a coherent paper: its abstract and full text are two different papers, so the abstract's claims have no supporting body.

Pith tools