REVIEW 8 cited by
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks. MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of 24 popular LVLMs, including both open-source and proprietary models, reveals significant challenges in multi-image comprehension, particularly in tasks involving spatial understanding. Even the most advanced models, such as GPT-4o, achieve only 55.7% accuracy on MMIU. Through multi-faceted analytical experiments, we identify key performance gaps and limitations, providing valuable insights for future model and data improvements. We aim for MMIU to advance the frontier of LVLM research and development, moving us toward achieving sophisticated multimodal multi-image user interactions.
Forward citations
Cited by 8 Pith papers
-
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
MMRB is the first benchmark combining multi-image inputs with chain-of-thought reasoning annotations, and its evaluation shows open-source MLLMs trail commercial models while multi-image reward models are unstable.
-
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
A hierarchical benchmark for multimodal models on human-centric visual understanding finds frontier models average under 60% and miss question-uncued visual evidence, with test-time scaling helping only marginally.
-
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
PeRL applies reinforcement learning to a vision-language model with image-order permutation and difficulty-based data filtering, improving multi-image reasoning while keeping single-image performance.
-
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
Context-to-Cue Direct Preference Optimization (CcDPO) reduces multi-image hallucinations in 7B multimodal LLMs by training on perturbed full-sequence captions and region-focused visual prompts, improving average multi...
-
Medical Large Vision Language Models with Multi-Image Visual Ability
Fine-tuning medical vision-language models on the Med-MIM multi-image instruction dataset improves their scores on the authors' multi-image benchmarks, but the held-in benchmark is drawn from the same data used for training.
-
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Multimodal context-learning benchmark CLBench-V separates grounding, information application, and knowledge acquisition; the best evaluated model scores 0.2847.
-
MANBench: Is Your Multimodal Model Smarter than Human?
A new bilingual 1,314-question benchmark finds the best multimodal model scores about 60%, below the average human score of 62%, beating humans only on knowledge and basic image-text tasks.
-
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
The submission cannot be reviewed as a coherent paper: its abstract and full text are two different papers, so the abstract's claims have no supporting body.
Discussion (0). Sign in to comment.