Pith. sign in

REVIEW 5 cited by

Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12742 v1 pith:PSQGL5I2 submitted 2024-06-18 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords modelsreasoninglanguagemulti-imagebenchmarkvisualvlmsgpt-4v
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The advancement of large language models (LLMs) has significantly broadened the scope of applications in natural language processing, with multi-modal LLMs extending these capabilities to integrate and interpret visual data. However, existing benchmarks for visual language models (VLMs) predominantly focus on single-image inputs, neglecting the crucial aspect of multi-image understanding. In this paper, we introduce a Multi-Image Relational Benchmark MIRB, designed to evaluate VLMs' ability to compare, analyze, and reason across multiple images. Our benchmark encompasses four categories: perception, visual world knowledge, reasoning, and multi-hop reasoning. Through a comprehensive evaluation of a wide range of open-source and closed-source models, we demonstrate that while open-source VLMs were shown to approach the performance of GPT-4V in single-image tasks, a significant performance gap remains in multi-image reasoning tasks. Our findings also reveal that even the state-of-the-art GPT-4V model struggles with our benchmark, underscoring the need for further research and development in this area. We believe our contribution of MIRB could serve as a testbed for developing the next-generation multi-modal models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A structured five-step reasoning template plus diverse-trajectory cold start and diversity-preserving two-stage RL lifts a 7B multimodal model to state-of-the-art multi-image reasoning on several benchmarks.

  2. FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark of real-world financial charts shows current vision-language models lag badly on questions that require reading values from chart axes.

  3. MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A self-supervised contrastive triplets plus weak-to-strong augmented GRPO training method improves multi-image reasoning in Qwen2.5-VL-7B.

  4. Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Context-to-Cue Direct Preference Optimization (CcDPO) reduces multi-image hallucinations in 7B multimodal LLMs by training on perturbed full-sequence captions and region-focused visual prompts, improving average multi...

  5. VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism

    cs.CV 2025-06 conditional novelty 5.0 of 10

    VReST combines Monte Carlo tree search with a self-reward signal inside a vision-language model to get higher accuracy than CoT, ToT, or voting baselines on MathVista, MathVision, and CharXiv, while spending several t...

Pith tools