Pith. sign in

REVIEW 4 cited by

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.15415 v1 pith:ZM2JPHV7 submitted 2025-04-21 cs.CV cs.CL

IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs

classification cs.CV cs.CL
keywords iv-benchreasoningvideoperceptiondataimage-groundedtasksbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Existing evaluation frameworks for Multimodal Large Language Models (MLLMs) primarily focus on image reasoning or general video understanding tasks, largely overlooking the significant role of image context in video comprehension. To bridge this gap, we propose IV-Bench, the first comprehensive benchmark for evaluating Image-Grounded Video Perception and Reasoning. IV-Bench consists of 967 videos paired with 2,585 meticulously annotated image-text queries across 13 tasks (7 perception and 6 reasoning tasks) and 5 representative categories. Extensive evaluations of state-of-the-art open-source (e.g., InternVL2.5, Qwen2.5-VL) and closed-source (e.g., GPT-4o, Gemini2-Flash and Gemini2-Pro) MLLMs demonstrate that current models substantially underperform in image-grounded video Perception and Reasoning, merely achieving at most 28.9% accuracy. Further analysis reveals key factors influencing model performance on IV-Bench, including inference pattern, frame number, and resolution. Additionally, through a simple data synthesis approach, we demonstratethe challenges of IV- Bench extend beyond merely aligning the data format in the training proecss. These findings collectively provide valuable insights for future research. Our codes and data are released in https://github.com/multimodal-art-projection/IV-Bench.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

    cs.CV 2026-03 unverdicted novelty 7.0

    SceneBench shows VLMs lose accuracy on scene-level questions in long videos due to forgetting, and Scene-RAG retrieval improves performance by 2.5%.

  2. VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

    cs.CL 2026-05 unverdicted novelty 5.0

    VISTA is a video synthesis framework that creates controllable egocentric videos of daily tasks with reactive and proactive agent intervention modes via causal reverse reasoning scripts.

  3. VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

    cs.CL 2026-05 reject novelty 5.0

    A staged, inspectable pipeline converts scenario descriptions into egocentric videos for assistant-AI training, but its claimed edge over one-pass baselines is an estimate, not a measured result.

  4. Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

    cs.CV 2026-03 unverdicted novelty 5.0

    SceneBench shows VLMs sharply lose accuracy on scene-level long-video questions, and Scene-RAG only partially mitigates that forgetting (+2.50%).