Pith. sign in

REVIEW 4 cited by

A Survey on Multimodal Benchmarks: In the Era of Large AI Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18142 v1 pith:GU6DHDWQ submitted 2024-09-21 cs.AI cs.MM

classification cs.AIcs.MM
keywords benchmarksmodelsmultimodalsurveyacrossanalysislargemllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid evolution of Multimodal Large Language Models (MLLMs) has brought substantial advancements in artificial intelligence, significantly enhancing the capability to understand and generate multimodal content. While prior studies have largely concentrated on model architectures and training methodologies, a thorough analysis of the benchmarks used for evaluating these models remains underexplored. This survey addresses this gap by systematically reviewing 211 benchmarks that assess MLLMs across four core domains: understanding, reasoning, generation, and application. We provide a detailed analysis of task designs, evaluation metrics, and dataset constructions, across diverse modalities. We hope that this survey will contribute to the ongoing advancement of MLLM research by offering a comprehensive overview of benchmarking practices and identifying promising directions for future work. An associated GitHub repository collecting the latest papers is available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code

    cs.CL 2025-06 conditional novelty 6.0 of 10

    WebUIBench is a 21,793-question benchmark that splits WebUI-to-Code into perception, HTML programming, and cross-modal understanding, and it ranks 29 multimodal LLMs on each sub-skill.

  2. Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A new 1,430-item multimodal benchmark shows that leading multimodal LLMs rarely notice small visual traps needed for commonsense safety reasoning, with top scores near 0.46 on a scale whose maximum is about 0.97.

  3. Unified Multimodal Understanding via Byte-Pair Visual Encoding

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Priority-guided byte-pair encoding of quantized image patches plus curriculum training yields an 8B discrete-token MLLM competitive with continuous-embedding models on VQA and multimodal benchmarks.

  4. MLLMs are Deeply Affected by Modality Bias

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A position paper with a case study showing that multimodal LLMs rely on language priors and underuse visual input, together with a research roadmap and calls for balanced training.

Pith tools