REVIEW 20 cited by
A Survey on Benchmarks of Multimodal Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multimodal Large Language Models (MLLMs) are gaining increasing popularity in both academia and industry due to their remarkable performance in various applications such as visual question answering, visual perception, understanding, and reasoning. Over the past few years, significant efforts have been made to examine MLLMs from multiple perspectives. This paper presents a comprehensive review of 200 benchmarks and evaluations for MLLMs, focusing on (1)perception and understanding, (2)cognition and reasoning, (3)specific domains, (4)key capabilities, and (5)other modalities. Finally, we discuss the limitations of the current evaluation methods for MLLMs and explore promising future directions. Our key argument is that evaluation should be regarded as a crucial discipline to support the development of MLLMs better. For more details, please visit our GitHub repository: https://github.com/swordlidev/Evaluation-Multimodal-LLMs-Survey.
Forward citations
Cited by 20 Pith papers
-
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
A new benchmark, TempVS, shows that state-of-the-art multimodal LLMs largely fail at multi-event temporal reasoning across image sequences, despite being able to ground individual events.
-
PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement
Existing MLLM unlearning methods reduce private-attribute leakage on entangled images but substantially harm co-occurring public figures and landmarks, with private knowledge often re-emerging after public finetuning.
-
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding
A structured five-step reasoning template plus diverse-trajectory cold start and diversity-preserving two-stage RL lifts a 7B multimodal model to state-of-the-art multi-image reasoning on several benchmarks.
-
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.
-
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
USB-SafeBench is a unified MLLM safety benchmark with 61 risk categories, 4 modality combinations, and dual-language vulnerability and oversensitivity tests.
-
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.
-
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
AvaMERG is a new text-speech-vision avatar benchmark for empathetic response generation, and the Empatheia system is claimed to outperform baselines on both textual and multimodal empathy tasks.
-
A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models
LoTbench, an interactive causality-aware benchmark built on Oogiri humor tasks, ranks multimodal LLMs and finds their creativity is moderately below human levels yet strongly correlated with general multimodal cogniti...
-
50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security Implications
A multimodal AI detector, DPGuard, combined with a unified 21-category taxonomy, claims state-of-the-art detection of deceptive UI patterns and finds them in 47% of popular websites and 24% of mobile screenshots.
-
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models
VARCO-VISION-14B is a Korean-English vision-language model that reports strong results among similar-size open models and introduces five Korean multimodal benchmarks.
-
Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains
Introduces ZNO-Vision, a 4,306-item Ukrainian multimodal exam benchmark, plus a translated VQA set and a 20-dish cuisine test, and finds only Gemini, Claude, and Qwen2-VL-72B beat the chance baseline.
-
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
SMART, an audio-enhanced MLLM with shot-aware token compression, reports new state-of-the-art moment retrieval accuracy on Charades-STA and QVHighlights.
-
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
A new multilingual, image-based Kangaroo math benchmark shows Gemini 2.0 Flash, Qwen-VL 2.5 72B, and GPT-4o lead current multimodal LLMs, but all remain far below human accuracy on visual math reasoning.
-
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
A new benchmark shows that GPT-4o and other multimodal LLMs perform near chance on ordering image events and far below humans on estimating time lapses.
-
Unified Multimodal Understanding via Byte-Pair Visual Encoding
Priority-guided byte-pair encoding of quantized image patches plus curriculum training yields an 8B discrete-token MLLM competitive with continuous-embedding models on VQA and multimodal benchmarks.
-
MLLMs are Deeply Affected by Modality Bias
A position paper with a case study showing that multimodal LLMs rely on language priors and underuse visual input, together with a research roadmap and calls for balanced training.
-
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
A pipeline that converts an image into a scene graph, embeds the graph chunks, retrieves the most relevant chunks, and prompts an LLM with them reports high VQA accuracy, but the comparison to MLLMs is not credible.
-
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.
-
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
VLAD combines contrastive vision-language alignment with hierarchical diffusion guidance and claims improved text-to-image generation, but the reported FID numbers in Table I do not support 'consistently outperforms a...
-
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
A survey tracing the evolution of visual question answering from 2015 CNN-LSTM models through attention mechanisms, modular networks, vision-language pretraining, and large multimodal models.
Discussion (0). Continue with ORCID to comment.