Pith. sign in

REVIEW 20 cited by

A Survey on Benchmarks of Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08632 v2 pith:HLAVBY5A submitted 2024-08-16 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords mllmsbenchmarksevaluationgithublanguagelargemodelsmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) are gaining increasing popularity in both academia and industry due to their remarkable performance in various applications such as visual question answering, visual perception, understanding, and reasoning. Over the past few years, significant efforts have been made to examine MLLMs from multiple perspectives. This paper presents a comprehensive review of 200 benchmarks and evaluations for MLLMs, focusing on (1)perception and understanding, (2)cognition and reasoning, (3)specific domains, (4)key capabilities, and (5)other modalities. Finally, we discuss the limitations of the current evaluation methods for MLLMs and explore promising future directions. Our key argument is that evaluation should be regarded as a crucial discipline to support the development of MLLMs better. For more details, please visit our GitHub repository: https://github.com/swordlidev/Evaluation-Multimodal-LLMs-Survey.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A new benchmark, TempVS, shows that state-of-the-art multimodal LLMs largely fail at multi-event temporal reasoning across image sequences, despite being able to ground individual events.

  2. PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Existing MLLM unlearning methods reduce private-attribute leakage on entangled images but substantially harm co-occurring public figures and landmarks, with private knowledge often re-emerging after public finetuning.

  3. Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A structured five-step reasoning template plus diverse-trajectory cold start and diversity-preserving two-stage RL lifts a 7B multimodal model to state-of-the-art multi-image reasoning on several benchmarks.

  4. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  5. USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

    cs.CR 2025-05 conditional novelty 6.0 of 10

    USB-SafeBench is a unified MLLM safety benchmark with 61 risk categories, 4 modality combinations, and dual-language vulnerability and oversensitivity tests.

  6. ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ZeroBench is a hand-built 100-question visual reasoning benchmark, adversarially filtered so every evaluated frontier LMM scored 0% at release.

  7. Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

    cs.MM 2025-02 conditional novelty 6.0 of 10

    AvaMERG is a new text-speech-vision avatar benchmark for empathetic response generation, and the Empatheia system is claimed to outperform baselines on both textual and multimodal empathy tasks.

  8. A Causality-aware Paradigm for Evaluating Creativity of Multimodal Large Language Models

    cs.AI 2025-01 conditional novelty 6.0 of 10

    LoTbench, an interactive causality-aware benchmark built on Oogiri humor tasks, ranks multimodal LLMs and finds their creativity is moderately below human levels yet strongly correlated with general multimodal cogniti...

  9. 50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security Implications

    cs.CR 2025-01 conditional novelty 6.0 of 10

    A multimodal AI detector, DPGuard, combined with a unified 21-category taxonomy, claims state-of-the-art detection of deceptive UI patterns and finds them in 47% of popular websites and 24% of mobile screenshots.

  10. VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    VARCO-VISION-14B is a Korean-English vision-language model that reports strong results among similar-size open models and introduces five Korean multimodal benchmarks.

  11. Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains

    cs.CL 2024-11 conditional novelty 6.0 of 10

    Introduces ZNO-Vision, a 4,306-item Ukrainian multimodal exam benchmark, plus a translated VQA set and a 20-dish cuisine test, and finds only Gemini, Claude, and Qwen2-VL-72B beat the chance baseline.

  12. SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM

    cs.CV 2025-11 conditional novelty 5.0 of 10

    SMART, an audio-enhanced MLLM with shot-aware token compression, reports new state-of-the-art moment retrieval accuracy on Charades-STA and QVHighlights.

  13. Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A new multilingual, image-based Kangaroo math benchmark shows Gemini 2.0 Flash, Qwen-VL 2.5 72B, and GPT-4o lead current multimodal LLMs, but all remain far below human accuracy on visual math reasoning.

  14. Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A new benchmark shows that GPT-4o and other multimodal LLMs perform near chance on ordering image events and far below humans on estimating time lapses.

  15. Unified Multimodal Understanding via Byte-Pair Visual Encoding

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Priority-guided byte-pair encoding of quantized image patches plus curriculum training yields an 8B discrete-token MLLM competitive with continuous-embedding models on VQA and multimodal benchmarks.

  16. MLLMs are Deeply Affected by Modality Bias

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A position paper with a case study showing that multimodal LLMs rely on language priors and underuse visual input, together with a research roadmap and calls for balanced training.

  17. Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering

    cs.CV 2024-12 reject novelty 4.0 of 10

    A pipeline that converts an image into a scene graph, embeds the graph chunks, retrieves the most relevant chunks, and prompts an LLM with them reports high VQA accuracy, but the comparison to MLLMs is not credible.

  18. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

  19. Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models

    cs.CV 2025-01 reject novelty 3.0 of 10

    VLAD combines contrastive vision-language alignment with hierarchical diffusion guidance and claims improved text-to-image generation, but the reported FID numbers in Table I do not support 'consistently outperforms a...

  20. The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering

    cs.CV 2025-01 reject novelty 2.0 of 10

    A survey tracing the evolution of visual question answering from 2015 CNN-LSTM models through attention mechanisms, modular networks, vision-language pretraining, and large multimodal models.

Pith tools