Pith. sign in

REVIEW 4 cited by

A Survey on Evaluation of Multimodal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.15769 v1 pith:5GUTQ2N6 submitted 2024-08-28 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords evaluationmllmsmllmcapabilitiesevaluategenerallanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) mimic human perception and reasoning system by integrating powerful Large Language Models (LLMs) with various modality encoders (e.g., vision, audio), positioning LLMs as the "brain" and various modality encoders as sensory organs. This framework endows MLLMs with human-like capabilities, and suggests a potential pathway towards achieving artificial general intelligence (AGI). With the emergence of all-round MLLMs like GPT-4V and Gemini, a multitude of evaluation methods have been developed to assess their capabilities across different dimensions. This paper presents a systematic and comprehensive review of MLLM evaluation methods, covering the following key aspects: (1) the background of MLLMs and their evaluation; (2) "what to evaluate" that reviews and categorizes existing MLLM evaluation tasks based on the capabilities assessed, including general multimodal recognition, perception, reasoning and trustworthiness, and domain-specific applications such as socioeconomic, natural sciences and engineering, medical usage, AI agent, remote sensing, video and audio processing, 3D point cloud analysis, and others; (3) "where to evaluate" that summarizes MLLM evaluation benchmarks into general and specific benchmarks; (4) "how to evaluate" that reviews and illustrates MLLM evaluation steps and metrics; Our overarching goal is to provide valuable insights for researchers in the field of MLLM evaluation, thereby facilitating the development of more capable and reliable MLLMs. We emphasize that evaluation should be regarded as a critical discipline, essential for advancing the field of MLLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

    cs.AI 2026-08 conditional novelty 6.0 of 10

    KnowHal is a new benchmark that jointly tests entity, attribute, relation, and knowledge hallucinations in multimodal language models using paired true/false questions on shared images.

  2. Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MMKC-Bench provides a human-verified benchmark of multimodal knowledge conflicts and shows that current LMMs prefer internal parametric knowledge over external evidence.

  3. Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Uncertainty-o estimates uncertainty in large multimodal models by perturbing prompts and computing entropy over semantically clustered answers, improving hallucination detection across five modalities.

  4. Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

    cs.CV 2025-09 reject novelty 4.0 of 10

    Using SigLIP retrieval to feed a Qwen2-VL or InternVL2 model with similar and dissimilar coordinates yields reported street-level accuracies of 23.2%, 17.1%, and 24.3% on IM2GPS, IM2GPS3k, and YFCC4k.

Pith tools