{"id":"7ad60ef5-8603-48c2-a09b-fd5f41881123","arxiv_id":"2411.15296","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.","lead":"This paper is a survey of how multimodal large language models are evaluated. It sorts more than a hundred benchmarks into capability categories, explains how benchmarks are built and scored, and highlights gaps in current evaluation practice.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey claims comprehensiveness, but it never states a benchmark selection protocol; Fig. 4 and Table 1 may overrepresent the authors' own benchmarks, so the taxonomy and field-level conclusions may not generalize.","rationale":"I agree with the reader's weakest_assumption: the absence of a stated benchmark selection protocol, combined with the authors' leadership of several prominently featured benchmarks, is the most load-bearing weakness of the survey's central claim. The survey does have independent value: it covers many important benchmarks, discusses construction pitfalls like data leakage and MCQ leakage, and describes toolkits in detail. However, the claim of comprehensiveness and systematicity cannot be evaluated without knowing how benchmarks were chosen. My proposed test would directly measure whether the selection is biased toward the authors' own benchmarks. If the test shows no overrepresentation, the concern is resolved and the survey can serve as a reliable map. If it shows significant overrepresentation, the authors should add a transparent selection protocol, disclose the potential conflict of interest in benchmark choice, and moderate the claim from 'comprehensive' to 'representative of a curated subset.' Since the reader already assigned a conditional verdict, my recommendation is unchanged: the survey should be accepted only after these methodological issues are addressed.","tokens_in":43511,"tokens_out":5479,"duration_ms":44458,"concrete_test":"Compile the universe of MLLM evaluation benchmarks from a neutral source, such as Papers with Code and the ACL Anthology, excluding the authors' own project page as the data source. For every benchmark, record whether it appears in Fig. 4 or Table 1 and whether it is authored by the survey's teams. Compare the inclusion rate of author-affiliated benchmarks to non-affiliated benchmarks using a Fisher exact test or a simple rate ratio. If author-affiliated benchmarks are included at a significantly higher rate (e.g., rate ratio > 2 or p < 0.05) or if a major category in Fig. 4 is dominated by author-affiliated benchmarks, then the selection is not representative and the comprehensiveness claim should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it presents a comprehensive and systematic survey of MLLM evaluation. The load-bearing assumption is that the benchmarks selected for Fig. 4 and Table 1 are a representative sample of the field. Section 3 introduces Fig. 4 as 'a quick scan' and Table 1 as 'statistics of some representative benchmarks,' but no inclusion or exclusion criteria are stated anywhere in the paper. The author footnote discloses that the authors lead the MME, MMBench, and LLaVA teams, and several of the most prominently featured benchmarks—MME, Video-MME, MME-RealWorld, MMBench, VLMEvalKit, LMMs-Eval—are their own. If this selection is skewed, the taxonomy in Fig. 4 and the aggregate conclusions in Sections 3.1–3.3 (e.g., 'open-source models have increasingly matched or even surpassed closed-source counterparts' in §3.1.1) could be artifacts of which benchmarks were chosen. Because no reproducible search and selection protocol is provided, the reader cannot distinguish a comprehensive map from a curated list. This is an internal methodological gap, not a disagreement with the field's consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a survey of evaluation methods and benchmarks for multimodal large language models (MLLMs). It organizes the benchmark landscape into a hierarchy of three top-level branches: foundational capabilities (comprehensive evaluation, OCR, chart/document understanding, mathematics, multidisciplinary knowledge, multilingual ability, instruction following, multi-round QA, multi-image understanding, interleaved content, high-resolution input, visual grounding, fine-grained perception, and video understanding), model self-analysis (hallucination, bias, safety, causation), and extended applications (medical imaging, emotion analysis, remote sensing, agents, code generation, GUI understanding, transfer capability, knowledge editing, embodied AI, and autonomous driving). It then discusses benchmark construction (Section 4), evaluation judges and metrics (Sections 5–6), four toolkits (Section 7), and future directions (Section 8). The paper's stated goal is to help researchers select and build benchmarks and to systematize MLLM evaluation, and its four-part structure matches the promises of the abstract.","tokens_in":43848,"tokens_out":15956,"duration_ms":129649,"significance":"The manuscript's main strength is its breadth and organization: roughly 200 benchmarks are arranged in a readable three-branch taxonomy, and the pipeline discussion in Sections 4–6 (construction, judge, metric, toolkit) gives practitioners concrete guidance. The discussion of multiple-choice-question leakage, data leakage, and vision-centric design in Section 4.3 distills recurring methodological lessons that are useful beyond any single benchmark. The companion project page (https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Benchmarks) makes the survey a living resource, which is a genuine practical asset. The survey does not introduce new theory or new measurements, which is appropriate for its genre; its validity depends on whether the displayed benchmark selection is representative of the field, and that is where the manuscript is currently weakest (see Major Comment 1).","major_comments":[{"comment":"The central claim of the paper is that it provides a comprehensive and systematic survey of MLLM evaluation, but the paper never states inclusion or exclusion criteria for the benchmarks shown in Fig. 4 and Table 1. The author footnote discloses that the authors lead the MME, MMBench, and LLaVA teams, and benchmarks or toolkits from those teams are among the most prominently featured (MME [24], MMBench [22], MME-RealWorld [35], Video-MME [87], MMBench-Video [91], VLMEvalKit, LMMs-Eval, and OpenCompass). Because the lists are presented without a protocol, the reader cannot determine whether the taxonomy in Fig. 4 and the aggregate conclusions in Sections 3.1.1 and 3.1.9 reflect the field or a curated subset of the literature; the concern is not the (disclosed) involvement of the authors in several featured projects but the fact that the comprehensiveness claim is unverifiable. A concrete instance is the claim in §3.1.1 that 'open-source models have increasingly matched or even surpassed closed-source counterparts,' which cites only [22], [24], [35], all from the authors' own teams. I note that the survey does include many third-party benchmarks (e.g., MathVista, MMMU, POPE, HallusionBench), so the selection is not exclusive; nevertheless, the paper should state how the benchmark sets were assembled (search scope, time window, selection criteria) and add a limitations statement, or explicitly reframe the selection as curated rather than comprehensive.","section":"§3, Fig. 4, Table 1, author footnote"},{"comment":"Several citation errors undermine the survey's reliability as a literature map. (a) The Fig. 2 caption attributes 'purely discrete modeling to achieve both understanding and generation' to reference [12], which is Lu et al., 'Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models'; that paper describes compositional prompting with LLMs, not discrete multimodal modeling, and the intended citation appears to be missing from the reference list. (b) Table 1 lists RefCOCO+ and RefCOCOg with citation [81] (Kazemzadeh et al., ReferItGame), while the text in §3.1.12 attributes them to [82] (Mao et al.); the RefCOCO family originates from [82], so the table entries should be corrected and made consistent with the text. (c) Several papers are assigned two reference numbers: [48] and [180] (AI2D), [70] and [210] (MIA-Bench), [87] and [178] (Video-MME), [104] and [204] (Bingo), and [109] and [186] (MHaluBench). (d) Reference [118] is used for what appear to be several distinct items: the OOD benchmarks in §3.2.3, the VLLM-safety-benchmark in §3.2.3, and VLAA in §3.3.7; if these are all from the same source, a cross-reference note is needed, otherwise separate citations are required. Because readers of a survey rely on the bibliography to locate the benchmarks discussed, this batch of errors should be corrected systematically before publication.","section":"§2.1 (Fig. 2 caption), Table 1, reference list"}],"minor_comments":[{"comment":"The manuscript contains numerous typos that should be fixed in a proofreading pass: 'extented applications' and 'counstruction' in the abstract, 'instrctions' in the Fig. 3 caption, 'benmarks' in §3.1.1, 'edites' in §4.1, 'lager' in §3.1.1, 'BLUE4' for BLEU4 in Fig. 6, and 'taxonmomy' and 'Recongi tion' in §8.1 and §8.4.","section":"Abstract, Fig. 3, §3.1.1, §4.1, Fig. 6, §8"},{"comment":"The benchmark from reference [29] (WildVision) is referred to inconsistently as 'WV-Bench' (Fig. 4, §3.1.1, §6.2.1) and 'WV-Arena' (§5.1); one name should be adopted throughout.","section":"§3.1.1, §5.1, Fig. 4"},{"comment":"In Table 1, the category labels 'EmbodiedAI' and 'EmbodiedAI(Video)' do not match the survey's own naming in the text ('Embodied AI', §3.3.9), and 'MMHAL-BENCH' should be 'MMHal-Bench' as in Fig. 4; the duplicated citation '[149], [149]' in Fig. 6's metric list should be de-duplicated.","section":"Table 1, Fig. 6"},{"comment":"The text refers to 'WCGB [145]' for webpage-to-code generation, but Fig. 4 and the reference list use 'Web2Code' for [145]; the names should be aligned, or a separate citation provided for WCGB if it is a different dataset.","section":"§3.3.5"},{"comment":"The mention of LiveBench in Section 7.2 has no citation and no reference-list entry; given that the surrounding paragraph is about LMMs-Eval, the connection to multimodal evaluation should also be made explicit.","section":"§7.2"},{"comment":"MMMU-Pro [62] appears in Fig. 4 but is not discussed in Section 3.1.5; since it is a notable recent robustness-oriented extension of MMMU, a one-sentence discussion would better match the survey's coverage claims.","section":"§3.1.5, Fig. 4"},{"comment":"The title 'MME-Survey' foregrounds a single benchmark from the authors' own team rather than the survey's general scope; a scope-reflecting title would match the content and the disclosure footnote more accurately.","section":"Title"}],"recommendation":"major_revision","confidential_remarks":"This is a credible survey with real practical value, and I expect that revision can address the concerns raised. The two load-bearing points are: (1) the absence of an explicit benchmark-selection protocol, which makes the 'comprehensive' claim hard to verify given the disclosed but heavy presence of the authors' own benchmarks and toolkits; and (2) the batch of citation errors (mis-attributed [12], RefCOCO+/g rows in Table 1, five duplicated reference numbers, and overloading of [118]), which for a survey directly affects its usability. The duplicate reference numbering suggests the final version was assembled quickly, so a careful proofreading pass is advisable. I would also flag that the title 'MME-Survey' may raise expectations about a single benchmark; the editor may want to discuss retitling with the authors. With the selection criteria stated and the bibliography corrected, the survey would be suitable for publication in this journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful map of the MLLM evaluation landscape, but it is a curated map, not a systematic one. The lack of a stated selection protocol and a handful of citation errors keep it from being the definitive reference it claims to be.\n\nWhat is new: the organizing taxonomy itself—foundational capabilities, model self-analysis, extended applications—and the four-part pipeline (benchmark construction, judge, metric, toolkit) are a sensible way to structure the field. Section 4 on construction is the most useful part; it lays out data collection methods, annotation strategies, and practical pitfalls like MCQ leakage and data contamination that are easy to overlook. Section 5's breakdown of human, LLM-based, and script-based judges is also clear and accurate. The coverage is broad: video, OCR, math, hallucination, medical, driving, agents—most of the benchmarks a researcher would need are here.\n\nWhere it is soft: the selection of benchmarks in Fig. 4 and Table 1 is never justified. The authors lead MME, Video-MME, MME-RealWorld, MMBench, VLMEvalKit, and LMMs-Eval, and those projects appear prominently. The footnote discloses this, which is good, but disclosure is not the same as a protocol. Without inclusion/exclusion criteria, \"representative benchmarks\" is doing a lot of work, and any field-level conclusion (e.g. open-source matching closed-source) inherits the selection bias. This is an internal methodological gap, not a fatal one—the taxonomy stands on its own—but it prevents the paper from being the systematic survey the abstract promises.\n\nThere are also mechanical errors that a reference work cannot afford. Section 2.1 cites [12] (Chameleon, compositional reasoning) for purely discrete multimodal modeling; the actual Chameleon paper is about plug-and-play reasoning, not discrete tokens. Table 1 lists RefCOCO+ and RefCOCOg with reference [81], but those are from [82]. Small, but they erode trust in a survey whose value is precisely its reliability.\n\nWho this is for: graduate students and researchers new to MLLM evaluation who need a bird's-eye view of the benchmark zoo. It is also a decent checklist for someone building a new benchmark. It is not a source to cite for specific benchmark statistics without checking the original.\n\nRecommendation: send it to review. A serious referee can push for a stated selection protocol and fix the citation errors; the core taxonomy and construction guidance are worth publishing.","headline":"A useful but curated map of MLLM evaluation; needs a stated selection protocol and citation fixes before it can be the definitive reference it claims to be.","tokens_in":44284,"tokens_out":2009,"would_cite":false,"duration_ms":17723,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps the entire landscape of how multimodal large language models are evaluated, from capability benchmarks to construction pipelines to scoring methods.","keywords":["multimodal large language models","MLLM evaluation","benchmark taxonomy","benchmark construction","evaluation metrics","evaluation toolkits","hallucination evaluation","vision-language models"],"falsifier":"Enumerate every MLLM evaluation paper released in a fixed window (for example, 2023–2024) using a neutral literature index, attempt to file each benchmark into the survey's three-branch taxonomy, and record any benchmark family that fits nowhere; if a substantial family such as audio-only or 3D evaluation is missing or misclassified, the survey's comprehensiveness claim is weakened.","tokens_in":43303,"feed_emoji":"🧭","tokens_out":7494,"duration_ms":64525,"temperature":0.7,"pith_summary":"This survey aims to organize the sprawling field of multimodal large language model (MLLM) evaluation into a single map with four coordinated parts. It claims that every current benchmark can be filed under three top-level branches: foundational capabilities (perception, OCR, charts, math, multilingual, instruction following, multi-round and multi-image understanding, video), model self-analysis (hallucination, bias, safety, causation), and extended applications (medicine, emotion, remote sensing, agents, code, GUI, transfer, knowledge editing, embodied AI, driving). It then lays out the pipeline for building a benchmark, the ways to judge outputs (human, LLM/MLLM, or script), the metric families that summarize scores, and the toolkits that run evaluations. If the map is accurate, a researcher can locate the right benchmark for a given question, avoid known construction pitfalls such as data leakage and text-only solvable questions, and see where the field's blind spots are.","feed_headline":"One survey maps every way multimodal LLMs are tested","feed_subtitle":"A three-branch taxonomy plus a build-and-measure pipeline helps researchers pick or create the right benchmark.","key_machinery":"The central object is the taxonomy-pipeline pair. The taxonomy (Fig. 4) sorts benchmarks into three top-level branches and many sub-branches, giving each benchmark a location in capability space so that a researcher can search the field by what they want to test. The pipeline (Fig. 6) organizes the choices a benchmark builder faces: where data comes from (existing datasets, modified data, or internet gathering), how QA pairs are annotated (automatic construction, LLM/MLLM prompting, or manual annotation), which judge evaluates responses (human, model, or script), which metric family summarizes results (deterministic or non-deterministic), and which toolkit executes the evaluation. The taxonomy does the work of making the field searchable, while the pipeline does the work of making new benchmarks constructible and comparable.","core_discovery":"The paper's central claim is that MLLM evaluation can be surveyed systematically along four dimensions: what capabilities are assessed, how benchmarks are built, how performance is measured, and where the next benchmarks should focus. Its main organizing device is a three-branch taxonomy of benchmarks, with sub-branches ranging from comprehensive evaluation and OCR to hallucination, safety, and autonomous driving. Its second device is a construction-and-measurement pipeline that runs from data collection through annotation to judge, metric, and toolkit. The paper introduces no new benchmark; its contribution is the synthesis itself, together with practical guidance for choosing among existing benchmarks and for building new ones that avoid known failure modes.","pith_inferences":["A reader should weigh the benchmark sample: several of the most prominently featured benchmarks were developed by this paper's own author group, and no inclusion or exclusion criteria for the survey are stated, so the field-level takeaways should be read as conditioned on that sample.","The taxonomy invites a testable extension: compute rank correlations of models across benchmark branches; if rankings diverge, 'MLLM capability' is not a single scalar and results should be reported per capability branch.","The reported gaps in audio and 3D evaluation suggest that an omni-modal benchmark suite reusing this taxonomy could reveal whether current models' cross-modal reasoning is general or mostly vision-language.","Since benchmarks are grouped by their declared capability rather than the skills actually required, a follow-up could re-annotate each benchmark by the minimal skill set needed to solve it, then redraw the taxonomy to test the map's validity."],"forward_implications":["A researcher can use the taxonomy as a checklist, locating the capability branch a new model claims to improve and selecting the corresponding benchmarks instead of relying on one aggregate leaderboard.","A benchmark builder can use the pipeline discussion to anticipate failure modes such as multiple-choice leakage, data contamination, and questions answerable without looking at the image.","The survey's gap analysis identifies where new benchmarks are most needed: instruction following, multi-turn dialogue, creativity, task-specific commercial applications, and audio and 3D modalities.","Because judge choice affects open-ended scores, results produced by different LLM judges or human judges are not directly comparable across papers.","The toolkit section implies that standardized evaluation infrastructure is becoming available, which should reduce the cost of reproducing and comparing MLLM results."],"supporting_citations":[{"why":"Anchors the comprehensive-evaluation branch and serves as the paper's recurring example of a manually annotated, script-scored yes/no benchmark.","marker":"[24]"},{"why":"Defines 20 ability dimensions and the CircularEval metric; used to illustrate comprehensive evaluation and choice-extraction by an LLM judge.","marker":"[22]"},{"why":"The largest manually annotated benchmark in the survey; grounds the high-resolution and real-world evaluation sub-branches.","marker":"[35]"},{"why":"Anchors the video-understanding branch and provides the example of a multimodal (frames, subtitles, audio) manually annotated benchmark.","marker":"[87]"},{"why":"Anchors the multidisciplinary and interleaved image-text branches and is used to discuss text-only solvability in multiple-choice questions.","marker":"[59]"},{"why":"Anchors the mathematical-reasoning branch and exemplifies assembling a new benchmark from existing datasets.","marker":"[53]"},{"why":"Anchors the hallucination branch and exemplifies the modify-existing-data route of benchmark construction.","marker":"[106]"},{"why":"One of the four toolkits detailed in the evaluation-toolkit section; load-bearing for the claim that standardized evaluation infrastructure exists.","marker":"[212]"},{"why":"The other general MLLM evaluation toolkit detailed, including the pruned LMMs-Eval lite subset-selection method.","marker":"[213]"}],"fun_headline_variants":["Survey maps the entire MLLM evaluation landscape","Four key dimensions for evaluating multimodal LLMs","A three-branch taxonomy plus a build-and-measure pipeline","Practical guide to choosing and creating MLLM benchmarks","Synthesis of MLLM evaluation: taxonomy, pipeline, and outlook"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's field-level conclusions rest on its benchmark sample being representative, yet it states no inclusion or exclusion criteria, and several of the most prominently featured benchmarks were developed by the authors themselves.","fun_headline_variants_meta":{"raw":{"variants":["Survey maps the entire MLLM evaluation landscape","Four key dimensions for evaluating multimodal LLMs","A three-branch taxonomy plus a build-and-measure pipeline","Practical guide to choosing and creating MLLM benchmarks","Synthesis of MLLM evaluation: taxonomy, pipeline, and outlook"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1381,"prompt_tokens":910,"completion_tokens":471,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":392}},"tokens_in":526,"tokens_out":471,"duration_ms":4676,"temperature":1.0,"reasoning_tokens":392,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:26:17.551440+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Enumerate every MLLM evaluation paper released in a fixed window (for example, 2023–2024) using a neutral literature index, attempt to file each benchmark into the survey's three-branch taxonomy, and record any benchmark family that fits nowhere; if a substantial family such as audio-only or 3D evaluation is missing or misclassified, the survey's comprehensiveness claim is weakened.","supporting_citations":[],"review_version":1}