{"id":"a2ee3d86-7ed3-4f92-8999-88c1d409c926","arxiv_id":"2505.03829","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of VideoLLM benchmarks and evaluation protocols that organizes known datasets and metrics, with proposed future benchmark designs.","lead":"This paper is a literature survey that reviews video benchmarks and evaluation methods for video large language models (VideoLLMs). It compares existing benchmarks and outlines future directions for evaluation, but introduces no new benchmark or model.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table II is the linchpin of the 'generational improvement' claim, but its numbers carry no per-row source or protocol; the trend is only as reliable as that undocumented aggregation.","rationale":"The reader's weakest assumption is exactly right: Table II is a load-bearing aggregation, and its provenance and protocol consistency are not established. I agree with the reader's classification of the paper as a survey rather than a falsifiable research claim. However, because Table II is the only quantitative support for the survey's most concrete assertion, the paper should not be treated as a reliable reference document as-is. The concern is testable and does not require assuming bad faith. If the protocol-controlled reconstruction confirms the trend, the concern is resolved; if it does not, Section IV-A is unsupported. A CONDITIONAL verdict captures this: the survey is acceptable only after the table is tied to primary sources and the consistency of the evaluation protocol is demonstrated.","tokens_in":19399,"tokens_out":6130,"duration_ms":66003,"concrete_test":"Reconstruct Table II row by row from primary sources. For each model, record the exact QA scores and generative-performance scores from its original paper or official repository, along with the protocol actually used: prompt template, GPT judge version, number of sampled videos per benchmark, and evaluation codebase. Then rerun the Section IV-A comparison using only rows whose protocols match exactly, recomputing whether MSVD-QA performance is monotone in model release order under protocol-controlled conditions. If the trend disappears or reverses, the survey's headline claim requires revision; if it survives, the missing citations can be added and the claim stands.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The survey's only empirical content is Table II, and Section IV-A's first and most emphasized trend ('clear generational improvement... newer models consistently outperforming their predecessors') is read directly off that table, e.g., Video-LLaMA 51.6 vs. IG-VLM 76.7 on MSVD-QA. Yet no row in Table II cites its source, and the text never states the evaluation protocol: which prompt template was used, which judge model assigned the generative scores, which subset of each QA benchmark was sampled, what decoding settings, or whether all models were evaluated with the same codebase. The table mixes results that originate from different papers and publication periods (Video-ChatGPT, then later entries such as PLLaVA and IG-VLM) and contains missing cells such as GPT4-V lacking MSVD-QA and MSRVTT-QA. If the later authors reproduced the Video-ChatGPT protocol with different video subsets or a different LLM judge, then cross-row ordering is not a valid basis for the temporal trend. The same provenance problem affects Table I, whose benchmark statistics also appear without per-benchmark citations in the table. This is not an accusation that the numbers are wrong; it is a statement that the central empirical assertion is currently not checkable from the manuscript.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of benchmarks and evaluation methodologies for Video Large Language Models (VideoLLMs). It inventories approximately 20 video QA and reasoning benchmarks in Table I (video counts, clip counts, durations, QA-pair counts), organizes evaluation practice into closed-set, open-set (LLM-judged), and specialized temporal/spatiotemporal categories in Section III, and aggregates zero-shot QA accuracies plus five generative-performance dimensions for roughly 28 models in Table II. Section IV reads performance trends and architectural observations off that aggregation, Sections V and VI propose challenges and future benchmark designs, and Section VII catalogues application domains. The paper explicitly positions itself as building on the prior survey of Tang et al. [12].","tokens_in":19622,"tokens_out":14929,"duration_ms":137117,"significance":"If the aggregation in Table II were properly documented and the factual entries of Table I corrected, this survey would be a useful reference for practitioners: it assembles a broad and recent benchmark inventory with duration statistics, articulates a workable three-way taxonomy of evaluation styles, and names concrete gaps (hallucination, long-form video, cross-modal integration, interactive evaluation). The paper makes no predictions and fits no parameters, so there is no fit-to-data circularity; its value is organizational. The contribution is modest relative to the acknowledged prior survey [12], and the claimed performance trends are currently not verifiable from the manuscript because Table II lacks per-row provenance and evaluation-protocol details, while at least one Table I entry (TGIF-QA) is internally inconsistent. If the provenance and factual issues are resolved, the benchmark inventory and taxonomy would be of genuine use to the VideoLLM community.","major_comments":[{"comment":"The central empirical content of the survey is Table II, and the first trend asserted in Section IV-A ('there is a clear generational improvement in model performance, with newer models consistently outperforming their predecessors') is read directly off that table, yet no row in Table II cites its source and the text nowhere states the evaluation protocol used for the generative scores (which judge LLM, which prompt template, which video-sampling scheme, which decoding settings, whether all rows were produced under a common codebase). The paper itself acknowledges in Section III-B that 'different evaluator models may produce different scores for the same response,' which makes the absence of judge-model reporting in Table II a direct threat to cross-row comparability: rows for the Video-ChatGPT-era models originate from [11], while later rows (PLLaVA, IG-VLM, ST-LLM) necessarily come from their own papers and may have been graded by different judges or prompts. Missing cells (e.g., GPT4-V with no MSVD-QA or MSRVTT-QA, Video-LLaMA 2 with no MSRVTT-QA) are unexplained. Please add a per-row source column, state the exact protocol (or explicitly assert that every row was produced under the [11] protocol), and discuss the missing cells; otherwise the Section IV-A trends are not checkable.","section":"Table II / Section III-B / Section IV-A"},{"comment":"The claim that newer models 'consistently outperform their predecessors' is contradicted by Table II itself. Within the same table, VideoGPT+, a later model, has the lowest temporal score (1.78, below Video-ChatGPT's 2.16); Video-LLaMA 2 scores 2.63 on temporal, exceeding the newer PLLaVA (2.33) and IG-VLM (2.34); and Chat-UniVi's temporal score (2.89) exceeds VideoChat2's (2.66). Monotonic improvement is at best visible in the MSVD-QA column (51.6 to 76.7), and the text should qualify the word 'consistently' accordingly. The same paragraph's closing claim that audio-visual models such as 'AV-LLM and AVicuna often show more balanced performance' is also not supported by the table: AVicuna (2.81/2.62/3.25/2.53/2.59) and AV-LLM (2.56/2.47/2.93/2.17/2.47) are neither the highest nor the most balanced rows, compared with, for example, LLaVA-NeXT-Video at 3.39/3.29/3.92/2.60/3.12.","section":"Section IV-A / Table II"},{"comment":"The TGIF-QA entry is internally inconsistent and factually wrong. Table I lists 8,506 QA pairs for 9,575 clips, and the prose repeats '9,575 short animated GIFs ... and 8,506 question-answer pairs'; the original TGIF-QA dataset contains on the order of 165k QA pairs across roughly 100k GIFs, and a QA count smaller than the clip count is implausible for a dataset with multiple questions per video. Additionally, several rows (MSVD-QA 504/13,157; MSRVTT-QA 2,990/72,821; NExT-QA 1,000/8,564; ActivityNet-QA 800/8,000) appear to be evaluation subsets or protocol-dependent splits rather than original dataset statistics, but Table I has no per-row source column and the text does not distinguish original-dataset statistics from evaluation subsets. Add per-row citations and a 'subset/protocol' flag where applicable.","section":"Table I / Section II-A"},{"comment":"The claims linking architecture class to evaluation results are asserted as if read off Table II, but the models invoked are largely absent from the table, and the table contradicts the stated patterns. VideoChat is classified as a Video Analyzer x LLM model in one paragraph and then as a hybrid (Analyzer + Embedder) model in the next ('Hybrid models that combine analyzer and embedder approaches, such as VideoChat and Vid2Seq'), an explicit internal contradiction. VTimeLLM is cited as exemplifying better temporal performance among Embedder x LLM models, yet its temporal score (2.49) is lower than its correctness (2.78), detail (3.10), and context (3.40) in Table II. ChatVideo, TimeChat, and Vid2Seq never appear in Table II. Either the claims should be derived strictly from the presented data, or each claim should be attributed to the specific analysis in [12] with the relevant section cited.","section":"Section IV-B / Table II"}],"minor_comments":[{"comment":"The itemized list is mis-numbered: it contains '(ii) Evidence grounding' followed by a second '(ii) Reasoning transparency' and a stray '(v).' with a period; renumber the list as (i)-(v).","section":"Section VI-F"},{"comment":"The phrase 'GPT-based metrics for MSVD-QA, MSRVTT-QA, and ActivityNet-QA' conflates the QA accuracy columns with the GPT-based generative scores; the QA columns are accuracy percentages, while GPT-based scoring applies to the five generative dimensions.","section":"Table II caption / Section III-B"},{"comment":"The table is described as 'sorted by temporal understanding performance,' but the ordering is violated by the PLLaVA (2.33) / IG-VLM (2.34) pair, which appears in decreasing order after the higher 2.34 of RED-VILLM; either re-sort the table or drop the sorting claim.","section":"Table II"},{"comment":"Model names are typeset inconsistently: 'Video LLaMA 2' vs 'Video-LLaMA', 'AVicuna' (presumably A-Vicuna), 'RED-VILLM' (presumably Red-VILLM), and spacing artifacts such as 'LLaV A' and 'V ALL Y'; normalize all model names across text, tables, and captions.","section":"Throughout"},{"comment":"Reference [26] (AutoEval-Video) lacks an arXiv identifier, and several references mix venue names with bare 'arXiv preprint' descriptors without consistent identifiers; add the missing arXiv IDs or DOIs throughout.","section":"References"},{"comment":"Average duration is listed as '-' for SOK-Bench and EditVid-QA with no explanation; either supply the values or state explicitly that duration is not reported for these benchmarks.","section":"Table I"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the survey's contribution overlaps heavily with the cited Tang et al. [12] and with several other VideoLLM surveys circulating in the arXiv pipeline; its distinctive value is the benchmark/evaluation-focused aggregation in Tables I and II. The main risk is provenance: Table II appears to assemble numbers from multiple primary papers (Video-ChatGPT's benchmark and later model papers), and the authors must be expected to supply per-row sources and a protocol statement in revision; I would suggest having a referee verify a random sample of Table II rows against the cited primary papers. I would also ask the authors to reconcile Table I with the original dataset papers, since the TGIF-QA entry is demonstrably incorrect. If the authors cannot provide provenance, the paper should be repositioned as a position paper rather than a survey, and the contribution claim should be scaled accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this survey does exactly what it says, organizing current VideoLLM benchmarks and evaluation methods, but the one empirical table behind its headline “generational improvement” claim has no sources and no stated evaluation protocol, so the trend is not checkable as written.\n\nWhat is genuinely useful: Table II aggregates zero-shot QA and generative performance scores for many VideoLLMs in one place, and the prose around closed-set versus open-set evaluation is a reasonable orientation for newcomers. The benchmark evolution narrative—short clips to long-form, general to specialized—is accurate at a high level. The list of future directions is broad, though so broad that much of it reads as a checklist.\n\nThe soft spots are real and load-bearing. Table II sorts models by temporal score and feeds the paper's main trend claim, yet no row cites its source paper, and the text never states which prompt template, judge model, video subset, or decoding settings were used. Different rows may come from different papers with different protocols. This is not an accusation that the numbers are wrong; it is that the central empirical assertion is currently unverifiable from the manuscript. The stress-test note lands. Table I also has an apparent error: TGIF-QA is listed with 8,506 QA pairs, which is far below the original dataset's size, and the text repeats the same number. That looks like a copy error. Some citations are off-topic (the first reference, Frozen in Time, is an odd fit for the video-challenges point), and the paper leans heavily on Tang et al. [12], which is fine but leaves little original content.\n\nThe paper makes no falsifiable claim and introduces no new benchmark, metric, or measurement, so it is not a research contribution. As a survey, it would be useful if the tables were trustworthy, and they are not yet.\n\nMy recommendation: this deserves a peer-review slot rather than a desk reject, but with a clear instruction to add per-cell sources and protocol details to both tables, correct the TGIF-QA numbers, and soften any claim that depends on the unverified aggregation. The organizational skeleton is decent, and a corrected version would be a citable orientation for someone entering the field. As is, I would not cite it.","headline":"A derivative but potentially useful survey whose main table—the one driving its central trend claim—has no per-row sources or protocol, making that claim uncheckable as written.","tokens_in":20100,"tokens_out":2410,"would_cite":false,"duration_ms":25277,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps how video-language models are benchmarked and evaluated, and argues that open-set, LLM-judged scoring now dominates the field.","keywords":["VideoLLM","video understanding benchmarks","video question answering","evaluation methodology","open-set evaluation","temporal reasoning","multimodal integration"],"falsifier":"Take any row in Table II and compare it with the cited source paper; for instance, check whether Video-LLaMA's MSVD-QA score is 51.6 in [9] under zero-shot open-ended question answering. If the original paper reports a different number, split, or scoring method, the across-model ordering on which the survey's main trend rests cannot be reconstructed from the evidence it cites.","tokens_in":19199,"feed_emoji":"🎬","tokens_out":6138,"duration_ms":58593,"temperature":0.7,"pith_summary":"The paper surveys how Video Large Language Models (VideoLLMs) are benchmarked and evaluated. It sets out to organise the field's benchmarks by size, duration, and task focus, and to classify evaluation into closed-set, open-set, and specialised temporal/spatiotemporal methods. The survey's main empirical observation, drawn from its own comparison table, is that newer VideoLLMs show a generational improvement in zero-shot question answering while still trailing a general-purpose multimodal model such as GPT-4V in generative dimensions like temporal understanding and context. A sympathetic reader comes away with a structured map of the terrain and a list of proposed next-generation benchmarks.","feed_headline":"Video language models improve, but temporal understanding still lags","feed_subtitle":"A survey of 20+ benchmarks maps how video models are scored, and where newer models still fall short.","key_machinery":"The organising device is a pair of tables. Table I inventories more than twenty benchmarks by video count, clip count, average duration, and question-answer pairs; Table II compiles model scores on five generative dimensions (correctness, detail, context, temporal, consistency) plus zero-shot accuracy on MSVD-QA, MSRVTT-QA, and ActivityNet-QA, with rows sorted by temporal score. The analytical mechanism that produces the trend claims is the open-set evaluation protocol introduced by Video-ChatGPT, in which a GPT-based judge scores free-form model responses on those five dimensions. The survey also adopts an architectural taxonomy involving Video Analyzer × LLM, Video Embedder × LLM, and hybrid designs, using it to explain why some models excel at factual accuracy and others at temporal reasoning.","core_discovery":"The central claim is that VideoLLM evaluation has shifted from closed-set multiple-choice QA toward open-set protocols in which a large language model grades free-form answers, and that benchmark design is evolving in step: from short clips and factual questions to long videos, temporal reasoning, and multimodal integration. Examining the compiled scores, the paper asserts a clear generational improvement in model performance, citing the move from Video-LLaMA at 51.6 on MSVD-QA to IG-VLM at 76.7, and identifies a persistent gap between specialised VideoLLMs and GPT-4V on generative performance dimensions. It also reports that performance drops as video length grows, and that audio-visual models show more balanced scores. On the basis of these observations, the paper proposes six benchmark designs for future evaluation: hierarchical understanding, multimodal integration, long-form narrative, interactive evaluation, robustness and adversarial testing, and explainability.","pith_inferences":["An implication the paper leaves implicit: the five-dimension scoring protocol likely carries judge-induced variance, so re-scoring one fixed set of model outputs with different judge models would quantify how much of Table II's ordering is an artifact of the evaluator.","A controlled extension the paper does not run: vary only the duration of the same narrative content across conditions; if accuracy declines smoothly with length, context retention is causal, whereas a step change would implicate specific architectural bottlenecks.","The survey aggregates rows from different source papers with unknown protocols; a natural editorial follow-up is to rebuild the comparison on a single execution harness before treating the generational trend as established.","The proposed explainability benchmark presupposes a shared standard for what counts as a good explanation; defining and validating that standard may be a prerequisite for the benchmark to be usable."],"forward_implications":["If the reported trends hold, comparisons between VideoLLMs should concentrate on long-video, temporal, and multimodal tasks, since basic factual QA no longer separates the top models.","Open-set LLM-judged scoring becomes the de facto currency for reporting VideoLLM capability, making the choice of judge model a source of cross-paper variation.","The proposed hierarchical and long-form narrative benchmarks would allow each capability to be measured separately, which could change how model strengths and weaknesses are ranked.","The persistence of a gap with general-purpose multimodal models suggests that video-specific architectures should direct their next steps at temporal and contextual integration rather than factual recognition.","Because longer videos consistently produce lower scores, context retention across extended durations is the binding constraint for real-world VideoLLM deployment."],"supporting_citations":[{"why":"Supplies the broader survey and architectural taxonomy that this paper builds on.","marker":"[12]"},{"why":"Introduces the open-end zero-shot QA and five-dimension generative evaluation protocol used in Table II.","marker":"[11]"},{"why":"Provides the MSVD-QA and MSRVTT-QA datasets whose scores populate Table II and statistics in Table I.","marker":"[13]"},{"why":"Contributes the TGIF-QA benchmark focused on repetition, state transitions, and frame-based QA.","marker":"[14]"},{"why":"Contributes the ActivityNet-QA long-video benchmark used in both tables.","marker":"[15]"},{"why":"Represents a recent comprehensive multi-modal video understanding benchmark.","marker":"[17]"},{"why":"Provides CinePile, a large long-video benchmark with 303,828 QA pairs supporting the long-form trend.","marker":"[20]"},{"why":"Provides InfiniBench, the very-long-video benchmark used to argue for long-form evaluation needs.","marker":"[21]"},{"why":"Supplies TempCompass, the specialised temporal-reasoning benchmark motivating dedicated evaluation.","marker":"[22]"},{"why":"Supplies Video-MME, supporting the claim that multimodal integration is an emerging benchmark direction.","marker":"[27]"}],"fun_headline_variants":["VideoLLM evaluation shifts from multiple choice to AI-graded answers","Survey: VideoLLMs gain, but long-video understanding still lags","New VideoLLM benchmarks push temporal and audio-visual reasoning","Six benchmark designs proposed to close VideoLLM evaluation gaps","Open-set grading and longer clips reshape VideoLLM testing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The trends in Section IV-A rest entirely on the scores compiled in Table II, but the paper does not disclose where those numbers come from, which prompts or judge models produced them, or whether every row was measured under the same protocol.","fun_headline_variants_meta":{"raw":{"variants":["VideoLLM evaluation shifts from multiple choice to AI-graded answers","Survey: VideoLLMs gain, but long-video understanding still lags","New VideoLLM benchmarks push temporal and audio-visual reasoning","Six benchmark designs proposed to close VideoLLM evaluation gaps","Open-set grading and longer clips reshape VideoLLM testing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001073,"raw_usage":{"total_tokens":4458,"prompt_tokens":875,"completion_tokens":3583,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":3495}},"tokens_in":491,"tokens_out":3583,"duration_ms":26769,"temperature":1.0,"reasoning_tokens":3495,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:05:32.037377+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any row in Table II and compare it with the cited source paper; for instance, check whether Video-LLaMA's MSVD-QA score is 51.6 in [9] under zero-shot open-ended question answering. If the original paper reports a different number, split, or scoring method, the across-model ordering on which the survey's main trend rests cannot be reconstructed from the evidence it cites.","supporting_citations":[{"cited_title":"Video question answering via gradually reﬁned attention o ver appearance and motion,","cited_arxiv_id":null,"evidence_quote":"Provides the MSVD-QA and MSRVTT-QA datasets whose scores populate Table II and statistics in Table I."},{"cited_title":"Tgif-qa: Toward spatio-temporal reasoning in visual question answering,","cited_arxiv_id":null,"evidence_quote":"Contributes the TGIF-QA benchmark focused on repetition, state transitions, and frame-based QA."},{"cited_title":"Activitynet-qa: A dataset for understanding complex web videos via question answering,","cited_arxiv_id":null,"evidence_quote":"Contributes the ActivityNet-QA long-video benchmark used in both tables."},{"cited_title":"Mvbench: A comprehensive multi-modal video understanding benchmark,","cited_arxiv_id":null,"evidence_quote":"Represents a recent comprehensive multi-modal video understanding benchmark."}],"review_version":1}