{"id":"d84a9f5c-b63b-4ec6-8e26-71303fb13b36","arxiv_id":"2505.23693","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new benchmark, VF-Eval, measures how well multimodal LLMs check, detect, and reason about errors in AI-generated videos, and shows frontier models remain far below human performance.","lead":"VF-Eval is a new benchmark that tests how well multimodal AI models can spot and explain mistakes in AI-generated videos, using 9,740 questions across four tasks. It finds that even the best model, GPT-4.1, scores about 52 percent versus 84 percent for humans, showing that AI video critics are still unreliable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Open-ended CV/RE scores are assigned by GPT-4.1-mini from text alone, with no video access and no reported judge calibration; the judge is also one of the evaluated models, so Table 3 rankings could reflect response-style preference rather than AIGC feedback ability.","rationale":"VF-Eval's central contribution is a measurement: frontier MLLMs provide unreliable feedback on AIGC videos, with GPT-4.1 at 51.6 vs 84.4 human. That measurement depends on the scores assigned to open-ended answers. The reader correctly identifies the GPT-4.1-mini judge as the weakest assumption. My reading of Appendix B.2 strengthens this: the judge is given no video, only question, response, and gold answer, so it cannot validate visual facts; it scores textual overlap. Because the judge belongs to the same model family as the top-scoring model and no calibration is reported, the CV/RE columns are not yet a validated measure. This is not an internal inconsistency in the benchmark, and the direction of the gap may survive; hence conditional acceptance remains appropriate. A secondary issue, that RePrompt uses human-revised prompts rather than MLLM-revised prompts, is real but does not threaten the benchmark's core claim; I therefore focus on the judge.","tokens_in":17678,"tokens_out":5379,"duration_ms":52806,"concrete_test":"Sample 300 open-ended items (150 CV, 150 RE) stratified by model and score bin. Have three human experts score each response against the gold answer using the same 0–2 rubric, and independently re-score with a second LLM judge (e.g., GPT-4o) using the same prompt but without video. Compute Cohen's kappa between GPT-4.1-mini and humans, and recompute per-model CV/RE means using human scores. If the judge–human agreement is below (say) 0.7, or if the ordering of GPT-4.1 relative to InternVL3-38B and Qwen2.5-VL-72B changes, then Eqs. 1/4 do not support the reported ranking and the benchmark should report human-judged scores or a calibrated judge.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. 1 and Eq. 4 compute CV and RE scores with GPT-4.1-mini, an LLM judge that is itself a row in Table 3. The scoring prompt (Fig. 12) gives the judge only the question, the model response, and the gold answer; it does not provide the video. Thus the judge cannot verify whether a response is visually correct; it can only reward textual similarity to the gold answer. This introduces two coupled threats to the central claim. First, a model whose answer is correct but differently worded, or that identifies a valid error not mentioned in the gold answer, will be scored 0 or 1, so the reported CV/RE accuracy is not a clean measure of feedback quality. Second, any stylistic bias of GPT-4.1-mini—toward its own output style or toward verbosity—will differentially affect models; because GPT-4.1 and GPT-4.1-mini share a family, the top ranking of GPT-4.1 in the overall score may be partly an artifact. No inter-judge agreement, human calibration of judge scores, or per-model judge score analysis is reported. Given that CV and RE account for 1,982 of the 9,740 items and are the only tasks using this judge, the headline human–model gap (84.4 vs 51.6) is not fully established until this evaluation channel is validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VF-Eval, a benchmark for evaluating multimodal large language models (MLLMs) on their ability to generate feedback on AI-generated videos (AIGC). The benchmark comprises four tasks: coherence validation (CV), error awareness (EA), error type detection (ED), and reasoning evaluation (RE), built over 9,740 question-answer pairs on videos from several text-to-video models. The authors evaluate 13 frontier MLLMs and report that even the best model, GPT-4.1, reaches only 51.6% overall accuracy against an 84.4% human baseline. They also present a RePrompt experiment intended to show that aligning MLLMs with human preferences can benefit video generation. The central claims are that VF-Eval is a valid benchmark for a previously underexplored capability and that current MLLMs are unreliable as feedback providers for AIGC videos.","tokens_in":17995,"tokens_out":7997,"duration_ms":72146,"significance":"If the benchmark is valid, it fills a genuine gap: existing video-understanding benchmarks focus on natural videos, while AIGC videos have distinct artifacts and are increasingly important. The dataset construction includes human annotation, expert validation, and public data/code releases, which are concrete strengths. The four-task design with fine-grained reasoning subcategories is a useful contribution to the community. However, the validity of the headline performance numbers depends on the evaluation protocol, particularly the use of an unvalidated LLM judge for open-ended responses and the all-positive design of the error-awareness task. These issues must be resolved before the benchmark can be considered a reliable measurement instrument.","major_comments":[{"comment":"The CV and RE scores are computed by GPT-4.1-mini, a judge that is prompted with only the question, the model response, and the gold answer, with no video access. This means the judge cannot verify visual correctness; it can only reward textual similarity to the gold answer. A model that gives a correct but differently-worded answer, or that identifies a valid error not listed in the gold answer, will be scored 0 or 1, making the reported CV/RE accuracy an indirect measure of textual overlap rather than a direct measure of feedback quality. Furthermore, GPT-4.1-mini is itself one of the rows in Table 3, so the rankings could partially reflect response-style preference within the GPT-4.1 family rather than AIGC feedback ability. The paper reports no inter-judge agreement, no human calibration of judge scores, and no per-model score distribution analysis. Because CV and RE account for 1,982 of 9,740 items and heavily influence the overall score, the headline gap between humans (84.4%) and GPT-4.1 (51.6%) is not fully established. Please either validate the judge (e.g., human agreement on a subset, comparison with another judge, or analysis of judge score distributions) or reformulate the open-ended evaluation to be less dependent on a single uncalibrated LLM.","section":"§3.1, Eq. (1) and Eq. (4); Appendix B.2 (Fig. 12)"},{"comment":"The Error Awareness task is all-positive: the paper states that all Yes-Or-No questions are designed with 'Yes' as the correct answer. Under this design, accuracy equals the model's 'Yes' rate, and a model that always answers 'Yes' would achieve 100%. The comparison to a 50% 'random guess' baseline in Table 3 and the interpretation in §4.2 that performance 'worse than random guessing' indicates a bias toward perceiving videos as normal conflates label imbalance with perceptual bias. A model with a generic tendency to answer 'No' would score below 50% even if it had perfect video understanding of the presented content. The intentional setup may be defensible as a probe of bias, but the current reporting and interpretation are misleading. Please reframe the result by reporting the model's 'No' rate, comparing against an all-Yes baseline, or including a subset of negative examples to separate label bias from genuine perceptual error.","section":"§3.3; §4.2; Table 3, EA columns"},{"comment":"The evaluated models receive different numbers of input frames (ranging from 2 to 16), with the choice made based on the model's context window. For example, GPT-4.1 receives 16 frames while llama3-llava-next-8b receives 2 frames. This introduces a confound: a model's performance may partly reflect how much of the video it sees rather than its intrinsic capability. The paper does not analyze this effect, yet it uses the resulting scores to rank models and to claim that a scaling law applies in §4.2. Please report results with a matched frame count for a subset of models, or at least provide an analysis of how performance varies with frame count and discuss the limitation explicitly.","section":"§4.1; Table 6"},{"comment":"The RePrompt experiment compares videos generated from the original prompts with videos generated from prompts revised by human annotators. This is a human-in-the-loop study, not an evaluation of MLLM feedback. The conclusion that 'aligning MLLMs more closely with human preferences can benefit video generation' is not directly supported because the revised prompts were produced by humans, not by MLLMs. To support the claim, the experiment should include a condition where MLLM-generated feedback is used to revise prompts, or the claims should be scaled back to what the experiment actually demonstrates (that human-revised prompts can improve some aspects of generated videos).","section":"§5.3; Table 4"}],"minor_comments":[{"comment":"The paper states that '2,395 question-answer pairs are corrected' and then refers to 'the low percentage of revisions' in the same paragraph. The corrected fraction is approximately 24.6% of the 9,740 total pairs, which is not obviously low. Please clarify how this number should be interpreted or revise the wording.","section":"§3.3"},{"comment":"The description of Error Type Detection says it 'intends to identify all the errors present in the AIGC video' and gives example questions such as 'Select the choices that reflect...', which suggests multi-select answers. However, Eq. (3) defines the score as an exact match I(y_i = ŷ_i) and Table 3 reports a 25% random-guess baseline, which is consistent with single-answer multiple-choice. Please clarify the answer format (single-answer vs. multi-select) and specify the scoring accordingly.","section":"§3.1 and §3.3; Eq. (3)"},{"comment":"The frame sampling procedure is not described: are frames uniformly sampled, randomly sampled, or keyframe-based? This information is needed for reproducibility, especially because the number of frames varies across models.","section":"§4.1; Table 6"},{"comment":"The win rates in Table 4 are all close to 50% (50.7% to 57.6%), but no confidence intervals or significance tests are reported. Please report statistical significance or at least the number of pairwise comparisons used for each aspect.","section":"§5.3; Table 4"},{"comment":"The prompt for Reasoning Evaluation is labeled '[CV_SHOT_PROMPT]', which is the same label used for Coherence Validation. Please rename it to '[RE_SHOT_PROMPT]' to avoid confusion. ","section":"Appendix B.2 (Fig. 9 and Fig. 12)"},{"comment":"The benchmark name is spelled inconsistently: 'VF-EVAL' in the abstract and some equations, 'VF-Eval' in the title and elsewhere. Please standardize the spelling. Also, in the introduction the sentence 'that significantly from those found in traditional video content' appears to be missing a verb (e.g., 'differ').","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The benchmark is timely and the dataset release is a positive asset. The major concerns are about the evaluation protocol rather than the dataset construction; in particular, the unvalidated GPT-4.1-mini judge for open-ended tasks and the all-positive Error Awareness set. Both are fixable, but the current version's headline numbers are not yet fully reliable. I would also gently suggest that the RePrompt section overclaims what the experiment shows, which the authors should address editorially as well as technically."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful benchmark resource - a new dataset of 9,740 QA pairs on AIGC videos, four tasks, code and data released, 13 MLLMs evaluated. The core finding, that frontier models are far below human on AIGC video feedback, is probably true in direction. But two design decisions muddy the numbers, and one experiment doesn't test what the abstract says it tests.\n\nWhat's new: VF-Eval is the first benchmark I've seen that combines coherence validation, error awareness, error type detection, and reasoning evaluation specifically on AI-generated videos. The dataset is fairly large, covers multiple T2V models (Pika, Kling, OpenSora, etc.), and has human validation with reported inter-annotator agreement. The model evaluation itself is broad and the error analysis is thoughtful. That is real work.\n\nThe soft spots: (1) All yes/no questions in Error Awareness have 'Yes' as the correct answer. So the 'worse than random' result on EA (GPT-4.1 at 39.7% quality, 24% CP) is not evidence that models 'perceive the video as normal' - it's evidence that they say 'No' more than half the time. Without negative examples, the task measures response bias, not error detection. (2) Open-ended CV and RE scores are assigned by GPT-4.1-mini from text alone, with no video access and no reported judge calibration. The judge is also one of the evaluated models, so rankings on those 1,982 items could reflect response-style preference. This affects about 20% of the benchmark and is part of the headline 51.6 vs 84.4 gap. (3) The RePrompt experiment has humans revise prompts; it does not compare MLLM feedback to human feedback. The conclusion that 'aligning MLLMs with human preferences can benefit video generation' is an extrapolation, not a result.\n\nMinor: Table 2 says 1,836 yes/no questions while the text says 1,826. Human baseline methodology is under-specified. These are fixable.\n\nFor a referee: the dataset and protocol are worth engaging with, and a good referee would ask for the EA negative-example fix, judge validation, and a repositioned RePrompt section. I'd send it to review, with major revisions.","headline":"Useful AIGC-video feedback benchmark, but the headline numbers are partly artifacts of an all-positive yes/no split and a text-only judge.","tokens_in":18503,"tokens_out":4886,"would_cite":true,"duration_ms":46044,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VF-Eval benchmarks four feedback skills on AI-generated video and finds even the best model, GPT-4.1, scores 51.6% overall versus an 84.4% human baseline.","keywords":["AIGC video","multimodal large language model","video understanding benchmark","feedback generation","coherence validation","error detection","reasoning evaluation","video generation"],"falsifier":"Re-score a random sample of the open-ended responses (coherence validation and reasoning evaluation) with independent human raters instead of GPT-4.1-mini, and compare the resulting model-human gap to the reported 51.6% versus 84.4%; if the gap narrows substantially, the headline result is an artifact of the automated judge rather than a true capability gap.","tokens_in":17506,"feed_emoji":"🎬","tokens_out":6788,"duration_ms":52010,"temperature":0.7,"pith_summary":"VF-Eval is a benchmark that tests whether multimodal large language models can give useful feedback on AI-generated (AIGC) videos, through four tasks: checking whether a video matches its generation prompt, saying whether an error is present, naming the error type, and answering reasoning questions about the video. The paper evaluates 13 models and reports that the best one, GPT-4.1, reaches 51.6% overall accuracy against an 84.4% human baseline. The authors argue that current models are therefore unreliable as automated critics for AI video generation, and that a further experiment, RePrompt, shows aligning model feedback with human preferences can improve regenerated videos. A sympathetic reader would take the central claim to be that the ability to critique AIGC videos is a distinct, underdeveloped capability worth benchmarking.","feed_headline":"Best AI model scores 51.6%, humans 84.4%","feed_subtitle":"New VF-Eval benchmark tests four feedback skills on AI-generated video; even the best model trails humans by 32.8 points.","key_machinery":"The load-bearing object is the VF-Eval dataset itself: 9,740 question-answer pairs built from AIGC videos generated by proprietary models (Pika, Kling, Pixeldance, Gen-3) and open-source models (T2V-Turbo-V2, plus videos from LaVie and OpenSora). The four tasks are the mechanism that operationalizes 'feedback': Yes/No questions for error awareness, multiple-choice for error type detection, and open-ended questions for coherence validation and reasoning evaluation, the latter judged by GPT-4.1-mini against human-written reference answers. The RePrompt experiment, where humans revise the original generation prompt and the revised prompt is used to regenerate the video, is the mechanism connecting benchmark performance to the practical goal of improving AIGC video generation.","core_discovery":"On its own terms, VF-Eval claims that frontier multimodal LLMs cannot yet reliably interpret AI-generated videos well enough to serve as feedback providers. The benchmark's four tasks—coherence validation, error awareness, error type detection, and reasoning evaluation—each measure a component of that ability, and the reported results show a wide gap: the strongest model, GPT-4.1, scores 51.6% overall, while humans score 84.4%. The paper also reports that open-source models are competitive with proprietary ones, that models often fail by relying on textual cues or commonsense rather than actual video content, and that a prompt-revision experiment (RePrompt) yields modest win rates for human-revised prompts over original ones, suggesting that better alignment with human feedback could make MLLM feedback useful for video generation.","pith_inferences":["We infer that, because every Error Awareness question has 'Yes' as the correct answer, a trivial always-yes baseline would score 100% on that task; the reported sub-random performance means the task primarily measures a response bias toward declaring videos error-free, not necessarily the absence of error-detection skill.","We infer that if the GPT-4.1-mini judge is found unreliable, the open-ended scores should be re-estimated with human adjudication or a panel of judges, and the benchmark's conclusions about coherence validation and reasoning would then need revision.","We infer a natural extension, not explored in the paper: testing image-to-video and audio-video AIGC, where error types such as audio-visual mismatch would likely create new failure modes.","We infer from the RePrompt win rates that a model scoring well on VF-Eval could act as an automated critic inside a generation pipeline, reducing human involvement in iterative prompt refinement."],"forward_implications":["Directly using MLLM feedback in AIGC video quality assessment pipelines is likely to produce inaccurate results until models improve.","The benchmark provides a training signal: open-source models are close to proprietary ones, so fine-tuning on VF-Eval has room to close part of the gap.","Combining MLLMs with computer-vision auxiliary methods should improve feedback precision, as the paper suggests.","RePrompt indicates that human-aligned prompt revision can improve subject consistency and aesthetic quality in regenerated videos, though gains in image and background quality are small.","The four-task structure gives a reusable protocol for measuring whether a model can critique AI-generated video, not just answer questions about natural video."],"supporting_citations":[{"why":"Provides the Videophy train split from which LaVie and OpenSora AIGC videos were collected for VF-Eval.","marker":"(Bansal et al., 2024)"},{"why":"LaVie is an open-source text-to-video model whose generated videos form part of the benchmark's AIGC content.","marker":"(Wang et al., 2023c)"},{"why":"OpenSora is an open-source text-to-video model whose generated videos are included in VF-Eval.","marker":"(Zheng et al., 2024)"},{"why":"T2V-Turbo-V2 is the open-source video generation model used to create AIGC videos for the dataset.","marker":"(Li et al., 2024b)"},{"why":"Q-Bench is an existing benchmark for MLLM-based low-level video/image quality assessment that VF-Eval extends and contrasts with.","marker":"(Wu et al., 2024a)"},{"why":"Video-MME is a natural-video understanding benchmark used as a comparison point to show AIGC-specific gaps.","marker":"(Fu et al., 2024)"},{"why":"VideoRepair is cited as prior evidence that integrating MLLM feedback into generation pipelines yields video quality gains, supporting RePrompt.","marker":"(Lee et al., 2024)"},{"why":"VBench represents traditional AIGC video quality assessment methods whose scores fall short of pinpointing human-preference divergences.","marker":"(Huang et al., 2024)"}],"fun_headline_variants":["MLLMs score 51.6% on AI-video feedback; humans hit 84.4%","New benchmark exposes MLLM blind spots on AIGC video feedback","Even GPT-4.1 fails at interpreting AI-generated videos","MLLMs lag humans by 33 points on AIGC video feedback tasks","VF-Eval: Frontier models struggle to critique AI videos"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The open-ended scores assume that GPT-4.1-mini, the LLM used to grade answers against human references, produces scores that faithfully reflect answer quality; if that judge is biased, noisy, or lenient, the reported gap between models and humans would be distorted.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs score 51.6% on AI-video feedback; humans hit 84.4%","New benchmark exposes MLLM blind spots on AIGC video feedback","Even GPT-4.1 fails at interpreting AI-generated videos","MLLMs lag humans by 33 points on AIGC video feedback tasks","VF-Eval: Frontier models struggle to critique AI videos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2893,"prompt_tokens":906,"completion_tokens":1987,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":1886}},"tokens_in":522,"tokens_out":1987,"duration_ms":12392,"temperature":1.0,"reasoning_tokens":1886,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:39:19.608623+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score a random sample of the open-ended responses (coherence validation and reasoning evaluation) with independent human raters instead of GPT-4.1-mini, and compare the resulting model-human gap to the reported 51.6% versus 84.4%; if the gap narrows substantially, the headline result is an artifact of the automated judge rather than a true capability gap.","supporting_citations":[],"review_version":1}