{"id":"2d323438-b2fc-43df-8e67-f309e257bac2","arxiv_id":"2411.17221","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new large benchmark of AI-generated videos with human ratings across four dimensions, plus an LMM-based video quality assessor that outperforms prior metrics.","lead":"Researchers introduce AIGVQA-DB, a set of 36,576 AI-generated videos from 15 text-to-video models, rated on four quality dimensions with both scores and pair comparisons. They also present AIGV-Assessor, a multimodal model that predicts video quality scores and preferences, reporting state-of-the-art results on this and existing benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pairwise ground truth is built from 3-annotator majority votes with no agreement or label-noise analysis, so the pair-ranking SOTA claim and human-preference alignment may be overstated.","rationale":"The reader's weakest assumption is exactly the one I find most load-bearing, so I mark agreement as 'agree' and leave the verdict CONDITIONAL unchanged. The dataset construction is a serious empirical effort, and the MOS side uses 20 annotators per video, which is more robust; the external correlation results in Tables 4-6 are useful evidence even though several gains are small and no confidence intervals are reported. However, the novel pair-ranking contribution is only as strong as the pair labels, and the manuscript does not supply the agreement statistics needed to show those labels are reliable. The proposed test, fresh annotations plus agreement analysis, would settle whether the reported pair accuracies reflect human preference or partly reflect fitting noisy labels. Until then, CONDITIONAL is the right verdict.","tokens_in":34318,"tokens_out":5897,"duration_ms":57460,"concrete_test":"On a random subset of, say, 500 pairs per dimension, have 5+ fresh annotators re-judge the same pairs and compute (i) Fleiss' kappa or pairwise agreement for the original 3-annotator labels and (ii) a new high-confidence majority label. Then recompute AIGV-Assessor's pair accuracy against the high-confidence labels only, and compare it with human self-consistency, i.e., the agreement between the original majority and the fresh majority. If the original agreement is low, or if the model's accuracy on unanimous/high-confidence pairs is close to or below human self-consistency, the pair-ranking SOTA claim should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (Sections 1 and 5.2) includes state-of-the-art pair ranking on AIGVQA-DB. That claim depends on the 30,000 pair labels produced by three annotators per pair and resolved by majority vote (Section 3.3). No inter-rater agreement, per-annotator accuracy, or label-noise statistics are reported; Appendix C.3 asserts consensus rather than measuring it. With only three binary votes, the majority label has a non-negligible error rate: if each annotator agrees with the true preference 80% of the time, the majority label is correct only about 90% of the time; at 70% individual accuracy it is correct about 78% of the time. Because the same noisy labels are used both as the stage-3 cross-entropy training target (Appendix E, Eq. 3) and as the evaluation ground truth for Table 3, the reported pair accuracy is accuracy against this noisy proxy, not against human preference. The external benchmarks provide only MOS, so they do not independently validate pair ranking. The protocol description is also internally unclear: Appendix C.3 says ties among three binary judgments are resolved by 'average opinions,' which is undefined for an odd number of annotators. These issues make the pair-ranking half of the central claim the least secure part of the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AIGVQA-DB, a large-scale text-to-video quality assessment dataset of 36,576 AI-generated videos from 15 generation models and 1,048 prompts, annotated with MOS and pairwise comparisons across four perceptual dimensions (static quality, temporal smoothness, dynamic degree, text-video correspondence). It also proposes AIGV-Assessor, an LMM-based model that fuses spatial (InternViT) and temporal (SlowFast) features, produces quality-level descriptions and regression scores, and is fine-tuned for pairwise preference prediction. The authors report state-of-the-art performance on AIGVQA-DB and on four external AIGVQA datasets (LGVQ, FETV, T2VQA-DB, GAIA).","tokens_in":82,"tokens_out":3664,"duration_ms":62694,"significance":"If the results are reliable, the dataset is a substantial community resource: 36,576 videos, 370k expert ratings, 30,000 annotated pairs, and a four-dimension annotation protocol go well beyond existing AIGVQA benchmarks in scale and scope. The proposed architecture is plausible, and the evaluation across five datasets with many baselines is commendable. The authors also release the dataset and code, which supports reproducibility. The central claim, however, rests on the quality of the pairwise labels and on the statistical strength of the reported gains; both need strengthening before the state-of-the-art claim can be accepted.","major_comments":[{"comment":"The pairwise ground truth is produced by three annotators per pair with majority voting, but no inter-rater agreement statistic, per-annotator accuracy, or label-noise analysis is reported. The description in Appendix C.3 that ties are resolved by 'the average opinions of the three participants' is undefined for an odd number of binary votes. Since the same pairwise labels are used both as the stage-3 training target (Appendix E, Eq. 3) and as the evaluation ground truth for the Pair Acc column of Table 3, the reported pair accuracy is accuracy against this noisy proxy. Please report inter-rater agreement, the distribution of majority margins (2-1 vs 3-0), a label-noise analysis, and, if possible, an external pairwise dataset for validation.","section":"Section 3.3 and Appendix C.3"},{"comment":"All results are averages over ten random splits, yet no standard deviations, confidence intervals, or significance tests are provided. Several headline gains over strong baselines are small: on LGVQ temporal smoothness AIGV-Assessor improves SRCC from 0.893 to 0.900 and PLCC from 0.907 to 0.920; on T2VQA-DB SRCC improves from 0.7965 to 0.8131. Because the differences are on the order of 0.007-0.02, significance testing is required to support the claim of 'state-of-the-art performance' over the best existing methods.","section":"Section 5.1 and Tables 3-6"},{"comment":"The dataset description is internally inconsistent. The abstract and introduction state 15 text-to-video models, but Section 3.1 says the MOS subset uses 12 generative models and the pair-comparison subset uses 12 generative models. Table 9 lists 15 models, with Gen-2, MoonValley, and Sora appearing only in the MOS subset and MorphStudio only in the pair subset; it is unclear how the 576 MOS videos and 36,000 pair videos combine into the stated total of 36,576, and whether the MOS subset overlaps with the pair subset. Please clarify the exact composition of each subset and the total number of distinct videos.","section":"Section 3.1, Appendix D.1, and Table 9"},{"comment":"The pairwise comparison training stage uses a 'judge network inspired by LPIPS' whose architecture, parameterization, and training status (frozen or trainable) are not specified. This makes it impossible to determine how much of the Pair Acc gain comes from the judge network versus the quality regression head. Please provide the full architecture, input features, and training details of the judge network.","section":"Section 4.2 and Appendix E"},{"comment":"For methods that do not have a dedicated pair-ranking head, it is unclear how Pair Acc is computed: whether each pair is classified by comparing the predicted scores of the two videos, or by feeding the pair into the model. This affects the fairness of the comparison, since AIGV-Assessor has an extra pairwise training stage. Please state the pair-evaluation protocol for all baselines and for AIGV-Assessor.","section":"Table 3 evaluation protocol"}],"minor_comments":[{"comment":"Reference [16] is cited for 'ITU-R BT.500-14' but the listed reference is a paper on confusing image quality assessment, not the ITU recommendation; please cite the correct standard.","section":"References"},{"comment":"Reference [68] is cited as InternViT and InternVL2-8B, but the reference is a paper about Charxiv chart understanding; this appears to be the wrong citation for the vision backbone and LLM used in the method.","section":"References"},{"comment":"The label 'Temporal smooothness' in Figure 5(a) contains a typo ('smooothness').","section":"Figure 5"},{"comment":"Several entries in Table 3 have formatting errors, such as '55.08%0.4594 0.4701' in the BVQA row and '0.8489' appearing in the simpleVQA row in a position that is inconsistent with the other rows; please correct the table formatting and verify the reported values.","section":"Table 3"},{"comment":"The heading 'Annotaion Criteria' contains a typo; it should read 'Annotation Criteria'.","section":"Appendix C.1"},{"comment":"The text-video correspondence criterion contains a double comma in the description of the 'Bad' level ('either missing or incorrectly represented, , resulting').","section":"Figure 17"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is potentially valuable and the paper is generally well-structured. The main risk is that the pairwise-ranking claim is evaluated on labels whose reliability is not quantified and whose tie-handling rule is undefined. The authors should also address the lack of significance testing, which is particularly important given the small numerical margins over strong baselines. The internal inconsistencies in subset composition and the incorrect backbone citation should be fixed before this can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe dataset is the real contribution here, and it is a good one. AIGVQA-DB — 36,576 videos, 15 T2V models, 1,048 prompts, four perceptual dimensions, and both MOS and pairwise labels — is a genuine step up from T2VQA-DB, LGVQ, FETV, and GAIA. The annotation effort is substantial, and the four-dimension breakdown (static quality, temporal smoothness, dynamic degree, TV correspondence) is useful. The paper also does a thorough job benchmarking existing methods and analyzing model strengths and weaknesses across prompt categories. That alone justifies a serious look.\n\nThe model, AIGV-Assessor, is a reasonable engineering combination: InternViT + SlowFast, projection layers, InternLM2 with LoRA, regression head, and a pairwise judge network. The ablations show each component adds something. So the method is credible, if not surprising.\n\nNow the soft spots. The pairwise ground truth is the least secure part. Each pair gets three binary votes, majority decides, and there is no inter-rater agreement, no label-noise analysis, and the tie-resolution sentence in Appendix C.3 (\"average opinions of the three participants\") is undefined. Since those same labels are training signal and evaluation ground truth, the pair-ranking SOTA claim is weaker than the MOS claim. The stress-test arithmetic is fair: at 80% per-annotator accuracy the majority label is only about 90% right, and the paper never gives us the per-annotator number.\n\nOther issues: no confidence intervals or significance tests despite ten random splits, and some reported gains are under a percent on FETV. The abstract says 370k ratings but Section 3.3 adds up to 406,080. The MOS subset is built from 48 Sora prompts only, which limits prompt diversity for that half. And the InternVL2 citation is wrong — reference [68] is Charxiv, not the model. None of these are fatal, but they need fixing.\n\nWho should read this: anyone building or evaluating T2V metrics. The benchmark will likely become a reference point once released. I would send it to review, with a request for annotation agreement statistics, corrected counts, significance testing, and a softened pair-ranking claim. It is a serious paper with a fixable weak spot.","headline":"A large, genuinely useful AIGVQA benchmark with a credible but unremarkable assessor; the pairwise-label noise and missing statistics make the pair-ranking SOTA claim the weak point.","tokens_in":35138,"tokens_out":2487,"would_cite":true,"duration_ms":23221,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AIGV-Assessor claims to outperform all existing methods for predicting both MOS and pairwise preferences on AI-generated video quality, across four dimensions.","keywords":["text-to-video generation","video quality assessment","large multimodal models","pairwise preference","mean opinion score","benchmark dataset","spatiotemporal features"],"falsifier":"Re-annotate a random subset of about 1,000 pairs with at least ten annotators each; if AIGV-Assessor's pair accuracy on this higher-confidence ground truth is substantially below the level it achieves on the original three-annotator majority labels, the claimed alignment with human preference would be overstated.","tokens_in":34044,"feed_emoji":"🎬","tokens_out":6167,"duration_ms":47689,"temperature":0.7,"pith_summary":"The paper tries to establish that AI-generated video quality can be measured automatically in a way that matches human preference across four distinct axes: static quality, temporal smoothness, dynamic degree, and text-video correspondence. To do this, it builds AIGVQA-DB, a dataset of 36,576 videos from 15 text-to-video generators and 1,048 prompts, annotated with both mean opinion scores and 30,000 pairwise comparisons. On top of the dataset, it proposes AIGV-Assessor, an LMM-based model that reads spatiotemporal visual features along with the prompt, outputs a quality-level description and a numerical score, and is fine-tuned to decide which of two videos is better. The paper reports that AIGV-Assessor outperforms existing scoring and evaluation methods on all four dimensions and on several other AIGV quality benchmarks. If correct, this gives the field a more human-aligned automatic evaluator for text-to-video models.","feed_headline":"AIGV-Assessor leads on all four AI-video quality axes","feed_subtitle":"A 36,576-video dataset plus an LMM scorer makes text-to-video quality checks track human preference more closely.","key_machinery":"AIGV-Assessor combines a 2D vision encoder (InternViT) for per-frame spatial content and a 3D SlowFast network for motion, projects both into the language space of an LLM (InternLM2-Chat-8B), and uses the LLM's hidden states for quality regression. The model is trained in three stages: aligning visual tokens with language, fine-tuning with LoRA and an L1 loss against MOS, and adding a pairwise comparison stage that uses an LPIPS-inspired judge network with cross-entropy loss. This design lets the same model produce quality-level text, numerical scores, and pairwise preferences.","core_discovery":"The central claim is that AIGV-Assessor achieves state-of-the-art performance for both MOS prediction and pair ranking on AIGVQA-DB, beating prior handcrafted, deep-learning, vision-language, and LMM-based methods on all four evaluated dimensions. The paper further claims that the model generalizes: it posts the best correlations on the LGVQ, FETV, T2VQA-DB, and GAIA benchmarks, and its predicted model rankings overlap most closely with ground-truth rankings among the compared evaluators.","pith_inferences":["If the pair labels are reliable, the three-annotator majority-vote design could be adopted by future dataset builders as a cheaper alternative to full MOS annotation for preference modeling.","The same spatiotemporal-plus-LMM architecture could be applied to other generative media, such as image-to-video or audio-driven video, where distinct quality axes also matter.","The prompt categorization could be used to build targeted stress tests for text-to-video models, testing specific failure modes like event-order violations or camera-view control.","The model's reliance on paired fine-tuning suggests that adding more pairwise data, or using synthetic pairs from strong generators, may further improve alignment with human preference."],"forward_implications":["Text-to-video models can now be ranked automatically on four separate quality axes, not just a single aggregate score.","The pairwise comparison subset offers a training signal that is complementary to MOS, reducing ambiguity of absolute ratings on high-quality content.","The dataset's prompt categorization (spatial content, temporal content, attribute control, complexity) allows diagnosing which video generators fail on which prompt types.","Existing AIGVQA benchmarks without pairwise data can still be evaluated by the model, giving cross-dataset comparisons.","The reported gains on external benchmarks suggest the method transfers beyond its own training distribution."],"supporting_citations":[{"why":"Supplies the InternViT vision encoder and InternVL2 LLM backbone used by AIGV-Assessor.","marker":"[68]"},{"why":"Supplies the SlowFast 3D network used to extract temporal motion features.","marker":"[17]"},{"why":"Supplies the LoRA fine-tuning strategy for adapting the vision encoder and LLM.","marker":"[23]"},{"why":"Supplies the LPIPS-inspired judge network design for pairwise quality comparison.","marker":"[81]"},{"why":"One of the external AIGVQA benchmarks used to validate the model's generalization.","marker":"[84]"},{"why":"Provides another external benchmark and the prompt categorization principles adopted for AIGVQA-DB.","marker":"[42]"},{"why":"Provides the T2VQA-DB benchmark used for cross-dataset evaluation.","marker":"[30]"},{"why":"Provides the GAIA benchmark used for evaluating action-quality assessment in AI-generated videos.","marker":"[11]"}],"fun_headline_variants":["AIGV-Assessor tops all four AI-video quality metrics","36K-video benchmark propels AIGV-Assessor to SOTA","New LMM video quality scorer beats the field","AIGV-Assessor bests prior methods on every quality axis","Text-to-video quality: AIGV-Assessor sets new record"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pairwise labels are trustworthy enough to serve as both training signal and evaluation ground truth, but each pair is judged by only three annotators with no reported agreement or noise analysis.","fun_headline_variants_meta":{"raw":{"variants":["AIGV-Assessor tops all four AI-video quality metrics","36K-video benchmark propels AIGV-Assessor to SOTA","New LMM video quality scorer beats the field","AIGV-Assessor bests prior methods on every quality axis","Text-to-video quality: AIGV-Assessor sets new record"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1738,"prompt_tokens":897,"completion_tokens":841,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":751}},"tokens_in":513,"tokens_out":841,"duration_ms":7324,"temperature":1.0,"reasoning_tokens":751,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:24:04.901096+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random subset of about 1,000 pairs with at least ten annotators each; if AIGV-Assessor's pair accuracy on this higher-confidence ground truth is substantially below the level it achieves on the original three-annotator majority labels, the claimed alignment with human preference would be overstated.","supporting_citations":[{"cited_title":"Slowfast networks for video recognition","cited_arxiv_id":null,"evidence_quote":"Supplies the SlowFast 3D network used to extract temporal motion features."},{"cited_title":"The unreasonable effectiveness of deep features as a perceptual metric","cited_arxiv_id":null,"evidence_quote":"Supplies the LPIPS-inspired judge network design for pairwise quality comparison."},{"cited_title":"Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation","cited_arxiv_id":null,"evidence_quote":"Provides another external benchmark and the prompt categorization principles adopted for AIGVQA-DB."}],"review_version":1}