Pith. sign in

REVIEW 5 major objections 6 minor 120 references

This paper claims that a single mixture-of-experts model can rate, compare, and answer questions about AI-generated human-centric videos across three quality dimensions, and that the accompanying dataset is the largest of its kind.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 20:02 UTC pith:MHGXVTQA

load-bearing objection HVEval+ is a substantial dataset extension, but the paper under-validates the three-way dimension split and overstates the model's margin over fine-tuned baselines. the 5 major comments →

arxiv 2607.16742 v1 pith:MHGXVTQA submitted 2026-07-18 cs.CV

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

classification cs.CV
keywords video quality assessmenttext-to-video generationhuman-centric videodatasetpairwise preferencemixture of expertsmultimodal large language modelbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper is trying to establish two things: that HVEval+ is the largest holistic quality assessment dataset for AI-generated human-centric videos, and that a single mixture-of-experts model, MoE-Rater, can handle all assessment tasks—multi-dimensional quality rating, pairwise comparison, and category-specific question answering—within one unified framework. To support this, the authors extended an earlier dataset with 60,000 pairwise preference judgments alongside 60,000 mean opinion scores and 20,000 yes/no Q&A pairs, all across three dimensions: spatial quality, temporal quality, and text-video correspondence. They then built MoE-Rater on a multimodal large language model with task-routed projectors and LoRA experts, trained in three stages. If the claims hold, the dataset gives the field a large preference-aligned benchmark, and the model provides a single tool that can both score videos and rank the T2V models that made them.

Core claim

The central claim is that HVEval+ is the largest holistic quality assessment dataset for AI-generated human-centric videos, comprising 1k prompts organized into 7 major categories, 20k videos from 24 text-to-video models, 60k MOSs and 60k preference pairs across three dimensions, and 20k category-specific Q&A pairs. The companion claim is that MoE-Rater, an MLLM-based all-in-one method, achieves superior performance on both HVEval+ and Human-AGVQA, with near-1.0 rank correlation to human judgments when ordering T2V models and strong zero-shot transfer to a related dataset.

What carries the argument

The carrying mechanism is the MoE-Rater routing architecture combined with HVEval+'s paired annotation design. S-MoPE and T-MoPE are mixtures of task-specific MLP projectors that route spatial and temporal features to top-k experts via gating networks; MoLE applies the same idea to LoRA experts inside the LLM. A shared set of task embeddings drives all gating networks, and a three-stage training strategy—task-aware pre-training, task-specific adaptation, and adaptive routing optimization—lets experts specialize and then recombine. On the dataset side, 60,000 same-prompt video pairs convert relative preferences into training signal.

Load-bearing premise

The load-bearing premise is that the 33 annotators actually rated spatial quality, temporal quality, and text-video correspondence as three separate, non-overlapping things—but the paper reports no inter-rater reliability or cross-dimension confusion analysis, so if annotators conflated the dimensions, the MOSs and preference labels become noisy and the benchmark's value drops.

What would settle it

Re-analyze the raw per-subject scores: compute per-dimension inter-rater agreement using a standard coefficient (e.g., alpha or kappa) and the correlation between the same subject's ratings across dimensions on the same videos. If within-dimension agreement is low (below 0.6) or between-dimension subject correlations rival within-dimension ones, the three-way separation is unsupported and the central benchmark claim collapses. A small dimension-classification test—asking new annotators to label which dimension a given distortion belongs to—would also settle it.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • HVEval+ provides a large multi-dimensional benchmark; methods trained or fine-tuned on it substantially outperform generic video-quality models on AI-generated human-centric videos.
  • A single MoE-Rater handles rating, comparison, and Q&A, so T2V model evaluation no longer requires separate pipelines for each task.
  • MoE-Rater's predicted scores reproduce human T2V model rankings with near-1.0 correlation, making automated model comparison viable.
  • Zero-shot MoE-Rater beats most fine-tuned methods on Human-AGVQA, suggesting the dataset teaches transferable quality assessment skills.
  • The 60k preference pairs provide training signal for reward or judgment models that could be used to optimize T2V generators.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The pairwise preference structure could be repurposed as a reward model for optimizing T2V generation, since same-prompt preference pairs are exactly the input reward models are trained on; the paper notes this potential but does not train such a model.
  • If the three dimensions are truly separable in the annotations, the dataset becomes a diagnostic tool: a video with high spatial but low temporal scores points to motion smoothness as the failure mode, which could guide model-specific improvements.
  • A testable extension is to measure inter-rater agreement per dimension and cross-dimension confusion; the paper does not report these, so dimension cleanliness remains an open empirical question.
  • The three-stage expert-routing recipe may generalize to other multi-task MLLM settings beyond video quality, but that is speculation beyond the paper's scope.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper presents HVEval+, an extension of the authors' earlier HVEval dataset for quality assessment of AI-generated human-centric videos. HVEval+ contains 1,000 prompts across a 7-category taxonomy, 20,000 videos from 24 T2V models, 60,000 MOSs, 60,000 pairwise preference labels across three dimensions (spatial quality, temporal quality, text-video correspondence), and 20,000 category-specific Q&A pairs. The paper also proposes MoE-Rater, an MLLM-based all-in-one model with Mixture of Projector Experts (S-MoPE/T-MoPE) and Mixture of LoRA Experts (MoLE), trained in three stages to unify rating, pairwise comparison, and question answering. Experiments on HVEval+ and zero-shot/fine-tuned evaluations on Human-AGVQA are reported, with MoE-Rater achieving the best results on most metrics.

Significance. If the annotation dimensions are reliable, HVEval+ would be a valuable large-scale multi-dimensional benchmark for AI-generated human-centric video quality assessment, and MoE-Rater would be a useful unified framework for rating, comparison, and QA. The prompt-based train/test split avoids content leakage, and the cross-dataset zero-shot evaluation on Human-AGVQA is a genuine strength that provides independent grounding. The subjective protocol follows standard outlier-rejection practice. However, the load-bearing assumption that annotators can reliably separate the three quality dimensions is not validated by any inter-rater reliability or agreement analysis, which makes the dataset's multi-dimensional claims and the model's dimension-specific gains conditional on an untested premise.

major comments (5)
  1. [§III-C/III-D] The load-bearing assumption of the dataset is that 33 annotators can reliably separate spatial quality, temporal quality, and text-video correspondence. The subjective-study description reports only a 15-item screening test, per-batch outlier rejection, and ITU-R BT.500 outlier removal; no inter-rater reliability (e.g., ICC, Krippendorff's α), no per-dimension agreement, and no cross-dimension confusion analysis are reported. Section VI acknowledges annotation uncertainty but does not quantify it. If annotators conflated dimensions, the 60k MOSs and preference labels are noisy, and the multi-dimensional gains of MoE-Rater in Tables III/IV/VII could reflect fitting to shared variance rather than true dimensional understanding. Please add per-dimension agreement statistics, a dimension-separability analysis (e.g., correlation or confusion matrix of per-video dimension MOSs), and a statemen
  2. [§III-B/III-D] The pairwise preference labels and MOSs were collected from the same videos, yet the paper never checks whether the two protocols are internally consistent: for a given pair and dimension, does the video with the higher MOS also win the majority preference vote? Without this consistency check, both annotation layers remain unvalidated. Please report per-dimension agreement between MOS differences and preference outcomes, and specify the number of ratings per video/pair and the majority-vote rule.
  3. [§V-A/Table III] The 'fine-tuned on HVEval+' baselines (†) are used to support the claim of superiority, but no fine-tuning protocol is described for them. Differences in training data, number of steps, loss, and task mixture could explain part of the gap (e.g., spatial SRCC 0.891 vs 0.860 against Qwen2.5-VL†). Please specify the exact protocol and, ideally, use the same three-stage recipe for baselines or state why that is not feasible.
  4. [§V-B/Table VI] Table VI reports 'SRCC to Human' values approaching 0.99, but the text does not state whether the per-model average scores are computed on the training split, test split, or the full dataset. Because MoE-Rater is trained on 80% of HVEval+ videos, ranking correlations on training data would be inflated by design. Please recompute and report ranking correlations on the prompt-disjoint test split only.
  5. [§V-B/Tables III-V] No confidence intervals or significance tests are reported for SRCC/PLCC/accuracy. Several margins are modest (e.g., correspondence SRCC 0.787 vs 0.755 for InternVL2.5†), so some 'superior' claims may be within noise. Add bootstrap confidence intervals or paired significance tests for the key comparisons.
minor comments (6)
  1. [Table I] HVEval+ is described as an extension of HVEval, but the table reports 80k MOS for HVEval and 60k MOS for HVEval+; clarify the relation and annotation counts.
  2. [§IV-B] The equations define features for one video; it is unclear how two input videos are processed for pairwise comparison (concatenation, separate placeholders, fusion).
  3. [§III-C] The pass threshold for the 15-item training test and the number of rejected subjects are not reported.
  4. [General] No dataset availability statement or download link is provided; for a dataset paper this should be added.
  5. [Figures 4/5] Figure 4 caption reports only spatial and text-video correspondence scores while Figure 5 includes temporal; make the reported dimensions consistent across figures.
  6. [Table VII] The row configuration of the ablation table is hard to read; mark each component explicitly for every row.

Circularity Check

0 steps flagged

No significant circularity: HVEval+ and MoE-Rater are standard supervised benchmarking with an independent cross-dataset zero-shot evaluation.

full rationale

The paper's central claims are that HVEval+ is a large multi-dimensional dataset and that MoE-Rater performs well on it and on Human-AGVQA. Training and evaluating a model on the same benchmark is standard supervised practice, not circular, and the paper compares against other methods trained/fine-tuned on the same HVEval+ data. The cross-dataset zero-shot evaluation on Human-AGVQA (Table V) provides an external check: MoE-Rater is trained only on HVEval+ and evaluated on Human-AGVQA without fine-tuning, so the reported transfer performance is not a re-statement of its training targets. The self-citation of HVEval [11] is transparent provenance for the dataset extension and is not used to justify the model's superiority or to exclude alternatives; no uniqueness theorem or load-bearing argument is imported from prior work by the same authors. The three annotation dimensions are defined by independent human rating instructions, not in terms of MoE-Rater outputs, and no equation equates a predicted quantity to a fitted parameter. The limitation section's acknowledgment of annotation uncertainty is a validity concern, not circularity. Overall, no load-bearing step reduces by construction to its own inputs, so the circularity score is 0.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The central claims rest on the assumption that human annotations are reliable and dimensionally separable, and that the prompt/model coverage is representative. The model introduces two new architectural modules (MoPE, MoLE) with hand-chosen hyperparameters (N=7, k=3, LoRA rank 8) whose choices are not systematically justified. No external benchmarks for the modules are used beyond the authors' own dataset.

free parameters (4)
  • Number of experts N = 7
    Chosen by hand in Section V-A; no sensitivity analysis provided.
  • Top-k routing k = 3
    Chosen by hand in Section V-A; ablation does not vary k.
  • LoRA rank = 8
    Standard choice for LoRA; no tuning reported.
  • Task embedding dimension = 256
    Set in Section V-A without justification or sensitivity analysis.
axioms (4)
  • domain assumption The three quality dimensions (spatial, temporal, correspondence) are well-defined and consistently interpreted by annotators.
    The entire dataset assumes annotators can distinguish these dimensions; no inter-rater reliability is reported.
  • domain assumption The prompt taxonomy (7 major categories, 22 subcategories) comprehensively covers human-centric T2V scenarios.
    The taxonomy is designed by the authors and used to generate prompts; if it is not representative, the dataset coverage is skewed.
  • domain assumption ITU-R BT.500 outlier rejection is valid for this annotation protocol.
    The paper applies a standard method, but the protocol differs from typical television-quality tests; its suitability is assumed.
  • domain assumption Pairwise preferences and majority voting yield reliable ground truth labels.
    The paper does not analyze annotator agreement on pairwise choices or Q&A answers.
invented entities (2)
  • Mixture of Projector Experts (MoPE) no independent evidence
    purpose: Task-specific projection of spatial and temporal features into the LLM input space.
    The module is evaluated only through ablation on the same dataset; no external or theoretical evidence is provided.
  • Mixture of LoRA Experts (MoLE) no independent evidence
    purpose: Task-specific low-rank adaptation of the LLM in a unified multi-task model.
    Similar to MoPE, its effectiveness is shown via ablation on HVEval+ only.

pith-pipeline@v1.3.0-alltime-deepseek · 33180 in / 9409 out tokens · 89859 ms · 2026-08-01T20:02:57.667127+00:00 · methodology

0 comments
read the original abstract

AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring the importance of effective quality assessment for such videos. To this end, we extend our previous dataset HVEval with pairwise preference annotations, resulting in HVEval+, the largest holistic quality assessment dataset for AI-generated human-centric videos, which comprises 1k prompts based on a comprehensive taxonomy, 20k videos generated by 24 text-to-video (T2V) models, and extensive human annotations, including 60k mean opinion scores (MOSs) and 60k preference pairs across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), as well as 20k category-specific question-answer (Q&A) pairs. Along with the HVEval+ dataset, we further propose MoE-Rater, a Mixture-of-Experts (MoE)-inspired and multimodal large language model (MLLM)-based all-in-one method that supports multi-dimensional quality rating, multi-dimensional pairwise comparison, and category-specific question answering within a single model. Specifically, we introduce Mixture of Projector Experts (MoPE) and Mixture of LoRA Experts (MoLE), together with a three-stage training strategy consisting of task-aware pre-training, task-specific adaptation, and adaptive routing optimization, to effectively unify multiple tasks, resulting in superior performance on both HVEval+ and Human-AGVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the HVEval+ dataset and the MoE-Rater method in advancing AI-generated video quality assessment and further facilitating the evaluation and optimization of T2V models.

Figures

Figures reproduced from arXiv: 2607.16742 by Guangtao Zhai, Huiyu Duan, Patrick Le Callet, Sijing Wu, Xiongkuo Min, Yucheng Zhu, Yunhao Li.

Figure 1
Figure 1. Figure 1: Illustration of typical distortions in AI-generated human-centric videos. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the prompt suite statistics. Left: taxonomy of the human [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Dataset construction pipeline. (a) We first generate 1k prompts based [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Sample video frames generated by 24 T2V models using the prompt “Two parents and their kids are building a snowman together in a snowy park”. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Examples from the proposed HVEval+ dataset. (a)-(d) illustrate the good and bad cases across four evaluation dimensions: spatial quality, temporal [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: MOS distribution histograms and kernel density curves for spatial [PITH_FULL_IMAGE:figures/full_fig_p005_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: MOS distributions of (a) spatial quality, (b) temporal quality, and (c) text-video correspondence across the seven prompt categories. [PITH_FULL_IMAGE:figures/full_fig_p006_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Comparison of average spatial quality, temporal quality, text-video correspondence, and category-specific accuracy across the 22 prompt subcategories. [PITH_FULL_IMAGE:figures/full_fig_p006_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Comparison of the 24 T2V models in terms of average spatial quality, temporal quality, text-video correspondence, category-specific accuracy, and [PITH_FULL_IMAGE:figures/full_fig_p006_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Overview of MoE-Rater. MoE-Rater serves as an all-in-one model that supports all evaluation functions, including single-video quality scoring across [PITH_FULL_IMAGE:figures/full_fig_p007_10.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

120 extracted references · 29 linked inside Pith

  1. [1]

    Hunyuanvideo: A systematic framework for large video generative models,

    W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhanget al., “Hunyuanvideo: A systematic framework for large video generative models,”arXiv preprint arXiv:2412.03603,

  2. [2]

    Wan: Open and advanced large-scale video generative models,

    T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yanget al., “Wan: Open and advanced large-scale video generative models,”arXiv preprint arXiv:2503.20314, 2025. I, II-A, II-A, II, III-B, VI

  3. [3]

    Magi-1: Autoregressive video generation at scale,

    H. Teng, H. Jia, L. Sun, L. Li, M. Li, M. Tang, S. Han, T. Zhang, W. Zhang, W. Luoet al., “Magi-1: Autoregressive video generation at scale,”arXiv preprint arXiv:2505.13211, 2025. I, II-A, II-A, II, III-B, VI

  4. [4]

    Team, “Sora,” https://openai.com/sora/, 2025

    O. Team, “Sora,” https://openai.com/sora/, 2025. I, II, III-B, VI

  5. [5]

    Pixverse,

    PixVerse, “Pixverse,” https://app.pixverse.ai, 2025. I, II, III-B, VI

  6. [6]

    I, II, III-B, VI

    Pika, “Pika,” https://www.pika.art, 2025. I, II, III-B, VI

  7. [7]

    Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation,

    H. Li, M. Xu, Y . Zhan, S. Mu, J. Li, K. Cheng, Y . Chen, T. Chen, M. Ye, J. Wanget al., “Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 7752–7762. I

  8. [8]

    Humanomni: A large vision-speech language model for human-centric video understanding,

    J. Zhao, Q. Yang, Y . Peng, D. Bai, S. Yao, B. Sun, X. Chen, S. Fu, X. Wei, L. Boet al., “Humanomni: A large vision-speech language model for human-centric video understanding,”arXiv preprint arXiv:2501.15111, 2025. I

  9. [9]

    Singinghead: A large-scale 4d dataset for singing head animation,

    S. Wu, Y . Li, W. Zhang, J. Jia, Y . Zhu, Y . Yan, G. Zhai, and X. Yang, “Singinghead: A large-scale 4d dataset for singing head animation,” IEEE Transactions on Multimedia, 2025. I

  10. [10]

    Human-activity agv quality assessment: A benchmark dataset and an objective evaluation metric,

    Z. Zhang, W. Sun, X. Li, Y . Li, Q. Ge, J. Jia, Z. Zhang, Z. Ji, F. Sun, S. Juiet al., “Human-activity agv quality assessment: A benchmark dataset and an objective evaluation metric,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 6771–6780. I, I, II-B, II-C, V, V-A, V-A, V-B, V-B, V-B

  11. [11]

    Hveval: Towards unified evaluation of human-centric video generation and understand- ing,

    S. Wu, Y . Li, H. Duan, Y . Jiang, Y . Zhu, and G. Zhai, “Hveval: Towards unified evaluation of human-centric video generation and understand- ing,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 13 376–13 383. I, I, II-C

  12. [12]

    Is this generated person existed in real-world? fine-grained detecting and calibrating abnormal human-body,

    Z. Wang, Q. Ma, W. Wan, H. Li, K. Wang, and Y . Tian, “Is this generated person existed in real-world? fine-grained detecting and calibrating abnormal human-body,”arXiv preprint arXiv:2411.14205,

  13. [13]

    Improved techniques for training gans,

    T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,”Advances in neural information processing systems, vol. 29, 2016. I, II-B, II-B

  14. [14]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium,

    M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,”Advances in neural information processing systems, vol. 30, 2017. I, II-B, II-B

  15. [15]

    Towards accurate generative models of video: A new metric & challenges,

    T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michal- ski, and S. Gelly, “Towards accurate generative models of video: A new metric & challenges,”arXiv preprint arXiv:1812.01717, 2018. I, II-B, II-B

  16. [16]

    Vbench: Comprehensive benchmark suite for video generative models,

    Z. Huang, Y . He, J. Yu, F. Zhang, C. Si, Y . Jiang, Y . Zhang, T. Wu, Q. Jin, N. Chanpaisitet al., “Vbench: Comprehensive benchmark suite for video generative models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 807–21 818. I, I, II-B, II-B, III, V

  17. [17]

    Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness,

    D. Zheng, Z. Huang, H. Liu, K. Zou, Y . He, F. Zhang, Y . Zhang, J. He, W.-S. Zheng, Y . Qiaoet al., “Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness,”arXiv preprint arXiv:2503.21755, 2025. I, I, II-B, II-B

  18. [18]

    T2v- compbench: A comprehensive benchmark for compositional text-to- video generation,

    K. Sun, K. Huang, X. Liu, Y . Wu, Z. Xu, Z. Li, and X. Liu, “T2v- compbench: A comprehensive benchmark for compositional text-to- video generation,”arXiv preprint arXiv:2407.14505, 2024. I, I, II-B, II-B

  19. [19]

    Subjective-aligned dataset and metric for text-to-video quality assessment,

    T. Kou, X. Liu, Z. Zhang, C. Li, H. Wu, X. Min, G. Zhai, and N. Liu, “Subjective-aligned dataset and metric for text-to-video quality assessment,” inProceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 7793–7802. I, I, II-B, II-B, II-C, V

  20. [20]

    Benchmarking aigc video quality assessment: A dataset and unified model,

    Z. Zhang, X. Li, W. Sun, J. Jia, X. Min, Z. Zhang, C. Li, Z. Chen, P. Wang, Z. Jiet al., “Benchmarking aigc video quality assessment: A dataset and unified model,”arXiv preprint arXiv:2407.21408, 2024. I, I, II-B, II-B, II-C, V

  21. [21]

    Aigv-assessor: benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,

    J. Wang, H. Duan, G. Zhai, J. Wang, and X. Min, “Aigv-assessor: benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 18 869–18 880. I, II-B, II-B, II-C

  22. [22]

    Evalcrafter: Benchmarking and evaluating large video generation models,

    Y . Liu, X. Cun, X. Liu, X. Wang, Y . Zhang, H. Chen, Y . Liu, T. Zeng, R. Chan, and Y . Shan, “Evalcrafter: Benchmarking and evaluating large video generation models,” inProceedings of the IEEE/CVF Conference 13 on Computer Vision and Pattern Recognition, 2024, pp. 22 139–22 149. I, V

  23. [23]

    Fetv: A benchmark for fine-grained evaluation of open-domain text-to- video generation,

    Y . Liu, L. Li, S. Ren, R. Gao, S. Li, S. Chen, X. Sun, and L. Hou, “Fetv: A benchmark for fine-grained evaluation of open-domain text-to- video generation,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 352–62 387, 2023. I

  24. [24]

    Evaluating and improving compositional text-to-visual generation,

    B. Li, Z. Lin, D. Pathak, J. Li, Y . Fei, K. Wu, X. Xia, P. Zhang, G. Neubig, and D. Ramanan, “Evaluating and improving compositional text-to-visual generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5290–5301. I

  25. [25]

    F-bench: Rethinking human preference evaluation metrics for benchmarking face generation, customization, and restoration,

    L. Liu, H. Duan, Q. Hu, L. Yang, C. Cai, T. Ye, H. Liu, X. Zhang, and G. Zhai, “F-bench: Rethinking human preference evaluation metrics for benchmarking face generation, customization, and restoration,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 10 982–10 994. I

  26. [26]

    Multi- dimensional text-to-face image quality assessment using llm: Database and method,

    Y . Gao, X. Min, J. Han, Y . Cao, S. Wu, Y . Dou, and G. Zhai, “Multi- dimensional text-to-face image quality assessment using llm: Database and method,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 6948–6957. I

  27. [27]

    Aghi-qa: A subjective-aligned dataset and metric for ai- generated human images,

    Y . Li, S. Wu, W. Sun, Z. Zhang, Y . Zhu, Z. Zhang, H. Duan, X. Min, and G. Zhai, “Aghi-qa: A subjective-aligned dataset and metric for ai- generated human images,”IEEE Transactions on Circuits and Systems for Video Technology, 2025. I

  28. [28]

    Dreambench++: A human- aligned benchmark for personalized image generation,

    Y . Peng, Y . Cui, H. Tang, Z. Qi, R. Dong, J. Bai, C. Han, Z. Ge, X. Zhang, and S.-T. Xia, “Dreambench++: A human- aligned benchmark for personalized image generation,”arXiv preprint arXiv:2406.16855, 2024. I, III-A, III-A

  29. [29]

    Chatgpt-4o,

    OpenAI, “Chatgpt-4o,” https://openai.com, 2025. I, III-A, III-A

  30. [30]

    Animatediff-lightning: Cross-model diffusion distillation,

    S. Lin and X. Yang, “Animatediff-lightning: Cross-model diffusion distillation,”arXiv preprint arXiv:2403.12706, 2024. II-A, II-A, II, III-B, VI

  31. [31]

    Animatelcm: Computation-efficient personalized style video generation without personalized video data,

    F.-Y . Wang, Z. Huang, W. Bian, X. Shi, K. Sun, G. Song, Y . Liu, and H. Li, “Animatelcm: Computation-efficient personalized style video generation without personalized video data,” inSIGGRAPH Asia 2024 Technical Communications, 2024, pp. 1–5. II-A, II-A, II, III-B, VI

  32. [32]

    Magictime: Time-lapse video generation models as metamorphic simulators,

    S. Yuan, J. Huang, Y . Shi, Y . Xu, R. Zhu, B. Lin, X. Cheng, L. Yuan, and J. Luo, “Magictime: Time-lapse video generation models as metamorphic simulators,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. II-A, II-A, II, III-B, VI

  33. [33]

    Modelscope text-to-video technical report,

    J. Wang, H. Yuan, D. Chen, Y . Zhang, X. Wang, and S. Zhang, “Modelscope text-to-video technical report,”arXiv preprint arXiv:2308.06571, 2023. II-A, II-A, II, III-B, VI

  34. [34]

    Show-1: Marrying pixel and latent diffusion models for text-to-video generation,

    D. J. Zhang, J. Z. Wu, J.-W. Liu, R. Zhao, L. Ran, Y . Gu, D. Gao, and M. Z. Shou, “Show-1: Marrying pixel and latent diffusion models for text-to-video generation,”International Journal of Computer Vision, pp. 1–15, 2024. II-A, II-A, II, III-B, VI

  35. [35]

    T2v-turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design,

    J. Li, Q. Long, J. Zheng, X. Gao, R. Piramuthu, W. Chen, and W. Y . Wang, “T2v-turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design,”arXiv preprint arXiv:2410.05677, 2024. II-A, II-A, II, III-B, VI

  36. [36]

    Videocrafter2: Overcoming data limitations for high-quality video diffusion models,

    H. Chen, Y . Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y . Shan, “Videocrafter2: Overcoming data limitations for high-quality video diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 7310–7320. II-A, II-A, II, III-B, VI

  37. [37]

    Zeroscope,

    S. Sterling, “Zeroscope,” https://huggingface.co/cerspense/zeroscope v2 576w, 2023. II-A, II-A, II, III-B, VI

  38. [38]

    Cogvideox: Text-to- video diffusion models with an expert transformer,

    Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y . Yang, W. Hong, X. Zhang, G. Fenget al., “Cogvideox: Text-to- video diffusion models with an expert transformer,”arXiv preprint arXiv:2408.06072, 2024. II-A, II-A, II, III-B, VI

  39. [39]

    Ltx-video: Realtime video latent diffusion,

    Y . HaCohen, N. Chiprut, B. Brazowski, D. Shalem, D. Moshe, E. Richardson, E. Levin, G. Shiran, N. Zabari, O. Gordon et al., “Ltx-video: Realtime video latent diffusion,”arXiv preprint arXiv:2501.00103, 2024. II-A, II-A, II, III-B, VI

  40. [40]

    Latte: Latent diffusion transformer for video generation,

    X. Ma, Y . Wang, G. Jia, X. Chen, Z. Liu, Y .-F. Li, C. Chen, and Y . Qiao, “Latte: Latent diffusion transformer for video generation,” arXiv preprint arXiv:2401.03048, 2024. II-A, II-A, II, III-B, VI

  41. [41]

    Step-video-t2v technical report: The practice, challenges, and future of video foundation model,

    G. Ma, H. Huang, K. Yan, L. Chen, N. Duan, S. Yin, C. Wan, R. Ming, X. Song, X. Chenet al., “Step-video-t2v technical report: The practice, challenges, and future of video foundation model,”arXiv preprint arXiv:2502.10248, 2025. II-A, II-A, II, III-B, VI

  42. [42]

    Mochi 1,

    G. Team, “Mochi 1,” https://github.com/genmoai/models, 2024. II-A, II-A, II, III-B, VI

  43. [43]

    Skyreels-v2: Infinite-length film generative model,

    G. Chen, D. Lin, J. Yang, C. Lin, J. Zhu, M. Fan, H. Zhang, S. Chen, Z. Chen, C. Maet al., “Skyreels-v2: Infinite-length film generative model,”arXiv preprint arXiv:2504.13074, 2025. II-A, II-A, II, III-B, VI

  44. [44]

    Cofnet: contrastive object-aware fusion using box-level masks for multispectral object detection,

    M. Zhou, Y . Li, G. Yang, X. Wei, H. Pu, J. Luo, and W. Jia, “Cofnet: contrastive object-aware fusion using box-level masks for multispectral object detection,”IEEE Transactions on Multimedia, 2025. II-B

  45. [45]

    Afes: Attention-based feature excitation and sorting for action recognition,

    M. Zhou, J. Li, X. Wei, J. Luo, H. Pu, W. Wang, J. He, and Z. Shang, “Afes: Attention-based feature excitation and sorting for action recognition,”IEEE Transactions on Consumer Electronics, 2025. II-B

  46. [46]

    An end-to-end blind image quality assessment method using a recurrent network and self-attention,

    M. Zhou, X. Lan, X. Wei, X. Liao, Q. Mao, Y . Li, C. Wu, T. Xiang, and B. Fang, “An end-to-end blind image quality assessment method using a recurrent network and self-attention,”IEEE Transactions on Broadcasting, vol. 69, no. 2, pp. 369–377, 2022. II-C

  47. [47]

    Attentional feature fusion for end-to-end blind image quality assessment,

    M. Zhou, S. Lang, T. Zhang, X. Liao, Z. Shang, T. Xiang, and B. Fang, “Attentional feature fusion for end-to-end blind image quality assessment,”IEEE Transactions on Broadcasting, vol. 69, no. 1, pp. 144–152, 2022. II-C

  48. [48]

    Multilevel feature fusion for end-to-end blind image quality assessment,

    X. Lan, M. Zhou, X. Xu, X. Wei, X. Liao, H. Pu, J. Luo, T. Xiang, B. Fang, and Z. Shang, “Multilevel feature fusion for end-to-end blind image quality assessment,”IEEE Transactions on Broadcasting, vol. 69, no. 3, pp. 801–811, 2023. II-C

  49. [49]

    Blind image quality assessment based on perceptual comparison,

    A. Li, J. Wu, Y . Liu, L. Li, W. Dong, and G. Shi, “Blind image quality assessment based on perceptual comparison,”IEEE Transactions on Multimedia, vol. 26, pp. 9671–9682, 2024. II-C

  50. [50]

    Graph-represented distribution similarity index for full-reference image quality assess- ment,

    W. Shen, M. Zhou, J. Luo, Z. Li, and S. Kwong, “Graph-represented distribution similarity index for full-reference image quality assess- ment,”IEEE Transactions on Image Processing, vol. 33, pp. 3075– 3089, 2024. II-C

  51. [51]

    Bridging the synthetic-to-authentic gap: Distortion-guided unsupervised domain adaptation for blind image quality assessment,

    A. Li, J. Wu, Y . Liu, and L. Li, “Bridging the synthetic-to-authentic gap: Distortion-guided unsupervised domain adaptation for blind image quality assessment,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 28 422–28 431. II-C

  52. [52]

    Blind image quality assessment: Exploring content fidelity perceptibility via quality adversarial learning,

    M. Zhou, W. Shen, X. Wei, J. Luo, F. Jia, X. Zhuang, and W. Jia, “Blind image quality assessment: Exploring content fidelity perceptibility via quality adversarial learning,”International Journal of Computer Vision, vol. 133, no. 6, pp. 3242–3258, 2025. II-C

  53. [53]

    Image quality assessment: Investigating causal perceptual effects with abductive counterfactual inference,

    W. Shen, M. Zhou, Y . Chen, X. Wei, Y . Feng, H. Pu, and W. Jia, “Image quality assessment: Investigating causal perceptual effects with abductive counterfactual inference,” inProceedings of the computer vision and pattern recognition conference, 2025, pp. 17 990–17 999. II-C

  54. [54]

    No-reference image quality assessment: Exploring intrinsic distortion characteristics via generative noise estimation with mamba,

    X. Lan, W. Xian, M. Zhou, J. Yan, X. Wei, J. Luo, W. Jia, and S. Kwong, “No-reference image quality assessment: Exploring intrinsic distortion characteristics via generative noise estimation with mamba,” IEEE Transactions on Circuits and Systems for Video Technology,

  55. [55]

    Towards syn-to-real iqa: A novel perspective on reshaping synthetic data distributions,

    A. Li, J. Wu, Y . Liu, L. Li, and W. Dong, “Towards syn-to-real iqa: A novel perspective on reshaping synthetic data distributions,”arXiv preprint arXiv:2601.00225, 2026. II-C

  56. [56]

    Uni-iqa: A unified approach for mutual promotion of natural and screen content image quality assessment,

    M. Song, C. Chen, W. Song, and Y . Fang, “Uni-iqa: A unified approach for mutual promotion of natural and screen content image quality assessment,”IEEE Transactions on Circuits and Systems for Video Technology, 2025. II-C

  57. [57]

    Beyond pixels: text-guided deep insights into graphic design image aesthetics,

    G. Shi, L. Li, and M. Song, “Beyond pixels: text-guided deep insights into graphic design image aesthetics,”Journal of Electronic Imaging, vol. 33, no. 5, pp. 053 059–053 059, 2024. II-C

  58. [58]

    A deep learning based no- reference quality assessment model for ugc videos,

    W. Sun, X. Min, W. Lu, and G. Zhai, “A deep learning based no- reference quality assessment model for ugc videos,” inProceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 856–865. II-C, II-C, III, V

  59. [59]

    Analysis of video quality datasets via design of minimalistic video quality models,

    W. Sun, W. Wen, X. Min, L. Lan, G. Zhai, and K. Ma, “Analysis of video quality datasets via design of minimalistic video quality models,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. II-C, II-C, III

  60. [60]

    Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,

    H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,” inEuropean conference on computer vision. Springer, 2022, pp. 538–554. II-C, II-C, III, V

  61. [61]

    Dynamic expert- knowledge ensemble for generalizable video quality assessment,

    P. Chen, L. Li, H. Li, J. Wu, W. Dong, and G. Shi, “Dynamic expert- knowledge ensemble for generalizable video quality assessment,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 6, pp. 2577–2589, 2022. II-C, II-C

  62. [62]

    Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,

    H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,” inProceedings of 14 the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20 144–20 154. II-C, II-C, III, V

  63. [63]

    Q-align: Teaching lmms for visual scoring via discrete text-defined levels,

    H. Wu, Z. Zhang, W. Zhang, C. Chen, L. Liao, C. Li, Y . Gao, A. Wang, E. Zhang, W. Sunet al., “Q-align: Teaching lmms for visual scoring via discrete text-defined levels,”arXiv preprint arXiv:2312.17090, 2023. II-C, II-C, III, V

  64. [64]

    Kvq: Kwai video quality assessment for short-form videos,

    Y . Lu, X. Li, Y . Pei, K. Yuan, Q. Xie, Y . Qu, M. Sun, C. Zhou, and Z. Chen, “Kvq: Kwai video quality assessment for short-form videos,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 25 963–25 973. II-C, II-C, II-C, III

  65. [65]

    Scaling and masking: A new paradigm of data sampling for image and video quality assess- ment,

    Y . Liu, Y . Quan, G. Xiao, A. Li, and J. Wu, “Scaling and masking: A new paradigm of data sampling for image and video quality assess- ment,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 4, 2024, pp. 3792–3801. II-C, II-C

  66. [66]

    Video quality assessment for online processing: From spatial to temporal sampling,

    J. Yan, L. Wu, Y . Fang, X. Liu, X. Xia, and W. Liu, “Video quality assessment for online processing: From spatial to temporal sampling,” IEEE Transactions on Circuits and Systems for Video Technology,

  67. [67]

    Pea265: Perceptual assessment of video compression artifacts,

    L. Lin, S. Yu, L. Zhou, W. Chen, T. Zhao, and Z. Wang, “Pea265: Perceptual assessment of video compression artifacts,”IEEE Transac- tions on Circuits and Systems for Video Technology, vol. 30, no. 11, pp. 3898–3910, 2020. II-C

  68. [68]

    Patch- vq:’patching up’the video quality problem,

    Z. Ying, M. Mandal, D. Ghadiyaram, and A. Bovik, “Patch- vq:’patching up’the video quality problem,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 14 019–14 029. II-C, V

  69. [69]

    Fvq: A large-scale dataset and an lmm-based method for face video quality assessment,

    S. Wu, Y . Li, Z. Xu, Y . Gao, H. Duan, W. Sun, and G. Zhai, “Fvq: A large-scale dataset and an lmm-based method for face video quality assessment,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 6928–6937. II-C

  70. [70]

    Rgc-vqa: An exploration database for robotic-generated video quality assessment,

    J. Jin, J. Ying, H. Duan, L. Yang, S. Wu, Y . Li, Y . Zheng, X. Min, and G. Zhai, “Rgc-vqa: An exploration database for robotic-generated video quality assessment,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 13 413–13 420. II-C

  71. [71]

    Ges-qa: A multidimensional quality assessment dataset for audio-to-3d gesture generation,

    Z. Gao, Y . Li, S. Wu, Y . Cao, H. Duan, and G. Zhai, “Ges-qa: A multidimensional quality assessment dataset for audio-to-3d gesture generation,”arXiv preprint arXiv:2508.12020, 2025. II-C

  72. [72]

    Sfqa: A comprehensive perceptual quality assessment dataset for singing face generation,

    Z. Gao, Y . Li, S. Wu, Y . Zhu, H. Duan, and G. Zhai, “Sfqa: A comprehensive perceptual quality assessment dataset for singing face generation,”arXiv preprint arXiv:2601.20385, 2026. II-C

  73. [73]

    Two-level approach for no-reference consumer video quality assessment,

    J. Korhonen, “Two-level approach for no-reference consumer video quality assessment,”IEEE Transactions on Image Processing, vol. 28, no. 12, pp. 5923–5938, 2019. II-C, V

  74. [74]

    Rapique: Rapid and accurate video quality prediction of user generated content,

    Z. Tu, X. Yu, Y . Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Rapique: Rapid and accurate video quality prediction of user generated content,”IEEE Open Journal of Signal Processing, vol. 2, pp. 425–440, 2021. II-C, V

  75. [75]

    Ugc- vqa: Benchmarking blind video quality assessment for user generated content,

    Z. Tu, Y . Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Ugc- vqa: Benchmarking blind video quality assessment for user generated content,”IEEE Transactions on Image Processing, vol. 30, pp. 4449– 4464, 2021. II-C, V

  76. [76]

    Learning gen- eralized spatial-temporal deep feature representation for no-reference video quality assessment,

    B. Chen, L. Zhu, G. Li, F. Lu, H. Fan, and S. Wang, “Learning gen- eralized spatial-temporal deep feature representation for no-reference video quality assessment,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 4, pp. 1903–1916, 2021. II-C, III

  77. [77]

    An end-to-end no-reference video quality assessment method with hierarchical spatiotemporal feature representation,

    W. Shen, M. Zhou, X. Liao, W. Jia, T. Xiang, B. Fang, and Z. Shang, “An end-to-end no-reference video quality assessment method with hierarchical spatiotemporal feature representation,”IEEE Transactions on Broadcasting, vol. 68, no. 3, pp. 651–660, 2022. II-C

  78. [78]

    Blind quality assessment of wide-angle videos based on deformation representation learning and multi-dimensional feature fusion,

    B. Hu, W. Wang, L. Li, L. He, W. Lu, and X. Gao, “Blind quality assessment of wide-angle videos based on deformation representation learning and multi-dimensional feature fusion,”IEEE Transactions on Circuits and Systems for Video Technology, 2025. II-C

  79. [79]

    Mi3s: a multimodal large language model assisted quality assessment framework for ai-generated talking heads,

    Y . Zhou, Z. Zhang, S. Wu, J. Jia, Y . Jiang, W. Sun, X. Liu, X. Min, and G. Zhai, “Mi3s: a multimodal large language model assisted quality assessment framework for ai-generated talking heads,”Information Processing & Management, vol. 63, no. 1, p. 104321, 2026. II-C

  80. [80]

    Objective quality assessment of ai- generated content videos with transformation consistency focus,

    X. Xu, L. Wu, J. Yan, and Y . Fang, “Objective quality assessment of ai- generated content videos with transformation consistency focus,”IEEE Transactions on Circuits and Systems for Video Technology, 2026. II-C

Showing first 80 references.