REVIEW 5 major objections 6 minor 120 references
This paper claims that a single mixture-of-experts model can rate, compare, and answer questions about AI-generated human-centric videos across three quality dimensions, and that the accompanying dataset is the largest of its kind.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 20:02 UTC pith:MHGXVTQA
load-bearing objection HVEval+ is a substantial dataset extension, but the paper under-validates the three-way dimension split and overstates the model's margin over fine-tuned baselines. the 5 major comments →
Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that HVEval+ is the largest holistic quality assessment dataset for AI-generated human-centric videos, comprising 1k prompts organized into 7 major categories, 20k videos from 24 text-to-video models, 60k MOSs and 60k preference pairs across three dimensions, and 20k category-specific Q&A pairs. The companion claim is that MoE-Rater, an MLLM-based all-in-one method, achieves superior performance on both HVEval+ and Human-AGVQA, with near-1.0 rank correlation to human judgments when ordering T2V models and strong zero-shot transfer to a related dataset.
What carries the argument
The carrying mechanism is the MoE-Rater routing architecture combined with HVEval+'s paired annotation design. S-MoPE and T-MoPE are mixtures of task-specific MLP projectors that route spatial and temporal features to top-k experts via gating networks; MoLE applies the same idea to LoRA experts inside the LLM. A shared set of task embeddings drives all gating networks, and a three-stage training strategy—task-aware pre-training, task-specific adaptation, and adaptive routing optimization—lets experts specialize and then recombine. On the dataset side, 60,000 same-prompt video pairs convert relative preferences into training signal.
Load-bearing premise
The load-bearing premise is that the 33 annotators actually rated spatial quality, temporal quality, and text-video correspondence as three separate, non-overlapping things—but the paper reports no inter-rater reliability or cross-dimension confusion analysis, so if annotators conflated the dimensions, the MOSs and preference labels become noisy and the benchmark's value drops.
What would settle it
Re-analyze the raw per-subject scores: compute per-dimension inter-rater agreement using a standard coefficient (e.g., alpha or kappa) and the correlation between the same subject's ratings across dimensions on the same videos. If within-dimension agreement is low (below 0.6) or between-dimension subject correlations rival within-dimension ones, the three-way separation is unsupported and the central benchmark claim collapses. A small dimension-classification test—asking new annotators to label which dimension a given distortion belongs to—would also settle it.
If this is right
- HVEval+ provides a large multi-dimensional benchmark; methods trained or fine-tuned on it substantially outperform generic video-quality models on AI-generated human-centric videos.
- A single MoE-Rater handles rating, comparison, and Q&A, so T2V model evaluation no longer requires separate pipelines for each task.
- MoE-Rater's predicted scores reproduce human T2V model rankings with near-1.0 correlation, making automated model comparison viable.
- Zero-shot MoE-Rater beats most fine-tuned methods on Human-AGVQA, suggesting the dataset teaches transferable quality assessment skills.
- The 60k preference pairs provide training signal for reward or judgment models that could be used to optimize T2V generators.
Where Pith is reading between the lines
- The pairwise preference structure could be repurposed as a reward model for optimizing T2V generation, since same-prompt preference pairs are exactly the input reward models are trained on; the paper notes this potential but does not train such a model.
- If the three dimensions are truly separable in the annotations, the dataset becomes a diagnostic tool: a video with high spatial but low temporal scores points to motion smoothness as the failure mode, which could guide model-specific improvements.
- A testable extension is to measure inter-rater agreement per dimension and cross-dimension confusion; the paper does not report these, so dimension cleanliness remains an open empirical question.
- The three-stage expert-routing recipe may generalize to other multi-task MLLM settings beyond video quality, but that is speculation beyond the paper's scope.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents HVEval+, an extension of the authors' earlier HVEval dataset for quality assessment of AI-generated human-centric videos. HVEval+ contains 1,000 prompts across a 7-category taxonomy, 20,000 videos from 24 T2V models, 60,000 MOSs, 60,000 pairwise preference labels across three dimensions (spatial quality, temporal quality, text-video correspondence), and 20,000 category-specific Q&A pairs. The paper also proposes MoE-Rater, an MLLM-based all-in-one model with Mixture of Projector Experts (S-MoPE/T-MoPE) and Mixture of LoRA Experts (MoLE), trained in three stages to unify rating, pairwise comparison, and question answering. Experiments on HVEval+ and zero-shot/fine-tuned evaluations on Human-AGVQA are reported, with MoE-Rater achieving the best results on most metrics.
Significance. If the annotation dimensions are reliable, HVEval+ would be a valuable large-scale multi-dimensional benchmark for AI-generated human-centric video quality assessment, and MoE-Rater would be a useful unified framework for rating, comparison, and QA. The prompt-based train/test split avoids content leakage, and the cross-dataset zero-shot evaluation on Human-AGVQA is a genuine strength that provides independent grounding. The subjective protocol follows standard outlier-rejection practice. However, the load-bearing assumption that annotators can reliably separate the three quality dimensions is not validated by any inter-rater reliability or agreement analysis, which makes the dataset's multi-dimensional claims and the model's dimension-specific gains conditional on an untested premise.
major comments (5)
- [§III-C/III-D] The load-bearing assumption of the dataset is that 33 annotators can reliably separate spatial quality, temporal quality, and text-video correspondence. The subjective-study description reports only a 15-item screening test, per-batch outlier rejection, and ITU-R BT.500 outlier removal; no inter-rater reliability (e.g., ICC, Krippendorff's α), no per-dimension agreement, and no cross-dimension confusion analysis are reported. Section VI acknowledges annotation uncertainty but does not quantify it. If annotators conflated dimensions, the 60k MOSs and preference labels are noisy, and the multi-dimensional gains of MoE-Rater in Tables III/IV/VII could reflect fitting to shared variance rather than true dimensional understanding. Please add per-dimension agreement statistics, a dimension-separability analysis (e.g., correlation or confusion matrix of per-video dimension MOSs), and a statemen
- [§III-B/III-D] The pairwise preference labels and MOSs were collected from the same videos, yet the paper never checks whether the two protocols are internally consistent: for a given pair and dimension, does the video with the higher MOS also win the majority preference vote? Without this consistency check, both annotation layers remain unvalidated. Please report per-dimension agreement between MOS differences and preference outcomes, and specify the number of ratings per video/pair and the majority-vote rule.
- [§V-A/Table III] The 'fine-tuned on HVEval+' baselines (†) are used to support the claim of superiority, but no fine-tuning protocol is described for them. Differences in training data, number of steps, loss, and task mixture could explain part of the gap (e.g., spatial SRCC 0.891 vs 0.860 against Qwen2.5-VL†). Please specify the exact protocol and, ideally, use the same three-stage recipe for baselines or state why that is not feasible.
- [§V-B/Table VI] Table VI reports 'SRCC to Human' values approaching 0.99, but the text does not state whether the per-model average scores are computed on the training split, test split, or the full dataset. Because MoE-Rater is trained on 80% of HVEval+ videos, ranking correlations on training data would be inflated by design. Please recompute and report ranking correlations on the prompt-disjoint test split only.
- [§V-B/Tables III-V] No confidence intervals or significance tests are reported for SRCC/PLCC/accuracy. Several margins are modest (e.g., correspondence SRCC 0.787 vs 0.755 for InternVL2.5†), so some 'superior' claims may be within noise. Add bootstrap confidence intervals or paired significance tests for the key comparisons.
minor comments (6)
- [Table I] HVEval+ is described as an extension of HVEval, but the table reports 80k MOS for HVEval and 60k MOS for HVEval+; clarify the relation and annotation counts.
- [§IV-B] The equations define features for one video; it is unclear how two input videos are processed for pairwise comparison (concatenation, separate placeholders, fusion).
- [§III-C] The pass threshold for the 15-item training test and the number of rejected subjects are not reported.
- [General] No dataset availability statement or download link is provided; for a dataset paper this should be added.
- [Figures 4/5] Figure 4 caption reports only spatial and text-video correspondence scores while Figure 5 includes temporal; make the reported dimensions consistent across figures.
- [Table VII] The row configuration of the ablation table is hard to read; mark each component explicitly for every row.
Circularity Check
No significant circularity: HVEval+ and MoE-Rater are standard supervised benchmarking with an independent cross-dataset zero-shot evaluation.
full rationale
The paper's central claims are that HVEval+ is a large multi-dimensional dataset and that MoE-Rater performs well on it and on Human-AGVQA. Training and evaluating a model on the same benchmark is standard supervised practice, not circular, and the paper compares against other methods trained/fine-tuned on the same HVEval+ data. The cross-dataset zero-shot evaluation on Human-AGVQA (Table V) provides an external check: MoE-Rater is trained only on HVEval+ and evaluated on Human-AGVQA without fine-tuning, so the reported transfer performance is not a re-statement of its training targets. The self-citation of HVEval [11] is transparent provenance for the dataset extension and is not used to justify the model's superiority or to exclude alternatives; no uniqueness theorem or load-bearing argument is imported from prior work by the same authors. The three annotation dimensions are defined by independent human rating instructions, not in terms of MoE-Rater outputs, and no equation equates a predicted quantity to a fitted parameter. The limitation section's acknowledgment of annotation uncertainty is a validity concern, not circularity. Overall, no load-bearing step reduces by construction to its own inputs, so the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- Number of experts N =
7
- Top-k routing k =
3
- LoRA rank =
8
- Task embedding dimension =
256
axioms (4)
- domain assumption The three quality dimensions (spatial, temporal, correspondence) are well-defined and consistently interpreted by annotators.
- domain assumption The prompt taxonomy (7 major categories, 22 subcategories) comprehensively covers human-centric T2V scenarios.
- domain assumption ITU-R BT.500 outlier rejection is valid for this annotation protocol.
- domain assumption Pairwise preferences and majority voting yield reliable ground truth labels.
invented entities (2)
-
Mixture of Projector Experts (MoPE)
no independent evidence
-
Mixture of LoRA Experts (MoLE)
no independent evidence
read the original abstract
AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring the importance of effective quality assessment for such videos. To this end, we extend our previous dataset HVEval with pairwise preference annotations, resulting in HVEval+, the largest holistic quality assessment dataset for AI-generated human-centric videos, which comprises 1k prompts based on a comprehensive taxonomy, 20k videos generated by 24 text-to-video (T2V) models, and extensive human annotations, including 60k mean opinion scores (MOSs) and 60k preference pairs across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), as well as 20k category-specific question-answer (Q&A) pairs. Along with the HVEval+ dataset, we further propose MoE-Rater, a Mixture-of-Experts (MoE)-inspired and multimodal large language model (MLLM)-based all-in-one method that supports multi-dimensional quality rating, multi-dimensional pairwise comparison, and category-specific question answering within a single model. Specifically, we introduce Mixture of Projector Experts (MoPE) and Mixture of LoRA Experts (MoLE), together with a three-stage training strategy consisting of task-aware pre-training, task-specific adaptation, and adaptive routing optimization, to effectively unify multiple tasks, resulting in superior performance on both HVEval+ and Human-AGVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the HVEval+ dataset and the MoE-Rater method in advancing AI-generated video quality assessment and further facilitating the evaluation and optimization of T2V models.
Figures
Reference graph
Works this paper leans on
-
[1]
Hunyuanvideo: A systematic framework for large video generative models,
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhanget al., “Hunyuanvideo: A systematic framework for large video generative models,”arXiv preprint arXiv:2412.03603,
-
[2]
Wan: Open and advanced large-scale video generative models,
T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yanget al., “Wan: Open and advanced large-scale video generative models,”arXiv preprint arXiv:2503.20314, 2025. I, II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2025
-
[3]
Magi-1: Autoregressive video generation at scale,
H. Teng, H. Jia, L. Sun, L. Li, M. Li, M. Tang, S. Han, T. Zhang, W. Zhang, W. Luoet al., “Magi-1: Autoregressive video generation at scale,”arXiv preprint arXiv:2505.13211, 2025. I, II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2025
-
[4]
Team, “Sora,” https://openai.com/sora/, 2025
O. Team, “Sora,” https://openai.com/sora/, 2025. I, II, III-B, VI
2025
-
[5]
Pixverse,
PixVerse, “Pixverse,” https://app.pixverse.ai, 2025. I, II, III-B, VI
2025
-
[6]
I, II, III-B, VI
Pika, “Pika,” https://www.pika.art, 2025. I, II, III-B, VI
2025
-
[7]
Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation,
H. Li, M. Xu, Y . Zhan, S. Mu, J. Li, K. Cheng, Y . Chen, T. Chen, M. Ye, J. Wanget al., “Openhumanvid: A large-scale high-quality dataset for enhancing human-centric video generation,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 7752–7762. I
2025
-
[8]
Humanomni: A large vision-speech language model for human-centric video understanding,
J. Zhao, Q. Yang, Y . Peng, D. Bai, S. Yao, B. Sun, X. Chen, S. Fu, X. Wei, L. Boet al., “Humanomni: A large vision-speech language model for human-centric video understanding,”arXiv preprint arXiv:2501.15111, 2025. I
Pith/arXiv arXiv 2025
-
[9]
Singinghead: A large-scale 4d dataset for singing head animation,
S. Wu, Y . Li, W. Zhang, J. Jia, Y . Zhu, Y . Yan, G. Zhai, and X. Yang, “Singinghead: A large-scale 4d dataset for singing head animation,” IEEE Transactions on Multimedia, 2025. I
2025
-
[10]
Human-activity agv quality assessment: A benchmark dataset and an objective evaluation metric,
Z. Zhang, W. Sun, X. Li, Y . Li, Q. Ge, J. Jia, Z. Zhang, Z. Ji, F. Sun, S. Juiet al., “Human-activity agv quality assessment: A benchmark dataset and an objective evaluation metric,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 6771–6780. I, I, II-B, II-C, V, V-A, V-A, V-B, V-B, V-B
2025
-
[11]
Hveval: Towards unified evaluation of human-centric video generation and understand- ing,
S. Wu, Y . Li, H. Duan, Y . Jiang, Y . Zhu, and G. Zhai, “Hveval: Towards unified evaluation of human-centric video generation and understand- ing,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 13 376–13 383. I, I, II-C
2025
-
[12]
Z. Wang, Q. Ma, W. Wan, H. Li, K. Wang, and Y . Tian, “Is this generated person existed in real-world? fine-grained detecting and calibrating abnormal human-body,”arXiv preprint arXiv:2411.14205,
-
[13]
Improved techniques for training gans,
T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Radford, and X. Chen, “Improved techniques for training gans,”Advances in neural information processing systems, vol. 29, 2016. I, II-B, II-B
2016
-
[14]
Gans trained by a two time-scale update rule converge to a local nash equilibrium,
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,”Advances in neural information processing systems, vol. 30, 2017. I, II-B, II-B
2017
-
[15]
Towards accurate generative models of video: A new metric & challenges,
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michal- ski, and S. Gelly, “Towards accurate generative models of video: A new metric & challenges,”arXiv preprint arXiv:1812.01717, 2018. I, II-B, II-B
Pith/arXiv arXiv 2018
-
[16]
Vbench: Comprehensive benchmark suite for video generative models,
Z. Huang, Y . He, J. Yu, F. Zhang, C. Si, Y . Jiang, Y . Zhang, T. Wu, Q. Jin, N. Chanpaisitet al., “Vbench: Comprehensive benchmark suite for video generative models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 807–21 818. I, I, II-B, II-B, III, V
2024
-
[17]
Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness,
D. Zheng, Z. Huang, H. Liu, K. Zou, Y . He, F. Zhang, Y . Zhang, J. He, W.-S. Zheng, Y . Qiaoet al., “Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness,”arXiv preprint arXiv:2503.21755, 2025. I, I, II-B, II-B
Pith/arXiv arXiv 2025
-
[18]
T2v- compbench: A comprehensive benchmark for compositional text-to- video generation,
K. Sun, K. Huang, X. Liu, Y . Wu, Z. Xu, Z. Li, and X. Liu, “T2v- compbench: A comprehensive benchmark for compositional text-to- video generation,”arXiv preprint arXiv:2407.14505, 2024. I, I, II-B, II-B
Pith/arXiv arXiv 2024
-
[19]
Subjective-aligned dataset and metric for text-to-video quality assessment,
T. Kou, X. Liu, Z. Zhang, C. Li, H. Wu, X. Min, G. Zhai, and N. Liu, “Subjective-aligned dataset and metric for text-to-video quality assessment,” inProceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 7793–7802. I, I, II-B, II-B, II-C, V
2024
-
[20]
Benchmarking aigc video quality assessment: A dataset and unified model,
Z. Zhang, X. Li, W. Sun, J. Jia, X. Min, Z. Zhang, C. Li, Z. Chen, P. Wang, Z. Jiet al., “Benchmarking aigc video quality assessment: A dataset and unified model,”arXiv preprint arXiv:2407.21408, 2024. I, I, II-B, II-B, II-C, V
Pith/arXiv arXiv 2024
-
[21]
Aigv-assessor: benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,
J. Wang, H. Duan, G. Zhai, J. Wang, and X. Min, “Aigv-assessor: benchmarking and evaluating the perceptual quality of text-to-video generation with lmm,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 18 869–18 880. I, II-B, II-B, II-C
2025
-
[22]
Evalcrafter: Benchmarking and evaluating large video generation models,
Y . Liu, X. Cun, X. Liu, X. Wang, Y . Zhang, H. Chen, Y . Liu, T. Zeng, R. Chan, and Y . Shan, “Evalcrafter: Benchmarking and evaluating large video generation models,” inProceedings of the IEEE/CVF Conference 13 on Computer Vision and Pattern Recognition, 2024, pp. 22 139–22 149. I, V
2024
-
[23]
Fetv: A benchmark for fine-grained evaluation of open-domain text-to- video generation,
Y . Liu, L. Li, S. Ren, R. Gao, S. Li, S. Chen, X. Sun, and L. Hou, “Fetv: A benchmark for fine-grained evaluation of open-domain text-to- video generation,”Advances in Neural Information Processing Systems, vol. 36, pp. 62 352–62 387, 2023. I
2023
-
[24]
Evaluating and improving compositional text-to-visual generation,
B. Li, Z. Lin, D. Pathak, J. Li, Y . Fei, K. Wu, X. Xia, P. Zhang, G. Neubig, and D. Ramanan, “Evaluating and improving compositional text-to-visual generation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 5290–5301. I
2024
-
[25]
F-bench: Rethinking human preference evaluation metrics for benchmarking face generation, customization, and restoration,
L. Liu, H. Duan, Q. Hu, L. Yang, C. Cai, T. Ye, H. Liu, X. Zhang, and G. Zhai, “F-bench: Rethinking human preference evaluation metrics for benchmarking face generation, customization, and restoration,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 10 982–10 994. I
2025
-
[26]
Multi- dimensional text-to-face image quality assessment using llm: Database and method,
Y . Gao, X. Min, J. Han, Y . Cao, S. Wu, Y . Dou, and G. Zhai, “Multi- dimensional text-to-face image quality assessment using llm: Database and method,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 6948–6957. I
2025
-
[27]
Aghi-qa: A subjective-aligned dataset and metric for ai- generated human images,
Y . Li, S. Wu, W. Sun, Z. Zhang, Y . Zhu, Z. Zhang, H. Duan, X. Min, and G. Zhai, “Aghi-qa: A subjective-aligned dataset and metric for ai- generated human images,”IEEE Transactions on Circuits and Systems for Video Technology, 2025. I
2025
-
[28]
Dreambench++: A human- aligned benchmark for personalized image generation,
Y . Peng, Y . Cui, H. Tang, Z. Qi, R. Dong, J. Bai, C. Han, Z. Ge, X. Zhang, and S.-T. Xia, “Dreambench++: A human- aligned benchmark for personalized image generation,”arXiv preprint arXiv:2406.16855, 2024. I, III-A, III-A
Pith/arXiv arXiv 2024
-
[29]
Chatgpt-4o,
OpenAI, “Chatgpt-4o,” https://openai.com, 2025. I, III-A, III-A
2025
-
[30]
Animatediff-lightning: Cross-model diffusion distillation,
S. Lin and X. Yang, “Animatediff-lightning: Cross-model diffusion distillation,”arXiv preprint arXiv:2403.12706, 2024. II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2024
-
[31]
Animatelcm: Computation-efficient personalized style video generation without personalized video data,
F.-Y . Wang, Z. Huang, W. Bian, X. Shi, K. Sun, G. Song, Y . Liu, and H. Li, “Animatelcm: Computation-efficient personalized style video generation without personalized video data,” inSIGGRAPH Asia 2024 Technical Communications, 2024, pp. 1–5. II-A, II-A, II, III-B, VI
2024
-
[32]
Magictime: Time-lapse video generation models as metamorphic simulators,
S. Yuan, J. Huang, Y . Shi, Y . Xu, R. Zhu, B. Lin, X. Cheng, L. Yuan, and J. Luo, “Magictime: Time-lapse video generation models as metamorphic simulators,”IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. II-A, II-A, II, III-B, VI
2025
-
[33]
Modelscope text-to-video technical report,
J. Wang, H. Yuan, D. Chen, Y . Zhang, X. Wang, and S. Zhang, “Modelscope text-to-video technical report,”arXiv preprint arXiv:2308.06571, 2023. II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2023
-
[34]
Show-1: Marrying pixel and latent diffusion models for text-to-video generation,
D. J. Zhang, J. Z. Wu, J.-W. Liu, R. Zhao, L. Ran, Y . Gu, D. Gao, and M. Z. Shou, “Show-1: Marrying pixel and latent diffusion models for text-to-video generation,”International Journal of Computer Vision, pp. 1–15, 2024. II-A, II-A, II, III-B, VI
2024
-
[35]
J. Li, Q. Long, J. Zheng, X. Gao, R. Piramuthu, W. Chen, and W. Y . Wang, “T2v-turbo-v2: Enhancing video generation model post-training through data, reward, and conditional guidance design,”arXiv preprint arXiv:2410.05677, 2024. II-A, II-A, II, III-B, VI
arXiv 2024
-
[36]
Videocrafter2: Overcoming data limitations for high-quality video diffusion models,
H. Chen, Y . Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y . Shan, “Videocrafter2: Overcoming data limitations for high-quality video diffusion models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 7310–7320. II-A, II-A, II, III-B, VI
2024
-
[37]
Zeroscope,
S. Sterling, “Zeroscope,” https://huggingface.co/cerspense/zeroscope v2 576w, 2023. II-A, II-A, II, III-B, VI
2023
-
[38]
Cogvideox: Text-to- video diffusion models with an expert transformer,
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y . Yang, W. Hong, X. Zhang, G. Fenget al., “Cogvideox: Text-to- video diffusion models with an expert transformer,”arXiv preprint arXiv:2408.06072, 2024. II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2024
-
[39]
Ltx-video: Realtime video latent diffusion,
Y . HaCohen, N. Chiprut, B. Brazowski, D. Shalem, D. Moshe, E. Richardson, E. Levin, G. Shiran, N. Zabari, O. Gordon et al., “Ltx-video: Realtime video latent diffusion,”arXiv preprint arXiv:2501.00103, 2024. II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2024
-
[40]
Latte: Latent diffusion transformer for video generation,
X. Ma, Y . Wang, G. Jia, X. Chen, Z. Liu, Y .-F. Li, C. Chen, and Y . Qiao, “Latte: Latent diffusion transformer for video generation,” arXiv preprint arXiv:2401.03048, 2024. II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2024
-
[41]
Step-video-t2v technical report: The practice, challenges, and future of video foundation model,
G. Ma, H. Huang, K. Yan, L. Chen, N. Duan, S. Yin, C. Wan, R. Ming, X. Song, X. Chenet al., “Step-video-t2v technical report: The practice, challenges, and future of video foundation model,”arXiv preprint arXiv:2502.10248, 2025. II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2025
-
[42]
Mochi 1,
G. Team, “Mochi 1,” https://github.com/genmoai/models, 2024. II-A, II-A, II, III-B, VI
2024
-
[43]
Skyreels-v2: Infinite-length film generative model,
G. Chen, D. Lin, J. Yang, C. Lin, J. Zhu, M. Fan, H. Zhang, S. Chen, Z. Chen, C. Maet al., “Skyreels-v2: Infinite-length film generative model,”arXiv preprint arXiv:2504.13074, 2025. II-A, II-A, II, III-B, VI
Pith/arXiv arXiv 2025
-
[44]
Cofnet: contrastive object-aware fusion using box-level masks for multispectral object detection,
M. Zhou, Y . Li, G. Yang, X. Wei, H. Pu, J. Luo, and W. Jia, “Cofnet: contrastive object-aware fusion using box-level masks for multispectral object detection,”IEEE Transactions on Multimedia, 2025. II-B
2025
-
[45]
Afes: Attention-based feature excitation and sorting for action recognition,
M. Zhou, J. Li, X. Wei, J. Luo, H. Pu, W. Wang, J. He, and Z. Shang, “Afes: Attention-based feature excitation and sorting for action recognition,”IEEE Transactions on Consumer Electronics, 2025. II-B
2025
-
[46]
An end-to-end blind image quality assessment method using a recurrent network and self-attention,
M. Zhou, X. Lan, X. Wei, X. Liao, Q. Mao, Y . Li, C. Wu, T. Xiang, and B. Fang, “An end-to-end blind image quality assessment method using a recurrent network and self-attention,”IEEE Transactions on Broadcasting, vol. 69, no. 2, pp. 369–377, 2022. II-C
2022
-
[47]
Attentional feature fusion for end-to-end blind image quality assessment,
M. Zhou, S. Lang, T. Zhang, X. Liao, Z. Shang, T. Xiang, and B. Fang, “Attentional feature fusion for end-to-end blind image quality assessment,”IEEE Transactions on Broadcasting, vol. 69, no. 1, pp. 144–152, 2022. II-C
2022
-
[48]
Multilevel feature fusion for end-to-end blind image quality assessment,
X. Lan, M. Zhou, X. Xu, X. Wei, X. Liao, H. Pu, J. Luo, T. Xiang, B. Fang, and Z. Shang, “Multilevel feature fusion for end-to-end blind image quality assessment,”IEEE Transactions on Broadcasting, vol. 69, no. 3, pp. 801–811, 2023. II-C
2023
-
[49]
Blind image quality assessment based on perceptual comparison,
A. Li, J. Wu, Y . Liu, L. Li, W. Dong, and G. Shi, “Blind image quality assessment based on perceptual comparison,”IEEE Transactions on Multimedia, vol. 26, pp. 9671–9682, 2024. II-C
2024
-
[50]
Graph-represented distribution similarity index for full-reference image quality assess- ment,
W. Shen, M. Zhou, J. Luo, Z. Li, and S. Kwong, “Graph-represented distribution similarity index for full-reference image quality assess- ment,”IEEE Transactions on Image Processing, vol. 33, pp. 3075– 3089, 2024. II-C
2024
-
[51]
Bridging the synthetic-to-authentic gap: Distortion-guided unsupervised domain adaptation for blind image quality assessment,
A. Li, J. Wu, Y . Liu, and L. Li, “Bridging the synthetic-to-authentic gap: Distortion-guided unsupervised domain adaptation for blind image quality assessment,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 28 422–28 431. II-C
2024
-
[52]
Blind image quality assessment: Exploring content fidelity perceptibility via quality adversarial learning,
M. Zhou, W. Shen, X. Wei, J. Luo, F. Jia, X. Zhuang, and W. Jia, “Blind image quality assessment: Exploring content fidelity perceptibility via quality adversarial learning,”International Journal of Computer Vision, vol. 133, no. 6, pp. 3242–3258, 2025. II-C
2025
-
[53]
Image quality assessment: Investigating causal perceptual effects with abductive counterfactual inference,
W. Shen, M. Zhou, Y . Chen, X. Wei, Y . Feng, H. Pu, and W. Jia, “Image quality assessment: Investigating causal perceptual effects with abductive counterfactual inference,” inProceedings of the computer vision and pattern recognition conference, 2025, pp. 17 990–17 999. II-C
2025
-
[54]
No-reference image quality assessment: Exploring intrinsic distortion characteristics via generative noise estimation with mamba,
X. Lan, W. Xian, M. Zhou, J. Yan, X. Wei, J. Luo, W. Jia, and S. Kwong, “No-reference image quality assessment: Exploring intrinsic distortion characteristics via generative noise estimation with mamba,” IEEE Transactions on Circuits and Systems for Video Technology,
-
[55]
Towards syn-to-real iqa: A novel perspective on reshaping synthetic data distributions,
A. Li, J. Wu, Y . Liu, L. Li, and W. Dong, “Towards syn-to-real iqa: A novel perspective on reshaping synthetic data distributions,”arXiv preprint arXiv:2601.00225, 2026. II-C
arXiv 2026
-
[56]
Uni-iqa: A unified approach for mutual promotion of natural and screen content image quality assessment,
M. Song, C. Chen, W. Song, and Y . Fang, “Uni-iqa: A unified approach for mutual promotion of natural and screen content image quality assessment,”IEEE Transactions on Circuits and Systems for Video Technology, 2025. II-C
2025
-
[57]
Beyond pixels: text-guided deep insights into graphic design image aesthetics,
G. Shi, L. Li, and M. Song, “Beyond pixels: text-guided deep insights into graphic design image aesthetics,”Journal of Electronic Imaging, vol. 33, no. 5, pp. 053 059–053 059, 2024. II-C
2024
-
[58]
A deep learning based no- reference quality assessment model for ugc videos,
W. Sun, X. Min, W. Lu, and G. Zhai, “A deep learning based no- reference quality assessment model for ugc videos,” inProceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 856–865. II-C, II-C, III, V
2022
-
[59]
Analysis of video quality datasets via design of minimalistic video quality models,
W. Sun, W. Wen, X. Min, L. Lan, G. Zhai, and K. Ma, “Analysis of video quality datasets via design of minimalistic video quality models,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. II-C, II-C, III
2024
-
[60]
Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,
H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,” inEuropean conference on computer vision. Springer, 2022, pp. 538–554. II-C, II-C, III, V
2022
-
[61]
Dynamic expert- knowledge ensemble for generalizable video quality assessment,
P. Chen, L. Li, H. Li, J. Wu, W. Dong, and G. Shi, “Dynamic expert- knowledge ensemble for generalizable video quality assessment,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 6, pp. 2577–2589, 2022. II-C, II-C
2022
-
[62]
Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,
H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,” inProceedings of 14 the IEEE/CVF International Conference on Computer Vision, 2023, pp. 20 144–20 154. II-C, II-C, III, V
2023
-
[63]
Q-align: Teaching lmms for visual scoring via discrete text-defined levels,
H. Wu, Z. Zhang, W. Zhang, C. Chen, L. Liao, C. Li, Y . Gao, A. Wang, E. Zhang, W. Sunet al., “Q-align: Teaching lmms for visual scoring via discrete text-defined levels,”arXiv preprint arXiv:2312.17090, 2023. II-C, II-C, III, V
Pith/arXiv arXiv 2023
-
[64]
Kvq: Kwai video quality assessment for short-form videos,
Y . Lu, X. Li, Y . Pei, K. Yuan, Q. Xie, Y . Qu, M. Sun, C. Zhou, and Z. Chen, “Kvq: Kwai video quality assessment for short-form videos,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 25 963–25 973. II-C, II-C, II-C, III
2024
-
[65]
Scaling and masking: A new paradigm of data sampling for image and video quality assess- ment,
Y . Liu, Y . Quan, G. Xiao, A. Li, and J. Wu, “Scaling and masking: A new paradigm of data sampling for image and video quality assess- ment,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 4, 2024, pp. 3792–3801. II-C, II-C
2024
-
[66]
Video quality assessment for online processing: From spatial to temporal sampling,
J. Yan, L. Wu, Y . Fang, X. Liu, X. Xia, and W. Liu, “Video quality assessment for online processing: From spatial to temporal sampling,” IEEE Transactions on Circuits and Systems for Video Technology,
-
[67]
Pea265: Perceptual assessment of video compression artifacts,
L. Lin, S. Yu, L. Zhou, W. Chen, T. Zhao, and Z. Wang, “Pea265: Perceptual assessment of video compression artifacts,”IEEE Transac- tions on Circuits and Systems for Video Technology, vol. 30, no. 11, pp. 3898–3910, 2020. II-C
2020
-
[68]
Patch- vq:’patching up’the video quality problem,
Z. Ying, M. Mandal, D. Ghadiyaram, and A. Bovik, “Patch- vq:’patching up’the video quality problem,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 14 019–14 029. II-C, V
2021
-
[69]
Fvq: A large-scale dataset and an lmm-based method for face video quality assessment,
S. Wu, Y . Li, Z. Xu, Y . Gao, H. Duan, W. Sun, and G. Zhai, “Fvq: A large-scale dataset and an lmm-based method for face video quality assessment,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 6928–6937. II-C
2025
-
[70]
Rgc-vqa: An exploration database for robotic-generated video quality assessment,
J. Jin, J. Ying, H. Duan, L. Yang, S. Wu, Y . Li, Y . Zheng, X. Min, and G. Zhai, “Rgc-vqa: An exploration database for robotic-generated video quality assessment,” inProceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 13 413–13 420. II-C
2025
-
[71]
Ges-qa: A multidimensional quality assessment dataset for audio-to-3d gesture generation,
Z. Gao, Y . Li, S. Wu, Y . Cao, H. Duan, and G. Zhai, “Ges-qa: A multidimensional quality assessment dataset for audio-to-3d gesture generation,”arXiv preprint arXiv:2508.12020, 2025. II-C
arXiv 2025
-
[72]
Sfqa: A comprehensive perceptual quality assessment dataset for singing face generation,
Z. Gao, Y . Li, S. Wu, Y . Zhu, H. Duan, and G. Zhai, “Sfqa: A comprehensive perceptual quality assessment dataset for singing face generation,”arXiv preprint arXiv:2601.20385, 2026. II-C
arXiv 2026
-
[73]
Two-level approach for no-reference consumer video quality assessment,
J. Korhonen, “Two-level approach for no-reference consumer video quality assessment,”IEEE Transactions on Image Processing, vol. 28, no. 12, pp. 5923–5938, 2019. II-C, V
2019
-
[74]
Rapique: Rapid and accurate video quality prediction of user generated content,
Z. Tu, X. Yu, Y . Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Rapique: Rapid and accurate video quality prediction of user generated content,”IEEE Open Journal of Signal Processing, vol. 2, pp. 425–440, 2021. II-C, V
2021
-
[75]
Ugc- vqa: Benchmarking blind video quality assessment for user generated content,
Z. Tu, Y . Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Ugc- vqa: Benchmarking blind video quality assessment for user generated content,”IEEE Transactions on Image Processing, vol. 30, pp. 4449– 4464, 2021. II-C, V
2021
-
[76]
Learning gen- eralized spatial-temporal deep feature representation for no-reference video quality assessment,
B. Chen, L. Zhu, G. Li, F. Lu, H. Fan, and S. Wang, “Learning gen- eralized spatial-temporal deep feature representation for no-reference video quality assessment,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 4, pp. 1903–1916, 2021. II-C, III
1903
-
[77]
An end-to-end no-reference video quality assessment method with hierarchical spatiotemporal feature representation,
W. Shen, M. Zhou, X. Liao, W. Jia, T. Xiang, B. Fang, and Z. Shang, “An end-to-end no-reference video quality assessment method with hierarchical spatiotemporal feature representation,”IEEE Transactions on Broadcasting, vol. 68, no. 3, pp. 651–660, 2022. II-C
2022
-
[78]
Blind quality assessment of wide-angle videos based on deformation representation learning and multi-dimensional feature fusion,
B. Hu, W. Wang, L. Li, L. He, W. Lu, and X. Gao, “Blind quality assessment of wide-angle videos based on deformation representation learning and multi-dimensional feature fusion,”IEEE Transactions on Circuits and Systems for Video Technology, 2025. II-C
2025
-
[79]
Mi3s: a multimodal large language model assisted quality assessment framework for ai-generated talking heads,
Y . Zhou, Z. Zhang, S. Wu, J. Jia, Y . Jiang, W. Sun, X. Liu, X. Min, and G. Zhai, “Mi3s: a multimodal large language model assisted quality assessment framework for ai-generated talking heads,”Information Processing & Management, vol. 63, no. 1, p. 104321, 2026. II-C
2026
-
[80]
Objective quality assessment of ai- generated content videos with transformation consistency focus,
X. Xu, L. Wu, J. Yan, and Y . Fang, “Objective quality assessment of ai- generated content videos with transformation consistency focus,”IEEE Transactions on Circuits and Systems for Video Technology, 2026. II-C
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.