REVIEW 4 major objections 6 minor 34 references
Leading video generators still fall far short of film-grade craft when judged on professional cinematic language rather than web-style plausibility.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 20:09 UTC pith:I7GYHDQZ
load-bearing objection Solid film-grade T2V/R2V infrastructure: real multi-shot prompts, a usable taxonomy, and clear field gaps—use the diagnostics more than the tight Overall margins. the 4 major comments →
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When text-to-video and reference-to-video models are scored on academy-aligned cinematic criteria using prompts reverse-engineered from real award-winning films, overall scores stay well below saturation, an automatic expert-grade evaluator reproduces human model rankings at Spearman 0.95–0.96, and the decisive gaps concentrate in dynamic aesthetics and multi-shot cinematic language rather than static image quality.
What carries the argument
FilmBench’s three-level Cinematic Language taxonomy (3 axes, 12 components, 35+3 sub-metrics) paired with FilmOps, an open suite of specialized operators that map video into professional labels for shot scale, camera movement, composition, and related craft dimensions that feed the automatic scorer.
Load-bearing premise
That matching human model rankings on roughly thirty percent of the prompts, with every fine-grained sub-metric weighted equally, is enough to treat the automatic judge as a faithful proxy for professional film craft—even where agreement is weaker on editing and audio.
What would settle it
Have independent film professionals score the same generated videos on the full taxonomy; if their model ranking diverges sharply from the automatic leaderboard, or if top models reach near-ceiling scores under fixed prompts and rubrics, the claim that current generators remain far from film-grade fails.
If this is right
- Web-style benchmarks that already look saturated will not predict which models can execute a professional multi-shot shot list.
- Multi-shot staging, motivated camera moves, and performance become first-class evaluation and training targets rather than side effects of longer clips.
- Raw reference fidelity and overall cinematic quality are separable skills; leading one need not mean leading the other.
- Fine-grained sub-metric championships expose complementary model strengths that a single overall score conceals.
- Open cinematic operators let others build custom film-grade judges without relying only on generic multimodal scorers.
Where Pith is reading between the lines
- Training that only maximizes generic aesthetics and caption alignment is likely to leave the multi-shot and performance gaps largely untouched.
- A natural next stress test is end-to-end storyboard-to-sequence evaluation against a full director shot list, not only reverse-engineered clip prompts.
- Field-wide collapse on prop-reference fidelity points to object-level identity binding as a distinct unsolved piece of controllable generation.
- Because weaker models lose far more points on multi-shot prompts, scaling single-take quality alone will not close the craft gap.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FilmBench, a T2V/R2V benchmark for cinematic video generation. Prompts are reverse-engineered from award-winning film clips selected by professional directors (515 T2V + 654 R2V prompts, 1,056 of 1,169 multi-shot, spanning 20 genres and Chinese/International markets), scored on a three-level taxonomy (3 L1 axes, 12 L2 components, 35+3 L3 sub-metrics) co-designed with Beijing Film Academy faculty and a film studio, by an in-house automatic evaluation agent built on an open-sourced Cinematic Language operator suite (FilmOps). Validated against 90 professional raters on ~300 prompts, the agent's model-level ranking matches human rankings at Spearman ρ = 0.95 (T2V) / 0.96 (R2V). Benchmarking 9 T2V and 7 R2V models, the authors report no saturation (tops of 88.93/86.66), a field-wide dynamic-aesthetics bottleneck, a 7.9-point average single→multi-shot drop (up to 22.8 for the weakest model), stable rankings across markets/genres/reference types, and a "championship mismatch" showing no model leads all sub-metrics.
Significance. If the validation holds up, FilmBench would be a genuinely useful addition to video-generation evaluation: it supplies (i) a prompt set anchored to verified professional footage rather than web/LLM templates, with mostly multi-shot shot lists; (ii) a hierarchical, expert-vetted taxonomy that resolves differences (camera work, editing, performance) that saturating web-style benchmarks miss; (iii) a human-agreement study over 90 film-trained raters; and (iv) an open-source operator suite (FilmOps) with trained weights, plus per-dimension diagnostics (dynamic-aesthetics bottleneck, single→multi-shot degradation, championship mismatch) that are concrete, falsifiable findings about current models. The leaderboard is far from ceiling, which is itself informative. The main caveat on impact is that the full judge agent — the component producing all reported scores — is not released, so third parties cannot yet reproduce the benchmark numbers.
major comments (4)
- [§5.2, Table 2; §4.2] The central practical claim — that the automatic Overall score reproduces expert judgment (ρ = 0.95/0.96) — is validated only at the model level, over 9 (T2V) and 7 (R2V) systems, on ~30% of prompts. Meanwhile Table 2 shows three components with markedly lower agreement (editing 0.60, audio quality 0.68, audio coherence 0.68) that are equal-weighted into Overall via the §4.2 aggregation (roughly 1/12 + 1/12 of Aesthetic-Quality-plus-Instruction-Following weight and 1/7 of Temporal Continuity, ~16% of Overall by my count). The head of the leaderboard is tight: Seedance 2.0 vs HappyHorse 1.1 is 88.93 vs 87.42 (T2V) and 86.66 vs 85.51 (R2V), so a ~1.5-point systematic judge error concentrated in the weak-agreement components could plausibly reorder the leaders even with ρ = 0.95 overall, especially if the correlation is driven by the large gap to Hailuo (68.94). This is checkable within the
- [§4, §4.1 vs. contribution (iii)] The paper's third contribution is an 'expert-grade automatic evaluation agent', but only the FilmOps operator core (classification standard, weights, inference scripts) is open-sourced; the judge model that converts operator outputs and video into the 1–5 sub-metric scores — the component that actually produces every number in Tables 7–10 — is neither released nor specified (which backbone judge model, what prompting/scoring protocol, how operator labels are combined with judge scores). As written, the benchmark's scores are not reproducible by anyone outside the collaboration, which undercuts the release claim in the abstract and contributions. Please either release the judge configuration/weights or describe it precisely enough to reimplement, and state explicitly in §4 what is and is not public.
- [§4.2] The aggregation rule treats all L3 sub-metrics as equally important, maps the 1–5 rubric linearly to 0–100 (1→0), and excludes N/A audio dimensions symmetrically. None of these choices is justified or ablated, and each is consequential: the linear map makes a score of 1 equivalent to complete absence (0), so aggregate differences partly reflect rubric anchoring rather than capability; and the symmetric-N/A handling means audio-capable and non-audio models are compared on different effective metric sets, yet appear in the same Overall ranking (Hailuo's exclusion from the variance analysis for 'audio outliers' suggests this asymmetry is not benign). Please add a short ablation or at least a justification: e.g., Overall recomputed with audio models only, and sensitivity to the anchor mapping.
- [§5.2] The human-validation protocol needs more detail to support the ρ claims. With 90 raters and 2,878 model–prompt pairs, the paper does not state how many raters scored each pair, what the inter-rater reliability is (e.g., Krippendorff's α or per-component IRR), or whether raters used the same 5-point anchor rubrics as the judge. If human–human agreement on editing/audio is itself low, the ρ = 0.60–0.68 there may reflect intrinsic subjectivity rather than judge failure — an important distinction for interpreting Table 2. Please report IRR and rater assignment details, and clarify whether the 90 raters (all from the co-designing institution) constitute an independent validation or a within-school consistency check; a small held-out panel from outside the academy/studio would substantially strengthen the external-validity claim.
minor comments (6)
- [Table 2] The phrase 'separated by a mid-rule' in the Table 2 caption is unclear; presumably a visual separator between high- and low-agreement components. Please reword.
- [§3.1] Prompt drafting relies on Gemini 3.1 Pro (§3.1), a third-party model; please note the version/date and acknowledge the dependence, since prompt quality — and hence benchmark content — partially inherits that model's reverse-captioning behavior.
- [§5.9] The multi-shot drop (7.9 avg, up to 22.8) is compelling, but single-shot T2V n = 113 vs multi-shot n = 402 come from different source clips; the comparison is cross-sectional, not paired. A caveat, or a matched-clip subset analysis, would prevent over-reading the drop as purely compositional difficulty.
- [Figures 5–14] Several figures (5, 6, 10–14) carry embedded model-name labels inside bars and small insets; in the compiled PDF some axis text is hard to read at column width. Consider larger fonts or moving labels to legends.
- [Broader Impact] The Broader Impact statement is thin for a benchmark built on award-winning film clips: please address licensing/copyright status of the released prompts and any frame assets derived from copyrighted films.
- [§5.7] The 'championship' counting (18/35 etc.) is a nice diagnostic, but with near-tied models a per-sub-metric win threshold of zero makes counts noisy; a sentence on ties/significance would help.
Circularity Check
No significant circularity: FilmBench is an empirical benchmark whose rankings and gaps are external measurements, not results forced by definition or self-citation.
full rationale
FilmBench does not claim a first-principles derivation or a fitted-parameter prediction. Its load-bearing chain is constructive and empirical: (i) directors select award-winning clips and reverse-engineer multi-shot prompts; (ii) a three-level Cinematic Language taxonomy is co-designed with Beijing Film Academy faculty; (iii) an automatic agent (FilmOps operators plus judge) scores L3 sub-metrics; (iv) model-level Spearman agreement is measured against independent professional human raters on ~30% of prompts (ρ=0.95 T2V / 0.96 R2V); (v) leading T2V/R2V systems are scored under that fixed protocol. Human agreement is an external check that could have failed (and is weaker on editing and audio). Prompts are anchored to verified film clips, not to the generators under test. Shared academy criteria between taxonomy design and FilmOps is methodological alignment, not a reduction of a claimed prediction to its inputs. No uniqueness theorem, ansatz, or self-citation chain forces the leaderboard or the multi-shot/dynamic-aesthetics findings. Score 0 with empty steps is the proportionate finding.
Axiom & Free-Parameter Ledger
free parameters (3)
- Equal L3→L1 and equal L1→Overall weights =
uniform weights
- Linear map from 1–5 rubric to 0–100 =
1→0, 2→25, 3→50, 4→75, 5→100
- Human-agreement subsample size (~30% / 300 prompts) =
~300 prompts, 2878 model–prompt pairs
axioms (5)
- domain assumption Beijing Film Academy / studio Cinematic Language teaching categories are the correct professional standard for judging generated film craft.
- domain assumption Reverse-engineered structured prompts (Gemini draft + director refinement) faithfully encode the cinematic intent of the source award-winning clips.
- ad hoc to paper Model-level Spearman correlation of automatic vs human L2 means is a sufficient validation target for leaderboard trustworthiness.
- ad hoc to paper Not-applicable audio dimensions can be excluded symmetrically without distorting Overall comparisons across audio-capable and non-audio models.
- standard math Standard classification / MLLM operator training and macro-F1 evaluation practices transfer to cinematic label prediction across live-action and animation.
invented entities (3)
-
FilmBench taxonomy (3 L1 / 12 L2 / 35+3 L3)
independent evidence
-
FilmOps cinematic language operator suite
independent evidence
-
In-house expert-grade automatic evaluation agent
no independent evidence
read the original abstract
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Figures
Reference graph
Works this paper leans on
-
[1]
Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models, 2024
Fan Bao, Chendong Xiang, Gang Yue, Guande He, Hongzhou Zhu, Kaiwen Zheng, Min Zhao, Shilong Liu, Yaole Wang, and Jun Zhu. Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models, 2024. URLhttps://arxiv.org/abs/2405.04233
Pith/arXiv arXiv 2024
-
[2]
TiViBench: Benchmarking think-in-video reasoning for video generation
Harold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu, Zihan Wang, Sirui Chen, Wenkai Cheng, Kanghao Chen, Hongfei Zhang, Zixin Zhang, Rongjin Guo, Yu Cheng, and Ying-Cong Chen. TiViBench: Benchmarking think-in-video reasoning for video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11403–11413,
-
[3]
Rethinking video generation model for the embodied world
Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, and Daquan Zhou. Rethinking video generation model for the embodied world. InProceedings of the International Conference on Machine Learning (ICML), 2026. URLhttps://openreview.net/forum? id=p5QSlnwume
2026
-
[4]
Veo: A text-to-video generation system
Google DeepMind. Veo: A text-to-video generation system. Technical report, Google, 2025. URL https://deepmind.google/discover/blog/veo-2/
2025
-
[5]
Video-Bench: Human-aligned video generation benchmark
Hui Han, Siyuan Li, Jiaqi Chen, Yiwen Yuan, Yuling Wu, Yufan Deng, Chak Tou Leong, Hanwen Du, Junchen Fu, Youhua Li, Jie Zhang, Chi Zhang, Li-jia Li, and Yongxin Ni. Video-Bench: Human-aligned video generation benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18858–18868,
-
[6]
Happyhorse 1.0: Core capabilities and structural features for sota video generation
HappyHorse AI Team. Happyhorse 1.0: Core capabilities and structural features for sota video generation. Technical report, HappyHorse AI, 2026. URLhttps://happy-horse.art/features
2026
-
[7]
Happyhorse 1.1: Advancing multimodal integration and ai-driven video production
HappyHorse AI Team. Happyhorse 1.1: Advancing multimodal integration and ai-driven video production. Technical report, HappyHorse AI, 2026. URLhttps://happy-horse.art/happy-horse-1-1-ai
2026
-
[8]
VideoScore: Building automatic metrics to simulate fine-grained human feedback for video generation
Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Yuchen Lin, and Wenhu Chen. VideoScore: Building automatic metrics to simulate fine-grained human feedback for video generation. InPr...
2024
-
[9]
VBench: Comprehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative models. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2024
-
[10]
Karthik Inbasekar, Guy Rom, and Omer Shlomovits. WorldJen: An end-to-end multi-dimensional benchmark for generative video models.arXiv preprint arXiv:2605.03475, 2026. doi: 10.48550/arXiv. 2605.03475. URLhttps://arxiv.org/abs/2605.03475
-
[11]
VGA-Bench: A unified benchmark and multi-model framework for video aesthetics and generation quality evaluation
Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, and Xin Jin. VGA-Bench: A unified benchmark and multi-model framework for video aesthetics and generation quality evaluation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 30457–30466, 2026....
2026
-
[12]
Kling-omni technical report: A unified framework for video generation, editing, and reasoning
Kling Team and Kuaishou Technology AI Team. Kling-omni technical report: A unified framework for video generation, editing, and reasoning. Technical Report arXiv:2512.16776, Kuaishou Technology, 2025. URLhttps://arxiv.org/abs/2512.16776
Pith/arXiv arXiv 2025
-
[13]
Xiaofeng Li, Leyi Sheng, Zhen Sun, Zongmin Zhang, Jiaheng Wei, and Xinlei He. IP-Bench: Benchmark for image protection methods in image-to-video generation scenarios.arXiv preprint arXiv:2603.26154,
-
[14]
VMBench: A benchmark for perception-aligned video motion generation
Xinran Ling, Chen Zhu, Meiqi Wu, Hangyu Li, Xiaokun Feng, Cundian Yang, Aiming Hao, Jiashu Zhu, Ji- ahong Wu, and Xiangxiang Chu. VMBench: A benchmark for perception-aligned video motion generation. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 13087– 13098, 2025. URL https://openaccess.thecvf.com/content/ICCV2025...
2025
-
[15]
FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation
Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation. InAdvances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, volume 36, 2023. URL https://proceedings.neurips.cc/paper_files/paper/2023/hash/ c48...
2023
-
[16]
URLhttps://arxiv.org/abs/2603.26154
doi: 10.48550/arXiv.2603.26154. URLhttps://arxiv.org/abs/2603.26154
-
[17]
Minimax hailuo 2.3: A new level of complex video performance & media agent
MiniMax AI Team. Minimax hailuo 2.3: A new level of complex video performance & media agent. MiniMax News Release, 2025. URLhttps://www.minimax.io/news/minimax-hailuo-23
2025
-
[18]
SVBench: Evaluation of video generation models on social reasoning
Wenshuo Peng, Gongxuan Wang, Tianmeng Yang, Chuanhao Li, Xiaojie Xu, Hui He, and Kaipeng Zhang. SVBench: Evaluation of video generation models on social reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 32872– 32881, 2026. URL https://openaccess.thecvf.com/content/CVPR2026/html/Peng_SVBench_ Evalu...
2026
-
[19]
SLVMEval: Synthetic meta evaluation benchmark for text-to-long video generation
Ryosuke Matsuda, Keito Kudo, Haruto Yoshida, Nobuyuki Shimizu, and Jun Suzuki. SLVMEval: Synthetic meta evaluation benchmark for text-to-long video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7784–7794, 2026. URLhttps: //openaccess.thecvf.com/content/CVPR2026/html/Matsuda_SLVMEval_Synthetic...
2026
-
[20]
Seedance 2.0: Advancing video generation for world complexity, 2026
Team Seedance, De Chen, Liyang Chen, et al. Seedance 2.0: Advancing video generation for world complexity, 2026. URLhttps://arxiv.org/abs/2604.14148
Pith/arXiv arXiv 2026
-
[21]
MSVBench: Towards human-level evaluation of multi-shot video generation
Haoyuan Shi, Yunxin Li, Nanhao Deng, Zhenran Xu, Xinyu Chen, Longyue Wang, Baotian Hu, and Min Zhang. MSVBench: Towards human-level evaluation of multi-shot video generation. InFindings of the Association for Computational Linguistics: ACL 2026, pages 24034–24058, 2026. doi: 10.18653/v1/2026. findings-acl.1203. URLhttps://aclanthology.org/2026.findings-acl.1203/
doi:10.18653/v1/2026 2026
-
[22]
ConsistI2V: Enhancing visual consistency for image-to-video generation.Transactions on Machine Learning Research,
Weiming Ren, Huan Yang, Ge Zhang, Cong Wei, Xinrun Du, Wenhao Huang, and Wenhu Chen. ConsistI2V: Enhancing visual consistency for image-to-video generation.Transactions on Machine Learning Research,
-
[23]
CineTechBench: A benchmark for cinematographic technique under- standing and generation
Xinran Wang, Songyu Xu, Xiangxuan Shan, Yuxuan Zhang, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, and Zhanyu Ma. CineTechBench: A benchmark for cinematographic technique under- standing and generation. InAdvances in Neural Information Processing Systems (NeurIPS), 2025. URL https://arxiv.org/abs/2505.15145
Pith/arXiv arXiv 2025
-
[24]
Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, and Robert Geirhos. Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025
Pith/arXiv arXiv 2025
-
[25]
MovieBench: A hierarchical movie level dataset for long video generation
Weijia Wu, Mingyu Liu, Zeyu Zhu, Xi Xia, Haoen Feng, Wen Wang, Kevin Qinghong Lin, Chunhua Shen, and Mike Zheng Shou. MovieBench: A hierarchical movie level dataset for long video generation. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 28984–28994, 2025. URL https: //openaccess.thecvf.com/content/CVP...
2025
-
[26]
T2V-CompBench: A comprehensive benchmark for compositional text-to-video generation
Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2V-CompBench: A comprehensive benchmark for compositional text-to-video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8406–8416, 2025. URLhttps: //openaccess.thecvf.com/content/CVPR2025/html/Sun_T2V-CompBench_A_C...
2025
-
[27]
ChronoMagic-Bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation
Shenghai Yuan, Jinfa Huang, Yongqi Xu, Yaoyang Liu, Shaofeng Zhang, Yujun Shi, Rui- jie Zhu, Xinhua Cheng, Jiebo Luo, and Li Yuan. ChronoMagic-Bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation. InAdvances in Neural In- formation Processing Systems (NeurIPS) Datasets and Benchmarks Track, volume 37, pages 21236–21270, 202...
2024
-
[28]
Ailing Zhang, Lina Lei, Dehong Kong, Zhixin Wang, Jiaqi Xu, Fenglong Song, Chun-Le Guo, Chang Liu, Fan Li, and Jie Chen. UI2V-Bench: An understanding-based image-to-video generation benchmark.arXiv preprint arXiv:2509.24427, 2025. URLhttps://arxiv.org/abs/2509.24427
arXiv 2025
-
[29]
Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Lulu Gu, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, and Ziwei Liu. VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025. URL https://arxiv.org/abs/2503. 21755
Pith/arXiv arXiv 2025
-
[30]
Grok imagine video 1.5: Native audio generation and sota image-to-video workflow
xAI Team. Grok imagine video 1.5: Native audio generation and sota image-to-video workflow. xAI Official Blog, 2026. URLhttps://x.ai/news/grok-imagine-video-1-5
2026
-
[34]
IF” = Instruction Following, “TC
Ziwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang, Yuqing Yang, Qi Dai, Lili Qiu, and Chong Luo. A VGen-Bench: A task-driven benchmark for multi-granular evaluation of text-to-audio-video generation. InProceedings of the International Conference on Machine Learning (ICML), 2026. URL https: //openreview.net/forum?id=aJdgt8xDMy. A Qualitative Evaluation Examples...
2026
-
[2024]
URLhttps://openreview.net/forum?id=vqniLmUDvj
-
[2025]
URL https://openaccess.thecvf.com/content/CVPR2025/html/Han_Video-Bench_ Human-Aligned_Video_Generation_Benchmark_CVPR_2025_paper.html
-
[2026]
URL https://openaccess.thecvf.com/content/CVPR2026/html/Chen_TiViBench_ Benchmarking_Think-in-Video_Reasoning_for_Video_Generation_CVPR_2026_paper.html
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.