Pith. sign in

REVIEW 4 major objections 6 minor 34 references

Leading video generators still fall far short of film-grade craft when judged on professional cinematic language rather than web-style plausibility.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 20:09 UTC pith:I7GYHDQZ

load-bearing objection Solid film-grade T2V/R2V infrastructure: real multi-shot prompts, a usable taxonomy, and clear field gaps—use the diagnostics more than the tight Overall margins. the 4 major comments →

arxiv 2607.24241 v2 pith:I7GYHDQZ submitted 2026-07-27 cs.CV cs.AI

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

classification cs.CV cs.AI
keywords video generationcinematic languagetext-to-videoreference-to-videobenchmarkmulti-shot videoautomatic evaluationfilm aesthetics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most video-generation benchmarks still test basic visual plausibility with web-sourced prompts and generic scorers. FilmBench rebuilds the test around the cinematic language used in real film production: prompts reverse-engineered from award-winning multi-shot clips across twenty genres, and a three-level taxonomy of instruction following, temporal continuity, and aesthetic quality with dozens of craft sub-metrics. An automatic judge whose core operators are open-sourced recovers expert model rankings at very high correlation, yet no model approaches the ceiling. The field is bottlenecked on dynamic aesthetics—believable motion, performance, and camera work—and drops sharply when a prompt requires planning and cutting multiple shots instead of rendering a single take. The practical claim is that today’s leaderboards overstate readiness for professional cinematic creation.

Core claim

When text-to-video and reference-to-video models are scored on academy-aligned cinematic criteria using prompts reverse-engineered from real award-winning films, overall scores stay well below saturation, an automatic expert-grade evaluator reproduces human model rankings at Spearman 0.95–0.96, and the decisive gaps concentrate in dynamic aesthetics and multi-shot cinematic language rather than static image quality.

What carries the argument

FilmBench’s three-level Cinematic Language taxonomy (3 axes, 12 components, 35+3 sub-metrics) paired with FilmOps, an open suite of specialized operators that map video into professional labels for shot scale, camera movement, composition, and related craft dimensions that feed the automatic scorer.

Load-bearing premise

That matching human model rankings on roughly thirty percent of the prompts, with every fine-grained sub-metric weighted equally, is enough to treat the automatic judge as a faithful proxy for professional film craft—even where agreement is weaker on editing and audio.

What would settle it

Have independent film professionals score the same generated videos on the full taxonomy; if their model ranking diverges sharply from the automatic leaderboard, or if top models reach near-ceiling scores under fixed prompts and rubrics, the claim that current generators remain far from film-grade fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Web-style benchmarks that already look saturated will not predict which models can execute a professional multi-shot shot list.
  • Multi-shot staging, motivated camera moves, and performance become first-class evaluation and training targets rather than side effects of longer clips.
  • Raw reference fidelity and overall cinematic quality are separable skills; leading one need not mean leading the other.
  • Fine-grained sub-metric championships expose complementary model strengths that a single overall score conceals.
  • Open cinematic operators let others build custom film-grade judges without relying only on generic multimodal scorers.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Training that only maximizes generic aesthetics and caption alignment is likely to leave the multi-shot and performance gaps largely untouched.
  • A natural next stress test is end-to-end storyboard-to-sequence evaluation against a full director shot list, not only reverse-engineered clip prompts.
  • Field-wide collapse on prop-reference fidelity points to object-level identity binding as a distinct unsolved piece of controllable generation.
  • Because weaker models lose far more points on multi-shot prompts, scaling single-take quality alone will not close the craft gap.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces FilmBench, a T2V/R2V benchmark for cinematic video generation. Prompts are reverse-engineered from award-winning film clips selected by professional directors (515 T2V + 654 R2V prompts, 1,056 of 1,169 multi-shot, spanning 20 genres and Chinese/International markets), scored on a three-level taxonomy (3 L1 axes, 12 L2 components, 35+3 L3 sub-metrics) co-designed with Beijing Film Academy faculty and a film studio, by an in-house automatic evaluation agent built on an open-sourced Cinematic Language operator suite (FilmOps). Validated against 90 professional raters on ~300 prompts, the agent's model-level ranking matches human rankings at Spearman ρ = 0.95 (T2V) / 0.96 (R2V). Benchmarking 9 T2V and 7 R2V models, the authors report no saturation (tops of 88.93/86.66), a field-wide dynamic-aesthetics bottleneck, a 7.9-point average single→multi-shot drop (up to 22.8 for the weakest model), stable rankings across markets/genres/reference types, and a "championship mismatch" showing no model leads all sub-metrics.

Significance. If the validation holds up, FilmBench would be a genuinely useful addition to video-generation evaluation: it supplies (i) a prompt set anchored to verified professional footage rather than web/LLM templates, with mostly multi-shot shot lists; (ii) a hierarchical, expert-vetted taxonomy that resolves differences (camera work, editing, performance) that saturating web-style benchmarks miss; (iii) a human-agreement study over 90 film-trained raters; and (iv) an open-source operator suite (FilmOps) with trained weights, plus per-dimension diagnostics (dynamic-aesthetics bottleneck, single→multi-shot degradation, championship mismatch) that are concrete, falsifiable findings about current models. The leaderboard is far from ceiling, which is itself informative. The main caveat on impact is that the full judge agent — the component producing all reported scores — is not released, so third parties cannot yet reproduce the benchmark numbers.

major comments (4)
  1. [§5.2, Table 2; §4.2] The central practical claim — that the automatic Overall score reproduces expert judgment (ρ = 0.95/0.96) — is validated only at the model level, over 9 (T2V) and 7 (R2V) systems, on ~30% of prompts. Meanwhile Table 2 shows three components with markedly lower agreement (editing 0.60, audio quality 0.68, audio coherence 0.68) that are equal-weighted into Overall via the §4.2 aggregation (roughly 1/12 + 1/12 of Aesthetic-Quality-plus-Instruction-Following weight and 1/7 of Temporal Continuity, ~16% of Overall by my count). The head of the leaderboard is tight: Seedance 2.0 vs HappyHorse 1.1 is 88.93 vs 87.42 (T2V) and 86.66 vs 85.51 (R2V), so a ~1.5-point systematic judge error concentrated in the weak-agreement components could plausibly reorder the leaders even with ρ = 0.95 overall, especially if the correlation is driven by the large gap to Hailuo (68.94). This is checkable within the
  2. [§4, §4.1 vs. contribution (iii)] The paper's third contribution is an 'expert-grade automatic evaluation agent', but only the FilmOps operator core (classification standard, weights, inference scripts) is open-sourced; the judge model that converts operator outputs and video into the 1–5 sub-metric scores — the component that actually produces every number in Tables 7–10 — is neither released nor specified (which backbone judge model, what prompting/scoring protocol, how operator labels are combined with judge scores). As written, the benchmark's scores are not reproducible by anyone outside the collaboration, which undercuts the release claim in the abstract and contributions. Please either release the judge configuration/weights or describe it precisely enough to reimplement, and state explicitly in §4 what is and is not public.
  3. [§4.2] The aggregation rule treats all L3 sub-metrics as equally important, maps the 1–5 rubric linearly to 0–100 (1→0), and excludes N/A audio dimensions symmetrically. None of these choices is justified or ablated, and each is consequential: the linear map makes a score of 1 equivalent to complete absence (0), so aggregate differences partly reflect rubric anchoring rather than capability; and the symmetric-N/A handling means audio-capable and non-audio models are compared on different effective metric sets, yet appear in the same Overall ranking (Hailuo's exclusion from the variance analysis for 'audio outliers' suggests this asymmetry is not benign). Please add a short ablation or at least a justification: e.g., Overall recomputed with audio models only, and sensitivity to the anchor mapping.
  4. [§5.2] The human-validation protocol needs more detail to support the ρ claims. With 90 raters and 2,878 model–prompt pairs, the paper does not state how many raters scored each pair, what the inter-rater reliability is (e.g., Krippendorff's α or per-component IRR), or whether raters used the same 5-point anchor rubrics as the judge. If human–human agreement on editing/audio is itself low, the ρ = 0.60–0.68 there may reflect intrinsic subjectivity rather than judge failure — an important distinction for interpreting Table 2. Please report IRR and rater assignment details, and clarify whether the 90 raters (all from the co-designing institution) constitute an independent validation or a within-school consistency check; a small held-out panel from outside the academy/studio would substantially strengthen the external-validity claim.
minor comments (6)
  1. [Table 2] The phrase 'separated by a mid-rule' in the Table 2 caption is unclear; presumably a visual separator between high- and low-agreement components. Please reword.
  2. [§3.1] Prompt drafting relies on Gemini 3.1 Pro (§3.1), a third-party model; please note the version/date and acknowledge the dependence, since prompt quality — and hence benchmark content — partially inherits that model's reverse-captioning behavior.
  3. [§5.9] The multi-shot drop (7.9 avg, up to 22.8) is compelling, but single-shot T2V n = 113 vs multi-shot n = 402 come from different source clips; the comparison is cross-sectional, not paired. A caveat, or a matched-clip subset analysis, would prevent over-reading the drop as purely compositional difficulty.
  4. [Figures 5–14] Several figures (5, 6, 10–14) carry embedded model-name labels inside bars and small insets; in the compiled PDF some axis text is hard to read at column width. Consider larger fonts or moving labels to legends.
  5. [Broader Impact] The Broader Impact statement is thin for a benchmark built on award-winning film clips: please address licensing/copyright status of the released prompts and any frame assets derived from copyrighted films.
  6. [§5.7] The 'championship' counting (18/35 etc.) is a nice diagnostic, but with near-tied models a per-sub-metric win threshold of zero makes counts noisy; a sentence on ties/significance would help.

Circularity Check

0 steps flagged

No significant circularity: FilmBench is an empirical benchmark whose rankings and gaps are external measurements, not results forced by definition or self-citation.

full rationale

FilmBench does not claim a first-principles derivation or a fitted-parameter prediction. Its load-bearing chain is constructive and empirical: (i) directors select award-winning clips and reverse-engineer multi-shot prompts; (ii) a three-level Cinematic Language taxonomy is co-designed with Beijing Film Academy faculty; (iii) an automatic agent (FilmOps operators plus judge) scores L3 sub-metrics; (iv) model-level Spearman agreement is measured against independent professional human raters on ~30% of prompts (ρ=0.95 T2V / 0.96 R2V); (v) leading T2V/R2V systems are scored under that fixed protocol. Human agreement is an external check that could have failed (and is weaker on editing and audio). Prompts are anchored to verified film clips, not to the generators under test. Shared academy criteria between taxonomy design and FilmOps is methodological alignment, not a reduction of a claimed prediction to its inputs. No uniqueness theorem, ansatz, or self-citation chain forces the leaderboard or the multi-shot/dynamic-aesthetics findings. Score 0 with empty steps is the proportionate finding.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 3 invented entities

As a benchmark paper, load-bearing content is mostly domain design choices and measurement conventions rather than physical free parameters. The claim that FilmBench measures ‘film-grade craft’ rests on academy cinematic-language categories, equal metric aggregation, reverse-engineering fidelity, and treating specialized operators plus an in-house judge as expert proxies.

free parameters (3)
  • Equal L3→L1 and equal L1→Overall weights = uniform weights
    Every L3 sub-metric is treated as equally important; Overall is the mean of three L1 axes. No learned or expert-elicited weights; ranking can shift under alternate weightings.
  • Linear map from 1–5 rubric to 0–100 = 1→0, 2→25, 3→50, 4→75, 5→100
    Scores are mapped 1→0 … 5→100 before aggregation; spacing assumes equal perceptual steps.
  • Human-agreement subsample size (~30% / 300 prompts) = ~300 prompts, 2878 model–prompt pairs
    Model-level ρ is computed on a chosen subset of prompts and L2 means; subset selection affects the reported 0.95/0.96 figures.
axioms (5)
  • domain assumption Beijing Film Academy / studio Cinematic Language teaching categories are the correct professional standard for judging generated film craft.
    Taxonomy co-design (§3.2, Figure 1, Table 1) treats this academy system as ground truth for L1–L3 dimensions.
  • domain assumption Reverse-engineered structured prompts (Gemini draft + director refinement) faithfully encode the cinematic intent of the source award-winning clips.
    Pipeline in §3.1 / Figure 2 is the sole bridge from real films to machine-readable prompts.
  • ad hoc to paper Model-level Spearman correlation of automatic vs human L2 means is a sufficient validation target for leaderboard trustworthiness.
    §5.2 reports ρ on rankings of model means rather than clip-level absolute score calibration for every L3.
  • ad hoc to paper Not-applicable audio dimensions can be excluded symmetrically without distorting Overall comparisons across audio-capable and non-audio models.
    §4.2 aggregation rule; affects cross-model fairness when audio L3s are dropped.
  • standard math Standard classification / MLLM operator training and macro-F1 evaluation practices transfer to cinematic label prediction across live-action and animation.
    FilmOps training setup (§4.1) relies on ordinary supervised vision/MLLM methods plus practitioner QC.
invented entities (3)
  • FilmBench taxonomy (3 L1 / 12 L2 / 35+3 L3) independent evidence
    purpose: Operationalize professional cinematic judgment into scorable sub-metrics for T2V and R2V.
    New hierarchical metric system co-designed with academy/studio; not a physical entity but a postulated measurement ontology the paper’s claims depend on.
  • FilmOps cinematic language operator suite independent evidence
    purpose: Map frames/shots to industry-aligned labels (shot scale, composition, angle, tone, layout, camera movement) for automatic judging.
    Trained specialist operators introduced because generic MLLMs misjudge professional categories; weights released for external use.
  • In-house expert-grade automatic evaluation agent no independent evidence
    purpose: Produce 1–5 scores on every L3 dimension combining FilmOps, expert annotation models, and a judge model.
    Central scoring engine behind leaderboards; only core operators are open-sourced, so the full agent is a paper-specific construct validated mainly via internal human agreement.

pith-pipeline@v1.2.0-grok45-kimik3 · 81280 in / 3787 out tokens · 87183 ms · 2026-07-31T20:09:27.708140+00:00 · methodology

0 comments
read the original abstract

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.

Figures

Figures reproduced from arXiv: 2607.24241 by Bing Zhao, Chenglong Huang, Chongxiao Wang, Fanshu Ding, Fei Ding, Guangzheng Hu, Han Wu, Hengxia Qiang, Hong Qi, Hua Li, Hu Wei, Jie Tian, Jingjing Chen, Jingjing Fan, Jing Li, Jinlin Wang, Jinyang Zhen, Lin Qu, Mingshuang Tang, Niantong Li, Peng Han, Shengyi Wang, Weibin Chen, Weixu Qiao, Xiaoqian Zhu, Xiaotong Lv, Yanhao Wu, Yushu Wang, Zhong Li, Zimeng Li.

Figure 1
Figure 1. Figure 1: FilmBench evaluation taxonomy: 3 L1 axes, 12 L2 components, 35+3 (R2V-only) L3 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The director-driven reverse-engineering pipeline that turns award-winning film clips into [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Fine-grained, per-dimension FilmBench evaluation on a multi-shot reference-scene (R2V) [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Multi-shot T2V example (a sci-fi mech battle): Seedance 2.0 (86.11) vs. Grok Imagine Video [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Overall FilmBench scores (0–100) per model; model names are printed inside each bar and [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Cross-model variance per L3 sub-metric (main bars, colored by axis), with L2 inset (upper [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Per-axis rankings for T2V (left, 9 models) and R2V (right, 7 models). Each model is scored [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Per-model rankings within the L2 components most related to audiovisual language: [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Per-model rankings for the three highest-variance L3 sub-metrics (computed over merged [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Overall ranking split by market, Chinese Films vs. International Films; left: T2V [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Overall score heatmap across the top 4 genres by sample size. Y-axis: genres (with sample [PITH_FULL_IMAGE:figures/full_fig_p014_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Left: T2V single-shot (hatched) vs multi-shot (solid) mean scores (0–100) per model, [PITH_FULL_IMAGE:figures/full_fig_p015_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: T2V (left, 9 models) and R2V (right, 7 models): dialogue (hatched) vs. action (solid) [PITH_FULL_IMAGE:figures/full_fig_p016_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: R2V visual-following (reference-fidelity) ranking by reference type (columns: scene / [PITH_FULL_IMAGE:figures/full_fig_p016_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Cinematic Language L3 sub-metric cross-model means (0–100), T2V (hatched) vs. R2V [PITH_FULL_IMAGE:figures/full_fig_p017_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: R2V reference-character example in 3D animation: Seedance 2.0 (85.43) vs. Grok Imag [PITH_FULL_IMAGE:figures/full_fig_p021_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: R2V reference-prop example in 3D animation: Seedance 2.0 (82.75) vs. Kling 3.0 Omni [PITH_FULL_IMAGE:figures/full_fig_p021_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: T2V cross-model variance per L3 sub-metric, with L2 inset (upper right). [PITH_FULL_IMAGE:figures/full_fig_p022_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: R2V cross-model variance per L3 sub-metric, with L2 inset (upper right). [PITH_FULL_IMAGE:figures/full_fig_p023_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: T2V per-model L3 heatmap (35 sub-metrics [PITH_FULL_IMAGE:figures/full_fig_p026_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: R2V per-model L3 heatmap (38 sub-metrics [PITH_FULL_IMAGE:figures/full_fig_p027_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: T2V per-model L2 heatmap (12 components × 9 models) [PITH_FULL_IMAGE:figures/full_fig_p028_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: R2V per-model L2 heatmap (13 components × 7 models; adds the Visual Following component). Category inventories. Shot scale (8): extreme close-up, close-up, close shot, medium shot, medium full shot, full shot, long shot, extreme long shot. Composition (12): center, rule of thirds, horizontal, vertical, symmetric, framing, scattered, leading lines, diagonal, oblique, triangular, depth of field. Viewing ang… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 1 canonical work pages

  1. [1]

    Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models, 2024

    Fan Bao, Chendong Xiang, Gang Yue, Guande He, Hongzhou Zhu, Kaiwen Zheng, Min Zhao, Shilong Liu, Yaole Wang, and Jun Zhu. Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models, 2024. URLhttps://arxiv.org/abs/2405.04233

  2. [2]

    TiViBench: Benchmarking think-in-video reasoning for video generation

    Harold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu, Zihan Wang, Sirui Chen, Wenkai Cheng, Kanghao Chen, Hongfei Zhang, Zixin Zhang, Rongjin Guo, Yu Cheng, and Ying-Cong Chen. TiViBench: Benchmarking think-in-video reasoning for video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11403–11413,

  3. [3]

    Rethinking video generation model for the embodied world

    Yufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li, Ruoqing Hu, Yufei Ding, Yiming Zou, Yan Zeng, and Daquan Zhou. Rethinking video generation model for the embodied world. InProceedings of the International Conference on Machine Learning (ICML), 2026. URLhttps://openreview.net/forum? id=p5QSlnwume

  4. [4]

    Veo: A text-to-video generation system

    Google DeepMind. Veo: A text-to-video generation system. Technical report, Google, 2025. URL https://deepmind.google/discover/blog/veo-2/

  5. [5]

    Video-Bench: Human-aligned video generation benchmark

    Hui Han, Siyuan Li, Jiaqi Chen, Yiwen Yuan, Yuling Wu, Yufan Deng, Chak Tou Leong, Hanwen Du, Junchen Fu, Youhua Li, Jie Zhang, Chi Zhang, Li-jia Li, and Yongxin Ni. Video-Bench: Human-aligned video generation benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18858–18868,

  6. [6]

    Happyhorse 1.0: Core capabilities and structural features for sota video generation

    HappyHorse AI Team. Happyhorse 1.0: Core capabilities and structural features for sota video generation. Technical report, HappyHorse AI, 2026. URLhttps://happy-horse.art/features

  7. [7]

    Happyhorse 1.1: Advancing multimodal integration and ai-driven video production

    HappyHorse AI Team. Happyhorse 1.1: Advancing multimodal integration and ai-driven video production. Technical report, HappyHorse AI, 2026. URLhttps://happy-horse.art/happy-horse-1-1-ai

  8. [8]

    VideoScore: Building automatic metrics to simulate fine-grained human feedback for video generation

    Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bohan Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Bill Yuchen Lin, and Wenhu Chen. VideoScore: Building automatic metrics to simulate fine-grained human feedback for video generation. InPr...

  9. [9]

    VBench: Comprehensive benchmark suite for video generative models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Comprehensive benchmark suite for video generative models. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...

  10. [10]

    WorldJen: An end-to-end multi-dimensional benchmark for generative video models.arXiv preprint arXiv:2605.03475, 2026

    Karthik Inbasekar, Guy Rom, and Omer Shlomovits. WorldJen: An end-to-end multi-dimensional benchmark for generative video models.arXiv preprint arXiv:2605.03475, 2026. doi: 10.48550/arXiv. 2605.03475. URLhttps://arxiv.org/abs/2605.03475

  11. [11]

    VGA-Bench: A unified benchmark and multi-model framework for video aesthetics and generation quality evaluation

    Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, and Xin Jin. VGA-Bench: A unified benchmark and multi-model framework for video aesthetics and generation quality evaluation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 30457–30466, 2026....

  12. [12]

    Kling-omni technical report: A unified framework for video generation, editing, and reasoning

    Kling Team and Kuaishou Technology AI Team. Kling-omni technical report: A unified framework for video generation, editing, and reasoning. Technical Report arXiv:2512.16776, Kuaishou Technology, 2025. URLhttps://arxiv.org/abs/2512.16776

  13. [13]

    IP-Bench: Benchmark for image protection methods in image-to-video generation scenarios.arXiv preprint arXiv:2603.26154,

    Xiaofeng Li, Leyi Sheng, Zhen Sun, Zongmin Zhang, Jiaheng Wei, and Xinlei He. IP-Bench: Benchmark for image protection methods in image-to-video generation scenarios.arXiv preprint arXiv:2603.26154,

  14. [14]

    VMBench: A benchmark for perception-aligned video motion generation

    Xinran Ling, Chen Zhu, Meiqi Wu, Hangyu Li, Xiaokun Feng, Cundian Yang, Aiming Hao, Jiashu Zhu, Ji- ahong Wu, and Xiangxiang Chu. VMBench: A benchmark for perception-aligned video motion generation. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 13087– 13098, 2025. URL https://openaccess.thecvf.com/content/ICCV2025...

  15. [15]

    FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation

    Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. FETV: A benchmark for fine-grained evaluation of open-domain text-to-video generation. InAdvances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, volume 36, 2023. URL https://proceedings.neurips.cc/paper_files/paper/2023/hash/ c48...

  16. [16]

    URLhttps://arxiv.org/abs/2603.26154

    doi: 10.48550/arXiv.2603.26154. URLhttps://arxiv.org/abs/2603.26154

  17. [17]

    Minimax hailuo 2.3: A new level of complex video performance & media agent

    MiniMax AI Team. Minimax hailuo 2.3: A new level of complex video performance & media agent. MiniMax News Release, 2025. URLhttps://www.minimax.io/news/minimax-hailuo-23

  18. [18]

    SVBench: Evaluation of video generation models on social reasoning

    Wenshuo Peng, Gongxuan Wang, Tianmeng Yang, Chuanhao Li, Xiaojie Xu, Hui He, and Kaipeng Zhang. SVBench: Evaluation of video generation models on social reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 32872– 32881, 2026. URL https://openaccess.thecvf.com/content/CVPR2026/html/Peng_SVBench_ Evalu...

  19. [19]

    SLVMEval: Synthetic meta evaluation benchmark for text-to-long video generation

    Ryosuke Matsuda, Keito Kudo, Haruto Yoshida, Nobuyuki Shimizu, and Jun Suzuki. SLVMEval: Synthetic meta evaluation benchmark for text-to-long video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7784–7794, 2026. URLhttps: //openaccess.thecvf.com/content/CVPR2026/html/Matsuda_SLVMEval_Synthetic...

  20. [20]

    Seedance 2.0: Advancing video generation for world complexity, 2026

    Team Seedance, De Chen, Liyang Chen, et al. Seedance 2.0: Advancing video generation for world complexity, 2026. URLhttps://arxiv.org/abs/2604.14148

  21. [21]

    MSVBench: Towards human-level evaluation of multi-shot video generation

    Haoyuan Shi, Yunxin Li, Nanhao Deng, Zhenran Xu, Xinyu Chen, Longyue Wang, Baotian Hu, and Min Zhang. MSVBench: Towards human-level evaluation of multi-shot video generation. InFindings of the Association for Computational Linguistics: ACL 2026, pages 24034–24058, 2026. doi: 10.18653/v1/2026. findings-acl.1203. URLhttps://aclanthology.org/2026.findings-acl.1203/

  22. [22]

    ConsistI2V: Enhancing visual consistency for image-to-video generation.Transactions on Machine Learning Research,

    Weiming Ren, Huan Yang, Ge Zhang, Cong Wei, Xinrun Du, Wenhao Huang, and Wenhu Chen. ConsistI2V: Enhancing visual consistency for image-to-video generation.Transactions on Machine Learning Research,

  23. [23]

    CineTechBench: A benchmark for cinematographic technique under- standing and generation

    Xinran Wang, Songyu Xu, Xiangxuan Shan, Yuxuan Zhang, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, and Zhanyu Ma. CineTechBench: A benchmark for cinematographic technique under- standing and generation. InAdvances in Neural Information Processing Systems (NeurIPS), 2025. URL https://arxiv.org/abs/2505.15145

  24. [24]

    Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025

    Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, and Robert Geirhos. Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025

  25. [25]

    MovieBench: A hierarchical movie level dataset for long video generation

    Weijia Wu, Mingyu Liu, Zeyu Zhu, Xi Xia, Haoen Feng, Wen Wang, Kevin Qinghong Lin, Chunhua Shen, and Mike Zheng Shou. MovieBench: A hierarchical movie level dataset for long video generation. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 28984–28994, 2025. URL https: //openaccess.thecvf.com/content/CVP...

  26. [26]

    T2V-CompBench: A comprehensive benchmark for compositional text-to-video generation

    Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2V-CompBench: A comprehensive benchmark for compositional text-to-video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8406–8416, 2025. URLhttps: //openaccess.thecvf.com/content/CVPR2025/html/Sun_T2V-CompBench_A_C...

  27. [27]

    ChronoMagic-Bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation

    Shenghai Yuan, Jinfa Huang, Yongqi Xu, Yaoyang Liu, Shaofeng Zhang, Yujun Shi, Rui- jie Zhu, Xinhua Cheng, Jiebo Luo, and Li Yuan. ChronoMagic-Bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation. InAdvances in Neural In- formation Processing Systems (NeurIPS) Datasets and Benchmarks Track, volume 37, pages 21236–21270, 202...

  28. [28]

    UI2V-Bench: An understanding-based image-to-video generation benchmark.arXiv preprint arXiv:2509.24427, 2025

    Ailing Zhang, Lina Lei, Dehong Kong, Zhixin Wang, Jiaqi Xu, Fenglong Song, Chun-Le Guo, Chang Liu, Fan Li, and Jie Chen. UI2V-Bench: An understanding-based image-to-video generation benchmark.arXiv preprint arXiv:2509.24427, 2025. URLhttps://arxiv.org/abs/2509.24427

  29. [29]

    VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025

    Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Lulu Gu, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, and Ziwei Liu. VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025. URL https://arxiv.org/abs/2503. 21755

  30. [30]

    Grok imagine video 1.5: Native audio generation and sota image-to-video workflow

    xAI Team. Grok imagine video 1.5: Native audio generation and sota image-to-video workflow. xAI Official Blog, 2026. URLhttps://x.ai/news/grok-imagine-video-1-5

  31. [34]

    IF” = Instruction Following, “TC

    Ziwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang, Yuqing Yang, Qi Dai, Lili Qiu, and Chong Luo. A VGen-Bench: A task-driven benchmark for multi-granular evaluation of text-to-audio-video generation. InProceedings of the International Conference on Machine Learning (ICML), 2026. URL https: //openreview.net/forum?id=aJdgt8xDMy. A Qualitative Evaluation Examples...

  32. [2024]

    URLhttps://openreview.net/forum?id=vqniLmUDvj

  33. [2025]

    URL https://openaccess.thecvf.com/content/CVPR2025/html/Han_Video-Bench_ Human-Aligned_Video_Generation_Benchmark_CVPR_2025_paper.html

  34. [2026]

    URL https://openaccess.thecvf.com/content/CVPR2026/html/Chen_TiViBench_ Benchmarking_Think-in-Video_Reasoning_for_Video_Generation_CVPR_2026_paper.html