Pith. sign in

REVIEW 3 major objections 5 minor 5 cited by

T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Ten leading text-to-video models all score below 0.60 on every physical-law category in a new human-evaluated benchmark, with the best overall average at 0.42.

desk verdict The qualitative finding is almost certainly right, but the abstract overclaims below-0.60 across every law category, and the missing inter-annotator agreement makes the precise numeric scores noisier than the paper admits. read the letter →

arxiv 2505.00337 v1 pith:FW4BIQPI submitted 2025-05-01 cs.LG cs.AIcs.CLcs.CV

classification cs.LGcs.AIcs.CLcs.CV
keywords T2VPhysBenchtext-to-videogenerationphysicalconsistencyhumanevaluationNewtonianmechanicsconservationlawscounterfactualpromptsbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces a benchmark, T2VPhysBench, that tests whether text-to-video models obey twelve fundamental physical laws drawn from Newtonian mechanics, conservation principles, and phenomenological effects. Ten leading open-source and commercial models are scored by three human annotators on 84 prompts, and the paper reports that every model averages below 0.60 in each law category, with the best overall average at 0.42. It also finds that adding explicit hints about the target law rarely improves compliance and often lowers scores, and that when models are asked to generate physically impossible events they frequently comply. For a sympathetic reader the takeaway is that current video generators match prompts aesthetically but do not reason about the physical world.

What carries the argument

The load-bearing instrument is T2VPhysBench itself: a set of 84 prompts, seven per law, anchored to named laws rather than everyday scenario descriptions, plus a four-level human rating scale (0.0, 0.25, 0.5, 1.0) applied by three independent annotators to every video. The counterfactual arm replaces the prompts with physically impossible versions of the same scenarios, so that a model with genuine physical reasoning would be expected to produce videos that violate the named law. This design is what lets the paper attribute low scores to missing physical reasoning rather than to aesthetic quality or instruction-following failures.

What would settle it

Re-annotate the generated video set with two independent panels, or use a physics-simulation checker that measures, for example, the ball's vertical acceleration in the throw prompts; if panel scores diverge strongly or the simulated trajectory disagrees with human ratings on a large share of clips, the central below-0.60 finding is rater-dependent rather than a property of the models.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that state-of-the-art text-to-video systems do not internally represent basic physics: across all ten models, all three law families, and all twelve laws, average human-rated compliance stays below 0.60, and conservation laws are the hardest, topping out at 0.29. Prompt engineering cannot close the gap, because naming the law or spelling out the mechanism does not reliably raise scores and sometimes lowers them. The counterfactual study shows that models will generate impossible outcomes when instructed, which the paper takes as evidence that normal-looking compliance is surface pattern matching rather than physical understanding.

Load-bearing premise

The numerical conclusions rest on three annotators' four-level ratings being a reliable, unbiased measure of physical correctness, and the paper reports no inter-annotator agreement or rater-error analysis; if the ratings are noisy or biased, every score in the benchmark table shifts.

Editorial extensions

If this is right

  • If the scores are taken at face value, no current model can be trusted for safety-relevant video generation in robotics, autonomous driving, or scientific visualization.
  • Adding law-specific hints is not a workable remedy; the paper's ablation predicts that prompt refinement alone will not make these architectures physics-aware.
  • Conservation laws are systematically harder than Newton's laws or phenomena, so progress on physical consistency should be measured per law family rather than by a single average.
  • Counterfactual compliance scores being low means that instruction-following masks physical understanding; this predicts that a model's apparent realism and its physical competence can decouple.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An extension the authors leave implicit: the same 84-prompt protocol could be run with longer videos and higher resolutions to test whether physics failures persist with more temporal context, since the current 4-to-6-second clips may understate or overstate the gap.
  • One consequence not drawn in the paper is that trajectory-level automated checks, such as fitting projectile motion or collision velocities from the generated frames, could complement human ratings and turn the benchmark into a scalable regression test.
  • A testable prediction from the counterfactual results is that fine-tuning on physics-annotated data would improve standard-prompt scores faster than it improves counterfactual-prompt scores, because models can memorize canonical scenarios without acquiring transferable physical rules.
  • The per-law asymmetry may reflect training-data frequency more than physical complexity, since common scenarios like throwing a ball are scored better than rare ones like gyroscope motion; if so, data rebalancing would be the first lever rather than new physics-specific architectures.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces T2VPhysBench, a human-annotated benchmark for evaluating whether text-to-video generation models respect fundamental physical laws. The benchmark comprises 12 laws grouped into Newtonian principles, conservation principles, and phenomenon principles, with 84 prompts per model and 10 evaluated models spanning open and closed systems. Three studies are reported: (1) an overall compliance assessment summarized in Table 2, (2) a hint-level ablation varying prompt specificity, and (3) a counterfactual robustness test in Table 3. The headline claim is that all models score below 0.60 on average in each law category and that, collectively, models reveal a reliance on surface-pattern matching rather than genuine physical reasoning. The manuscript also presents observations on category-level difficulty, hint ineffectiveness, and counterfactual failure, and concludes with suggestions for physics-aware video generation.

Significance. If its empirical results are reliable, T2VPhysBench would be a useful community resource. Its strengths include the systematic coverage of twelve named physical laws, the inclusion of both open-source and commercial models, a human evaluation protocol that goes beyond automatic pixel-level metrics, and the use of a hint ablation and a counterfactual probe to interrogate failure modes. These are genuinely useful design choices, and the qualitative finding that current text-to-video models frequently violate basic physics is plausible and consistent with prior human-evaluated benchmarks such as VideoPhy. However, the quantitative claims rest entirely on a small human-rating protocol with no reported inter-annotator reliability or released annotation data, and the paper's headline 'below 0.60 in each law category' is contradicted by its own Table 2. The benchmark's contribution is therefore currently significant in conception but not yet established in its quantitative specifics.

major comments (3)
  1. [Section 3.3, Tables 2 and 3] The central quantitative result is not reproducible as reported. Section 3.3 states that three annotators independently assign a four-level score, and Section 4.1 averages across prompts and annotators, but the paper reports no inter-annotator agreement statistic (e.g., Cohen's kappa or Krippendorff's alpha), no per-item variance, no annotator-level score distributions, and the raw annotations are not released. With only 7 prompts per law and 3 annotators per cell, differences such as Kling 0.35 versus Mochi-1 0.34 in Table 2 cannot be distinguished from rater noise or systematic annotator bias. In addition, the rubric maps the ordinal levels 0.0, 0.25, 0.5, and 1.0 to equal intervals without validation, and the boundary between 'fails to demonstrate the intended physical behavior' (0.0) and 'clear violation of the law' (0.25) is especially susceptible to arbitrary thresholding in ambiguous videos. Please publish the annotation data and agreement statistics, or the numerical scores, model rankings, and category-level orderings should be treated as provisional.
  2. [Abstract and Section 4.1, Table 2] The abstract's claim that 'all models score below 0.60 on average in each law category' is internally contradicted by Table 2, where Qingying receives 0.63 on Phenomenon Principles. Observation 4.1 narrows the claim to 'basic Newtonian and conservation laws,' but the abstract and the contribution bullet still assert the stronger statement over 'each law category' and 'every law category.' This is a factual inconsistency in the paper's headline result, not merely a wording preference. Either the claim must be restricted to the laws for which it holds, as Observation 4.1 does, or the data must be revisited; in all cases the abstract, the contributions list, and Observation 4.1 must be made mutually consistent.
  3. [Section 4.3, Table 3] The counterfactual robustness study's interpretation is not supported by its design. The premise that a model with genuine physical reasoning 'should understand how to generate videos that violate some specific physical laws' is an unargued assumption: a model that generates physically plausible behavior even when instructed to produce an impossibility could instead be exhibiting a beneficial physics prior, and the counterfactual prompts vary in how unambiguously they specify the required violation. Moreover, the Section 3.3 rubric is defined for adherence to the target law, with Level 4 meaning the video 'fully and accurately conforms to the law,' so applying the same scale to measure deliberate violations leaves the scores with unclear semantics under counterfactual instructions. Consequently, Observations 4.5 and 4.6, which conclude that models 'demonstrate an inability to understand impossible physics' and that their compliance 'is rooted in memorized patterns,' overreach beyond the evidence. The experiment needs a counterfactual-specific scoring rubric or a stated control condition, and the corresponding conclusions should be softened accordingly.
minor comments (5)
  1. [Section 4.3] In the paragraph following Table 3, the reference 'Table 4.1' should be 'Table 2', which is the actual table containing the overall compliance scores.
  2. [Section 4.3] There are multiple typos and grammatical errors in this section, including 'vaccum', 'without and force', 'liinear motions', and 'Threrefore'; these should be corrected.
  3. [Figures 3-5] Figures 3 through 5 appear in the manuscript as unreadable character-encoding sequences (for example, '/uni00000031/uni00000048/...'), and the hint-level results are reported only as figures with no accompanying table, per-condition standard errors, or number of videos; please replace them with legible figures or provide the underlying per-law numerical results so that Observation 4.4 can be verified.
  4. [Section 3.1 and Appendix A] The models are evaluated at different resolutions and durations (e.g., Mochi-1 at 480p, LTX Video at 512p, Sora at 720p, and SD Video at 4 seconds), so resolution and clip length are potential confounds when comparing physics compliance across models; the authors should either match generation settings where possible or report scores separately by configuration.
  5. [Section 3.2 and Contributions] The full set of 84 prompts is not provided in the paper or appendix; please include the complete prompt list so the benchmark is reproducible. In addition, the text alternates between 'first-principled' and 'first-principles', and the contributions bullet 'a first first-principled benchmark' contains a typo that should be fixed.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: T2VPhysBench is an empirical human-evaluation benchmark; its headline scores are summary statistics of annotations, not derivations from fitted inputs or self-citations.

full rationale

This paper reports an empirical benchmark and does not contain a derivation chain that could reduce to its own inputs. The central claim, that models score below 0.60 on physical law compliance, is an arithmetic summary of Table 2, which is produced by averaging four-level human ratings over prompts and annotators as described in Section 3.3. The mapping from rating levels to scores in [0,1] is a measurement convention, not a function defined in terms of the target conclusion, so it is not self-definitional. No parameter is fitted to a subset of data and then renamed as a prediction; the experiments directly measure generated videos against a rubric. The only self-citation, [GHH+25], appears in the related-work discussion of object counting benchmarks and is not load-bearing for any claim in this paper. The counterfactual study in Section 4.3 rests on an assumption about what genuine physical reasoning should produce, which may be debatable, but it is not circular: the premise does not define the observed scores into existence. The abstract's statement that 'all models score below 0.60 on average in each law category' is internally inconsistent with Table 2, where Qingying scores 0.63 on Phenomenon Principles, and Observation 4.1 narrows the claim to 'basic Newtonian and conservation laws'; however, this is an accuracy or consistency problem, not circularity. Appendix B explicitly acknowledges that 'our study is entirely empirical,' confirming that there is no formal derivation whose conclusion is assumed among its premises. No self-definitional step, fitted input called a prediction, load-bearing self-citation, imported uniqueness theorem, ansatz smuggled via citation, or renaming of a known result was found. The appropriate finding is therefore no significant circularity, with score 0.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The benchmark introduces no physical constants or fitted model parameters, but it relies on a hand-chosen measurement mapping and several implicit domain assumptions. The main quantitative output (average compliance scores between 0.19 and 0.42) is an artifact of the arbitrary level-to-score mapping as much as of model behavior.

free parameters (1)
  • Human rating score mapping = Level 1=0.0, Level 2=0.25, Level 3=0.5, Level 4=1.0
    The mapping from ordinal quality labels to numbers is chosen by hand with arbitrary equal spacing. All average scores in Tables 2 and 3 are linear functions of these values, so the mapping directly determines every reported number. The scale assumes the gap between 'clear violation' and 'minor inaccuracy' matches other adjacent gaps, which is untested.
assumptions (4)
  • domain assumption The ratings of three annotators are a reliable and unbiased measure of physical correctness.
    Section 3.3 describes the rating protocol but reports no inter-annotator agreement; all numeric claims depend on this assumption.
  • domain assumption A model with genuine physical reasoning should intentionally generate physically impossible videos when such counterfactual behavior is requested.
    Section 4.3 uses low counterfactual compliance as evidence of pattern memorization without validating this diagnostic assumption.
  • ad hoc to paper Each of the 84 prompts isolates exactly one target physical law.
    Prompts like 'A person quickly pull out the paper pressed under the water bottle' involve multiple physical mechanisms (friction, inertia, rotation), yet each is scored against a single law; no validation of this mapping is provided.
  • domain assumption The 12 chosen laws and 7 prompts per law constitute a representative first-principles coverage of physics for video generation.
    The selection is ad hoc; no sampling method or completeness argument is given for the law or prompt set.

how reviews work

0 comments
Cite this review

Pith. "Pith review of T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation." pith.science (2026). https://pith.science/paper/FW4BIQPI

@misc{pith2026250500337,
  author       = {Pith},
  title        = {Pith review of: T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FW4BIQPI}},
  note         = {Machine review of arXiv:2505.00337}
}
read the original abstract

Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art creation and user engagement online. Yet, despite these advancements, their ability to respect fundamental physical laws remains largely untested: many outputs still violate basic constraints such as rigid-body collisions, energy conservation, and gravitational dynamics, resulting in unrealistic or even misleading content. Existing physical-evaluation benchmarks typically rely on automatic, pixel-level metrics applied to simplistic, life-scenario prompts, and thus overlook both human judgment and first-principles physics. To fill this gap, we introduce \textbf{T2VPhysBench}, a first-principled benchmark that systematically evaluates whether state-of-the-art text-to-video systems, both open-source and commercial, obey twelve core physical laws including Newtonian mechanics, conservation principles, and phenomenological effects. Our benchmark employs a rigorous human evaluation protocol and includes three targeted studies: (1) an overall compliance assessment showing that all models score below 0.60 on average in each law category; (2) a prompt-hint ablation revealing that even detailed, law-specific hints fail to remedy physics violations; and (3) a counterfactual robustness test demonstrating that models often generate videos that explicitly break physical rules when so instructed. The results expose persistent limitations in current architectures and offer concrete insights for guiding future research toward truly physics-aware video generation.

Figures

Figures reproduced from arXiv: 2505.00337 by the authors.

Figure 1
Figure 1. All 12 physical laws evaluated in this benchmark, illustrated with video examples from various text-to-video models. In this benchmark, we address the problem of enforcing physical constraints using a first￾4 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Prompt and Video Examples with Different Hint Levels. Based on our previous findings in Section 4.1, we have observed that most text-to-video models fail to generate videos that comply with physical laws. To show that this inherent limitation is non-trivial and cannot be resolved simply through prompt improvements, in this study we explore a simple but critical problem: can text-to-video models follow physical const… view at source ↗
Figure 3
Figure 3. Ablation Study of Different Hint Level Prompts in Newton Principles. Conservation of Energy Conservation of Mass Conservation of Momentum Conservation of Angular Momentum 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Score 0.25 0.25 0.12 0.56 0.44 0.50 0.12 0.25 0.31 0.44 0.12 0.44 Impact of Different Hint Level Prompt in Conservation Principles Inital Prompt First-Level Hint Prompt Second-Level Hint Prompt [PITH_FULL_IM… view at source ↗
Figures from the paper (27 more)
Figure 4
Figure 4. Figure 4: Ablation Study of Different Hint Level Prompts in Conservation Principles. Hooke's Law Snell's Law Law of Reflection Bernoulli's Principle 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Score 0.50 0.38 0.19 0.31 0.44 0.31 0.25 0.25 0.38 0.38 0.25 0.31 Impact of Different Hint…
Figure 5
Figure 5. Figure 5: Ablation Study of Different Hint Level Prompts in Phenomenon Principles. 8 [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Examples of counterfactual prompts in this benchmark, along with generated example videos from all ten text-to-video models. 9 [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Results of Generating Videos Following Newton’s First Law. 20 [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: Results of Generating Videos Following Newton’s First Law. 21 [PITH_FULL_IMAGE:figures/full_fig_p022_8.png]
Figure 9
Figure 9. Figure 9: Results of Generating Videos Following Newton’s Second Law. 22 [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]
Figure 10
Figure 10. Figure 10: Results of Generating Videos Following Newton’s Second Law. 23 [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: Results of Generating Videos Following Newton’s Third Law. 24 [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]
Figure 12
Figure 12. Figure 12: Results of Generating Videos Following Newton’s Third Law. 25 [PITH_FULL_IMAGE:figures/full_fig_p026_12.png]
Figure 13
Figure 13. Figure 13: Results of Generating Videos Following Law of Universal Gravitation. 26 [PITH_FULL_IMAGE:figures/full_fig_p027_13.png]
Figure 14
Figure 14. Figure 14: Results of Generating Videos Following Law of Universal Gravitation. 27 [PITH_FULL_IMAGE:figures/full_fig_p028_14.png]
Figure 15
Figure 15. Figure 15: Results of Generating Videos Following Conservation of Energy. 28 [PITH_FULL_IMAGE:figures/full_fig_p029_15.png]
Figure 16
Figure 16. Figure 16: Results of Generating Videos Following Conservation of Energy. 29 [PITH_FULL_IMAGE:figures/full_fig_p030_16.png]
Figure 17
Figure 17. Figure 17: Results of Generating Videos Following Conservation of Mass. 30 [PITH_FULL_IMAGE:figures/full_fig_p031_17.png]
Figure 18
Figure 18. Figure 18: Results of Generating Videos Following Conservation of Mass. 31 [PITH_FULL_IMAGE:figures/full_fig_p032_18.png]
Figure 19
Figure 19. Figure 19: Results of Generating Videos Following Conservation of Momentum. 32 [PITH_FULL_IMAGE:figures/full_fig_p033_19.png]
Figure 20
Figure 20. Figure 20: Results of Generating Videos Following Conservation of Momentum. 33 [PITH_FULL_IMAGE:figures/full_fig_p034_20.png]
Figure 21
Figure 21. Figure 21: Results of Generating Videos Following Conservation of Angular Momentum. 34 [PITH_FULL_IMAGE:figures/full_fig_p035_21.png]
Figure 22
Figure 22. Figure 22: Results of Generating Videos Following Conservation of Angular Momentum. 35 [PITH_FULL_IMAGE:figures/full_fig_p036_22.png]
Figure 23
Figure 23. Figure 23: Results of Generating Videos Following Hooke’s Law. 36 [PITH_FULL_IMAGE:figures/full_fig_p037_23.png]
Figure 24
Figure 24. Figure 24: Results of Generating Videos Following Hooke’s Law. 37 [PITH_FULL_IMAGE:figures/full_fig_p038_24.png]
Figure 25
Figure 25. Figure 25: Results of Generating Videos Following Snell’s Law. 38 [PITH_FULL_IMAGE:figures/full_fig_p039_25.png]
Figure 26
Figure 26. Figure 26: Results of Generating Videos Following Snell’s Law. 39 [PITH_FULL_IMAGE:figures/full_fig_p040_26.png]
Figure 27
Figure 27. Figure 27: Results of Generating Videos Following Law of Reflection. 40 [PITH_FULL_IMAGE:figures/full_fig_p041_27.png]
Figure 28
Figure 28. Figure 28: Results of Generating Videos Following Law of Reflection. 41 [PITH_FULL_IMAGE:figures/full_fig_p042_28.png]
Figure 29
Figure 29. Figure 29: Results of Generating Videos Following Bernoulli’s Principle. 42 [PITH_FULL_IMAGE:figures/full_fig_p043_29.png]
Figure 30
Figure 30. Figure 30: Results of Generating Videos Following Bernoulli’s Principle. 43 [PITH_FULL_IMAGE:figures/full_fig_p044_30.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Detecting AI-Generated Video: A Vision-Language Dual-View Survey

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AIGC-V detection should be treated as factual fidelity verification and organized by a four-layer vision-language dual-view taxonomy spanning cues, motion, cross-modal consistency, and world-level reasoning.

  2. "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.

  3. RoboScape: Physics-informed Embodied World Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RoboScape jointly learns RGB video, depth, and keypoint-token consistency in one autoregressive world model, improving video quality, geometry, action control, synthetic-data policy training, and policy evaluation for...

  4. Minimalist Softmax Attention Provably Learns Constrained Boolean Functions

    cs.LG 2025-05 reject novelty 5.0 of 10

    With teacher forcing that reveals pairwise products of the relevant bits, one gradient step lets a single-head attention recover the support of a k-bit AND/OR; the paper's claimed end-to-end hardness lower bound is in...

  5. Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse

    cs.LG 2025-05 reject novelty 4.0 of 10

    A residual self-attention network with all weight entries bounded by a small η can be approximated by one layer to error O(η)‖X‖∞, so skip connections do not prevent layer collapse.

Reference graph

Works this paper leans on

23 extracted references · 6 canonical work pages · cited by 5 Pith papers

  1. [1]

    Cosmos world foundation model platform for physical ai

    [AAB+25] Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575 ,

  2. [4]

    Conditional gan with discriminative filter generation for text-to-video synthe- sis

    [BMB+19] Yogesh Balaji, Martin Renqiang Min, Bing Bai, Rama Chellappa, and Hans Peter Graf. Conditional gan with discriminative filter generation for text-to-video synthe- sis. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, pages 1995–2001. International Joint Conferences on Artificial Intelligence...

  3. [5]

    Tc-bench: Benchmarking temporal compositionality in text-to-video and image-to-video generation

    [FLS+24] Weixi Feng, Jiachen Li, Michael Saxon, Tsu-jui Fu, Wenhu Chen, and William Yang Wang. Tc-bench: Benchmarking temporal compositionality in text-to-video and image-to-video generation. arXiv preprint arXiv:2406.08656 ,

  4. [9]

    Genai-bench: A holistic benchmark for com- positional text-to-visual generation

    [LLP+24] Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Emily Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. Genai-bench: A holistic benchmark for com- positional text-to-visual generation. In Synthetic Data for Computer Vision Work- shop@ CVPR 2024 ,

  5. [10]

    Exploring the evolution of physics cognition in video generation: A survey

    [LWW+25] Minghui Lin, Xiang Wang, Yishan Wang, Shu Wang, Fengqi Dai, Pengxiang Ding, Cunxiang Wang, Zhengrong Zuo, Nong Sang, Siteng Huang, et al. Exploring the evolution of physics cognition in video generation: A survey. arXiv preprint arXiv:2503.21765,

  6. [11]

    Do generative video models learn physical principles from watching videos? arXiv preprint arXiv:2501.09038,

    [MCS+25] Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, and Robert Geirhos. Do generative video models learn physical principles from watching videos? arXiv preprint arXiv:2501.09038,

  7. [12]

    Towards world simulator: Craft- ing physical commonsense-based benchmark for video generation

    [MLT+24] Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo. Towards world simulator: Craft- ing physical commonsense-based benchmark for video generation. arXiv preprint arXiv:2410.05363,

  8. [13]

    Phybench: A phys- ical commonsense benchmark for evaluating text-to-image models

    [MSL+24] Fanqing Meng, Wenqi Shao, Lixin Luo, Yahong Wang, Yiran Chen, Quanfeng Lu, Yue Yang, Tianshuo Yang, Kaipeng Zhang, Yu Qiao, et al. Phybench: A phys- ical commonsense benchmark for evaluating text-to-image models. arXiv preprint arXiv:2406.11802,

Show all 23 references
  1. [14]

    Dreamingv2: Reinforcement learning with discrete world models without reconstruction

    [OT22] Masashi Okada and Tadahiro Taniguchi. Dreamingv2: Reinforcement learning with discrete world models without reconstruction. In 2022 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS) , pages 985–991. IEEE,

  2. [19]

    Modelscope text-to-video technical report

    [WYC+23] Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571 ,

  3. [20]

    VideoCLIP: Contrastive pre- training for zero-shot video-text understanding

    [XGH+21] Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. VideoCLIP: Contrastive pre- training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in...

  4. [21]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    [YTZ+24] Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072,

  5. [22]

    In Section A, we present the details of each evaluated model

    16 Appendix Roadmap. In Section A, we present the details of each evaluated model. In Section B, we discuss the limitations of this paper. In Section C, we illustrate the societal implications of this work. In Section D, we show detailed video examples. A Implementation Detail...

  6. [23]

    It uses Deepseek-R1 for prompt enhancement, and provides aspect ratio choices including 16:9, 21:9, 4:3, 1:1, 3:4, and 9:16

    It comes in four variants: Video S2.0, Video S2.0 Pro, Video P2.0 Pro, and Video 1.2. It uses Deepseek-R1 for prompt enhancement, and provides aspect ratio choices including 16:9, 21:9, 4:3, 1:1, 3:4, and 9:16. Video S2.0, Video S2.0 Pro, and Video P2.0 Pro are able to create ...

  7. [1995]

    Wisa: World simulator assistant for physics-aware text-to-video generation

    [WMC+25] Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, et al. Wisa: World simulator assistant for physics-aware text-to-video generation. arXiv preprint arXiv:2503.08153 ,

  8. [2016]

    T2v-compbench: A comprehensive benchmark for compositional text-to-video generation

    [SHL+24] Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2v-compbench: A comprehensive benchmark for compositional text-to-video generation. arXiv preprint arXiv:2407.14505 ,

  9. [2019]

    Learning a driving simulator

    [SH16] Eder Santana and George Hotz. Learning a driving simulator. arXiv preprint arXiv:1608.01230,

  10. [2020]

    Ltx-video: Realtime video latent diffusion

    [HCB+24] Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Ei- tan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103 ,

  11. [2021]

    Physics informed deep learning (part i): Data-driven solutions of nonlinear partial differential equations

    [RPK17] Maziar Raissi, Paris Perdikaris, and George Em Karniadakis. Physics informed deep learning (part i): Data-driven solutions of nonlinear partial differential equations. arXiv preprint arXiv:1711.10561 ,

  12. [2022]

    World models.arXiv preprint arXiv:1803.10122,

    [HS18] David Ha and J¨ urgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122,

  13. [2023]

    Videophy: Eval- uating physical commonsense for video generation

    [BLX+25] Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover. Videophy: Eval- uating physical commonsense for video generation. In Workshop on Video-Language Models @ NeurIPS 2024 ,

  14. [2024]

    Can you count to nine? a human evaluation benchmark for counting limits in modern text-to-video models

    [GHH+25] Xuyang Guo, Zekai Huang, Jiayan Huo, Yingyu Liang, Zhenmei Shi, Zhao Song, and Jiahao Zhang. Can you count to nine? a human evaluation benchmark for counting limits in modern text-to-video models. arXiv preprint arXiv:2504.04051 ,

  15. [2025]

    Stable video diffusion: Scaling latent video diffusion models to large datasets

    [BDK+23] Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kil- ian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127,

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.