REVIEW 3 major objections 5 minor 5 cited by
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Ten leading text-to-video models all score below 0.60 on every physical-law category in a new human-evaluated benchmark, with the best overall average at 0.42.
desk verdict The qualitative finding is almost certainly right, but the abstract overclaims below-0.60 across every law category, and the missing inter-annotator agreement makes the precise numeric scores noisier than the paper admits. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is T2VPhysBench itself: a set of 84 prompts, seven per law, anchored to named laws rather than everyday scenario descriptions, plus a four-level human rating scale (0.0, 0.25, 0.5, 1.0) applied by three independent annotators to every video. The counterfactual arm replaces the prompts with physically impossible versions of the same scenarios, so that a model with genuine physical reasoning would be expected to produce videos that violate the named law. This design is what lets the paper attribute low scores to missing physical reasoning rather than to aesthetic quality or instruction-following failures.
What would settle it
Re-annotate the generated video set with two independent panels, or use a physics-simulation checker that measures, for example, the ball's vertical acceleration in the throw prompts; if panel scores diverge strongly or the simulated trajectory disagrees with human ratings on a large share of clips, the central below-0.60 finding is rater-dependent rather than a property of the models.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that state-of-the-art text-to-video systems do not internally represent basic physics: across all ten models, all three law families, and all twelve laws, average human-rated compliance stays below 0.60, and conservation laws are the hardest, topping out at 0.29. Prompt engineering cannot close the gap, because naming the law or spelling out the mechanism does not reliably raise scores and sometimes lowers them. The counterfactual study shows that models will generate impossible outcomes when instructed, which the paper takes as evidence that normal-looking compliance is surface pattern matching rather than physical understanding.
Load-bearing premise
The numerical conclusions rest on three annotators' four-level ratings being a reliable, unbiased measure of physical correctness, and the paper reports no inter-annotator agreement or rater-error analysis; if the ratings are noisy or biased, every score in the benchmark table shifts.
Editorial extensions
If this is right
- If the scores are taken at face value, no current model can be trusted for safety-relevant video generation in robotics, autonomous driving, or scientific visualization.
- Adding law-specific hints is not a workable remedy; the paper's ablation predicts that prompt refinement alone will not make these architectures physics-aware.
- Conservation laws are systematically harder than Newton's laws or phenomena, so progress on physical consistency should be measured per law family rather than by a single average.
- Counterfactual compliance scores being low means that instruction-following masks physical understanding; this predicts that a model's apparent realism and its physical competence can decouple.
Reading between the lines
- An extension the authors leave implicit: the same 84-prompt protocol could be run with longer videos and higher resolutions to test whether physics failures persist with more temporal context, since the current 4-to-6-second clips may understate or overstate the gap.
- One consequence not drawn in the paper is that trajectory-level automated checks, such as fitting projectile motion or collision velocities from the generated frames, could complement human ratings and turn the benchmark into a scalable regression test.
- A testable prediction from the counterfactual results is that fine-tuning on physics-annotated data would improve standard-prompt scores faster than it improves counterfactual-prompt scores, because models can memorize canonical scenarios without acquiring transferable physical rules.
- The per-law asymmetry may reflect training-data frequency more than physical complexity, since common scenarios like throwing a ball are scored better than rare ones like gyroscope motion; if so, data rebalancing would be the first lever rather than new physics-specific architectures.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces T2VPhysBench, a human-annotated benchmark for evaluating whether text-to-video generation models respect fundamental physical laws. The benchmark comprises 12 laws grouped into Newtonian principles, conservation principles, and phenomenon principles, with 84 prompts per model and 10 evaluated models spanning open and closed systems. Three studies are reported: (1) an overall compliance assessment summarized in Table 2, (2) a hint-level ablation varying prompt specificity, and (3) a counterfactual robustness test in Table 3. The headline claim is that all models score below 0.60 on average in each law category and that, collectively, models reveal a reliance on surface-pattern matching rather than genuine physical reasoning. The manuscript also presents observations on category-level difficulty, hint ineffectiveness, and counterfactual failure, and concludes with suggestions for physics-aware video generation.
Significance. If its empirical results are reliable, T2VPhysBench would be a useful community resource. Its strengths include the systematic coverage of twelve named physical laws, the inclusion of both open-source and commercial models, a human evaluation protocol that goes beyond automatic pixel-level metrics, and the use of a hint ablation and a counterfactual probe to interrogate failure modes. These are genuinely useful design choices, and the qualitative finding that current text-to-video models frequently violate basic physics is plausible and consistent with prior human-evaluated benchmarks such as VideoPhy. However, the quantitative claims rest entirely on a small human-rating protocol with no reported inter-annotator reliability or released annotation data, and the paper's headline 'below 0.60 in each law category' is contradicted by its own Table 2. The benchmark's contribution is therefore currently significant in conception but not yet established in its quantitative specifics.
major comments (3)
- [Section 3.3, Tables 2 and 3] The central quantitative result is not reproducible as reported. Section 3.3 states that three annotators independently assign a four-level score, and Section 4.1 averages across prompts and annotators, but the paper reports no inter-annotator agreement statistic (e.g., Cohen's kappa or Krippendorff's alpha), no per-item variance, no annotator-level score distributions, and the raw annotations are not released. With only 7 prompts per law and 3 annotators per cell, differences such as Kling 0.35 versus Mochi-1 0.34 in Table 2 cannot be distinguished from rater noise or systematic annotator bias. In addition, the rubric maps the ordinal levels 0.0, 0.25, 0.5, and 1.0 to equal intervals without validation, and the boundary between 'fails to demonstrate the intended physical behavior' (0.0) and 'clear violation of the law' (0.25) is especially susceptible to arbitrary thresholding in ambiguous videos. Please publish the annotation data and agreement statistics, or the numerical scores, model rankings, and category-level orderings should be treated as provisional.
- [Abstract and Section 4.1, Table 2] The abstract's claim that 'all models score below 0.60 on average in each law category' is internally contradicted by Table 2, where Qingying receives 0.63 on Phenomenon Principles. Observation 4.1 narrows the claim to 'basic Newtonian and conservation laws,' but the abstract and the contribution bullet still assert the stronger statement over 'each law category' and 'every law category.' This is a factual inconsistency in the paper's headline result, not merely a wording preference. Either the claim must be restricted to the laws for which it holds, as Observation 4.1 does, or the data must be revisited; in all cases the abstract, the contributions list, and Observation 4.1 must be made mutually consistent.
- [Section 4.3, Table 3] The counterfactual robustness study's interpretation is not supported by its design. The premise that a model with genuine physical reasoning 'should understand how to generate videos that violate some specific physical laws' is an unargued assumption: a model that generates physically plausible behavior even when instructed to produce an impossibility could instead be exhibiting a beneficial physics prior, and the counterfactual prompts vary in how unambiguously they specify the required violation. Moreover, the Section 3.3 rubric is defined for adherence to the target law, with Level 4 meaning the video 'fully and accurately conforms to the law,' so applying the same scale to measure deliberate violations leaves the scores with unclear semantics under counterfactual instructions. Consequently, Observations 4.5 and 4.6, which conclude that models 'demonstrate an inability to understand impossible physics' and that their compliance 'is rooted in memorized patterns,' overreach beyond the evidence. The experiment needs a counterfactual-specific scoring rubric or a stated control condition, and the corresponding conclusions should be softened accordingly.
minor comments (5)
- [Section 4.3] In the paragraph following Table 3, the reference 'Table 4.1' should be 'Table 2', which is the actual table containing the overall compliance scores.
- [Section 4.3] There are multiple typos and grammatical errors in this section, including 'vaccum', 'without and force', 'liinear motions', and 'Threrefore'; these should be corrected.
- [Figures 3-5] Figures 3 through 5 appear in the manuscript as unreadable character-encoding sequences (for example, '/uni00000031/uni00000048/...'), and the hint-level results are reported only as figures with no accompanying table, per-condition standard errors, or number of videos; please replace them with legible figures or provide the underlying per-law numerical results so that Observation 4.4 can be verified.
- [Section 3.1 and Appendix A] The models are evaluated at different resolutions and durations (e.g., Mochi-1 at 480p, LTX Video at 512p, Sora at 720p, and SD Video at 4 seconds), so resolution and clip length are potential confounds when comparing physics compliance across models; the authors should either match generation settings where possible or report scores separately by configuration.
- [Section 3.2 and Contributions] The full set of 84 prompts is not provided in the paper or appendix; please include the complete prompt list so the benchmark is reproducible. In addition, the text alternates between 'first-principled' and 'first-principles', and the contributions bullet 'a first first-principled benchmark' contains a typo that should be fixed.
Circularity Check
No circularity: T2VPhysBench is an empirical human-evaluation benchmark; its headline scores are summary statistics of annotations, not derivations from fitted inputs or self-citations.
full rationale
This paper reports an empirical benchmark and does not contain a derivation chain that could reduce to its own inputs. The central claim, that models score below 0.60 on physical law compliance, is an arithmetic summary of Table 2, which is produced by averaging four-level human ratings over prompts and annotators as described in Section 3.3. The mapping from rating levels to scores in [0,1] is a measurement convention, not a function defined in terms of the target conclusion, so it is not self-definitional. No parameter is fitted to a subset of data and then renamed as a prediction; the experiments directly measure generated videos against a rubric. The only self-citation, [GHH+25], appears in the related-work discussion of object counting benchmarks and is not load-bearing for any claim in this paper. The counterfactual study in Section 4.3 rests on an assumption about what genuine physical reasoning should produce, which may be debatable, but it is not circular: the premise does not define the observed scores into existence. The abstract's statement that 'all models score below 0.60 on average in each law category' is internally inconsistent with Table 2, where Qingying scores 0.63 on Phenomenon Principles, and Observation 4.1 narrows the claim to 'basic Newtonian and conservation laws'; however, this is an accuracy or consistency problem, not circularity. Appendix B explicitly acknowledges that 'our study is entirely empirical,' confirming that there is no formal derivation whose conclusion is assumed among its premises. No self-definitional step, fitted input called a prediction, load-bearing self-citation, imported uniqueness theorem, ansatz smuggled via citation, or renaming of a known result was found. The appropriate finding is therefore no significant circularity, with score 0.
Assumptions & free parameters
free parameters (1)
- Human rating score mapping =
Level 1=0.0, Level 2=0.25, Level 3=0.5, Level 4=1.0
assumptions (4)
- domain assumption The ratings of three annotators are a reliable and unbiased measure of physical correctness.
- domain assumption A model with genuine physical reasoning should intentionally generate physically impossible videos when such counterfactual behavior is requested.
- ad hoc to paper Each of the 84 prompts isolates exactly one target physical law.
- domain assumption The 12 chosen laws and 7 prompts per law constitute a representative first-principles coverage of physics for video generation.
Cite this review
Pith. "Pith review of T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation." pith.science (2026). https://pith.science/paper/FW4BIQPI
@misc{pith2026250500337,
author = {Pith},
title = {Pith review of: T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/FW4BIQPI}},
note = {Machine review of arXiv:2505.00337}
}
read the original abstract
Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following, and have become central to digital art creation and user engagement online. Yet, despite these advancements, their ability to respect fundamental physical laws remains largely untested: many outputs still violate basic constraints such as rigid-body collisions, energy conservation, and gravitational dynamics, resulting in unrealistic or even misleading content. Existing physical-evaluation benchmarks typically rely on automatic, pixel-level metrics applied to simplistic, life-scenario prompts, and thus overlook both human judgment and first-principles physics. To fill this gap, we introduce \textbf{T2VPhysBench}, a first-principled benchmark that systematically evaluates whether state-of-the-art text-to-video systems, both open-source and commercial, obey twelve core physical laws including Newtonian mechanics, conservation principles, and phenomenological effects. Our benchmark employs a rigorous human evaluation protocol and includes three targeted studies: (1) an overall compliance assessment showing that all models score below 0.60 on average in each law category; (2) a prompt-hint ablation revealing that even detailed, law-specific hints fail to remedy physics violations; and (3) a counterfactual robustness test demonstrating that models often generate videos that explicitly break physical rules when so instructed. The results expose persistent limitations in current architectures and offer concrete insights for guiding future research toward truly physics-aware video generation.
Figures
Figures from the paper (27 more)
Forward citations
Cited by 5 Pith papers
-
Detecting AI-Generated Video: A Vision-Language Dual-View Survey
AIGC-V detection should be treated as factual fidelity verification and organized by a four-layer vision-language dual-view taxonomy spanning cues, motion, cross-modal consistency, and world-level reasoning.
-
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.
-
RoboScape: Physics-informed Embodied World Model
RoboScape jointly learns RGB video, depth, and keypoint-token consistency in one autoregressive world model, improving video quality, geometry, action control, synthetic-data policy training, and policy evaluation for...
-
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
With teacher forcing that reveals pairwise products of the relevant bits, one gradient step lets a single-head attention recover the support of a k-bit AND/OR; the paper's claimed end-to-end hardness lower bound is in...
-
Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse
A residual self-attention network with all weight entries bounded by a small η can be approximated by one layer to error O(η)‖X‖∞, so skip connections do not prevent layer collapse.
Reference graph
Works this paper leans on
-
[1]
Cosmos world foundation model platform for physical ai
[AAB+25] Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575 ,
-
[4]
Conditional gan with discriminative filter generation for text-to-video synthe- sis
[BMB+19] Yogesh Balaji, Martin Renqiang Min, Bing Bai, Rama Chellappa, and Hans Peter Graf. Conditional gan with discriminative filter generation for text-to-video synthe- sis. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, pages 1995–2001. International Joint Conferences on Artificial Intelligence...
work page 1995
-
[5]
Tc-bench: Benchmarking temporal compositionality in text-to-video and image-to-video generation
[FLS+24] Weixi Feng, Jiachen Li, Michael Saxon, Tsu-jui Fu, Wenhu Chen, and William Yang Wang. Tc-bench: Benchmarking temporal compositionality in text-to-video and image-to-video generation. arXiv preprint arXiv:2406.08656 ,
-
[9]
Genai-bench: A holistic benchmark for com- positional text-to-visual generation
[LLP+24] Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Emily Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. Genai-bench: A holistic benchmark for com- positional text-to-visual generation. In Synthetic Data for Computer Vision Work- shop@ CVPR 2024 ,
work page 2024
-
[10]
Exploring the evolution of physics cognition in video generation: A survey
[LWW+25] Minghui Lin, Xiang Wang, Yishan Wang, Shu Wang, Fengqi Dai, Pengxiang Ding, Cunxiang Wang, Zhengrong Zuo, Nong Sang, Siteng Huang, et al. Exploring the evolution of physics cognition in video generation: A survey. arXiv preprint arXiv:2503.21765,
-
[11]
[MCS+25] Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, and Robert Geirhos. Do generative video models learn physical principles from watching videos? arXiv preprint arXiv:2501.09038,
-
[12]
Towards world simulator: Craft- ing physical commonsense-based benchmark for video generation
[MLT+24] Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo. Towards world simulator: Craft- ing physical commonsense-based benchmark for video generation. arXiv preprint arXiv:2410.05363,
-
[13]
Phybench: A phys- ical commonsense benchmark for evaluating text-to-image models
[MSL+24] Fanqing Meng, Wenqi Shao, Lixin Luo, Yahong Wang, Yiran Chen, Quanfeng Lu, Yue Yang, Tianshuo Yang, Kaipeng Zhang, Yu Qiao, et al. Phybench: A phys- ical commonsense benchmark for evaluating text-to-image models. arXiv preprint arXiv:2406.11802,
Show all 23 references
-
[14]
Dreamingv2: Reinforcement learning with discrete world models without reconstruction
[OT22] Masashi Okada and Tadahiro Taniguchi. Dreamingv2: Reinforcement learning with discrete world models without reconstruction. In 2022 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS) , pages 985–991. IEEE,
2022
-
[19]
Modelscope text-to-video technical report
[WYC+23] Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571 ,
-
[20]
VideoCLIP: Contrastive pre- training for zero-shot video-text understanding
[XGH+21] Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. VideoCLIP: Contrastive pre- training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in...
2021
-
[21]
Cogvideox: Text-to-video diffusion models with an expert transformer
[YTZ+24] Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072,
-
[22]
In Section A, we present the details of each evaluated model
16 Appendix Roadmap. In Section A, we present the details of each evaluated model. In Section B, we discuss the limitations of this paper. In Section C, we illustrate the societal implications of this work. In Section D, we show detailed video examples. A Implementation Detail...
2024
-
[23]
It uses Deepseek-R1 for prompt enhancement, and provides aspect ratio choices including 16:9, 21:9, 4:3, 1:1, 3:4, and 9:16
It comes in four variants: Video S2.0, Video S2.0 Pro, Video P2.0 Pro, and Video 1.2. It uses Deepseek-R1 for prompt enhancement, and provides aspect ratio choices including 16:9, 21:9, 4:3, 1:1, 3:4, and 9:16. Video S2.0, Video S2.0 Pro, and Video P2.0 Pro are able to create ...
2023
-
[1995]
Wisa: World simulator assistant for physics-aware text-to-video generation
[WMC+25] Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, et al. Wisa: World simulator assistant for physics-aware text-to-video generation. arXiv preprint arXiv:2503.08153 ,
-
[2016]
T2v-compbench: A comprehensive benchmark for compositional text-to-video generation
[SHL+24] Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2v-compbench: A comprehensive benchmark for compositional text-to-video generation. arXiv preprint arXiv:2407.14505 ,
-
[2019]
Learning a driving simulator
[SH16] Eder Santana and George Hotz. Learning a driving simulator. arXiv preprint arXiv:1608.01230,
-
[2020]
Ltx-video: Realtime video latent diffusion
[HCB+24] Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Ei- tan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103 ,
-
[2021]
Physics informed deep learning (part i): Data-driven solutions of nonlinear partial differential equations
[RPK17] Maziar Raissi, Paris Perdikaris, and George Em Karniadakis. Physics informed deep learning (part i): Data-driven solutions of nonlinear partial differential equations. arXiv preprint arXiv:1711.10561 ,
-
[2022]
World models.arXiv preprint arXiv:1803.10122,
[HS18] David Ha and J¨ urgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122,
-
[2023]
Videophy: Eval- uating physical commonsense for video generation
[BLX+25] Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover. Videophy: Eval- uating physical commonsense for video generation. In Workshop on Video-Language Models @ NeurIPS 2024 ,
2024
-
[2024]
Can you count to nine? a human evaluation benchmark for counting limits in modern text-to-video models
[GHH+25] Xuyang Guo, Zekai Huang, Jiayan Huo, Yingyu Liang, Zhenmei Shi, Zhao Song, and Jiahao Zhang. Can you count to nine? a human evaluation benchmark for counting limits in modern text-to-video models. arXiv preprint arXiv:2504.04051 ,
-
[2025]
Stable video diffusion: Scaling latent video diffusion models to large datasets
[BDK+23] Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kil- ian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127,
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.