{"id":"68a1cd91-5714-4525-8682-a37349533ccb","arxiv_id":"2505.00337","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A human-scored benchmark of 84 physics-law prompts finds all ten tested text-to-video models average below 0.42, indicating weak physical consistency.","lead":"This paper evaluates ten text-to-video AI models against twelve basic physics laws using a human-scored benchmark, and finds that none of them pass. It also shows that adding physics hints to prompts rarely fixes the failures, and that models struggle to generate deliberately impossible scenes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central numeric claim is not reproducible: no inter-annotator agreement or rater-error model is reported, and the abstract's 'each law category' below-0.60 claim is internally contradicted by Table 2 (Qingying, Phenomenon: 0.63).","rationale":"The reader's weakest_assumption identifies the absence of inter-annotator agreement and rater-error modeling as the load-bearing weakness; my independent pass reaches the same conclusion. Every numeric comparison in Tables 2 and 3 is a mean of three annotators' ordinal ratings, and the paper gives no evidence that this measurement is stable or unbiased. If the rating scale is noisy, the central claim 'all models score below 0.60 in each law category' could be an artifact of arbitrary thresholds; if the annotators are systematically lenient or strict, the specific values and rankings would shift. I also note a separate internal inconsistency: the abstract claims 'each law category' while Table 2 shows Qingying at 0.63 on Phenomenon Principles. This does not by itself refute the qualitative finding that models fail basic physics, but it means the precise headline as stated in the abstract is not even consistent with the paper's own table. The counterfactual study's interpretive premise is also questionable, as the reader notes, but I do not treat it as the most load-bearing element because the main compliance claim does not depend on it. The appropriate verdict remains CONDITIONAL: the paper's qualitative conclusion is plausible and consistent with prior work, but the quantitative claims require reliability analysis, correction of the abstract, and release of artifacts before they can be accepted at face value.","tokens_in":21239,"tokens_out":5557,"duration_ms":56996,"concrete_test":"Compute Krippendorff's alpha (or Cohen's kappa) from the three raw annotator ratings for every prompt-video pair, and release the per-annotator, per-law score matrices. If alpha >= 0.8 for each law category and a re-computed category maximum stays below 0.60 after correcting or re-auditing the Qingying Phenomenon entry (currently 0.63), the central claim survives. If alpha < 0.7, or if any category maximum exceeds 0.60 under an alternative aggregation such as majority vote, the paper must downgrade the headline to a qualitative finding and add error bars, and the abstract should be amended to match Observation 4.1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that 'all models score below 0.60 on average in each law category' rests entirely on the human rating protocol described in Section 3.3. That section states that three annotators independently assign one of four levels, and Section 4.1 averages 'across all prompts and annotators,' but the paper reports no inter-annotator agreement (Cohen's kappa, Krippendorff's alpha), no per-item variance, no annotator-level score distributions, and no rater-error model. All entries in Tables 2 and 3 are means of four-level ordinal ratings over 7 prompts x 3 annotators per law. With no reliability evidence, differences as small as 0.01 (Kling 0.35 vs Mochi-1 0.34) and the claimed ordering of law categories cannot be distinguished from rater noise or systematic bias. The rubric also places 'fails to demonstrate the intended behavior' (0.0) adjacent to 'clear violation of the law' (0.25), inviting arbitrary thresholding in ambiguous videos. Independently, the abstract's specific claim is internally contradicted by Table 2: Qingying scores 0.63 on Phenomenon Principles, and while Observation 4.1 narrows the claim to 'basic Newtonian and conservation laws,' the abstract and the contribution bullet still assert 'each law category.' Thus the most precise version of the headline is not self-consistent as written. Until raw annotations and agreement statistics are released, the central quantitative claim is not reproducible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces T2VPhysBench, a human-annotated benchmark for evaluating whether text-to-video generation models respect fundamental physical laws. The benchmark comprises 12 laws grouped into Newtonian principles, conservation principles, and phenomenon principles, with 84 prompts per model and 10 evaluated models spanning open and closed systems. Three studies are reported: (1) an overall compliance assessment summarized in Table 2, (2) a hint-level ablation varying prompt specificity, and (3) a counterfactual robustness test in Table 3. The headline claim is that all models score below 0.60 on average in each law category and that, collectively, models reveal a reliance on surface-pattern matching rather than genuine physical reasoning. The manuscript also presents observations on category-level difficulty, hint ineffectiveness, and counterfactual failure, and concludes with suggestions for physics-aware video generation.","tokens_in":21492,"tokens_out":6128,"duration_ms":60050,"significance":"If its empirical results are reliable, T2VPhysBench would be a useful community resource. Its strengths include the systematic coverage of twelve named physical laws, the inclusion of both open-source and commercial models, a human evaluation protocol that goes beyond automatic pixel-level metrics, and the use of a hint ablation and a counterfactual probe to interrogate failure modes. These are genuinely useful design choices, and the qualitative finding that current text-to-video models frequently violate basic physics is plausible and consistent with prior human-evaluated benchmarks such as VideoPhy. However, the quantitative claims rest entirely on a small human-rating protocol with no reported inter-annotator reliability or released annotation data, and the paper's headline 'below 0.60 in each law category' is contradicted by its own Table 2. The benchmark's contribution is therefore currently significant in conception but not yet established in its quantitative specifics.","major_comments":[{"comment":"The central quantitative result is not reproducible as reported. Section 3.3 states that three annotators independently assign a four-level score, and Section 4.1 averages across prompts and annotators, but the paper reports no inter-annotator agreement statistic (e.g., Cohen's kappa or Krippendorff's alpha), no per-item variance, no annotator-level score distributions, and the raw annotations are not released. With only 7 prompts per law and 3 annotators per cell, differences such as Kling 0.35 versus Mochi-1 0.34 in Table 2 cannot be distinguished from rater noise or systematic annotator bias. In addition, the rubric maps the ordinal levels 0.0, 0.25, 0.5, and 1.0 to equal intervals without validation, and the boundary between 'fails to demonstrate the intended physical behavior' (0.0) and 'clear violation of the law' (0.25) is especially susceptible to arbitrary thresholding in ambiguous videos. Please publish the annotation data and agreement statistics, or the numerical scores, model rankings, and category-level orderings should be treated as provisional.","section":"Section 3.3, Tables 2 and 3"},{"comment":"The abstract's claim that 'all models score below 0.60 on average in each law category' is internally contradicted by Table 2, where Qingying receives 0.63 on Phenomenon Principles. Observation 4.1 narrows the claim to 'basic Newtonian and conservation laws,' but the abstract and the contribution bullet still assert the stronger statement over 'each law category' and 'every law category.' This is a factual inconsistency in the paper's headline result, not merely a wording preference. Either the claim must be restricted to the laws for which it holds, as Observation 4.1 does, or the data must be revisited; in all cases the abstract, the contributions list, and Observation 4.1 must be made mutually consistent.","section":"Abstract and Section 4.1, Table 2"},{"comment":"The counterfactual robustness study's interpretation is not supported by its design. The premise that a model with genuine physical reasoning 'should understand how to generate videos that violate some specific physical laws' is an unargued assumption: a model that generates physically plausible behavior even when instructed to produce an impossibility could instead be exhibiting a beneficial physics prior, and the counterfactual prompts vary in how unambiguously they specify the required violation. Moreover, the Section 3.3 rubric is defined for adherence to the target law, with Level 4 meaning the video 'fully and accurately conforms to the law,' so applying the same scale to measure deliberate violations leaves the scores with unclear semantics under counterfactual instructions. Consequently, Observations 4.5 and 4.6, which conclude that models 'demonstrate an inability to understand impossible physics' and that their compliance 'is rooted in memorized patterns,' overreach beyond the evidence. The experiment needs a counterfactual-specific scoring rubric or a stated control condition, and the corresponding conclusions should be softened accordingly.","section":"Section 4.3, Table 3"}],"minor_comments":[{"comment":"In the paragraph following Table 3, the reference 'Table 4.1' should be 'Table 2', which is the actual table containing the overall compliance scores.","section":"Section 4.3"},{"comment":"There are multiple typos and grammatical errors in this section, including 'vaccum', 'without and force', 'liinear motions', and 'Threrefore'; these should be corrected.","section":"Section 4.3"},{"comment":"Figures 3 through 5 appear in the manuscript as unreadable character-encoding sequences (for example, '/uni00000031/uni00000048/...'), and the hint-level results are reported only as figures with no accompanying table, per-condition standard errors, or number of videos; please replace them with legible figures or provide the underlying per-law numerical results so that Observation 4.4 can be verified.","section":"Figures 3-5"},{"comment":"The models are evaluated at different resolutions and durations (e.g., Mochi-1 at 480p, LTX Video at 512p, Sora at 720p, and SD Video at 4 seconds), so resolution and clip length are potential confounds when comparing physics compliance across models; the authors should either match generation settings where possible or report scores separately by configuration.","section":"Section 3.1 and Appendix A"},{"comment":"The full set of 84 prompts is not provided in the paper or appendix; please include the complete prompt list so the benchmark is reproducible. In addition, the text alternates between 'first-principled' and 'first-principles', and the contributions bullet 'a first first-principled benchmark' contains a typo that should be fixed.","section":"Section 3.2 and Contributions"}],"recommendation":"major_revision","confidential_remarks":"The benchmark topic and design are well aligned with the journal's scope, and the paper has clear strengths: law-grounded prompts, a diverse model set, and three complementary studies. The main risk is empirical reproducibility: the central numeric claims depend entirely on a three-annotator protocol with no released annotations and no agreement measures. I would make public release of per-video annotations and inter-annotator statistics a condition of acceptance, and require that the abstract's headline claim be reconciled with Table 2. The counterfactual study's interpretation also needs revision, but that is fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful benchmark paper with a real internal inconsistency. The abstract claims all models score below 0.60 on average in each law category, but Table 2 shows Qingying at 0.63 on Phenomenon Principles. The qualitative conclusion — that current text-to-video models are bad at basic physics — survives that contradiction; the precise numeric claim does not.\n\nWhat's new: the 12-law structure, the hint-level ablation, and the counterfactual stress test. The hint study is the strongest part; it shows that naming the law or spelling out the mechanism rarely helps and sometimes hurts, which is a non-obvious and useful result. The counterfactual study is suggestive, though the premise that a physics-aware model should be able to deliberately violate laws is an assumption, not a validated diagnostic. I'd frame it as exploratory rather than proof of pattern matching.\n\nWhere it's soft: (1) No inter-annotator agreement. Three annotators independently rate on a four-level scale, but there is no kappa, no per-item variance, no rater model. With seven prompts per law and three raters, the 0.01 differences between Kling (0.35) and Mochi-1 (0.34) are indistinguishable from noise. (2) No error bars or significance tests anywhere in the tables. (3) The abstract overclaim noted above. (4) No released artifacts — no full prompt list, no raw annotations, no videos — so the benchmark is not yet reproducible. Appendix B mentions the scalability burden of annotation but does not acknowledge the reliability gap. On the positive side, the citation pattern is fine: they engage with VideoPhy, Physics-IQ, and related work, and the self-citation to their counting benchmark is not problematic.\n\nThe central finding is probably right: other benchmarks point the same direction, and the raw averages in Table 2 are low across the board. So the soft spots are in the measurement instrument and the claims, not in the existence of the phenomenon.\n\nVerdict: this deserves peer review, but only after the authors fix the abstract, add annotation-reliability statistics, and release the artifacts. The benchmark has real value for the T2V evaluation community; without those fixes, the quantitative claims should not be taken at face value. I'd bring it to reading group as a case study in benchmark design and evaluation hygiene, but I wouldn't cite it in my own work until the data are available.","headline":"The qualitative finding is almost certainly right, but the abstract overclaims below-0.60 across every law category, and the missing inter-annotator agreement makes the precise numeric scores noisier than the paper admits.","tokens_in":22083,"tokens_out":2246,"would_cite":false,"duration_ms":23796,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ten leading text-to-video models all score below 0.60 on every physical-law category in a new human-evaluated benchmark, with the best overall average at 0.42.","keywords":["T2VPhysBench","text-to-video generation","physical consistency","human evaluation","Newtonian mechanics","conservation laws","counterfactual prompts","benchmark"],"falsifier":"Re-annotate the generated video set with two independent panels, or use a physics-simulation checker that measures, for example, the ball's vertical acceleration in the throw prompts; if panel scores diverge strongly or the simulated trajectory disagrees with human ratings on a large share of clips, the central below-0.60 finding is rater-dependent rather than a property of the models.","tokens_in":20982,"feed_emoji":"🎬","tokens_out":4361,"duration_ms":41609,"temperature":0.7,"pith_summary":"This paper introduces a benchmark, T2VPhysBench, that tests whether text-to-video models obey twelve fundamental physical laws drawn from Newtonian mechanics, conservation principles, and phenomenological effects. Ten leading open-source and commercial models are scored by three human annotators on 84 prompts, and the paper reports that every model averages below 0.60 in each law category, with the best overall average at 0.42. It also finds that adding explicit hints about the target law rarely improves compliance and often lowers scores, and that when models are asked to generate physically impossible events they frequently comply. For a sympathetic reader the takeaway is that current video generators match prompts aesthetically but do not reason about the physical world.","feed_headline":"Video AI models fail basic physics tests, benchmark finds","feed_subtitle":"Ten top text-to-video systems score below 0.60 on every law family; the best averages 0.42.","key_machinery":"The load-bearing instrument is T2VPhysBench itself: a set of 84 prompts, seven per law, anchored to named laws rather than everyday scenario descriptions, plus a four-level human rating scale (0.0, 0.25, 0.5, 1.0) applied by three independent annotators to every video. The counterfactual arm replaces the prompts with physically impossible versions of the same scenarios, so that a model with genuine physical reasoning would be expected to produce videos that violate the named law. This design is what lets the paper attribute low scores to missing physical reasoning rather than to aesthetic quality or instruction-following failures.","core_discovery":"On its own terms, the paper's central discovery is that state-of-the-art text-to-video systems do not internally represent basic physics: across all ten models, all three law families, and all twelve laws, average human-rated compliance stays below 0.60, and conservation laws are the hardest, topping out at 0.29. Prompt engineering cannot close the gap, because naming the law or spelling out the mechanism does not reliably raise scores and sometimes lowers them. The counterfactual study shows that models will generate impossible outcomes when instructed, which the paper takes as evidence that normal-looking compliance is surface pattern matching rather than physical understanding.","pith_inferences":["An extension the authors leave implicit: the same 84-prompt protocol could be run with longer videos and higher resolutions to test whether physics failures persist with more temporal context, since the current 4-to-6-second clips may understate or overstate the gap.","One consequence not drawn in the paper is that trajectory-level automated checks, such as fitting projectile motion or collision velocities from the generated frames, could complement human ratings and turn the benchmark into a scalable regression test.","A testable prediction from the counterfactual results is that fine-tuning on physics-annotated data would improve standard-prompt scores faster than it improves counterfactual-prompt scores, because models can memorize canonical scenarios without acquiring transferable physical rules.","The per-law asymmetry may reflect training-data frequency more than physical complexity, since common scenarios like throwing a ball are scored better than rare ones like gyroscope motion; if so, data rebalancing would be the first lever rather than new physics-specific architectures."],"forward_implications":["If the scores are taken at face value, no current model can be trusted for safety-relevant video generation in robotics, autonomous driving, or scientific visualization.","Adding law-specific hints is not a workable remedy; the paper's ablation predicts that prompt refinement alone will not make these architectures physics-aware.","Conservation laws are systematically harder than Newton's laws or phenomena, so progress on physical consistency should be measured per law family rather than by a single average.","Counterfactual compliance scores being low means that instruction-following masks physical understanding; this predicts that a model's apparent realism and its physical competence can decouple."],"supporting_citations":[{"why":"Supplies the human-evaluation protocol this benchmark adopts and is the prior physics benchmark it extends.","marker":"[BLX+25]"},{"why":"Is the prior physics benchmark that relies on automated metrics, which this paper argues are insufficient.","marker":"[MCS+25]"},{"why":"Documents the physics-cognition gap in video generation that motivates the study.","marker":"[LWW+25]"},{"why":"Is a physics-aware generation system whose claimed capability the benchmark challenges.","marker":"[WMC+25]"},{"why":"Provides an earlier physical-commonsense benchmark with scenario prompts that this paper contrasts with law-derived prompts.","marker":"[MLT+24]"},{"why":"Is one of the ten evaluated models, anchoring the closed-source commercial results.","marker":"[Ope24]"},{"why":"Is the open-source model that scores highest overall, anchoring the best-case result.","marker":"[Ali25]"},{"why":"Is another evaluated closed-source model that drops sharply under counterfactual prompts, supporting the pattern-matching claim.","marker":"[Kli24]"}],"fun_headline_variants":["AI videos flunk physics: all models score under 0.60","Text-to-video AI can't obey basic laws of physics","Physics benchmark: top video AIs fail all 12 laws","Even physics hints don't fix AI video models' errors","Video AIs break physics even when told the rules"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The numerical conclusions rest on three annotators' four-level ratings being a reliable, unbiased measure of physical correctness, and the paper reports no inter-annotator agreement or rater-error analysis; if the ratings are noisy or biased, every score in the benchmark table shifts.","fun_headline_variants_meta":{"raw":{"variants":["AI videos flunk physics: all models score under 0.60","Text-to-video AI can't obey basic laws of physics","Physics benchmark: top video AIs fail all 12 laws","Even physics hints don't fix AI video models' errors","Video AIs break physics even when told the rules"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000453,"raw_usage":{"total_tokens":2274,"prompt_tokens":936,"completion_tokens":1338,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":1267}},"tokens_in":552,"tokens_out":1338,"duration_ms":8035,"temperature":1.0,"reasoning_tokens":1267,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:44:32.047217+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the generated video set with two independent panels, or use a physics-simulation checker that measures, for example, the ball's vertical acceleration in the throw prompts; if panel scores diverge strongly or the simulated trajectory disagrees with human ratings on a large share of clips, the central below-0.60 finding is rater-dependent rather than a property of the models.","supporting_citations":[],"review_version":1}