Pith. sign in

REVIEW 2 major objections 4 minor 115 references

The paper argues that video models can generate visually plausible motion while failing to bind physical quantities, select the governing law, or follow law-consistent dynamics, and that Apple-π is the first benchmark to localize these fail

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 21:02 UTC pith:IETGGOCN

load-bearing objection A genuinely useful law-grounded benchmark with a smart three-stage protocol, but the headline Deduction bottleneck rests on a temporal-alignment assumption the authors flag but never stress-test. the 2 major comments →

arxiv 2607.16401 v1 pith:IETGGOCN submitted 2026-07-17 cs.CV

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

classification cs.CV
keywords video generationworld modelsphysical reasoning benchmarkclassical mechanicschain-of-frameslaw-grounded evaluationvideo understandingphysical intelligence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to test whether video-generation models reason about physics the way a scientist would — by perceiving quantities, formulating the governing law, and deducing law-consistent dynamics — rather than merely generating visually plausible motion. It introduces Apple-π, a benchmark built on Orchard, a 400-video dataset of classical-mechanics scenes, and evaluates every model through five subtracks that isolate Perception, Formulation, and Deduction. The central finding is that current video models remain far from reliable law-grounded world simulators: the strongest video model scores 0.473 overall, and even the strongest unified understanding-generation models hover around 0.40 on the Deduction stage. The paper argues that visual plausibility does not imply law grounding, and that explicit understanding modules and reasoning-oriented supervision are needed next. A sympathetic reader would care because the benchmark gives a way to localize where a model's physical reasoning breaks down instead of treating output plausibility as evidence of physical understanding.

Core claim

Apple-π's discovery: a video can look physically plausible while the model has failed to bind quantities, select the governing law, or obey it frame by frame. The benchmark makes each step testable — reading annotations, segmenting objects, choosing a law, predicting a state, generating dynamics — and on its 400 Orchard cases the best video model scores 0.473, the strongest unified model 0.704, with Deduction near 0.40 even for leaders. Stage-resolved scores reveal a Perception-to-Formulation-to-Deduction bottleneck, weaker multi-law than single-law performance, and lower real- than simulated-video scores. Large-scale video training supplies priors, not dependable law-grounded intelligence.

What carries the argument

The central mechanism is the Apple-π protocol itself: a chain-of-frames prompt that evolves an infographic-annotated first frame into a video, so the generated video serves as the model's visible reasoning trace. The protocol splits scientific reasoning into Perception (reading quantities or segmenting objects), Formulation (selecting a symbolic law or predicting a target state), and Deduction (generating the full trajectory), and scores each with a hybrid of multimodal-judge rubric scoring and physics-law-grounded objective measures such as velocity error, masked PSNR, and spatiotemporal IoU against law-predicted ground truth. Orchard's two-level taxonomy — single-law tasks for clean diagno

Load-bearing premise

The Deduction scoring assumes every generated video is a continuous trajectory spanning exactly the requested physical duration, so a model that returns a slower or longer clip than the prompt asks for is systematically penalized on frame-aligned physics metrics even when its motion is law-correct.

What would settle it

Re-score the Deduction subtrack after matching generated frames by the model's own output duration instead of forcing the decoded clip onto the requested physical interval; if video-model Deduction scores rise toward their Formulation scores, the reported stage bottleneck is substantially an artifact of temporal normalization. A second check: take a video model that internally computes trajectories from the annotated quantities and see whether its spatiotemporal IoU and velocity accuracy improve when the prompt's requested duration equals the model's native container length.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If Apple-π's finding holds, 'world model' claims for current video generators are overstated: generating plausible motion is not the same as generating law-consistent motion.
  • The Perception-to-Deduction bottleneck implies that improving video generation alone will not close the gap; models need explicit understanding modules and physical-reasoning data.
  • The multi-law deficit implies that future systems must carry state variables across law transitions, not just imitate isolated motion priors.
  • The Sim-to-Real gap implies that law-grounding must be robust to visual appearance, not just to physical equations.
  • The benchmark provides a reusable stage-resolved diagnostic: any new video model can be scored not only on whether it fails, but where it fails.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: The temporal-normalization rule (every decoded output is scored as covering exactly the requested duration) may be inflating the reported Deduction gap; a duration-invariant re-scoring would show how much of the bottleneck is an alignment artifact rather than a physics failure.
  • Inference: The same Perception–Formulation–Deduction protocol could be transplanted to other domains with closed-form or simulable ground truth (fluid flow, deformable bodies, electromagnetism); the benchmark's lasting contribution may be the diagnostic protocol rather than the mechanics dataset.
  • Inference: If explicit understanding is what lifts unified models, a natural next experiment is adding perception and formulation supervision to a video generator and measuring whether Deduction improves; the paper leaves this test open.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper introduces Apple-π, a benchmark for evaluating video-generation models as law-grounded physical reasoners. It contributes Orchard, a 400-video dataset of classical-mechanics scenarios with simulator, self-recorded, and internet sources; a three-stage protocol (Perception, Formulation, Deduction) with five subtracks; and a hybrid evaluation suite combining MLLM-judge scoring with physics-law-grounded objective metrics. Eleven models are benchmarked. The central finding is that current video models remain far from reliable law-grounded world simulators: the best video model scores 0.473 overall, while the strongest unified understanding-generation models reach 0.704. The paper also reports a Perception→Formulation→Deduction bottleneck, weaker multi-law state transfer, and a Sim-to-Real gap.

Significance. If the quantitative results are robust, Apple-π is a valuable diagnostic resource: it is one of the first benchmarks to decompose video-model evaluation into auditable reasoning stages and to anchor ground truth in explicit physical laws rather than in model outputs. Strengths include the law-first dataset design with instrumented simulator ground truth, the separation of single-law and multi-law cases, multi-pass annotation with reported inter-annotator agreement, a cross-check with a second MLLM judge, and protocol ablations showing that the infographic interface is not a shortcut. The main risk is the temporal-alignment assumption used for Deduction scoring, which is load-bearing for the headline bottleneck claim and needs a sensitivity analysis before the central conclusions can be fully trusted.

major comments (2)
  1. [Section D.2, Eqs. (2)–(3); Limitations H] The Deduction metrics are computed after resampling every generated clip to exactly the requested physical duration Tc via bf_gen=N_gen/Tc. A model that returns a longer/shorter container with physically correct but time-scaled motion is compared against the wrong GT frames, so Spatial IoU, Spatiotemporal IoU, and velocity accuracy (Eq. 8) can be systematically low. Because Deduction carries weight 0.60 in Eq. (12), the headline 'Deduction is the hardest stage' and the best-video-model score 0.473 rest on this assumption. The paper acknowledges the issue in Limitations H but does not quantify sensitivity. Please report the distribution of generated container lengths/fps, recompute Deduction scores under alternative alignment rules (raw duration, per-model container length, or time-invariant variants), and show that the stage ordering and model ranking persist. Without this evidence, the
  2. [Section 4.1, Table 12, Appendix D.1] For Deduction, unified understanding-generation models are evaluated on timestamped keyframes at Tc, while video models generate dense videos and are sampled after time rescaling. This is a different task: keyframe generation removes temporal coherence, motion smoothness, and frame-count constraints that video models face. The comparison supporting the 'explicit understanding before generation' takeaway (Takeaway 1) and the unified models' higher overall scores is therefore confounded by output format. Please add a controlled comparison (e.g., dense-video outputs from unified models, or keyframe-only evaluation for video models) or explicitly bound the effect of output format on the reported 0.473-vs-0.704 gap.
minor comments (4)
  1. [Table 2] Several cells are typeset without spaces (e.g., Wan2.2 row '0.2000.310'; HunyuanVideo row '0.1550.128'). Please fix so the table is readable.
  2. [Throughout] Scores are reported as point estimates. Given 400 cases and 3 rollouts, please include standard errors or bootstrap confidence intervals for key comparisons (model rankings, stage funnel, pillar/source differences).
  3. [Section D.2] In the example, after computing bf_gen, it would clarify the assumption to state explicitly that generated frame 96 corresponds to 4 s of raw time, not 2 s, so the reader sees the rescaling effect.
  4. [Limitations H] The temporal-alignment caveat appears only in the appendix. Since it directly affects the headline Deduction result, please reference it (or a short version of it) in Section 4.2 where the 0.473 figure is introduced.

Circularity Check

0 steps flagged

No significant circularity: Apple-PI is a benchmark whose ground truth comes from independent physics, not from the models being evaluated.

full rationale

Apple-PI is an empirical benchmark rather than a derivation. The Orchard ground-truth trajectories are generated from explicit Newtonian equations (e.g., h = 1/2 gt^2, r(t) = r0 + v0 t + 1/2 gt^2, momentum conservation with restitution e) and from simulator/engine state, independent of the evaluated models. Evaluation metrics compare model outputs with these externally fixed GT frames; no benchmark parameter (rubric weights, timestamps, formula choices) is fitted to model outputs. The 'prediction' being scored is the model's generated video, not a quantity derived from the benchmark's own scoring rule. Thus there is no fitted input renamed as prediction and no definitional equivalence. The paper's self-citations (e.g., ref [58] for VBVR-Dataset, refs [18,55] for chain-of-frames) are contextual and not load-bearing for the central finding: even if those citations were absent, the benchmark's construction and measurements stand on the independently specified physics. The only substantive caveat is the D.2 temporal normalization rule, which treats every generated video as covering exactly the requested physical duration [0, Tc] regardless of the provider's nominal container length; as the authors acknowledge in Limitations H, this assumes the generated clip represents requested physical time rather than the provider's nominal duration. That is a validity/robustness concern that could bias the reported Deduction scores, but it is not circular: the aligned frames are still compared against independent GT trajectories derived from physics, and the rule is applied uniformly rather than being fitted to model performance. No actual circular step was found.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 0 invented entities

The paper introduces no new physical entities. The main subjective degrees of freedom are the hand-chosen rubric weights and metric transforms, which are design parameters rather than fitted parameters. The central physical laws are standard external input, not derived from the benchmark. The more consequential assumptions are interpretational: chain-of-frames as reasoning trace, MLLM judge reliability, SAM3 mask quality, and single-law isolation in real-world clips.

free parameters (5)
  • Deduction score group weights = 0.20 integrity, 0.20 fidelity, 0.60 physics
    Hand-chosen in §E.3 to emphasize physics over visual quality; changes the absolute Deduction scores and could affect model ordering if altered.
  • Perception-Text group weights = 0.50 content, 0.30 layout, 0.20 style
    Hand-chosen in §E.1.1 with no empirical justification; content is weighted highest even though some content criteria require exact font matching.
  • Formulation-Text group weights = 0.20 option, 0.30 formula, 0.40 substitution, 0.10 presentation
    Hand-chosen in §E.1.1; substitution correctness is weighted highest, which may interact with OCR ability rather than pure physics reasoning.
  • Velocity accuracy transform = Svel = 1/(1+ev)
    Arbitrary monotone mapping from mean velocity error to [0,1] in §E.2; no principled calibration is given.
  • PSNR clipping = min(PSNR/40,1)
    Ad hoc normalization in §E.2; the 40 dB threshold is not derived from perceptual or physical considerations.
axioms (5)
  • standard math Newtonian mechanics equations (free fall, projectile, momentum conservation, etc.) are the correct ground-truth laws.
    Used throughout §3.1 and Table 5 to generate ground-truth trajectories and answer keys.
  • domain assumption A generated video can be interpreted as a visible reasoning trace (chain-of-frames).
    This is the core protocol assumption in §3.2.1, based on prior work [18,55]; the benchmark depends on this interpretability to claim it evaluates 'thinking with video'.
  • domain assumption MLLM judge (Gemini 3 Flash at temperature 0) reliably scores the rubric dimensions.
    Used for all subjective scores in §4.1 and Appendix E; cross-checked with Qwen3-VL, but no judge is guaranteed to be a perfect physical or perceptual oracle.
  • domain assumption SAM3 segmentation yields comparable motion masks for generated and GT videos.
    All Deduction IoU metrics in §E.2 depend on SAM3 masks being accurate for both generated and ground-truth clips.
  • domain assumption Single-law real-world cases approximately isolate one dominant physical principle.
    Self-recorded and internet-sourced clips may contain unmodeled effects such as air resistance or micro-friction; the paper acknowledges this in Limitations H.

pith-pipeline@v1.3.0-alltime-deepseek · 29721 in / 9763 out tokens · 126637 ms · 2026-08-01T21:02:27.282202+00:00 · methodology

0 comments
read the original abstract

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce Apple-PI, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-PI comprises three components. 1) Orchard: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. 2) Benchmark Protocol: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses chain-of-frames prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. 3) Evaluation Suite: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-PI as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.

Figures

Figures reproduced from arXiv: 2607.16401 by Hao Li, Kairui Hu, Lei Yang, Ruisi Wang, Runmao Yao, Shulin Tian, Weichen Fan, Yuhao Dong, Yukang Cao, Zhaoxi Chen, Zhongang Cai, Ziang Cao, Ziqi Huang, Ziwei Liu.

Figure 1
Figure 1. Figure 1: Apple-π at a glance. Apple-π benchmarks law-grounded physical intelligence in video models through chain-of-frames traces over Perception, Formulation, and Deduction. Built on the 400-video Orchard mechanics dataset, it combines a hybrid evaluation suite of MLLM-based and physics-law-grounded metrics to diagnose where reasoning fails. Abstract Modern video generation models are increasingly hailed as emerg… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of Orchard. Left: Per-task data-source breakdown across simulated, self￾recorded, and Internet-sourced videos. Right: Two-level task taxonomy organizing 400 cases into three single-law pillars and a multi-law composition branch. Data sources. Orchard draws from three complementary origins, with the per-task source break￾down shown in the left panel of [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Apple-π benchmark protocol. An infographic-annotated first frame and chain-of-frames prompt elicit five reasoning subtracks: Perception-Text for reading quantities, Perception-Graphic for segmenting objects, Formulation-Text for selecting laws, Formulation-Graphic for predicting target states, and Deduction for generating full dynamics. in total. All five subtracks share a common input format: an annotated… view at source ↗
Figure 4
Figure 4. Figure 4: Stage-, pillar-, and source-resolved analysis. (a,b) Stage-wise results show a decline from Perception to Formulation to Deduction. (c) Pillar-wise results compare single- and multi-law cases. (d) Source-wise results show the Sim-to-Real gap. show a clear difficulty order: Perception is easiest, Formulation is harder, and Deduction is hardest. This hierarchy follows the protocol: Perception mainly requires… view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative failure analysis. The model preserves annotations and objects, but misbinds the initial-velocity cue, yielding a wrong target state and trajectory. To complement the quantitative results, [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

115 extracted references · 34 linked inside Pith

  1. [1]

    Videophy: Evaluating physical commonsense for video generation, 2024

    Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover. Videophy: Evaluating physical commonsense for video generation, 2024. URLhttps://arxiv.org/abs/2406.03520

  2. [2]

    Videophy-2: A challenging action-centric physical commonsense evaluation in video generation, 2025

    Hritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg, Aditya Grover, and Kai-Wei Chang. Videophy-2: A challenging action-centric physical commonsense evaluation in video generation, 2025. URLhttps://arxiv.org/abs/2503.06800

  3. [3]

    Bear, Elias Wang, Damian Mrowca, Felix J

    Daniel M. Bear, Elias Wang, Damian Mrowca, Felix J. Binder, Hsiao-Yu Fish Tung, R. T. Pramod, Cameron Holdaway, Sirui Tao, Kevin Smith, and Fan-Yun Sun et al. Physion: Evaluating physical prediction from vision in humans and machines, 2022. URLhttps://arxiv.org/abs/2106.08261

  4. [4]

    Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments, 2025

    Florian Bordes, Quentin Garrido, Justine T Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux. Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments, 2025. URL https://arxiv.org/abs/2506.09849

  5. [5]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, and Eric Luhman et al. Video generation models as world simulators. 2024. URL https://openai.com/research/video-generation-models-as-world-simulators

  6. [6]

    Genie: Generative interactive environ- ments

    Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, and Chris Apps et al. Genie: Generative interactive environ- ments. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors,Proceedings o...

  7. [7]

    Physx-3d: Physical-grounded 3d asset generation,

    Ziang Cao, Zhaoxi Chen, Liang Pan, and Ziwei Liu. Physx-3d: Physical-grounded 3d asset generation,

  8. [8]

    Physx-anything: Simulation-ready physical 3d assets from single image, 2025

    Ziang Cao, Fangzhou Hong, Zhaoxi Chen, Liang Pan, and Ziwei Liu. Physx-anything: Simulation-ready physical 3d assets from single image, 2025. URLhttps://arxiv.org/abs/2511.13648

  9. [9]

    Physx-omni: Unified simulation-ready physical 3d generation for rigid, deformable, and articulated objects, 2026

    Ziang Cao, Yinghao Liu, Haitian Li, Runmao Yao, Fangzhou Hong, Zhaoxi Chen, Liang Pan, and Ziwei Liu. Physx-omni: Unified simulation-ready physical 3d generation for rigid, deformable, and articulated objects, 2026. URLhttps://arxiv.org/abs/2605.21572

  10. [10]

    Tivibench: Benchmarking think-in-video reasoning for video generative models, 2025

    Harold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu, Zihan Wang, Sirui Chen, Wenkai Cheng, Kanghao Chen, Hongfei Zhang, and Zixin Zhang et al. Tivibench: Benchmarking think-in-video reasoning for video generative models, 2025. URLhttps://arxiv.org/abs/2511.13704

  11. [11]

    Feng, and Yiannis Aloimonos

    Jingxi Chen, Zongxia Li, Zhichao Liu, Guangyao Shi, Xiyang Wu, Fuxiao Liu, Cornelia Fermuller, Brandon Y . Feng, and Yiannis Aloimonos. First frame is the place to go for video content customization,

  12. [12]

    Unit: Unified multimodal chain-of-thought test-time scaling, 2026

    Leon Liangyu Chen, Haoyu Ma, Zhipeng Fan, Ziqi Huang, Animesh Sinha, Xiaoliang Dai, Jialiang Wang, Zecheng He, Jianwei Yang, and Chunyuan Li et al. Unit: Unified multimodal chain-of-thought test-time scaling, 2026. URLhttps://arxiv.org/abs/2602.12279

  13. [13]

    Physbench: Benchmarking and enhancing vision-language models for physical world understanding, 2025

    Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Guizilini, and Yue Wang. Physbench: Benchmarking and enhancing vision-language models for physical world understanding, 2025. URL https://arxiv. org/abs/2501.16411

  14. [14]

    Emerging properties in unified multimodal pretraining, 2025

    Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, and Ziang Song et al. Emerging properties in unified multimodal pretraining, 2025. URL https://arxiv.org/abs/2505.14683

  15. [15]

    More: Motion-aware feed-forward 4d reconstruction transformer

    Juntong Fang, Zequn Chen, Weiqi Zhang, Donglin Di, Xuancheng Zhang, Chengmin Yang, and Yu-Shen Liu. More: Motion-aware feed-forward 4d reconstruction transformer. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 28914–28924, June 2026. 10

  16. [16]

    Seedance 1.0: Exploring the boundaries of video generation models, 2025

    Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, and Xiaojie Li et al. Seedance 1.0: Exploring the boundaries of video generation models, 2025. URL https://arxiv.org/abs/2506.09113

  17. [17]

    Intuitive physics understanding emerges from self-supervised pretraining on natural videos, 2025

    Quentin Garrido, Nicolas Ballas, Mahmoud Assran, Adrien Bardes, Laurent Najman, Michael Rabbat, Emmanuel Dupoux, and Yann LeCun. Intuitive physics understanding emerges from self-supervised pretraining on natural videos, 2025. URLhttps://arxiv.org/abs/2502.11831

  18. [18]

    Chain-of-frames: Advancing video understanding in multimodal llms via frame-aware reasoning, 2025

    Sara Ghazanfari, Francesco Croce, Nicolas Flammarion, Prashanth Krishnamurthy, Farshad Khorrami, and Siddharth Garg. Chain-of-frames: Advancing video understanding in multimodal llms via frame-aware reasoning, 2025. URLhttps://arxiv.org/abs/2506.00318

  19. [19]

    Nano banana 2: Combining pro capabilities with lightning-fast speed

    Google. Nano banana 2: Combining pro capabilities with lightning-fast speed. https://blog.google/ innovation-and-ai/technology/ai/nano-banana-2/ , February 2026. The Keyword. Accessed: 2026-05-03

  20. [20]

    phyworldbench

    Jing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan, Fangrui Zhu, Daniel Hong, Yue Fan, Qianqi Yan, Kaiwen Zhou, and Ming-Yu Liu et al. "phyworldbench": A comprehensive evaluation of physical realism in text-to-video models, 2026. URLhttps://arxiv.org/abs/2507.13428

  21. [21]

    Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark, 2025

    Ziyu Guo, Xinyan Chen, Renrui Zhang, Ruichuan An, Yu Qi, Dongzhi Jiang, Xiangtai Li, Manyuan Zhang, Hongsheng Li, and Pheng-Ann Heng. Are video models ready as zero-shot reasoners? an empirical study with the mme-cof benchmark, 2025. URLhttps://arxiv.org/abs/2510.26802

  22. [22]

    Recurrent world models facilitate policy evolution.Advances in neural information processing systems, 31, 2018

    David Ha and Jürgen Schmidhuber. Recurrent world models facilitate policy evolution.Advances in neural information processing systems, 31, 2018

  23. [23]

    Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022. URLhttps://arxiv.org/abs/2205.15868

  24. [24]

    Benchmarking scientific understanding and reasoning for video generation using videoscience-bench, 2025

    Lanxiang Hu, Abhilash Shankarampeta, Yixin Huang, Zilin Dai, Haoyang Yu, Yujie Zhao, Haoqiang Kang, Daniel Zhao, Tajana Rosing, and Hao Zhang. Benchmarking scientific understanding and reasoning for video generation using videoscience-bench, 2025. URLhttps://arxiv.org/abs/2512.02942

  25. [25]

    Visual sketchpad: Sketching as a visual chain of thought for multimodal language models

    Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Krishna. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Processing Systems, volume 37, p...

  26. [26]

    Vbench: Comprehensive benchmark suite for video generative models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, and Nattapol Chanpaisit et al. Vbench: Comprehensive benchmark suite for video generative models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 21807–21818, June 2024

  27. [27]

    Vchain: Chain-of-visual- thought for reasoning in video generation, 2025

    Ziqi Huang, Ning Yu, Gordon Chen, Haonan Qiu, Paul Debevec, and Ziwei Liu. Vchain: Chain-of-visual- thought for reasoning in video generation, 2025. URLhttps://arxiv.org/abs/2510.05094

  28. [28]

    Vbench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(3):3268–3285,

    Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, and Yuming Jiang et al. Vbench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(3):3268–3285,

  29. [29]

    Hy-world 2.0: A multi-modal world model for reconstructing, generating, and simulating 3d worlds, 2026

    Team HY-World, Chenjie Cao, Xuhui Zuo, Zhenwei Wang, Yisu Zhang, Junta Wu, Zhenyang Liu, Yuning Gong, Yang Liu, and Bo Yuan et al. Hy-world 2.0: A multi-modal world model for reconstructing, generating, and simulating 3d worlds, 2026. URLhttps://arxiv.org/abs/2604.14268

  30. [30]

    How far is video generation from world model: A physical law perspective

    Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, and Jiashi Feng. How far is video generation from world model: A physical law perspective. In Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu, editors,Proceedings of the 42nd International Conference on M...

  31. [31]

    Hunyuanvideo: A systematic framework for large video generative models, 2025

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, and Jianwei Zhang et al. Hunyuanvideo: A systematic framework for large video generative models, 2025. URLhttps://arxiv.org/abs/2412.03603. 11

  32. [32]

    doi: 10.1109/TPAMI.2025.3633890

  33. [33]

    A path towards autonomous machine intelligence version 0.9

    Yann LeCun et al. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27. 2022

  34. [34]

    Thinking in frames: How visual context and test-time scaling empower video reasoning, 2026

    Chengzu Li, Zanyi Wang, Jiaang Li, Yi Xu, Han Zhou, Huanyu Zhang, Ruichuan An, Dengyang Jiang, Zhaochong An, and Ivan Vuli´c et al. Thinking in frames: How visual context and test-time scaling empower video reasoning, 2026. URLhttps://arxiv.org/abs/2601.21037

  35. [35]

    Gonzalez et al

    Dacheng Li, Yunhao Fang, Yukang Chen, Shuo Yang, Shiyi Cao, Justin Wong, Michael Luo, Xiaolong Wang, Hongxu Yin, and Joseph E. Gonzalez et al. Worldmodelbench: Judging video generation models as world models, 2025. URLhttps://arxiv.org/abs/2502.20694

  36. [36]

    What about gravity in video generation? post-training newton’s laws with verifiable rewards, 2025

    Minh-Quan Le, Yuanzhi Zhu, Vicky Kalogeiton, and Dimitris Samaras. What about gravity in video generation? post-training newton’s laws with verifiable rewards, 2025. URL https://arxiv.org/abs/ 2512.00425

  37. [37]

    Video-t1: Test-time scaling for video generation

    Fangfu Liu, Hanyang Wang, Yimo Cai, Kaiyan Zhang, Xiaohang Zhan, and Yueqi Duan. Video-t1: Test-time scaling for video generation. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 18671–18681, October 2025

  38. [38]

    V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025

    Yang Luo, Xuanlei Zhao, Baijiong Lin, Lingting Zhu, Liyao Tang, Yuqi Liu, Ying-Cong Chen, Shengju Qian, Xin Wang, and Yang You. V-reasonbench: Toward unified reasoning benchmark suite for video generation models, 2025. URLhttps://arxiv.org/abs/2511.16668

  39. [39]

    Physicsmind: Sim and real mechanics benchmarking for physical reasoning and prediction in foundational vlms and world models, 2026

    Chak-Wing Mak, Guanyu Zhu, Boyi Zhang, Hongji Li, Xiaowei Chi, Kevin Zhang, Yichen Wu, Yangfan He, Chun-Kai Fan, and Wentao Lu et al. Physicsmind: Sim and real mechanics benchmarking for physical reasoning and prediction in foundational vlms and world models, 2026. URL https://arxiv.org/abs/ 2601.16007

  40. [40]

    Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026

    Yifan Li, Yukai Gu, Yingqian Min, Zikang Liu, Yifan Du, Kun Zhou, Min Yang, Wayne Xin Zhao, and Minghui Qiu. Beyond the last frame: Process-aware evaluation for generative video reasoning, 2026. URL https://arxiv.org/abs/2512.24952

  41. [41]

    Do generative video mod- els understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 948–958, March 2026

    Saman Motamed, Laura Culp, Kevin Swersky, Priyank Jaini, and Robert Geirhos. Do generative video mod- els understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 948–958, March 2026

  42. [42]

    Isaac Sim

    NVIDIA. Isaac Sim. URLhttps://github.com/isaac-sim/IsaacSim

  43. [43]

    Thinking with images, 2025

    OpenAI. Thinking with images, 2025. URL https://openai.com/index/thinking-with-images/ . Accessed: 2026-04-20

  44. [44]

    Towards world simulator: Crafting physical commonsense-based benchmark for video generation

    Fanqing Meng, Jiaqi Liao, Xinyu Tan, Quanfeng Lu, Wenqi Shao, Kaipeng Zhang, Yu Cheng, Dianqi Li, and Ping Luo. Towards world simulator: Crafting physical commonsense-based benchmark for video generation. In Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste-Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu, editors,Proceedings of th...

  45. [45]

    SenseNova-U1: Unifying multimodal understanding and generation with NEO-Unify architecture

    OpenSenseNova. SenseNova-U1: Unifying multimodal understanding and generation with NEO-Unify architecture. https://github.com/OpenSenseNova/SenseNova-U1, April 2026. GitHub repository. Accessed: 2026-05-03

  46. [46]

    Quantiphy: A quantitative benchmark evaluating physical reasoning abilities of vision-language models,

    Li Puyin, Tiange Xiang, Ella Mao, Shirley Wei, Xinye Chen, Adnan Masood, Li Fei-fei, and Ehsan Adeli. Quantiphy: A quantitative benchmark evaluating physical reasoning abilities of vision-language models,

  47. [47]

    Intphys: A framework and benchmark for visual intuitive physics reasoning, 2020

    Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux. Intphys: A framework and benchmark for visual intuitive physics reasoning, 2020. URLhttps://arxiv.org/abs/1803.07616

  48. [48]

    Introducing ChatGPT Images 2.0

    OpenAI. Introducing ChatGPT Images 2.0. https://openai.com/index/ introducing-chatgpt-images-2-0/ , April 2026. OpenAI product release. Accessed: 2026- 05-03

  49. [49]

    Seedance 2.0: Advancing video generation for world complexity,

    Team Seedance, De Chen, Liyang Chen, Xin Chen, Ying Chen, Zhuo Chen, Zhuowei Chen, Feng Cheng, Tianheng Cheng, and Yufeng Cheng et al. Seedance 2.0: Advancing video generation for world complexity,

  50. [50]

    Phyx: Does your model have the "wits" for physical reasoning?, 2025

    Hui Shen, Taiqiang Wu, Qi Han, Yunta Hsieh, Jizhou Wang, Yuyue Zhang, Yuxin Cheng, Zijian Hao, Yuansheng Ni, and Xin Wang et al. Phyx: Does your model have the "wits" for physical reasoning?, 2025. URLhttps://arxiv.org/abs/2505.15929

  51. [51]

    URLhttps://arxiv.org/abs/2512.19526

  52. [52]

    Kling-omni technical report, 2025

    Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du, Zipeng Feng, Kun Gai, Sainan Guo, Feng Han, Jingbin He, and Kang He et al. Kling-omni technical report, 2025. URLhttps://arxiv.org/abs/2512.16776

  53. [53]

    Seedance 1.5 pro: A native audio-visual joint generation foundation model, 2025

    Team Seedance, Heyi Chen, Siyan Chen, Xin Chen, Yanfei Chen, Ying Chen, Zhuo Chen, Feng Cheng, Tianheng Cheng, and Xinqi Cheng et al. Seedance 1.5 pro: A native audio-visual joint generation foundation model, 2025. URLhttps://arxiv.org/abs/2512.13507. 12

  54. [54]

    Cof-t2i: Video models as pure visual reasoners for text-to-image generation, 2026

    Chengzhuo Tong, Mingkun Chang, Shenglong Zhang, Yuran Wang, Cheng Liang, Zhizheng Zhao, Ruichuan An, Bohan Zeng, Yang Shi, and Yifan Dai et al. Cof-t2i: Video models as pure visual reasoners for text-to-image generation, 2026. URLhttps://arxiv.org/abs/2601.10061

  55. [55]

    URLhttps://arxiv.org/abs/2604.14148

  56. [56]

    Wan: Open and advanced large-scale video generative models, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, and Jianxiao Yang et al. Wan: Open and advanced large-scale video generative models, 2025. URL https://arxiv.org/abs/2503.20314

  57. [57]

    Worldplay: Towards long-term geometric consistency for real-time interactive world modeling, 2025

    Wenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu, Zehan Wang, Zhenwei Wang, Yunhong Wang, Jun Zhang, Tengfei Wang, and Chunchao Guo. Worldplay: Towards long-term geometric consistency for real-time interactive world modeling, 2025. URLhttps://arxiv.org/abs/2512.14614

  58. [58]

    A very big video reasoning suite, 2026

    Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thaddäus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, and Zelong Hong et al. A very big video reasoning suite, 2026. URL https: //arxiv.org/abs/2602.20159

  59. [59]

    Advancing open-source world models, 2026

    Robbyant Team, Zelin Gao, Qiuyu Wang, Yanhong Zeng, Jiapeng Zhu, Ka Leong Cheng, Yixuan Li, Hanlin Wang, Yinghao Xu, and Shuailei Ma et al. Advancing open-source world models, 2026. URL https://arxiv.org/abs/2601.20540

  60. [60]

    Univideo: Unified understanding, generation, and editing for videos, 2026

    Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, and Wenhu Chen. Univideo: Unified understanding, generation, and editing for videos, 2026. URL https://arxiv.org/ abs/2510.08377

  61. [61]

    Thinking with video: Video generation as a promising multimodal reasoning paradigm, 2025

    Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, and Yining Zheng et al. Thinking with video: Video generation as a promising multimodal reasoning paradigm, 2025. URLhttps://arxiv.org/abs/2511.04570

  62. [62]

    Video models are zero-shot learners and reasoners, 2025

    Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini, and Robert Geirhos. Video models are zero-shot learners and reasoners, 2025. URL https://arxiv.org/abs/2509.20328

  63. [63]

    Videoscene: Distilling video diffusion model to generate 3d scenes in one step

    Hanyang Wang, Fangfu Liu, Jiawei Chi, and Yueqi Duan. Videoscene: Distilling video diffusion model to generate 3d scenes in one step. In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16475–16485, 2025. doi: 10.1109/CVPR52734.2025.01536

  64. [64]

    Hunyuanvideo 1.5 technical report, 2025

    Bing Wu, Chang Zou, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Jack Peng, Jianbing Wu, Jiangfeng Xiong, and Jie Jiang et al. Hunyuanvideo 1.5 technical report, 2025. URL https://arxiv.org/abs/ 2511.18870

  65. [65]

    Demystifing video reasoning, 2026

    Ruisi Wang, Zhongang Cai, Fanyi Pu, Junxiang Xu, Wanqi Yin, Maijunxian Wang, Ran Ji, Chenyang Gu, Bo Li, and Ziqi Huang et al. Demystifing video reasoning, 2026. URL https://arxiv.org/abs/2603. 16870

  66. [66]

    Omni-worldbench: Towards a comprehensive interaction-centric evaluation for world models, 2026

    Meiqi Wu, Zhixin Cai, Fufangchen Zhao, Xiaokun Feng, Rujing Dang, Bingze Song, Ruitian Tian, Jiashu Zhu, Jiachen Lei, and Hao Dou et al. Omni-worldbench: Towards a comprehensive interaction-centric evaluation for world models, 2026. URLhttps://arxiv.org/abs/2603.22212

  67. [67]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors,Ad- vances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Asso...

  68. [68]

    Anchoreddream: Zero-shot 360° indoor scene generation from a single view via geometric grounding, 2026

    Runmao Yao, Junsheng Zhou, Zhen Dong, and Yu-Shen Liu. Anchoreddream: Zero-shot 360° indoor scene generation from a single view via geometric grounding, 2026. URL https://arxiv.org/abs/ 2601.16532

  69. [69]

    Veo 3.1 ingredients to video: More consistency, creativity and control

    Ricky Wong. Veo 3.1 ingredients to video: More consistency, creativity and control. https://blog. google/innovation-and-ai/technology/ai/veo-3-1-ingredients-to-video/ , January 2026. Google Blog, The Keyword. Accessed: 2026-05-03

  70. [70]

    Think in strokes, not pixels: Process-driven image generation via interleaved reasoning, 2026

    Lei Zhang, Junjiao Tian, Zhipeng Fan, Kunpeng Li, Jialiang Wang, Weifeng Chen, Markos Georgopoulos, Felix Juefei-Xu, Yuxiang Bao, and Julian McAuley et al. Think in strokes, not pixels: Process-driven image generation via interleaved reasoning, 2026. URLhttps://arxiv.org/abs/2604.04746

  71. [71]

    Omnigen2: Towards instruction-aligned multimodal generation, 2026

    Chenyuan Wu, Pengfei Zheng, Ruiran Yan, Shitao Xiao, Xin Luo, Yueze Wang, Wanli Li, Xiyan Jiang, Yexin Liu, and Junjie Zhou et al. Omnigen2: Towards instruction-aligned multimodal generation, 2026. URLhttps://arxiv.org/abs/2506.18871. 13

  72. [72]

    Nguyen, Baixuan Xu, Zhaowei Wang, Jiayang Cheng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, and Tianqing Fang et al

    Tianshi Zheng, Kelvin Kiu-Wai Tam, Newt Hue-Nam K. Nguyen, Baixuan Xu, Zhaowei Wang, Jiayang Cheng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, and Tianqing Fang et al. Newtonbench: Benchmarking generalizable scientific law discovery in llm agents, 2026. URL https://arxiv.org/abs/2510.07172

  73. [73]

    Cogvideox: Text-to-video diffusion models with an expert transformer, 2025

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, and Guanyu Feng et al. Cogvideox: Text-to-video diffusion models with an expert transformer, 2025. URLhttps://arxiv.org/abs/2408.06072

  74. [75]

    Tenenbaum

    Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. Clevrer: Collision events for video representation and reasoning, 2020. URL https: //arxiv.org/abs/1910.01442

  75. [77]

    Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness, 2025

    Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Lulu Gu, Yuanhan Zhang, Jingwen He, and Wei-Shi Zheng et al. Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness, 2025. URLhttps://arxiv.org/abs/2503.21755

  76. [79]

    Physinone: Visual physics learning and reasoning in one suite, 2026

    Siyuan Zhou, Hejun Wang, Hu Cheng, Jinxi Li, Dongsheng Wang, Junwei Jiang, Yixiao Jin, Jiayue Huang, Shiwei Mao, and Shangjia Liu et al. Physinone: Visual physics learning and reasoning in one suite, 2026. URLhttps://arxiv.org/abs/2604.09415. 14 Appendix Contents A Orchard Design Principles and Data Card 16 A.1 Design Principles . . . . . . . . . . . . . ...

  77. [80]

    the correct governing formula for the case

  78. [81]

    a confusing real formula that shares symbols with the annotations but does not govern the case; 18 Table 8: Examples of issues caught and corrected during simulator QA. Fix category Cases caught Physics-duration adjustment 23 Velocity-label theoretical-value override 10 Configuration re-recording due to camera, material, or position issue 30 Table 9: Inte...

  79. [82]

    an unrelated real formula from another mechanics family

  80. [83]

    v = X. ,→XX

    a fabricated formula that is syntactically plausible but not a valid physical law. This construction distinguishes law selection from superficial symbol matching. A model that chooses a formula merely because it contains the same variables should fail on the confusing distractor. B.5 Three-Pass Review Protocol Every case undergoes a 1+2 review protocol: o...

Showing first 80 references.