Pith. sign in

REVIEW 3 major objections 1 minor 1 cited by

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

T0 review · 3 major / 1 minor · reviewed 2026-06-29 · grok-4.3

Pith's one-line read Video diffusion models notice when time runs backward but do not grasp cause and effect like humans do.

desk verdict The paper's time-reversal benchmark for separating temporal perception from causality in video models has a practical setup but rests on unvalidated assumptions about counterfactuals and VLM labels. read the letter →

arxiv 2605.30346 v1 pith:N5A4EID6 submitted 2026-05-28 cs.CV

classification cs.CV
keywords videodiffusionmodelscausalityworldbenchmarkcounterfactualsarrowoftimeviolationexpectation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests whether video diffusion models function as world models by checking if they understand causality or only statistical timing patterns. It introduces YoCausal, which reverses real videos at no cost to create natural counterfactual examples and measures both time-direction awareness and causal reasoning separately. On 13 current models, time-direction detection turns out to be unrelated to actual causal understanding, leaving a clear shortfall compared with human performance. This distinction matters because models positioned as simulators of the physical world need reliable cause-effect reasoning to predict what happens next in new situations.

What carries the argument

YoCausal benchmark that creates natural counterfactuals by temporally reversing real videos, then computes Reverse Surprise Index for time-direction sensitivity and Causality Cognition Index to isolate genuine causal reasoning from temporal bias.

What would settle it

A model that scores equally on causal and non-causal subsets in the Causality Cognition Index or reaches human-level scores on both indices would contradict the reported gap between time perception and causal understanding.

Watch

Extended reading notes

Core claim

YoCausal is a two-level benchmark that first quantifies arrow-of-time perception through a Reverse Surprise Index based on denoising loss when videos are played backward, then applies a Causality Cognition Index that uses a vision-language model to split videos into causal and non-causal groups. Evaluation across 13 state-of-the-art video diffusion models shows that strong performance on the first index does not produce strong performance on the second, revealing that temporal pattern recognition alone does not deliver causal cognition and that current models remain far from human levels on real-world videos.

Load-bearing premise

Reversing real-world videos produces valid natural counterfactual samples, and a vision-language model can accurately separate causal from non-causal videos.

Editorial extensions

If this is right

  • Models can detect time reversal without acquiring causal reasoning.
  • Synthetic-data benchmarks may overlook real-world causal failures.
  • Current video diffusion models fall short of human causal cognition.
  • The two-level protocol can be extended to new models at low cost.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Improving the Causality Cognition Index could lead models to generate more physically consistent future frames.
  • The same reversal technique might expose causal gaps in other generative domains such as audio or 3D scenes.
  • Explicit causal objectives beyond standard diffusion training may be needed to close the human gap.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The paper introduces YoCausal, a two-level benchmark for evaluating causal understanding in video diffusion models (VDMs) inspired by the Violation of Expectation paradigm. It treats temporally reversed real-world videos as natural counterfactual samples, defines the Reverse Surprise Index (RSI) to quantify arrow-of-time perception via denoising loss, and the Causality Cognition Index (CCI) via VLM-based stratification into causal vs. non-causal subsets. Evaluation across 13 state-of-the-art VDMs concludes that arrow-of-time perception does not imply causal understanding and that a significant gap remains relative to human causal cognition.

Significance. If the premises hold, this provides a scalable, real-world, zero-cost protocol for disentangling temporal bias from causal reasoning in generative video models, extending cognitive science methods to assess progress toward world models. It offers falsifiable indices and highlights a dissociation that could guide future VDM development.

major comments (3)
  1. [Abstract] Abstract: Treating temporally reversed videos as 'natural counterfactual samples' is load-bearing for the central dissociation claim, yet reversal simultaneously violates multiple irreversible processes (entropy, gravity, friction) without corresponding to a targeted do-intervention or single-cause counterfactual in a causal graph; this risks conflating general physics-violation detection with causal reasoning.
  2. [Abstract] Abstract: CCI relies on an off-the-shelf VLM to partition videos into causal vs. non-causal subsets with no reported calibration against human judgments or formal causal criteria; without this, the reported gap between RSI and CCI may reflect VLM annotation artifacts rather than VDM causal understanding.
  3. [Abstract] Abstract: The evaluation on 13 VDMs reports a dissociation and gap to humans but supplies no quantitative details on CCI computation, error bars, dataset sizes, or controls for VLM bias, preventing verification that the data support the central claim.
minor comments (1)
  1. The abstract would benefit from explicit references to causal inference literature (e.g., Pearl's do-calculus) and prior VoE implementations to situate the protocol.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their constructive feedback. We address each major comment below with clarifications on our methodology and indicate planned revisions where appropriate.

read point-by-point responses
  1. Referee: Treating temporally reversed videos as 'natural counterfactual samples' is load-bearing for the central dissociation claim, yet reversal simultaneously violates multiple irreversible processes (entropy, gravity, friction) without corresponding to a targeted do-intervention or single-cause counterfactual in a causal graph; this risks conflating general physics-violation detection with causal reasoning.

    Authors: We agree that time reversal is not a targeted do-intervention on a single causal variable. Our method draws directly from the Violation of Expectation paradigm, using reversal to create scalable, real-world violations of expected physical dynamics rather than precise graph interventions. RSI quantifies detection of such violations as a necessary (but not sufficient) component of causal perception. We will revise the abstract and method sections to describe these as 'approximate natural counterfactuals' to prevent overstatement. revision: partial

  2. Referee: CCI relies on an off-the-shelf VLM to partition videos into causal vs. non-causal subsets with no reported calibration against human judgments or formal causal criteria; without this, the reported gap between RSI and CCI may reflect VLM annotation artifacts rather than VDM causal understanding.

    Authors: The concern is valid. The current manuscript applies an off-the-shelf VLM with prompts targeting agent-driven cause-effect relations but does not report human calibration. We will add a human validation study on a data subset, report agreement metrics, and include the exact stratification prompts and criteria in the revised version. revision: yes

  3. Referee: The evaluation on 13 VDMs reports a dissociation and gap to humans but supplies no quantitative details on CCI computation, error bars, dataset sizes, or controls for VLM bias, preventing verification that the data support the central claim.

    Authors: The full manuscript contains dataset sizes, the CCI formula, and per-model results. We agree that error bars, explicit dataset statistics, and VLM bias controls are insufficiently detailed. We will add a table with video counts, standard errors across VLM runs, and a discussion of bias mitigation in the revision. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: benchmark metrics are external evaluations, not self-derived

full rationale

The paper introduces RSI (denoising loss on time-reversed videos) and CCI (VLM-based stratification) purely as evaluation protocols applied to existing VDMs. Neither metric is obtained by fitting parameters to the target result, nor does any central claim reduce to a self-citation chain or definitional equivalence. The reported dissociation between arrow-of-time perception and causal understanding follows directly from applying these independent indices to 13 models; no derivation step equates output to input by construction. This is a standard empirical benchmark paper with no load-bearing self-referential steps.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the domain assumption that video reversal supplies valid counterfactuals and that VLM-based stratification isolates causal reasoning; no free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption Temporally reversing real-world videos at zero cost produces natural counterfactual samples suitable for testing causality
    Invoked to justify the core evaluation protocol in both Level 1 and Level 2.
  • domain assumption A VLM can accurately stratify videos into causal and non-causal subsets to disentangle causal reasoning from temporal bias
    Required for the CCI to isolate genuine causality understanding.

how reviews work

0 comments
Cite this review

Pith. "Pith review of YoCausal: How Far is Video Generation from World Model? A Causality Perspective." pith.science (2026). https://pith.science/paper/N5A4EID6

@misc{pith2026260530346,
  author       = {Pith},
  title        = {Pith review of: YoCausal: How Far is Video Generation from World Model? A Causality Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N5A4EID6}},
  note         = {Machine review of arXiv:2605.30346}
}
read the original abstract

As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data, limiting real-world generalization due to the sim-to-real gap. We present YoCausal, a two-level benchmark inspired by the Violation of Expectation (VoE) paradigm from cognitive science. By temporally reversing real-world videos at zero cost as natural counterfactual samples, YoCausal establishes an arbitrarily extensible evaluation protocol. Level 1 introduces the Reverse Surprise Index (RSI), quantifying arrow-of-time perception via denoising loss. Level 2 introduces the Causality Cognition Index (CCI), which leverages a VLM to stratify datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias. Evaluation of 13 state-of-the-art VDMs reveals that perceiving the arrow of time does not imply understanding causality, and a significant gap persists relative to human-level causal cognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PhiZero: A World Model Built Around Physical Language

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A self-supervised discrete physical-language bottleneck plus a VLM reasoner lets a world model predict state transitions before rendering video, improving physical coherence and enabling zero-shot motion transfer.

Reference graph

Works this paper leans on

134 extracted references · 61 canonical work pages · cited by 1 Pith paper

  1. [1]

    Abdi et al

    H. Abdi et al. The kendall rank correlation coefficient.Encyclopedia of measurement and statistics, 2:508–510, 2007

  2. [2]

    Cosmos World Foundation Model Platform for Physical AI

    N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P . Chattopadhyay, Y. Chen, Y. Cui, Y. Ding, et al. Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575, 2025

  3. [3]

    T. Ates, M. Ate¸ so˘ glu, Ç. Yi˘ git, I. Kesen, M. Kobas, E. Erdem, A. Erdem, T. Goksun, and D. Yuret. Craft: A benchmark for causal reasoning about forces and interactions. InFindings of the Association for Computational Linguistics: ACL 2022, pages 2602–2627, 2022

  4. [4]

    Z. Bai, H. Ci, and M. Z. Shou. Impossible videos.arXiv preprint arXiv:2503.14378, 2025

  5. [5]

    Baillargeon

    R. Baillargeon. Infants’ physical world.Current directions in psychological science, 13(3):89–94, 2004

  6. [6]

    Baillargeon, E

    R. Baillargeon, E. S. Spelke, and S. Wasserman. Object permanence in five-month-old infants.Cognition, 20(3):191–208, 1985

  7. [7]

    VideoPhy: Evaluating Physical Commonsense for Video Generation

    H. Bansal, Z. Lin, T. Xie, Z. Zong, M. Yarom, Y. Bitton, C. Jiang, Y. Sun, K.-W. Chang, and A. Grover. Videophy: Evaluating physical commonsense for video generation.arXiv preprint arXiv:2406.03520, 2024

  8. [8]

    arXiv preprint arXiv:2503.06800 (2025)

    H. Bansal, C. Peng, Y. Bitton, R. Goldenberg, A. Grover, and K.-W. Chang. Videophy-2: A challenging action-centric physical commonsense evaluation in video generation.arXiv preprint arXiv:2503.06800, 2025

Show all 134 references
  1. [9]

    Bar-Tal, H

    O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, G. Liu, A. Raj, et al. Lumiere: A space-time diffusion model for video generation. InSIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024

  2. [10]

    Baradel, N

    F. Baradel, N. Neverova, J. Mille, G. Mori, and C. Wolf. Cophy: Counterfactual learning of physical dynamics.arXiv preprint arXiv:1909.12000, 2019

  3. [11]

    P . W. Battaglia, J. B. Hamrick, and J. B. Tenenbaum. Simulation as an engine of physical scene understanding.Proceedings of the national academy of sciences, 110(45):18327–18332, 2013

  4. [12]

    D. M. Bear, E. Wang, D. Mrowca, F. J. Binder, H.-Y. F. Tung, R. Pramod, C. Holdaway, S. Tao, K. Smith, F.-Y. Sun, et al. Physion: Evaluating physical prediction from vision in humans and machines.arXiv preprint arXiv:2106.08261, 2021

  5. [13]

    Blattmann, T

    A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V . Voleti, A. Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. 23 Figure A.5 Scaling laws and generational trends in causal cognition...

  6. [14]

    Blattmann, R

    A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22563–22575, 2023

  7. [15]

    Bordes, Q

    F. Bordes, Q. Garrido, J. T. Kao, A. Williams, M. Rabbat, and E. Dupoux. Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environments.arXiv preprint arXiv:2506.09849, 2025

  8. [16]

    Brooks, B

    T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, et al. Video generation models as world simulators.OpenAI Blog, 1(8):1, 2024

  9. [17]

    Bruce, M

    J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al. Genie: Generative interactive environments. InForty-first International Conference on Machine Learning, 2024

  10. [18]

    M. Cai, R. Tan, J. Zhang, B. Zou, K. Zhang, F. Yao, F. Zhu, J. Gu, Y. Zhong, Y. Shang, et al. Tempo- ralbench: Benchmarking fine-grained temporal understanding for multimodal video models.arXiv preprint arXiv:2410.10818, 2024

  11. [19]

    Chandrasegaran, A

    K. Chandrasegaran, A. Gupta, L. M. Hadzic, T. Kota, J. He, C. Eyzaguirre, Z. Durante, M. Li, J. Wu, and L. Fei-Fei. Hourvideo: 1-hour video-language understanding.Advances in Neural Information Processing Systems, 37:53168–53197, 2024

  12. [20]

    Chao, W.-F

    C.-H. Chao, W.-F. Sun, B.-W. Cheng, Y.-C. Lo, C.-C. Chang, Y.-L. Liu, Y.-L. Chang, C.-P . Chen, and C.-Y. Lee. Denoising likelihood score matching for conditional score-based data generation.arXiv preprint arXiv:2203.14206, 2022

  13. [21]

    Y. Chen, J. Liu, X. Lin, and R. Tang. Countervqa: Evaluating and improving counterfactual reasoning in vision-language models for video understanding.arXiv preprint arXiv:2511.19923, 2025

  14. [22]

    H. Chi, H. Li, W. Yang, F. Liu, L. Lan, X. Ren, T. Liu, and B. Han. Unveiling causal reasoning in large language models: Reality or mirage?Advances in Neural Information Processing Systems, 37:96640–96670, 2024

  15. [23]

    Clark and P

    K. Clark and P . Jaini. Text-to-image diffusion models are zero shot classifiers.Advances in Neural Information Processing Systems, 36:58921–58937, 2023

  16. [24]

    Cores, M

    D. Cores, M. Dorkenwald, M. Mucientes, C. G. Snoek, and Y. M. Asano. Tvbench: Redesigning video-language evaluation. 2024

  17. [25]

    Croitoru, V

    F.-A. Croitoru, V . Hondru, R. T. Ionescu, and M. Shah. Diffusion models in vision: A survey.IEEE transactions on pattern analysis and machine intelligence, 45(9):10850–10869, 2023

  18. [26]

    Dasgupta, J

    A. Dasgupta, J. Duan, M. H. Ang Jr, and C. Tan. Avoe: a synthetic 3d dataset on understanding violation of expectation for artificial cognition.arXiv preprint arXiv:2110.05836, 2021

  19. [27]

    Didelez and I

    V . Didelez and I. Pigeot. Causality: models, reasoning, and inference, 2001

  20. [28]

    Y. Du, M. Yang, P . Florence, F. Xia, A. Wahid, B. Ichter, P . Sermanet, T. Yu, P . Abbeel, J. B. Tenenbaum, 24 et al. Video language planning.arXiv preprint arXiv:2310.10625, 2023

  21. [29]

    Dummett.Principles of electoral reform

    M. Dummett.Principles of electoral reform. Oxford University Press, 1997

  22. [30]

    P . Emerson. The original borda count and partial voting.Social Choice and Welfare, 40(2):353–358, 2013

  23. [31]

    Esser, J

    P . Esser, J. Chiu, P . Atighehchian, J. Granskog, and A. Germanidis. Structure and content-guided video synthesis with diffusion models. InProceedings of the IEEE/CVF international conference on computer vision, pages 7346–7356, 2023

  24. [32]

    A. Foss, C. Evans, S. Mitts, K. Sinha, A. Rizvi, and J. T. Kao. Causalvqa: A physically grounded causal reasoning benchmark for video models.arXiv preprint arXiv:2506.09943, 2025

  25. [33]

    C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, et al. Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition...

  26. [34]

    Gandhi, G

    K. Gandhi, G. Stojnic, B. M. Lake, and M. R. Dillon. Baby intuitions benchmark (bib): Discerning the goals, preferences, and actions of others.Advances in neural information processing systems, 34:9963–9976, 2021

  27. [35]

    Garrido, N

    Q. Garrido, N. Ballas, M. Assran, A. Bardes, L. Najman, M. Rabbat, E. Dupoux, and Y. LeCun. Intuitive physics understanding emerges from self-supervised pretraining on natural videos.arXiv preprint arXiv:2502.11831, 2025

  28. [36]

    S. Ge, A. Mahapatra, G. Parmar, J.-Y. Zhu, and J.-B. Huang. On the content bias in fréchet video distance. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7277–7288, 2024

  29. [37]

    Girdhar, M

    R. Girdhar, M. Singh, A. Brown, Q. Duval, S. Azadi, S. S. Rambhatla, A. Shah, X. Yin, D. Parikh, and I. Misra. Factorizing text-to-video generation by explicit image conditioning. InEuropean Conference on Computer Vision, pages 205–224. Springer, 2024

  30. [38]

    Y. Guo, C. Yang, A. Rao, Z. Liang, Y. Wang, Y. Qiao, M. Agrawala, D. Lin, and B. Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning.arXiv preprint arXiv:2307.04725, 2023

  31. [39]

    Gupta, L

    A. Gupta, L. Yu, K. Sohn, X. Gu, M. Hahn, F.-F. Li, I. Essa, L. Jiang, and J. Lezama. Photorealistic video generation with diffusion models. InEuropean Conference on Computer Vision, pages 393–411. Springer, 2024

  32. [40]

    Ha and J

    D. Ha and J. Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2(3):440, 2018

  33. [41]

    HaCohen, N

    Y. HaCohen, N. Chiprut, B. Brazowski, D. Shalem, D. Moshe, E. Richardson, E. Levin, G. Shiran, N. Zabari, O. Gordon, et al. Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024

  34. [42]

    Hafner, T

    D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. InInternational conference on machine learning, pages 2555–2565. PMLR, 2019

  35. [43]

    Hafner, T

    D. Hafner, T. Lillicrap, M. Norouzi, and J. Ba. Mastering atari with discrete world models.arXiv preprint arXiv:2010.02193, 2020

  36. [44]

    Hanyu, K

    N. Hanyu, K. Watanabe, and S. Kitazawa. Ready to detect a reversal of time’s arrow: a psychophysical study using short video clips in daily scenes.Royal Society open science, 10(4), 2023

  37. [45]

    J. Ho, A. Jain, and P . Abbeel. Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

  38. [46]

    Ho and T

    J. Ho and T. Salimans. Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022

  39. [47]

    J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet. Video diffusion models. Advances in neural information processing systems, 35:8633–8646, 2022

  40. [48]

    W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers.arXiv preprint arXiv:2205.15868, 2022

  41. [49]

    A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado. Gaia-1: A generative world model for autonomous driving.arXiv preprint arXiv:2309.17080, 2023

  42. [50]

    Huang, Y

    Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21807–21818, 2024

  43. [51]

    Huang, F

    Z. Huang, F. Zhang, X. Xu, Y. He, J. Yu, Z. Dong, Q. Ma, N. Chanpaisit, C. Si, Y. Jiang, et al. Vbench++: Comprehensive and versatile benchmark suite for video generative models.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 25

  44. [52]

    Hurst, A

    A. Hurst, A. Lerer, A. P . Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. Gpt-4o system card.arXiv preprint arXiv:2410.21276, 2024

  45. [53]

    Z. Jin, Y. Chen, F. Leeb, L. Gresele, O. Kamal, Z. Lyu, K. Blin, F. Gonzalez Adauto, M. Kleiman-Weiner, M. Sachan, et al. Cladder: Assessing causal reasoning in language models.Advances in Neural Information Processing Systems, 36:31038–31065, 2023

  46. [54]

    Z. Jin, J. Liu, Z. Lyu, S. Poff, M. Sachan, R. Mihalcea, M. Diab, and B. Schölkopf. Can large language models infer causation from correlation?arXiv preprint arXiv:2306.05836, 2023

  47. [55]

    B. Kang, Y. Yue, R. Lu, Z. Lin, Y. Zhao, K. Wang, G. Huang, and J. Feng. How far is video generation from world model: A physical law perspective.arXiv preprint arXiv:2411.02385, 2024

  48. [56]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

  49. [57]

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P . Natsev, et al. The kinetics human action video dataset.arXiv preprint arXiv:1705.06950, 2017

  50. [58]

    Kiciman, R

    E. Kiciman, R. Ness, A. Sharma, and C. Tan. Causal reasoning and large language models: Opening a new frontier for causality.Transactions on Machine Learning Research, 2023

  51. [59]

    Kingma, T

    D. Kingma, T. Salimans, B. Poole, and J. Ho. Variational diffusion models.Advances in neural information processing systems, 34:21696–21707, 2021

  52. [60]

    Kondratyuk, L

    D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, G. Schindler, R. Hornung, V . Birodkar, J. Yan, M.-C. Chiu, et al. Videopoet: A large language model for zero-shot video generation.arXiv preprint arXiv:2312.14125, 2023

  53. [61]

    W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. Hunyuan- video: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

  54. [62]

    B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman. Building machines that learn and think like people.Behavioral and brain sciences, 40:e253, 2017

  55. [63]

    D. Layzer. The arrow of time.Scientific American, 233(6):56–69, 1975

  56. [64]

    LeCun et al

    Y. LeCun et al. A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27.Open Review, 62(1):1–62, 2022

  57. [65]

    A. M. Leslie and S. Keeble. Do six-month-old infants perceive causality?Cognition, 25(3):265–288, 1987

  58. [66]

    A. C. Li, M. Prabhudesai, S. Duggal, E. Brown, and D. Pathak. Your diffusion model is secretly a zero-shot classifier. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 2206–2217, 2023

  59. [67]

    C. Li, O. Michel, X. Pan, S. Liu, M. Roberts, and S. Xie. Pisa experiments: Exploring physics post-training for video diffusion models by watching stuff drop.arXiv preprint arXiv:2503.09595, 2025

  60. [68]

    D. Li, Y. Fang, Y. Chen, S. Yang, S. Cao, J. Wong, M. Luo, X. Wang, H. Yin, J. E. Gonzalez, et al. Worldmodelbench: Judging video generation models as world models.arXiv preprint arXiv:2502.20694, 2025

  61. [69]

    J. Li, L. Niu, and L. Zhang. From representation to reasoning: Towards both evidence and common- sense reasoning for video question-answering. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 21273–21282, 2022

  62. [70]

    K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P . Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206, 2024

  63. [71]

    S. Li, L. Li, Y. Liu, S. Ren, Y. Liu, R. Gao, X. Sun, and L. Hou. Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models. InEuropean Conference on Computer Vision, pages 331–348. Springer, 2024

  64. [72]

    Y. Li, W. Tian, Y. Jiao, J. Chen, and Y.-G. Jiang. Eyes can deceive: Benchmarking counterfactual reasoning abilities of multi-modal large language models.arXiv preprint arXiv:2404.12966, 3, 2024

  65. [73]

    Liang, H

    Z. Liang, H. He, C. Yang, and B. Dai. Scaling laws for diffusion transformers.arXiv preprint arXiv:2410.08184, 2024

  66. [74]

    J. Lin, Y. Du, O. Watkins, D. Hafner, P . Abbeel, D. Klein, and A. Dragan. Learning to model the world with language.arXiv preprint arXiv:2308.01399, 2023

  67. [75]

    X. Liu, Z. Xu, M. Li, K. Wang, Y. J. Lee, and Y. Shang. Can world simulators reason? gen-vire: A generative visual reasoning benchmark.arXiv preprint arXiv:2511.13853, 2025. 26

  68. [76]

    Y. Liu, S. Li, Y. Liu, Y. Wang, S. Ren, L. Li, S. Chen, X. Sun, and L. Hou. Tempcompass: Do video llms really understand videos? InFindings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, 2024

  69. [77]

    Y. Liu, K. Zhang, Y. Li, Z. Yan, C. Gao, R. Chen, Z. Yuan, Y. Huang, H. Sun, J. Gao, et al. Sora: A review on background, technology, limitations, and opportunities of large vision models.arXiv preprint arXiv:2402.17177, 2024

  70. [78]

    Margoni, L

    F. Margoni, L. Surian, and R. Baillargeon. The violation-of-expectation paradigm: A conceptual overview.Psychological Review, 131(3):716, 2024

  71. [79]

    Matsuo, Y

    Y. Matsuo, Y. LeCun, M. Sahani, D. Precup, D. Silver, M. Sugiyama, E. Uchibe, and J. Morimoto. Deep learning, reinforcement learning, and world models.Neural Networks, 152:267–275, 2022

  72. [80]

    R. P . McDonald. Judea pearl. causality: Models, reasoning, and inference. cambridge: Cambridge university press. 384 pp., 2000, isbn 0521773628.Psychometrika, 67(2):321–322, 2002

  73. [81]

    F. Meng, J. Liao, X. Tan, W. Shao, Q. Lu, K. Zhang, Y. Cheng, D. Li, Y. Qiao, and P . Luo. Towards world simulator: Crafting physical commonsense-based benchmark for video generation.arXiv preprint arXiv:2410.05363, 2024

  74. [82]

    Michotte.The perception of causality

    A. Michotte.The perception of causality. Routledge, 2017

  75. [83]

    Misra, C

    I. Misra, C. L. Zitnick, and M. Hebert. Shuffle and learn: unsupervised learning using temporal order verification. InEuropean conference on computer vision, pages 527–544. Springer, 2016

  76. [84]

    Monfort, A

    M. Monfort, A. Andonian, B. Zhou, K. Ramakrishnan, S. A. Bargal, T. Yan, L. Brown, Q. Fan, D. Gutfreund, C. Vondrick, et al. Moments in time dataset: one million videos for event understanding. IEEE transactions on pattern analysis and machine intelligence, 42(2):502–508, 2019

  77. [85]

    Motamed, L

    S. Motamed, L. Culp, K. Swersky, P . Jaini, and R. Geirhos. Do generative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958, 2026

  78. [86]

    L. G. Neuberg. Causality: models, reasoning, and inference, by judea pearl, cambridge university press, 2000.Econometric Theory, 19(4):675–685, 2003

  79. [87]

    X. L. Ng, K. E. Ong, Q. Zheng, Y. Ni, S. Y. Yeo, and J. Liu. Animal kingdom: A large and diverse dataset for animal behavior understanding. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19023–19034, 2022

  80. [88]

    Patraucean, L

    V . Patraucean, L. Smaira, A. Gupta, A. Recasens, L. Markeeva, D. Banarse, S. Koppula, M. Malinowski, Y. Yang, C. Doersch, et al. Perception test: A diagnostic benchmark for multimodal video models. Advances in Neural Information Processing Systems, 36:42748–42761, 2023

  81. [89]

    Peebles and S

    W. Peebles and S. Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023

  82. [90]

    X. Peng, Z. Zheng, C. Shen, T. Young, X. Guo, B. Wang, H. Xu, H. Liu, M. Jiang, W. Li, et al. Open-sora 2.0: Training a commercial-level video generation model in $200 k.arXiv preprint arXiv:2503.09642, 2025

  83. [91]

    L. C. Pickup, Z. Pan, D. Wei, Y. Shih, C. Zhang, A. Zisserman, B. Scholkopf, and W. T. Freeman. Seeing the arrow of time. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2035–2042, 2014

  84. [92]

    L. S. Piloto, A. Weinstein, P . Battaglia, and M. Botvinick. Intuitive physics learning in a deep-learning model inspired by developmental psychology.Nature human behaviour, 6(9):1257–1267, 2022

  85. [93]

    Polyak, A

    A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, et al. Movie gen: A cast of media foundation models.arXiv preprint arXiv:2410.13720, 2024

  86. [94]

    Y. Qin, Z. Shi, J. Yu, X. Wang, E. Zhou, L. Li, Z. Yin, X. Liu, L. Sheng, J. Shao, et al. Worldsimbench: Towards video generation models as world simulators.arXiv preprint arXiv:2410.18072, 2024

  87. [95]

    Riochet, M

    R. Riochet, M. Y. Castro, M. Bernard, A. Lerer, R. Fergus, V . Izard, and E. Dupoux. Intphys: A framework and benchmark for visual intuitive physics reasoning.arXiv preprint arXiv:1803.07616, 2018

  88. [96]

    Riochet, M

    R. Riochet, M. Y. Castro, M. Bernard, A. Lerer, R. Fergus, V . Izard, and E. Dupoux. Intphys 2019: A benchmark for visual intuitive physics understanding.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9):5016–5025, 2021

  89. [97]

    T. Shu, A. Bhandwaldar, C. Gan, K. Smith, S. Liu, D. Gutfreund, E. Spelke, J. Tenenbaum, and T. Ullman. Agent: A benchmark for core psychological reasoning. InInternational conference on machine learning, pages 9614–9625. PMLR, 2021. 27

  90. [98]

    Singer, A

    U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al. Make-a-video: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792, 2022

  91. [99]

    Smith, L

    K. Smith, L. Mei, S. Yao, J. Wu, E. Spelke, J. Tenenbaum, and T. Ullman. Modeling expectation violation in intuitive physics with coarse probabilistic object representations.Advances in neural information processing systems, 32, 2019

  92. [100]

    Sohl-Dickstein, E

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. InInternational conference on machine learning, pages 2256–2265. pmlr, 2015

  93. [101]

    Y. Song, C. Durkan, I. Murray, and S. Ermon. Maximum likelihood training of score-based diffusion models.Advances in neural information processing systems, 34:1415–1428, 2021

  94. [102]

    Y. Song, J. Sohl-Dickstein, D. P . Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations.arXiv preprint arXiv:2011.13456, 2020

  95. [103]

    E. S. Spelke, K. Breinlinger, J. Macomber, and K. Jacobson. Origins of knowledge.Psychological review, 99(4):605, 1992

  96. [104]

    E. S. Spelke and K. D. Kinzler. Core knowledge.Developmental science, 10(1):89–96, 2007

  97. [105]

    A. E. Stahl and L. Feigenson. Observing the unexpected enhances infants’ learning and exploration. Science, 348(6230):91–94, 2015

  98. [106]

    G. Team. Mochi 1.https://github.com/genmoai/models, 2024

  99. [107]

    G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Mil- lican, et al. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805, 2023

  100. [108]

    Q. Team. Qwen3. 5-omni technical report.arXiv preprint arXiv:2604.15804, 2026

  101. [109]

    Teed and J

    Z. Teed and J. Deng. Raft: Recurrent all-pairs field transforms for optical flow. InEuropean conference on computer vision, pages 402–419. Springer, 2020

  102. [110]

    Téglás, E

    E. Téglás, E. Vul, V . Girotto, M. Gonzalez, J. B. Tenenbaum, and L. L. Bonatti. Pure reasoning in 12-month-old infants as probabilistic inference.science, 332(6033):1054–1059, 2011

  103. [111]

    T. D. Ullman, E. Spelke, P . Battaglia, and J. B. Tenenbaum. Mind games: Game engines as an architecture for intuitive physics.Trends in cognitive sciences, 21(9):649–665, 2017

  104. [112]

    Unterthiner, S

    T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly. Towards accurate generative models of video: A new metric & challenges.arXiv preprint arXiv:1812.01717, 2018

  105. [113]

    Valevski, Y

    D. Valevski, Y. Leviathan, M. Arar, and S. Fruchter. Diffusion models are real-time game engines. arXiv preprint arXiv:2408.14837, 2024

  106. [114]

    T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

  107. [115]

    J. Wang, H. Yuan, D. Chen, Y. Zhang, X. Wang, and S. Zhang. Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571, 2023

  108. [116]

    Z. Wang, S. Zhang, C. Tang, and K. Wang. Timecausality: Evaluating the causal ability in time dimension for vision language models.arXiv preprint arXiv:2505.15435, 2025

  109. [117]

    D. Wei, J. J. Lim, A. Zisserman, and W. T. Freeman. Learning and using the arrow of time. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8052–8060, 2018

  110. [118]

    J. Wu, J. J. Lim, H. Zhang, J. B. Tenenbaum, and W. T. Freeman. Physics 101: Learning physical object properties from unlabeled videos. InBMVC, volume 2, page 7, 2016

  111. [119]

    K. Wynn. Addition and subtraction by human infants.Nature, 358(6389):749–750, 1992

  112. [120]

    J. Xiao, X. Shang, A. Yao, and T.-S. Chua. Next-qa: Next phase of question-answering to explaining temporal actions. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786, 2021

  113. [121]

    Z. Xue, M. Luo, and K. Grauman. Seeing the arrow of time in large multimodal models.arXiv preprint arXiv:2506.03340, 2025

  114. [122]

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

  115. [123]

    M. Yang, Y. Du, K. Ghasemipour, J. Tompson, D. Schuurmans, and P . Abbeel. Learning interactive real-world simulators.arXiv preprint arXiv:2310.06114, 1(2):6, 2023

  116. [124]

    Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072, 2024. 28

  117. [125]

    K. Yi, C. Gan, Y. Li, P . Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum. Clevrer: Collision events for video representation and reasoning.arXiv preprint arXiv:1910.01442, 2019

  118. [126]

    T. Yin, Q. Zhang, R. Zhang, W. T. Freeman, F. Durand, E. Shechtman, and X. Huang. From slow bidirectional to fast autoregressive video diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22963–22974, 2025

  119. [127]

    Y. Yin, Y. Zhao, M. Zheng, K. Lin, J. Ou, R. Chen, V . S.-J. Huang, J. Wang, X. Tao, P . Wan, et al. Towards precise scaling laws for video diffusion transformers. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 18155–18165, 2025

  120. [128]

    J. Yuan, F. Pizzati, F. Pinto, L. Kunze, I. Laptev, P . Newman, P . Torr, and D. De Martini. Likephys: Evaluating intuitive physics understanding in video diffusion models via likelihood preference.arXiv preprint arXiv:2510.11512, 2025

  121. [129]

    S. Yuan, J. Huang, Y. Xu, Y. Liu, S. Zhang, Y. Shi, R.-J. Zhu, X. Cheng, J. Luo, and L. Yuan. Chronomagic- bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation.Advances in Neural Information Processing Systems, 37:21236–21270, 2024

  122. [130]

    Zeˇ cevi´ c, M

    M. Zeˇ cevi´ c, M. Willig, D. S. Dhami, and K. Kersting. Causal parrots: Large language models may talk causality but are not causal.arXiv preprint arXiv:2308.13067, 2023

  123. [131]

    Zhang, D

    C. Zhang, D. Cherniavskii, A. Tragoudaras, A. Vozikis, T. Nijdam, D. W. Prinzhorn, M. Bodracska, N. Sebe, A. Zadaianchuk, and E. Gavves. Morpheus: Benchmarking physical reasoning of video generative models with real physical experiments.arXiv preprint arXiv:2504.02918, 2025

  124. [132]

    Zheng, Z

    D. Zheng, Z. Huang, H. Liu, K. Zou, Y. He, F. Zhang, L. Gu, Y. Zhang, J. He, W.-S. Zheng, et al. Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025

  125. [133]

    Zheng, X

    Z. Zheng, X. Peng, T. Yang, C. Shen, S. Li, H. Liu, Y. Zhou, T. Li, and Y. You. Open-sora: Democratizing efficient video production for all.arXiv preprint arXiv:2412.20404, 2024

  126. [134]

    Z. Zhu, X. Wang, W. Zhao, C. Min, B. Li, N. Deng, M. Dou, Y. Wang, B. Shi, K. Wang, et al. Is sora a world simulator? a comprehensive survey on general world models and beyond.arXiv preprint arXiv:2405.03520, 2024. 29

Pith tools

Reviewed June 29, 2026 · model on record in the stance chip above.