Pith. sign in

REVIEW 3 major objections 5 minor 52 references

PhiZero learns a compact discrete physical language of state transitions from unlabeled video, reasons future world evolution in that space, then renders it as video.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 01:48 UTC pith:HHDQLKYJ

load-bearing objection Solid systems paper: discrete transition codes + reason-then-render beat direct pixel prediction on physics video benches; novelty is packaging, understanding eval is partly self-referential. the 3 major comments →

arxiv 2607.28624 v1 pith:HHDQLKYJ submitted 2026-07-30 cs.CV

PhiZero: A World Model Built Around Physical Language

classification cs.CV
keywords physical languageworld modelsvideo generationreason-then-renderstate transitionsmotion transferaction-conditioned simulationself-supervised video
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Mainstream video world models predict the next frames in pixel space, so the rules of how the world changes stay buried inside a huge visual generator and often break physics. PhiZero instead extracts a short discrete code—called physical language—from ordinary videos by forcing a bottleneck to explain only what changes between frames, given the first frame and a strong generative decoder. A language-model-style reasoner then predicts that code from the current image and an action intent; a diffusion renderer turns the code back into video. The claim is that making dynamics an explicit intermediate target yields more physically coherent generation and understanding, plus a transferable handle on motion that can be edited, controlled, and moved across appearances and bodies without paired training. A sympathetic reader cares because a reusable code for how the world evolves is the missing substrate for controllable simulators and Physical AI, not just prettier clips.

Core claim

A self-supervised discrete physical language of state-transition patterns, learned from in-the-wild video and used in a reason-then-render loop, models physically coherent world evolution better than direct pixel-space prediction and doubles as a controllable, appearance-disentangled interface for action-conditioned simulation and zero-shot motion transfer.

What carries the argument

Physical language: a length-256 discrete FSQ sequence of transition tokens produced by a transition-level Q-Former over adjacent video latents. It is the intermediate that the reasoner predicts and the diffusion decoder renders, separating dynamics inference from pixel synthesis.

Load-bearing premise

The first-frame-conditioned reconstruction bottleneck plus likelihoods over its codes really capture reusable physical state-change structure, not just compressible visual change patterns that happen to score well on the chosen physics tests.

What would settle it

Hold the first frame and action fixed, swap or scramble the physical-language sequence, and check whether rendered outcomes systematically violate the same physical laws the paper’s benchmarks reward; if scrambled codes still look physically fine or transfer fails when appearance is held fixed and only dynamics should move, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Future video world models can treat dynamics as an explicit discrete reasoning target rather than an implicit side-effect of pixel prediction.
  • The same physical-language codes support interactive rollouts, trajectory- and action-conditioned driving and robot simulation, and fine-grained control without redesigning the generator.
  • Motion patterns can be transferred zero-shot across object appearance, human-to-robot embodiments, and sim-to-real looks by editing only the first frame.
  • Physical plausibility can be scored by comparing reasoner likelihoods of candidate videos’ codes, enabling understanding benchmarks without separate classifiers.
  • Scaling the discrete transition vocabulary and reasoner with more motion-rich and simulator data should further improve physical fidelity.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the codes truly factor dynamics from appearance, abundant human video could become a low-cost teacher for robot policies via code-level retargeting rather than paired teleoperation.
  • Hierarchical or recurrent prediction over physical-language sequences is a natural next step for long-horizon consistency beyond fixed four-second clips.
  • Grounding subsets of the discrete symbols to measurable quantities (contact, force proxies, occlusion events) would test whether the language can move from empirical transitions toward interpretable physical variables.
  • The same bottleneck may serve as a drop-in dynamics head for existing VLMs, turning generation priors into pairwise physics discriminators without new labeled physics datasets.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. PhiZero proposes a physical world model organized around “physical language”: a compact discrete code of video state transitions learned by a self-supervised tokenizer (transition-level Q-Former + FSQ) that, with the first frame, conditions a pretrained diffusion decoder (reason-then-render; Eq. 1). A VLM-initialized reasoner then autoregressively predicts these codes from an image and textual action intent; the frozen decoder renders them to video. The paper reports leading or near-leading scores on Physics-IQ Verified, PhyGround, and WorldModelBench generation benchmarks, competitive pairwise results on IntPhys2/LikePhys/YoCausal via reasoner likelihoods over tokenizer codes, ablations of tokenizer/reasoner design, and demos of interactive/action-conditioned control and zero-shot motion and sim-to-real transfer by reusing codes under edited first frames.

Significance. If the factorization and discrete transition codes genuinely capture reusable world dynamics rather than decoder-friendly appearance change, the work offers a clear alternative to pure pixel-space world models and a practical interface for control and cross-embodiment transfer. Strengths include multi-benchmark generation evaluation (including rule-based comparison to real physical experiments on Physics-IQ Verified), explicit ablations (Tables 8–9), reconstruction efficiency vs. other compressed tokenizers (Table 7), and concrete transfer/control applications (Figs. 5–8). The contribution is primarily architectural and empirical rather than theoretical; its lasting value hinges on whether physical language is shown to be more than a strong motion bottleneck for Wan2.2-class decoders.

major comments (3)
  1. [Appendix B.2, Eq. (7); Tables 4–6] Appendix B.2 / Eq. (7): On IntPhys2, LikePhys, and YoCausal, “understanding” is defined as higher autoregressive log-likelihood of z = T_φ(V) under the reasoner trained to imitate the frozen tokenizer. Validity labels never enter training or scoring except through whatever statistics T_φ already packs. This is not an independent physics probe: if the bottleneck mainly stores compressible motion/appearance residuals that correlate with the benchmark construction, pairwise preference can improve without encoding physical laws. The generation wins do not close this gap. Please either (i) evaluate with an external physics judge or frozen independent scorer on the same pairs, (ii) corrupt dynamics while matching low-level motion energy and show likelihood drops selectively, or (iii) substantially soften claims that these tables demonstrate physical understanding.
  2. [Sec. 3.2, Eq. (1)–(4); Tables 1–3, 9] Sec. 3.2 and the central interpretation of “physical language”: the tokenizer is optimized for reconstruction with a strong pretrained diffusion prior, first-frame appearance anchoring, pure-noise warm-up, and heavy VLM-based data filtering (Fig. 3). Tables 1–3 and Fig. 4 show better physical outcomes than Wan2.2-5B and other video models, but they do not isolate the discrete transition codes from (a) the Wan2.2 decoder prior and LoRA finetuning, (b) the curated motion/physics-heavy corpus (including simulation), and (c) reasoner capacity. A load-bearing control is missing: e.g., the same reason-then-render pipeline with a matched continuous or permuted/codebook-shuffled bottleneck, or pixel/latent autoregression on the identical data and decoder budget. Without that, “explicit physical-language reasoning” remains under-identified relative to “better-conditioned video generation on filte
  3. [Sec. 4.5; Appendix C.1–C.2; Figs. 7–8] Table 9 and Sec. 4.5 applications: prompt-enhanced Wan2.2-5B reaches only 26.6 IQ-Score vs. 41.2 for PhiZero, which supports the pipeline, but action-conditioned driving/robotics and interactive rollouts appear to use domain-adapted tokenizers/reasoners (Appendix C.1) while zero-shot transfer uses brief source-domain tokenizer finetuning plus GPT-Image first-frame edits (Appendix C.2). The paper should state clearly which results are zero-shot general PhiZero vs. domain-adapted, report quantitative metrics (not only qualitative figures) for trajectory/action fidelity and transfer identity preservation, and avoid implying a single frozen physical language serves all demos without adaptation.
minor comments (5)
  1. [Sec. 4.1] Clarify FSQ vocabulary size: levels (8,5,5,5,5,5) give 8×5^5 = 25 000 symbols as stated; ensure consistency wherever “25K” and sequence length 256 are discussed relative to 33-frame / 9-latent encoding.
  2. [Table 2] Table 2: PhiZero’s General Quality (2.93) is below Veo3.1/Wan2.2-14B while Physics Score is best; briefly discuss the quality–physics tradeoff so Overall leadership is not over-read.
  3. [Sec. 4.4, Fig. 6] Fig. 6 UMAP clusters on driving/manipulation are suggestive but use domain-adapted tokenizers and selected categories; label that limitation in the caption.
  4. [Sec. 2; Appendix D] Related work (Sec. 2 and Appendix D) is thorough; a short explicit contrast table vs. latent-action models (Genie, LAPA, etc.) and appearance–motion tokenizers would help readers place the novelty claim.
  5. [Sec. 3.3; throughout] Typos/style: “T wo-stage” spacing (Table 9, Sec. 3.3); “insufficient”/“sufficient” ligature artifacts; footnote on “physical” vs physical laws is helpful—consider elevating one sentence into Sec. 1.

Circularity Check

1 steps flagged

No load-bearing circular derivation: PhiZero is an empirical reason-then-render system trained and judged against external data/benchmarks; understanding uses own-code likelihoods but external validity labels.

specific steps
  1. other [Appendix B.2, Eq. 7; Sec. 4.2 Video Understanding Results]
    "For each pair, we encode both videos into physical-language sequences, compute their log-likelihoods under the Physical Language Reasoner, and select the video with the higher likelihood as valid. ... s(V±, cpair) = Σ_j log pθ(z±_j | I0±, cpair, z±_<j)."

    Not circular by construction: external pair labels still decide correctness. Mild coupling only—the scored object is likelihood of the model's own tokenizer codes under a reasoner trained to imitate those codes—so 'physical understanding' is operationalized as in-stack preference for T_φ encodings of benchmark-valid clips, not an independent physics probe. This weakens interpretation of Tables 4–6 but does not make the reported accuracies identically equal to training objectives or fitted parameters.

full rationale

This is a systems/ML paper, not a first-principles derivation. Physical language is learned by self-supervised reconstruction (tokenizer + diffusion decoder, Eqs. 2–4), the reasoner is teacher-forced to imitate frozen tokenizer codes (Eqs. 5–6), and generation claims are scored on external suites (Physics-IQ Verified, PhyGround, WorldModelBench) with independent references or judges. That chain does not reduce a claimed prediction to its fitted inputs by construction. The only soft spot is the understanding protocol (Appendix B.2, Eq. 7): validity is ranked by autoregressive log-likelihood of z = T_φ(V) under the same reasoner trained to imitate T_φ. That couples the metric to the model stack and can reward decoder-friendly compressible change rather than laws—but pair labels (possible/impossible, valid/invalid, forward/reverse) still come from external benchmarks, so accuracy is not tautological (a random ranker stays near chance). No self-citation uniqueness theorem, ansatz smuggled as theorem, or renamed closed-form identity carries the central claim. Score 1 only for that mild metric–stack coupling, not for a forced result.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 2 invented entities

The central claim rests on engineering assumptions about representation learning and evaluation, not on formal physical axioms. Load-bearing choices include: first-frame appearance anchoring, discrete FSQ bottleneck sufficiency, pretrained diffusion/VLM priors as carriers of residual detail and commonsence, and benchmark proxies for ‘physical coherence.’ Free parameters are standard ML hyperparameters and codebook design; invented entity is ‘physical language’ as an empirical symbol system.

free parameters (6)
  • FSQ quantization levels (8,5,5,5,5,5) → ~25K vocab = (8,5,5,5,5,5); vocab ~25K
    Hand-chosen discrete bottleneck size that defines the physical-language alphabet; not derived from physics.
  • Transition Q-Former query count M=32 and sequence length 256 = M=32; N=(9-1)*32=256
    Architectural capacity choice controlling how many symbols represent each adjacent latent pair on 33-frame clips.
  • LoRA rank 32 on diffusion decoder = rank 32
    Adaptation capacity hyperparameter for the renderer.
  • FSQ entropy-loss weight 0.005 and REPA-loss weight 0.1 = 0.005 / 0.1
    Regularizer weights chosen during tokenizer training schedule (Table 12).
  • Data filtering thresholds and VLM judge criteria for motion/physics/aesthetics
    Curation knobs that select the 10K-hour, 5M-clip, and 1M-clip corpora; directly shape what counts as learnable ‘physical’ transitions.
  • Two-stage reasoner LR schedule and data mix (5M then ~1M with 200K sim) = 5e-5 then 1e-5; 50K+10K steps
    Training recipe choices that ablations show move Physics-IQ scores (Table 9).
axioms (6)
  • ad hoc to paper World evolution p(V,z|I0,c) factorizes as reason-then-render p(z|I0,c)p(V|I0,z) with z sufficient for dynamics (Eq. 1).
    Core modeling assumption introduced in §3.1; not forced by physics, chosen as architecture.
  • domain assumption Conditioning the decoder on the clean first frame lets a tight discrete bottleneck focus on state change rather than static appearance (§3.2).
    Standard content-motion separation bet; load-bearing for calling z ‘physical language’.
  • domain assumption Pretrained video diffusion priors can restore fine appearance so reconstruction loss supervises transition codes usefully (flow-matching objective Eq. 4; Wan2.2 init).
    Relies on external generative model quality; ablation without diffusion decoder hurts most (Table 8).
  • domain assumption Pretrained VLM commonsense is a good prior for autoregressive physical-language prediction from image+text intent (§3.3).
    Initialization and training target assume VLM semantics align with transition codes.
  • domain assumption Benchmark suites (Physics-IQ, PhyGround, WorldModelBench, IntPhys2, LikePhys, YoCausal) are adequate proxies for physical coherence and understanding (§4.2, App. B).
    External validity of the central empirical claim depends on these proxies and judges.
  • standard math Finite scalar quantization and standard optimizers/deep nets behave as usual.
    Ordinary ML background; FSQ cited from Mentzer et al.
invented entities (2)
  • Physical language (discrete FSQ state-transition symbol sequences) no independent evidence
    purpose: Serve as compact intermediate for reasoning, control, and appearance-disentangled transfer of world evolution.
    Named and centralized as the paper’s key object; explicitly not a symbolic law system (footnote p.2). Evidence is internal (reconstruction, clusters, transfer demos, benchmarks), not an independent physical measurement of the symbols.
  • Transition-level Q-Former over adjacent latent pairs no independent evidence
    purpose: Impose local temporal inductive bias when compressing videos into physical-language tokens.
    Architectural invention relative to a global Q-Former; supported only by ablation gains inside this paper.

pith-pipeline@v1.2.0-daily-grok45 · 29758 in / 4246 out tokens · 81904 ms · 2026-07-31T01:48:25.071321+00:00 · methodology

0 comments
read the original abstract

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.

Figures

Figures reproduced from arXiv: 2607.28624 by Lue Fan, Ruopeng Gao, Shuyao Shang, Tieniu Tan, Xu Chen, Yuqi Wang, Zhaoxiang Zhang.

Figure 1
Figure 1. Figure 1: We introduce PHIZERO, a physical world model that learns a compact discrete physical language from in-the-wild videos. It reasons world evolution in this space and renders the inferred transitions into videos, enabling physically realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer. ABSTRACT We introduce PHIZERO, a physical world model built ar… view at source ↗
Figure 2
Figure 2. Figure 2: The overall pipeline of PHIZERO. (a) The Physical Language Tokenizer uses a transition￾level Q-Former to compress video state transitions into a compact discrete physical language. A diffusion decoder reconstructs the video conditioned on the physical language and the first frame. (b) The Physical Language Reasoner uses an autoregressive VLM to predict a physical-language sequence from the first frame and … view at source ↗
Figure 3
Figure 3. Figure 3: Hierarchical data curation pipeline. We collect Internet videos to construct a 50K-hour pool of in-the-wild real-world videos and a 1K-hour pool of simulation videos. Progressive filtering yields 10K hours for tokenizer pretraining, 5M four-second clips for tokenizer SFT and reasoner pretraining, and 1M motion-rich four-second clips for reasoner SFT. Pure-noise Warm-up The pretrained diffusion decoder may … view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison on Physics-IQ Verified. Compared with the Wan2.2-5B baseline, PHIZERO better captures the physical consequences of interactions, including collision-induced object displacement, gravity-driven deformation, chain reactions, and dynamically changing shadows [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Transferability of Physical Language. We encode each source video’s state transition into a physical-language sequence and decode the same sequence conditioned on an edited first frame with a different visual appearance. The transferred videos preserve the source evolution across substantial changes in object and scene appearance. 0 2 4 6 8 1 2 3 4 5 2 4 6 8 Left turn Right turn Straight Stationary (a) Aut… view at source ↗
Figure 6
Figure 6. Figure 6: Semantic structure of physical language. We aggregate transition features into clip-level representations, reduce them to 20 dimensions using PCA, and project them into 3D using UMAP. The resulting embeddings naturally organize according to the underlying transition patterns, forming structured manifolds or clusters across the two domains [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: PHIZERO as a physically realistic, controllable, and interactive video world model. PHIZERO generates physically coherent world evolution, follows fine-grained action conditions, and supports interactive rollouts under sequential controls. ing PCA, and further project them into 3D using UMAP (McInnes et al., 2018). As shown in [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Zero-shot cross-embodiment and sim-to-real transfer with physical language. We en￾code source state transitions into physical-language sequences and render them conditioned on edited first frames specifying new embodiments or visual domains, enabling zero-shot cross-embodiment and sim-to-real transfer. Zero-Shot Cross-embodiment and Sim-to-real Transfer Because physical language disentan￾gles state transit… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 37 linked inside Pith

  1. [1]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774,

  2. [5]

    Latent action world models for control with unlabeled trajectories

    Marvin Alles, Xingyuan Zhang, Patrick van der Smagt, and Philip Becker-Ehmck. Latent action world models for control with unlabeled trajectories. arXiv preprint arXiv:2512.10016,

  3. [6]

    Visual imitation enables contextual humanoid control

    Arthur Allshire, Hongsuk Choi, Junyi Zhang, David McAllister, Anthony Zhang, Chung Min Kim, Trevor Darrell, Pieter Abbeel, Jitendra Malik, and Angjoo Kanazawa. Visual imitation enables contextual humanoid control. arXiv preprint arXiv:2505.03729,

  4. [7]

    V-jepa 2: Self-supervised video models enable understanding, prediction and planning

    Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Am- mar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-jepa 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985,

  5. [8]

    Qwen3-vl technical report

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025a. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Y ang, Z...

  6. [11]

    Motus: A unified latent action world model

    Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Y ao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model. arXiv preprint arXiv:2512.13030,

  7. [12]

    Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environ- ments

    Florian Bordes, Quentin Garrido, Justine T Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux. Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environ- ments. arXiv preprint arXiv:2506.09849,

  8. [13]

    Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems

    Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Y an Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, Xu Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. arXiv preprint arXiv:2503.06669, 2025a. Qingwen Bu, Y anting Y ang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Y ao, Ping Luo,...

  9. [15]

    Comphy: Compositional physical reasoning of objects and events from videos.arXiv preprint arXiv:2205.01089,

    Zhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba, Joshua B Tenenbaum, and Chuang Gan. Comphy: Compositional physical reasoning of objects and events from videos.arXiv preprint arXiv:2205.01089,

  10. [17]

    Adaworld: Learning adaptable world models with latent actions

    Shenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang, and Chuang Gan. Adaworld: Learning adaptable world models with latent actions. arXiv preprint arXiv:2503.18938,

  11. [18]

    Learning latent action world models in the wild

    Quentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas, Y ann LeCun, and Michael Rabbat. Learning latent action world models in the wild. arXiv preprint arXiv:2601.05230,

  12. [20]

    Genmo AI

    URL https://arxiv.org/abs/2312.11805. Genmo AI. Genmo AI Blog. https://www.genmo.ai/blog,

  13. [21]

    Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

    Yuwei Guo, Ceyuan Y ang, Anyi Rao, Zhengyang Liang, Y aohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725,

  14. [22]

    Ltx-2: Efficient joint audio-visual foundation model

    Y oav HaCohen, Benny Brazowski, Nisan Chiprut, Y aki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, et al. Ltx-2: Efficient joint audio-visual foundation model. arXiv preprint arXiv:2601.03233,

  15. [23]

    Lamo: Self-supervised latent motion priors for physical realism in video generation

    Bo Jiang, Depu Meng, Yihan Hu, Yichen Xie, Tianshuo Xu, and Wei Zhan. Lamo: Self-supervised latent motion priors for physical realism in video generation. arXiv preprint arXiv:2605.23878 , 2026a. Yuxin Jiang, Yuchao Gu, Ivor W Tsang, and Mike Zheng Shou. Olaf-world: Orienting latent actions for video world modeling. arXiv preprint arXiv:2602.10104, 2026b....

  16. [24]

    Hunyuanvideo: A systematic framework for large video generative models

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603,

  17. [25]

    Uni- formerv2: Spatiotemporal learning by arming image vits with video uniformer

    Kunchang Li, Y ali Wang, Yinan He, Yizhuo Li, Yi Wang, Limin Wang, and Yu Qiao. Uni- formerv2: Spatiotemporal learning by arming image vits with video uniformer. arXiv preprint arXiv:2211.09552,

  18. [26]

    Dicode: Diffusion-compressed deep to- kens for autoregressive video generation with language models

    Yizhuo Li, Yuying Ge, Yixiao Ge, Ying Shan, and Ping Luo. Dicode: Diffusion-compressed deep to- kens for autoregressive video generation with language models. arXiv preprint arXiv:2412.04446,

  19. [27]

    Sekai: A video dataset towards world exploration

    Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin, Ming Li, Shitian Zhao, Zhaopan Xu, Xinyue Li, Yukang Feng, Jianwen Sun, et al. Sekai: A video dataset towards world exploration. Advances in Neural Information Processing Systems, 38, 2026b. Jongbin Lim, Taeyun Ha, Mingi Choi, Jisoo Kim, Byungjun Kim, Subin Jeon, and Hanbyul Joo. Hrdexdb: A paired human-ro...

  20. [28]

    Open-sora plan: Open-source large video generation model

    21 Preprint Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Y ang Y e, Shenghai Yuan, Liuhan Chen, et al. Open-sora plan: Open-source large video generation model. arXiv preprint arXiv:2412.00131,

  21. [29]

    Phyground: Benchmarking physical reasoning in generative world models

    Juyi Lin, Arash Akbari, Yumei He, Lin Zhao, Haichao Zhang, Arman Akbari, Xingchen Xu, Zoe Y Lu, Enfu Nan, Hokin Deng, et al. Phyground: Benchmarking physical reasoning in generative world models. arXiv preprint arXiv:2605.10806,

  22. [30]

    Libero: Benchmarking knowledge transfer for lifelong robot learning

    Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36:44776–44791, 2023a. Huaize Liu, Wenzhang Sun, Qiyuan Zhang, Donglin Di, Biao Gong, Hao Li, Chen Wei, and Changqing Zou. Hi-vae: Efficient video autoenc...

  23. [31]

    Umap: Uniform manifold approximation and projection for dimension reduction

    Leland McInnes, John Healy, and James Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426,

  24. [33]

    Openvid-1m: A large-scale high-quality dataset for text-to-video generation

    Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Y ang, Zhijie Chen, Xiang Li, Jian Y ang, and Ying Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. In International Conference on Learning Representations , volume 2025, pp. 1045–1064,

  25. [34]

    Omniweaving: Towards unified video generation with free-form composition and reasoning

    Kaihang Pan, Qi Tian, Jianwei Zhang, Weijie Kong, Jiangfeng Xiong, Y anxin Long, Shixue Zhang, Haiyi Qiu, Tan Wang, Zheqi Lv, et al. Omniweaving: Towards unified video generation with free-form composition and reasoning. arXiv preprint arXiv:2603.24458,

  26. [35]

    Physics-iq verified

    22 Preprint Tim Rädsch, Yuki M Asano, Hilde Kuehne, Stefan Bauer, Priyank Jaini, Robert Geirhos, and Carsten T Lüth. Physics-iq verified. arXiv preprint arXiv:2606.18943,

  27. [36]

    Dismo: Disentangled motion representations for open-world motion transfer

    Thomas Ressler-Antal, Frank Fundel, Malek Ben Alaya, Stefan Andreas Baumann, Felix Krause, Ming Gui, and Björn Ommer. Dismo: Disentangled motion representations for open-world motion transfer. arXiv preprint arXiv:2511.23428,

  28. [37]

    Learning to act without actions

    Dominik Schmidt and Minqi Jiang. Learning to act without actions. In International Conference on Learning Representations, volume 2024, pp. 9379–9395,

  29. [38]

    Dynvla: Learning world dynamics for action reasoning in autonomous driving

    Shuyao Shang, Bing Zhan, Yunfei Y an, Yuqi Wang, Yingyan Li, Y asong An, Xiaoman Wang, Jierui Liu, Lu Hou, Lue Fan, et al. Dynvla: Learning world dynamics for action reasoning in autonomous driving. arXiv preprint arXiv:2603.11041,

  30. [39]

    Adaptive 1d video diffusion autoencoder

    Y ao Teng, Minxuan Lin, Xian Liu, Shuai Wang, Xiao Y ang, and Xihui Liu. Adaptive 1d video diffusion autoencoder. arXiv preprint arXiv:2602.04220,

  31. [40]

    Wan: Open and advanced large-scale video generative models

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haim- ing Zhao, Jianxiao Y ang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Y an, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fan...

  32. [41]

    Vtok: A unified video tokenizer with decoupled spatial-temporal latents

    Feng Wang, Yichun Shi, Ceyuan Y ang, Qiushan Guo, Jingxiang Sun, Alan Yuille, and Peng Wang. Vtok: A unified video tokenizer with decoupled spatial-temporal latents. arXiv preprint arXiv:2602.04202, 2026a. Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Jiasong Feng, Zhanjie Zhang, Wanyuan Pang, and Xiaodan Liang. Wisa: World simulator assistant for physics-aware te...

  33. [42]

    Co-evolving latent action world models

    Yucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao, Kaixin Wang, and Jiang Bian. Co-evolving latent action world models. arXiv preprint arXiv:2510.26433, 2025a. Yuchi Wang, Junliang Guo, Xinyi Xie, Tianyu He, Xu Sun, and Jiang Bian. Vidtwin: Video vae with decoupled structure and dynamics. InProceedings of the Computer Vision and Pattern Recognition Confe...

  34. [43]

    Pandora: Towards general world model with natural language actions and video states

    Jiannan Xiang, Guangyi Liu, Yi Gu, Qiyue Gao, Yuting Ning, Yuheng Zha, Zeyu Feng, Tianhua Tao, Shibo Hao, Y emin Shi, et al. Pandora: Towards general world model with natural language actions and video states. arXiv preprint arXiv:2406.09455,

  35. [44]

    Pan: A world model for general, interactable, and long-horizon world simulation

    Jiannan Xiang, Yi Gu, Zihan Liu, Zeyu Feng, Qiyue Gao, Yiyan Hu, Benhao Huang, Guangyi Liu, Yichi Y ang, Kun Zhou, et al. Pan: A world model for general, interactable, and long-horizon world simulation. arXiv preprint arXiv:2511.09057,

  36. [45]

    Y ocausal: How far is video generation from world model? a causality perspective

    Y ou-Zhe Xie, Yu-Hsuan Li, Jie- Ying Lee, Kaipeng Zhang, Yu-Lun Liu, and Zhixiang Wang. Y ocausal: How far is video generation from world model? a causality perspective. arXiv preprint arXiv:2605.30346,

  37. [46]

    Physalign: Physics-coherent image-to-video generation through feature and 3d representation alignment

    Zhexiao Xiong, Yizhi Song, Liu He, Wei Xiong, Yu Yuan, Feng Qiao, and Nathan Jacobs. Physalign: Physics-coherent image-to-video generation through feature and 3d representation alignment. arXiv preprint arXiv:2603.13770,

  38. [47]

    Ultravideo: High-quality uhd video dataset with comprehensive captions

    Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He, Yinan Chen, Yuxuan Cai, Y abiao Wang, Chengjie Wang, Y ong Liu, Xiangtai Li, et al. Ultravideo: High-quality uhd video dataset with comprehensive captions. arXiv preprint arXiv:2506.13691,

  39. [48]

    Rethinking video tokenization: A conditioned diffusion-based approach

    Nianzu Y ang, Pandeng Li, Liming Zhao, Y ang Li, Chen-Wei Xie, Y ehui Tang, Xudong Lu, Zhihang Liu, Yun Zheng, Yu Liu, et al. Rethinking video tokenization: A conditioned diffusion-based approach. arXiv preprint arXiv:2503.03708, 2025a. Sherry Y ang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Leslie Kaelbling, Dale Schu- urmans, and Pieter Abbeel. Le...

  40. [49]

    Vlipp: Towards physically plausible video generation with vision and language informed physical prior

    Xindi Y ang, Baolu Li, Yiming Zhang, Zhenfei Yin, Lei Bai, Liqian Ma, Zhiyong Wang, Jianfei Cai, Tien-Tsin Wong, Huchuan Lu, et al. Vlipp: Towards physically plausible video generation with vision and language informed physical prior. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 12360–12370, 2025b. 24 Preprint Zhuoyi Y a...

  41. [50]

    Clevrer: Collision events for video representation and reasoning

    Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum. Clevrer: Collision events for video representation and reasoning. arXiv preprint arXiv:1910.01442,

  42. [52]

    Representation alignment for generation: Training diffusion transformers is easier than you think

    Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. arXiv preprint arXiv:2410.06940, 2024a. Sihyun Yu, Weili Nie, De-An Huang, Boyi Li, Jinwoo Shin, et al. Efficient video diffusion models via content-frame mo...

  43. [53]

    Dila: Disentangled latent action world models

    Tianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang, and Si Wu. Dila: Disentangled latent action world models. arXiv preprint arXiv:2605.15725, 2026a. Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng, Xiangpeng Wan, Junchi Y an, and Yu Cheng. Videorepa: Learning physics for video generation through relational alignment with foundation models. Advan...

  44. [2018]

    Finite scalar quanti- zation: Vq-vae made simple

    Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. Finite scalar quanti- zation: Vq-vae made simple. In International Conference on Learning Representations , volume 2024, pp. 51772–51783,

  45. [2019]

    Language model beats diffusion– tokenizer is key to visual generation

    Lijun Yu, José Lezama, Nitesh B Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Y ong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, et al. Language model beats diffusion– tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737,

  46. [2020]

    Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling

    Boyu Chen, Yi Chen, Lu Qiu, Jerry Bai, Yuying Ge, and Yixiao Ge. Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling. arXiv preprint arXiv:2604.19734, 2026a. Weiliang Chen, Yuanhui Huang, Xuebo Wang, and Yueqi Duan. Tivtok: Broadcasting time-invariant tokens for scalable video tokenization. arXiv preprint arXi...

  47. [2021]

    Moalign: Motion-centric representation alignment for video diffusion models

    Aritra Bhowmik, Denis Korzhenkov, Cees GM Snoek, Amirhossein Habibian, and Mohsen Ghafoo- rian. Moalign: Motion-centric representation alignment for video diffusion models. arXiv preprint arXiv:2510.19022,

  48. [2022]

    Warp: Whole-body retargeting for learning from offline human demonstrations

    Zhenyang Chen, Chuizheng Kong, Chuye Zhang, Yuanshao Y ang, Lawrence Y Zhu, Shreyas Kousik, and Danfei Xu. Warp: Whole-body retargeting for learning from offline human demonstrations. arXiv preprint arXiv:2606.29940, 2026c. Hanchen Cui and Y ang Gao. A universal world model learned from large scale and diverse videos. In NeurIPS 2023 Foundation Models for...

  49. [2023]

    Cosmos world foundation model platform for physical ai

    Niket Agarwal, Arslan Ali, Maciej Bala, Y ogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chat- topadhyay, Y ongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575,

  50. [2024]

    Physion: Evaluating physical prediction from vision in humans and machines

    Daniel M Bear, Elias Wang, Damian Mrowca, Felix J Binder, Hsiao- Yu Fish Tung, RT Pramod, Cameron Holdaway, Sirui Tao, Kevin Smith, Fan- Yun Sun, et al. Physion: Evaluating physical prediction from vision in humans and machines. arXiv preprint arXiv:2106.08261,

  51. [2025]

    Cosmos 3: Omnimodal world models for physical ai

    Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Y ogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical ai. arXiv preprint arXiv:2606.02800,

  52. [2026]

    World simulation with video foundation models for physical ai

    Arslan Ali, Junjie Bai, Maciej Bala, Y ogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, et al. World simulation with video foundation models for physical ai. arXiv preprint arXiv:2511.00062,