REVIEW 3 major objections 5 minor 52 references
PhiZero learns a compact discrete physical language of state transitions from unlabeled video, reasons future world evolution in that space, then renders it as video.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 01:48 UTC pith:HHDQLKYJ
load-bearing objection Solid systems paper: discrete transition codes + reason-then-render beat direct pixel prediction on physics video benches; novelty is packaging, understanding eval is partly self-referential. the 3 major comments →
PhiZero: A World Model Built Around Physical Language
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A self-supervised discrete physical language of state-transition patterns, learned from in-the-wild video and used in a reason-then-render loop, models physically coherent world evolution better than direct pixel-space prediction and doubles as a controllable, appearance-disentangled interface for action-conditioned simulation and zero-shot motion transfer.
What carries the argument
Physical language: a length-256 discrete FSQ sequence of transition tokens produced by a transition-level Q-Former over adjacent video latents. It is the intermediate that the reasoner predicts and the diffusion decoder renders, separating dynamics inference from pixel synthesis.
Load-bearing premise
The first-frame-conditioned reconstruction bottleneck plus likelihoods over its codes really capture reusable physical state-change structure, not just compressible visual change patterns that happen to score well on the chosen physics tests.
What would settle it
Hold the first frame and action fixed, swap or scramble the physical-language sequence, and check whether rendered outcomes systematically violate the same physical laws the paper’s benchmarks reward; if scrambled codes still look physically fine or transfer fails when appearance is held fixed and only dynamics should move, the central claim fails.
If this is right
- Future video world models can treat dynamics as an explicit discrete reasoning target rather than an implicit side-effect of pixel prediction.
- The same physical-language codes support interactive rollouts, trajectory- and action-conditioned driving and robot simulation, and fine-grained control without redesigning the generator.
- Motion patterns can be transferred zero-shot across object appearance, human-to-robot embodiments, and sim-to-real looks by editing only the first frame.
- Physical plausibility can be scored by comparing reasoner likelihoods of candidate videos’ codes, enabling understanding benchmarks without separate classifiers.
- Scaling the discrete transition vocabulary and reasoner with more motion-rich and simulator data should further improve physical fidelity.
Where Pith is reading between the lines
- If the codes truly factor dynamics from appearance, abundant human video could become a low-cost teacher for robot policies via code-level retargeting rather than paired teleoperation.
- Hierarchical or recurrent prediction over physical-language sequences is a natural next step for long-horizon consistency beyond fixed four-second clips.
- Grounding subsets of the discrete symbols to measurable quantities (contact, force proxies, occlusion events) would test whether the language can move from empirical transitions toward interpretable physical variables.
- The same bottleneck may serve as a drop-in dynamics head for existing VLMs, turning generation priors into pairwise physics discriminators without new labeled physics datasets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. PhiZero proposes a physical world model organized around “physical language”: a compact discrete code of video state transitions learned by a self-supervised tokenizer (transition-level Q-Former + FSQ) that, with the first frame, conditions a pretrained diffusion decoder (reason-then-render; Eq. 1). A VLM-initialized reasoner then autoregressively predicts these codes from an image and textual action intent; the frozen decoder renders them to video. The paper reports leading or near-leading scores on Physics-IQ Verified, PhyGround, and WorldModelBench generation benchmarks, competitive pairwise results on IntPhys2/LikePhys/YoCausal via reasoner likelihoods over tokenizer codes, ablations of tokenizer/reasoner design, and demos of interactive/action-conditioned control and zero-shot motion and sim-to-real transfer by reusing codes under edited first frames.
Significance. If the factorization and discrete transition codes genuinely capture reusable world dynamics rather than decoder-friendly appearance change, the work offers a clear alternative to pure pixel-space world models and a practical interface for control and cross-embodiment transfer. Strengths include multi-benchmark generation evaluation (including rule-based comparison to real physical experiments on Physics-IQ Verified), explicit ablations (Tables 8–9), reconstruction efficiency vs. other compressed tokenizers (Table 7), and concrete transfer/control applications (Figs. 5–8). The contribution is primarily architectural and empirical rather than theoretical; its lasting value hinges on whether physical language is shown to be more than a strong motion bottleneck for Wan2.2-class decoders.
major comments (3)
- [Appendix B.2, Eq. (7); Tables 4–6] Appendix B.2 / Eq. (7): On IntPhys2, LikePhys, and YoCausal, “understanding” is defined as higher autoregressive log-likelihood of z = T_φ(V) under the reasoner trained to imitate the frozen tokenizer. Validity labels never enter training or scoring except through whatever statistics T_φ already packs. This is not an independent physics probe: if the bottleneck mainly stores compressible motion/appearance residuals that correlate with the benchmark construction, pairwise preference can improve without encoding physical laws. The generation wins do not close this gap. Please either (i) evaluate with an external physics judge or frozen independent scorer on the same pairs, (ii) corrupt dynamics while matching low-level motion energy and show likelihood drops selectively, or (iii) substantially soften claims that these tables demonstrate physical understanding.
- [Sec. 3.2, Eq. (1)–(4); Tables 1–3, 9] Sec. 3.2 and the central interpretation of “physical language”: the tokenizer is optimized for reconstruction with a strong pretrained diffusion prior, first-frame appearance anchoring, pure-noise warm-up, and heavy VLM-based data filtering (Fig. 3). Tables 1–3 and Fig. 4 show better physical outcomes than Wan2.2-5B and other video models, but they do not isolate the discrete transition codes from (a) the Wan2.2 decoder prior and LoRA finetuning, (b) the curated motion/physics-heavy corpus (including simulation), and (c) reasoner capacity. A load-bearing control is missing: e.g., the same reason-then-render pipeline with a matched continuous or permuted/codebook-shuffled bottleneck, or pixel/latent autoregression on the identical data and decoder budget. Without that, “explicit physical-language reasoning” remains under-identified relative to “better-conditioned video generation on filte
- [Sec. 4.5; Appendix C.1–C.2; Figs. 7–8] Table 9 and Sec. 4.5 applications: prompt-enhanced Wan2.2-5B reaches only 26.6 IQ-Score vs. 41.2 for PhiZero, which supports the pipeline, but action-conditioned driving/robotics and interactive rollouts appear to use domain-adapted tokenizers/reasoners (Appendix C.1) while zero-shot transfer uses brief source-domain tokenizer finetuning plus GPT-Image first-frame edits (Appendix C.2). The paper should state clearly which results are zero-shot general PhiZero vs. domain-adapted, report quantitative metrics (not only qualitative figures) for trajectory/action fidelity and transfer identity preservation, and avoid implying a single frozen physical language serves all demos without adaptation.
minor comments (5)
- [Sec. 4.1] Clarify FSQ vocabulary size: levels (8,5,5,5,5,5) give 8×5^5 = 25 000 symbols as stated; ensure consistency wherever “25K” and sequence length 256 are discussed relative to 33-frame / 9-latent encoding.
- [Table 2] Table 2: PhiZero’s General Quality (2.93) is below Veo3.1/Wan2.2-14B while Physics Score is best; briefly discuss the quality–physics tradeoff so Overall leadership is not over-read.
- [Sec. 4.4, Fig. 6] Fig. 6 UMAP clusters on driving/manipulation are suggestive but use domain-adapted tokenizers and selected categories; label that limitation in the caption.
- [Sec. 2; Appendix D] Related work (Sec. 2 and Appendix D) is thorough; a short explicit contrast table vs. latent-action models (Genie, LAPA, etc.) and appearance–motion tokenizers would help readers place the novelty claim.
- [Sec. 3.3; throughout] Typos/style: “T wo-stage” spacing (Table 9, Sec. 3.3); “insufficient”/“sufficient” ligature artifacts; footnote on “physical” vs physical laws is helpful—consider elevating one sentence into Sec. 1.
Circularity Check
No load-bearing circular derivation: PhiZero is an empirical reason-then-render system trained and judged against external data/benchmarks; understanding uses own-code likelihoods but external validity labels.
specific steps
-
other
[Appendix B.2, Eq. 7; Sec. 4.2 Video Understanding Results]
"For each pair, we encode both videos into physical-language sequences, compute their log-likelihoods under the Physical Language Reasoner, and select the video with the higher likelihood as valid. ... s(V±, cpair) = Σ_j log pθ(z±_j | I0±, cpair, z±_<j)."
Not circular by construction: external pair labels still decide correctness. Mild coupling only—the scored object is likelihood of the model's own tokenizer codes under a reasoner trained to imitate those codes—so 'physical understanding' is operationalized as in-stack preference for T_φ encodings of benchmark-valid clips, not an independent physics probe. This weakens interpretation of Tables 4–6 but does not make the reported accuracies identically equal to training objectives or fitted parameters.
full rationale
This is a systems/ML paper, not a first-principles derivation. Physical language is learned by self-supervised reconstruction (tokenizer + diffusion decoder, Eqs. 2–4), the reasoner is teacher-forced to imitate frozen tokenizer codes (Eqs. 5–6), and generation claims are scored on external suites (Physics-IQ Verified, PhyGround, WorldModelBench) with independent references or judges. That chain does not reduce a claimed prediction to its fitted inputs by construction. The only soft spot is the understanding protocol (Appendix B.2, Eq. 7): validity is ranked by autoregressive log-likelihood of z = T_φ(V) under the same reasoner trained to imitate T_φ. That couples the metric to the model stack and can reward decoder-friendly compressible change rather than laws—but pair labels (possible/impossible, valid/invalid, forward/reverse) still come from external benchmarks, so accuracy is not tautological (a random ranker stays near chance). No self-citation uniqueness theorem, ansatz smuggled as theorem, or renamed closed-form identity carries the central claim. Score 1 only for that mild metric–stack coupling, not for a forced result.
Axiom & Free-Parameter Ledger
free parameters (6)
- FSQ quantization levels (8,5,5,5,5,5) → ~25K vocab =
(8,5,5,5,5,5); vocab ~25K
- Transition Q-Former query count M=32 and sequence length 256 =
M=32; N=(9-1)*32=256
- LoRA rank 32 on diffusion decoder =
rank 32
- FSQ entropy-loss weight 0.005 and REPA-loss weight 0.1 =
0.005 / 0.1
- Data filtering thresholds and VLM judge criteria for motion/physics/aesthetics
- Two-stage reasoner LR schedule and data mix (5M then ~1M with 200K sim) =
5e-5 then 1e-5; 50K+10K steps
axioms (6)
- ad hoc to paper World evolution p(V,z|I0,c) factorizes as reason-then-render p(z|I0,c)p(V|I0,z) with z sufficient for dynamics (Eq. 1).
- domain assumption Conditioning the decoder on the clean first frame lets a tight discrete bottleneck focus on state change rather than static appearance (§3.2).
- domain assumption Pretrained video diffusion priors can restore fine appearance so reconstruction loss supervises transition codes usefully (flow-matching objective Eq. 4; Wan2.2 init).
- domain assumption Pretrained VLM commonsense is a good prior for autoregressive physical-language prediction from image+text intent (§3.3).
- domain assumption Benchmark suites (Physics-IQ, PhyGround, WorldModelBench, IntPhys2, LikePhys, YoCausal) are adequate proxies for physical coherence and understanding (§4.2, App. B).
- standard math Finite scalar quantization and standard optimizers/deep nets behave as usual.
invented entities (2)
-
Physical language (discrete FSQ state-transition symbol sequences)
no independent evidence
-
Transition-level Q-Former over adjacent latent pairs
no independent evidence
read the original abstract
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.
Figures
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774,
-
[5]
Latent action world models for control with unlabeled trajectories
Marvin Alles, Xingyuan Zhang, Patrick van der Smagt, and Philip Becker-Ehmck. Latent action world models for control with unlabeled trajectories. arXiv preprint arXiv:2512.10016,
-
[6]
Visual imitation enables contextual humanoid control
Arthur Allshire, Hongsuk Choi, Junyi Zhang, David McAllister, Anthony Zhang, Chung Min Kim, Trevor Darrell, Pieter Abbeel, Jitendra Malik, and Angjoo Kanazawa. Visual imitation enables contextual humanoid control. arXiv preprint arXiv:2505.03729,
-
[7]
V-jepa 2: Self-supervised video models enable understanding, prediction and planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Am- mar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-jepa 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985,
-
[8]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025a. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Y ang, Z...
-
[11]
Motus: A unified latent action world model
Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Y ao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model. arXiv preprint arXiv:2512.13030,
-
[12]
Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environ- ments
Florian Bordes, Quentin Garrido, Justine T Kao, Adina Williams, Michael Rabbat, and Emmanuel Dupoux. Intphys 2: Benchmarking intuitive physics understanding in complex synthetic environ- ments. arXiv preprint arXiv:2506.09849,
-
[13]
Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Y an Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, Xu Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. arXiv preprint arXiv:2503.06669, 2025a. Qingwen Bu, Y anting Y ang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Y ao, Ping Luo,...
-
[15]
Zhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba, Joshua B Tenenbaum, and Chuang Gan. Comphy: Compositional physical reasoning of objects and events from videos.arXiv preprint arXiv:2205.01089,
-
[17]
Adaworld: Learning adaptable world models with latent actions
Shenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang, and Chuang Gan. Adaworld: Learning adaptable world models with latent actions. arXiv preprint arXiv:2503.18938,
-
[18]
Learning latent action world models in the wild
Quentin Garrido, Tushar Nagarajan, Basile Terver, Nicolas Ballas, Y ann LeCun, and Michael Rabbat. Learning latent action world models in the wild. arXiv preprint arXiv:2601.05230,
-
[20]
URL https://arxiv.org/abs/2312.11805. Genmo AI. Genmo AI Blog. https://www.genmo.ai/blog,
-
[21]
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Y ang, Anyi Rao, Zhengyang Liang, Y aohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725,
-
[22]
Ltx-2: Efficient joint audio-visual foundation model
Y oav HaCohen, Benny Brazowski, Nisan Chiprut, Y aki Bitterman, Andrew Kvochko, Avishai Berkowitz, Daniel Shalem, Daphna Lifschitz, Dudu Moshe, Eitan Porat, et al. Ltx-2: Efficient joint audio-visual foundation model. arXiv preprint arXiv:2601.03233,
-
[23]
Lamo: Self-supervised latent motion priors for physical realism in video generation
Bo Jiang, Depu Meng, Yihan Hu, Yichen Xie, Tianshuo Xu, and Wei Zhan. Lamo: Self-supervised latent motion priors for physical realism in video generation. arXiv preprint arXiv:2605.23878 , 2026a. Yuxin Jiang, Yuchao Gu, Ivor W Tsang, and Mike Zheng Shou. Olaf-world: Orienting latent actions for video world modeling. arXiv preprint arXiv:2602.10104, 2026b....
-
[24]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603,
-
[25]
Uni- formerv2: Spatiotemporal learning by arming image vits with video uniformer
Kunchang Li, Y ali Wang, Yinan He, Yizhuo Li, Yi Wang, Limin Wang, and Yu Qiao. Uni- formerv2: Spatiotemporal learning by arming image vits with video uniformer. arXiv preprint arXiv:2211.09552,
-
[26]
Dicode: Diffusion-compressed deep to- kens for autoregressive video generation with language models
Yizhuo Li, Yuying Ge, Yixiao Ge, Ying Shan, and Ping Luo. Dicode: Diffusion-compressed deep to- kens for autoregressive video generation with language models. arXiv preprint arXiv:2412.04446,
-
[27]
Sekai: A video dataset towards world exploration
Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin, Ming Li, Shitian Zhao, Zhaopan Xu, Xinyue Li, Yukang Feng, Jianwen Sun, et al. Sekai: A video dataset towards world exploration. Advances in Neural Information Processing Systems, 38, 2026b. Jongbin Lim, Taeyun Ha, Mingi Choi, Jisoo Kim, Byungjun Kim, Subin Jeon, and Hanbyul Joo. Hrdexdb: A paired human-ro...
-
[28]
Open-sora plan: Open-source large video generation model
21 Preprint Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Y ang Y e, Shenghai Yuan, Liuhan Chen, et al. Open-sora plan: Open-source large video generation model. arXiv preprint arXiv:2412.00131,
-
[29]
Phyground: Benchmarking physical reasoning in generative world models
Juyi Lin, Arash Akbari, Yumei He, Lin Zhao, Haichao Zhang, Arman Akbari, Xingchen Xu, Zoe Y Lu, Enfu Nan, Hokin Deng, et al. Phyground: Benchmarking physical reasoning in generative world models. arXiv preprint arXiv:2605.10806,
-
[30]
Libero: Benchmarking knowledge transfer for lifelong robot learning
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36:44776–44791, 2023a. Huaize Liu, Wenzhang Sun, Qiyuan Zhang, Donglin Di, Biao Gong, Hao Li, Chen Wei, and Changqing Zou. Hi-vae: Efficient video autoenc...
-
[31]
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426,
-
[33]
Openvid-1m: A large-scale high-quality dataset for text-to-video generation
Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Y ang, Zhijie Chen, Xiang Li, Jian Y ang, and Ying Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. In International Conference on Learning Representations , volume 2025, pp. 1045–1064,
2025
-
[34]
Omniweaving: Towards unified video generation with free-form composition and reasoning
Kaihang Pan, Qi Tian, Jianwei Zhang, Weijie Kong, Jiangfeng Xiong, Y anxin Long, Shixue Zhang, Haiyi Qiu, Tan Wang, Zheqi Lv, et al. Omniweaving: Towards unified video generation with free-form composition and reasoning. arXiv preprint arXiv:2603.24458,
-
[35]
22 Preprint Tim Rädsch, Yuki M Asano, Hilde Kuehne, Stefan Bauer, Priyank Jaini, Robert Geirhos, and Carsten T Lüth. Physics-iq verified. arXiv preprint arXiv:2606.18943,
-
[36]
Dismo: Disentangled motion representations for open-world motion transfer
Thomas Ressler-Antal, Frank Fundel, Malek Ben Alaya, Stefan Andreas Baumann, Felix Krause, Ming Gui, and Björn Ommer. Dismo: Disentangled motion representations for open-world motion transfer. arXiv preprint arXiv:2511.23428,
-
[37]
Learning to act without actions
Dominik Schmidt and Minqi Jiang. Learning to act without actions. In International Conference on Learning Representations, volume 2024, pp. 9379–9395,
2024
-
[38]
Dynvla: Learning world dynamics for action reasoning in autonomous driving
Shuyao Shang, Bing Zhan, Yunfei Y an, Yuqi Wang, Yingyan Li, Y asong An, Xiaoman Wang, Jierui Liu, Lu Hou, Lue Fan, et al. Dynvla: Learning world dynamics for action reasoning in autonomous driving. arXiv preprint arXiv:2603.11041,
-
[39]
Adaptive 1d video diffusion autoencoder
Y ao Teng, Minxuan Lin, Xian Liu, Shuai Wang, Xiao Y ang, and Xihui Liu. Adaptive 1d video diffusion autoencoder. arXiv preprint arXiv:2602.04220,
-
[40]
Wan: Open and advanced large-scale video generative models
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haim- ing Zhao, Jianxiao Y ang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Y an, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fan...
-
[41]
Vtok: A unified video tokenizer with decoupled spatial-temporal latents
Feng Wang, Yichun Shi, Ceyuan Y ang, Qiushan Guo, Jingxiang Sun, Alan Yuille, and Peng Wang. Vtok: A unified video tokenizer with decoupled spatial-temporal latents. arXiv preprint arXiv:2602.04202, 2026a. Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Jiasong Feng, Zhanjie Zhang, Wanyuan Pang, and Xiaodan Liang. Wisa: World simulator assistant for physics-aware te...
-
[42]
Co-evolving latent action world models
Yucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao, Kaixin Wang, and Jiang Bian. Co-evolving latent action world models. arXiv preprint arXiv:2510.26433, 2025a. Yuchi Wang, Junliang Guo, Xinyi Xie, Tianyu He, Xu Sun, and Jiang Bian. Vidtwin: Video vae with decoupled structure and dynamics. InProceedings of the Computer Vision and Pattern Recognition Confe...
-
[43]
Pandora: Towards general world model with natural language actions and video states
Jiannan Xiang, Guangyi Liu, Yi Gu, Qiyue Gao, Yuting Ning, Yuheng Zha, Zeyu Feng, Tianhua Tao, Shibo Hao, Y emin Shi, et al. Pandora: Towards general world model with natural language actions and video states. arXiv preprint arXiv:2406.09455,
-
[44]
Pan: A world model for general, interactable, and long-horizon world simulation
Jiannan Xiang, Yi Gu, Zihan Liu, Zeyu Feng, Qiyue Gao, Yiyan Hu, Benhao Huang, Guangyi Liu, Yichi Y ang, Kun Zhou, et al. Pan: A world model for general, interactable, and long-horizon world simulation. arXiv preprint arXiv:2511.09057,
-
[45]
Y ocausal: How far is video generation from world model? a causality perspective
Y ou-Zhe Xie, Yu-Hsuan Li, Jie- Ying Lee, Kaipeng Zhang, Yu-Lun Liu, and Zhixiang Wang. Y ocausal: How far is video generation from world model? a causality perspective. arXiv preprint arXiv:2605.30346,
-
[46]
Zhexiao Xiong, Yizhi Song, Liu He, Wei Xiong, Yu Yuan, Feng Qiao, and Nathan Jacobs. Physalign: Physics-coherent image-to-video generation through feature and 3d representation alignment. arXiv preprint arXiv:2603.13770,
-
[47]
Ultravideo: High-quality uhd video dataset with comprehensive captions
Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He, Yinan Chen, Yuxuan Cai, Y abiao Wang, Chengjie Wang, Y ong Liu, Xiangtai Li, et al. Ultravideo: High-quality uhd video dataset with comprehensive captions. arXiv preprint arXiv:2506.13691,
-
[48]
Rethinking video tokenization: A conditioned diffusion-based approach
Nianzu Y ang, Pandeng Li, Liming Zhao, Y ang Li, Chen-Wei Xie, Y ehui Tang, Xudong Lu, Zhihang Liu, Yun Zheng, Yu Liu, et al. Rethinking video tokenization: A conditioned diffusion-based approach. arXiv preprint arXiv:2503.03708, 2025a. Sherry Y ang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Leslie Kaelbling, Dale Schu- urmans, and Pieter Abbeel. Le...
-
[49]
Vlipp: Towards physically plausible video generation with vision and language informed physical prior
Xindi Y ang, Baolu Li, Yiming Zhang, Zhenfei Yin, Lei Bai, Liqian Ma, Zhiyong Wang, Jianfei Cai, Tien-Tsin Wong, Huchuan Lu, et al. Vlipp: Towards physically plausible video generation with vision and language informed physical prior. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 12360–12370, 2025b. 24 Preprint Zhuoyi Y a...
2025
-
[50]
Clevrer: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum. Clevrer: Collision events for video representation and reasoning. arXiv preprint arXiv:1910.01442,
Pith/arXiv arXiv 1910
-
[52]
Representation alignment for generation: Training diffusion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. arXiv preprint arXiv:2410.06940, 2024a. Sihyun Yu, Weili Nie, De-An Huang, Boyi Li, Jinwoo Shin, et al. Efficient video diffusion models via content-frame mo...
Pith/arXiv arXiv 2024
-
[53]
Dila: Disentangled latent action world models
Tianqiu Zhang, Muyang Lyu, Yufan Zhang, Fang Fang, and Si Wu. Dila: Disentangled latent action world models. arXiv preprint arXiv:2605.15725, 2026a. Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng, Xiangpeng Wan, Junchi Y an, and Yu Cheng. Videorepa: Learning physics for video generation through relational alignment with foundation models. Advan...
-
[2018]
Finite scalar quanti- zation: Vq-vae made simple
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. Finite scalar quanti- zation: Vq-vae made simple. In International Conference on Learning Representations , volume 2024, pp. 51772–51783,
2024
-
[2019]
Language model beats diffusion– tokenizer is key to visual generation
Lijun Yu, José Lezama, Nitesh B Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Y ong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, et al. Language model beats diffusion– tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737,
-
[2020]
Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling
Boyu Chen, Yi Chen, Lu Qiu, Jerry Bai, Yuying Ge, and Yixiao Ge. Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling. arXiv preprint arXiv:2604.19734, 2026a. Weiliang Chen, Yuanhui Huang, Xuebo Wang, and Yueqi Duan. Tivtok: Broadcasting time-invariant tokens for scalable video tokenization. arXiv preprint arXi...
-
[2021]
Moalign: Motion-centric representation alignment for video diffusion models
Aritra Bhowmik, Denis Korzhenkov, Cees GM Snoek, Amirhossein Habibian, and Mohsen Ghafoo- rian. Moalign: Motion-centric representation alignment for video diffusion models. arXiv preprint arXiv:2510.19022,
-
[2022]
Warp: Whole-body retargeting for learning from offline human demonstrations
Zhenyang Chen, Chuizheng Kong, Chuye Zhang, Yuanshao Y ang, Lawrence Y Zhu, Shreyas Kousik, and Danfei Xu. Warp: Whole-body retargeting for learning from offline human demonstrations. arXiv preprint arXiv:2606.29940, 2026c. Hanchen Cui and Y ang Gao. A universal world model learned from large scale and diverse videos. In NeurIPS 2023 Foundation Models for...
Pith/arXiv arXiv 2023
-
[2023]
Cosmos world foundation model platform for physical ai
Niket Agarwal, Arslan Ali, Maciej Bala, Y ogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chat- topadhyay, Y ongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575,
-
[2024]
Physion: Evaluating physical prediction from vision in humans and machines
Daniel M Bear, Elias Wang, Damian Mrowca, Felix J Binder, Hsiao- Yu Fish Tung, RT Pramod, Cameron Holdaway, Sirui Tao, Kevin Smith, Fan- Yun Sun, et al. Physion: Evaluating physical prediction from vision in humans and machines. arXiv preprint arXiv:2106.08261,
-
[2025]
Cosmos 3: Omnimodal world models for physical ai
Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Y ogesh Balaji, Josh Bapst, et al. Cosmos 3: Omnimodal world models for physical ai. arXiv preprint arXiv:2606.02800,
-
[2026]
World simulation with video foundation models for physical ai
Arslan Ali, Junjie Bai, Maciej Bala, Y ogesh Balaji, Aaron Blakeman, Tiffany Cai, Jiaxin Cao, Tianshi Cao, Elizabeth Cha, Yu-Wei Chao, et al. World simulation with video foundation models for physical ai. arXiv preprint arXiv:2511.00062,
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.