REVIEW 3 major objections 6 minor 55 references
Geometry-aware egocentric video synthesis can expand scarce robot demos and raise real OOD manipulation success.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 13:31 UTC pith:6WBUGRLB
load-bearing objection Solid geometry-conditioned egocentric generator with a real robot data-mix win; the OOD lift is directionally credible but under-powered and not cleanly attributed to the two modules. the 3 major comments →
EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
EgoGenesis shows that an autoregressive video prior conditioned with Online Anchored Projective Memory and Action-3D RoPE can synthesize action-aligned egocentric rollouts whose visual and geometric quality is high enough that mixing equal amounts of real and generated trajectories measurably improves downstream world-action model generalization on held-out real-robot layouts and appearances.
What carries the argument
Two coupled conditioners: OAPM (an immutable first-frame 3D scene anchor plus a replace-only recent snapshot read by gated projective cross-attention) and A3D-RoPE (metric, camera-aware 3D rotary phases on skeleton/end-effector patches in skeleton-to-video cross-attention).
Load-bearing premise
Resimulating a real trajectory under edited looks while keeping the same action labels produces synthetic contact dynamics that help, rather than hurt, a real robot policy on new layouts.
What would settle it
Hold the downstream model and schedule fixed, train on 400 real plus 400 EgoGenesis videos versus 800 real only (or versus 400 real plus an ablated generator without OAPM/A3D-RoPE), and check whether OOD success on the eight-task single- and dual-arm suite still rises by the reported margins over 25 trials per task.
If this is right
- Scarce real egocentric demos can be doubled with appearance- and scene-edited resimulations without changing the action label space.
- Long autoregressive egocentric rollouts need both a frozen first-frame 3D anchor and online recent-state refresh to limit object and background drift.
- Metric 3D rotary encoding of end-effectors in cross-attention is a practical control path for hands, grippers, and dual-arm skeletons under a shared prior.
- Downstream bimanual policies stand to gain more from this style of augmentation than single-arm ones when real OOD data are limited.
Where Pith is reading between the lines
- If the dual-slot memory is doing most of the anti-drift work, similar anchored-plus-refresh memories may transfer to third-person robot world models that also suffer long-horizon identity collapse.
- Cross-embodiment retargeting (hand skeleton to gripper) in unseen rooms suggests a path to train one generator once and serve multiple robot morphologies from human video.
- Judge-based physical-faithfulness scores may become a cheap filter for which synthetic clips enter policy training, not only a generation leaderboard metric.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. EgoGenesis is an autoregressive egocentric video–action generator built on a pretrained DiT prior. It adds two geometry-aware conditioners: Online Anchored Projective Memory (OAPM), which keeps an immutable first-frame 3D scene anchor and periodically refreshes a recent slot via VGGT-Ω features and gated cross-attention, and Action-3D Rotary Position Embedding (A3D-RoPE), which injects camera-aware metric skeleton/end-effector coordinates into skeleton-to-video cross-attention. Under matched scene and action conditioning on held-out trajectories, the method reports best or near-best scores on fidelity, keypoint error, physical faithfulness, and consistency (Table 1), with component ablations and geometric-drift curves supporting both modules (Tables 2/4, Fig. 5). The central applied claim is that adding 400 EgoGenesis rollouts to 400 real trajectories, with downstream LingBot-VA architecture and schedule fixed, raises OOD real-robot success from 77% to 84% (single-arm) and 53% to 70% (dual-arm) on a Tianji M6 eight-task suite (Table 3; Supp. Table 6).
Significance. If the generation and data-augmentation results hold under tighter controls, the paper is a useful contribution to embodied data engines: it targets a concrete failure mode of egocentric generators (scene drift and weak metric action control) with two implementable modules, shows multi-metric gains over strong generic and egocentric baselines, and links synthetic rollouts to real-robot OOD generalization under a fixed policy recipe. Strengths include complementary-module ablations, long-horizon depth/camera drift measurements, embodiment-balanced training, and a separately initialized downstream WAM that does not inherit generator weights. The work is systems-empirical rather than theoretical; its value hinges on whether the robot lift can be attributed to the proposed geometry mechanisms and complementary variation rather than to evaluation noise or generic dataset enlargement.
major comments (3)
- [Section 5.5, Table 3] Section 5.5 and Table 3 (and Supp. Table 6): the headline claim that EgoGenesis-synthesized data “substantially improve downstream WAM generalization” is only partially isolated. The experiment correctly fixes LingBot-VA architecture and schedule and varies data mix (400 real vs 400 synth. vs 400+400), and synthetic rollouts are described as resimulations that preserve the action-label space while editing appearance/scene. There is, however, no control that pairs the same preserved action labels and edit protocol with a weaker generator (e.g., Wan2.2-5B-Control or RynnWorld-TeleOp from Table 1). Generation ablations (Table 2, Fig. 5) support OAPM/A3D-RoPE for PSNR/Kpt.Err/drift, but the robot table cannot yet attribute the OOD lift to those modules rather than to extra trajectories, appearance duplication, or training-set size. A weaker-generator augmentation arm—or an explicit down-weig
- [Table 3, Supp. Table 6, Figure 9] Table 3 / Supp. Table 6: statistical support for the OOD lifts is thin relative to how they are stated. Each task uses 25 trials (4-point granularity); suite aggregates are 100 trials with no confidence intervals, no multi-seed fine-tunes, and no hypothesis test. Under binomial sampling noise alone, the dual-arm 53%→70% change is on the order of a few standard errors, and the single-arm 77%→84% change is weaker. Stage-progress histograms in Fig. 9 are helpful but still single-run. Please report CIs or bootstrap intervals, at least two independent fine-tune seeds (or repeated trial blocks), and temper abstract/conclusion language to match the precision of a 25-trial-per-task protocol.
- [Section 5.5 (synthetic data pipeline)] Synthetic supervision pipeline (paragraph preceding Table 3; Section 5.5): the weakest load-bearing assumption is that resimulating real trajectories under edited appearance/scene while “preserving the action-label space” yields net-helpful, label-consistent contact dynamics for policy learning. The manuscript does not quantify residual contact/physics error on the synthetic set used for fine-tuning (e.g., Kpt.Err/Phys.Faith on the 400 augmented trajectories, or failure filters), nor how aggressive the appearance/layout edits are. Without that, gains could partly reflect benign visual diversity or, conversely, be limited by systematic contact mismatch. A short audit of synthetic label fidelity and edit distribution would substantially strengthen attribution of the OOD improvement to complementary variation.
minor comments (6)
- [Abstract, Figure 1, Conclusion] Notation inconsistency: the method is written as EGOGENESIS / EgoGenesis / \method, and WAM appears as “W AM” with a stray space in several places (abstract, Fig. 1 caption, conclusion). Normalize naming throughout.
- [Section 4, OAPM] Eqs. (4)–(6) and (10)–(12): refresh stride s_r is central to OAPM but never given a default value or sensitivity study in the main text; state the operating value used for Tables 1–3.
- [Eq. (8), Supplementary A] A3D-RoPE hyperparameters s=4 and κ=10^4 (Eq. 8) and tube radius r_0=0.10 (Supp. A) are fixed without ablation; a one-row sensitivity or justification would help reproducibility.
- [Table 1, Supplementary F] Phys.Faith relies on Kimi K2.7 with a 0–5 rubric (Supp. F). Disclose judge variance (repeated queries) or a small human correlation check so the metric is easier to interpret beside PSNR/Kpt.Err.
- [Figure 2, Section 2] Figure 2 and related-work framing: “overfitted hand” / gripper morphology failures are clear qualitatively; a quantitative embodiment-confusion rate on baselines would make the motivation sharper.
- [Conclusion, Section 5.1, Figure 5] Typos and wording: “spanding”→“expanding” (Conclusion); “momory slots”→“memory slots” (Section 5.1); “Pl ¨ucker” spacing; arXiv date “30 Jul 2026” looks like a metadata error.
Circularity Check
No derivation circularity: empirical systems paper with held-out generation metrics and separately initialized real-robot evaluation.
full rationale
EgoGenesis does not present a first-principles derivation whose outputs reduce to its inputs by construction. OAPM and A3D-RoPE are architectural conditioning mechanisms (dual-slot VGGT-Ω scene memory; metric 3D rotary phases in skeleton-to-video cross-attention) trained with a standard flow-matching objective on a source-balanced corpus; claims are empirical comparisons under matched scene/action conditioning (Table 1) and controlled ablations (Table 2, Fig. 5). Downstream WAM results fix LingBot-VA architecture and schedule, independently re-initialize from the official checkpoint (no weight inheritance from the generator), and measure physical task success on ID/OOD real-robot trials (Table 3, Supp. Table 6). Generation quality is scored against held-out real videos (PSNR/SSIM/LPIPS/Kpt.Err), not against quantities fitted from the same targets. Self-citations are ordinary related-work pointers, not load-bearing uniqueness theorems. Dual use of VGGT-Ω for OAPM features and for Depth/Cam-ERR ablations is an evaluation-bias concern, not a circular reduction of a claimed prediction. No self-definitional loop, fitted-input-as-prediction, or ansatz-smuggled-via-self-citation pattern is present.
Axiom & Free-Parameter Ledger
free parameters (5)
- A3D-RoPE scale s and base κ =
s=4, κ=10^4
- OAPM recent-memory refresh stride s_r
- Skeleton tube radius r_0 and adaptive edge width =
r_0=0.10
- Two-stage training schedule (6k+6k steps) and source mixture weights =
6k SFT + 6k AR; 210K clips
- Downstream WAM optimization hyperparameters =
lr=1e-5, 20000 steps, video CFG=5, action CFG=1
axioms (5)
- domain assumption A frozen pretrained video diffusion/flow prior (Wan-family) plus VAE latents is a suitable backbone for long egocentric manipulation synthesis.
- domain assumption VGGT-Ω 3D reconstruction features from anchor⊕recent RGB slots are an adequate scene memory for gated cross-attention without overwriting identity.
- domain assumption Rendered skeleton/EEF trajectories with camera intrinsics/extrinsics can be unprojected into a stable anchor-frame metric 3D coordinate field on latent patches.
- ad hoc to paper Preserving the action-label space while editing appearance/scene yields complementary supervision rather than conflicting visual-action pairs for policy learning.
- domain assumption Flow-matching / Euler integration of block-causal DiT with KV cache produces committed history suitable both as world-model rollouts and as training videos.
invented entities (2)
-
Online Anchored Projective Memory (OAPM)
no independent evidence
-
Action-3D Rotary Position Embedding (A3D-RoPE)
no independent evidence
read the original abstract
Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77\% to 84\% on single-arm tasks and from 53\% to 70\% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.
Figures
Reference graph
Works this paper leans on
-
[1]
Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems, 2025
AgiBot-World-Contributors. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems, 2025
2025
-
[2]
Black, Dimitrios Tzionas, and Victoria Fern ´andez Abrevaya
Rick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas, and Victoria Fern ´andez Abrevaya. Interdyn: Con- trollable interactive dynamics with video diffusion models, 2024
2024
-
[3]
Masked visual actions for unified world modeling, 2026
Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, and Jia-Bin Huang. Masked visual actions for unified world modeling, 2026
2026
-
[4]
Fan, et al
Mido Assran, Adrien Bardes, David P. Fan, et al. V-JEPA 2: Self-supervised video models enable understanding, predic- tion and planning, 2025
2025
-
[5]
HOT3D: Hand and object tracking in 3d from egocentric multi-view videos
Prithviraj Banerjee et al. HOT3D: Hand and object tracking in 3d from egocentric multi-view videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025
2025
-
[6]
Gen2Act: Human video gener- ation in novel scenarios enables generalizable robot manipu- lation, 2024
Homanga Bharadhwaj et al. Gen2Act: Human video gener- ation in novel scenarios enables generalizable robot manipu- lation, 2024
2024
-
[7]
Sta- ble video diffusion: Scaling latent video diffusion models to large datasets, 2023
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, et al. Sta- ble video diffusion: Scaling latent video diffusion models to large datasets, 2023
2023
-
[8]
VideoCrafter2: Overcoming data limitations for high-quality video diffusion models
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. VideoCrafter2: Overcoming data limitations for high-quality video diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
2024
-
[9]
HandsOnWorld: Unconstrained ego- centric video generation with camera-disentangled hand con- trol, 2026
Yushuo Chen, Xiaoyu Shi, Xintao Wu, Xintao Wang, Pengfei Wan, and Yebin Liu. HandsOnWorld: Unconstrained ego- centric video generation with camera-disentangled hand con- trol, 2026
2026
-
[10]
Semantically controllable augmentations for generaliz- able robot learning.The International Journal of Robotics Research, 2024
Zoey Qiuyu Chen, Zhao Mandi, Homanga Bharadhwaj, Mo- hit Sharma, Shuran Song, Abhishek Gupta, and Vikash Ku- mar. Semantically controllable augmentations for generaliz- able robot learning.The International Journal of Robotics Research, 2024
2024
-
[11]
Large video planner enables generalizable robot control, 2025
Yilun Du et al. Large video planner enables generalizable robot control, 2025. arXiv preprint
2025
-
[12]
TokenFlow: Consistent diffusion features for consistent video editing
Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. TokenFlow: Consistent diffusion features for consistent video editing. InProceedings of the IEEE/CVF International Conference on Computer Vision, 2023
2023
-
[13]
Byrne, et al
Kristen Grauman, Andrew Westbury, Eugene H. Byrne, et al. Ego4D: Around the world in 3,000 hours of egocentric video. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022
2022
-
[14]
Ego-Exo4D: Understanding skilled human activity from first- and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, et al. Ego-Exo4D: Understanding skilled human activity from first- and third-person perspectives. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
2024
-
[15]
Ctrl-world: A controllable generative world model for robot manipulation, 2025
Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, and Chelsea Finn. Ctrl-world: A controllable generative world model for robot manipulation, 2025
2025
-
[16]
Egosim: Egocentric world simulator for em- bodied interaction generation, 2026
Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu, Jiafei Cao, Xihui Liu, Ran Yi, Lizhuang Ma, Jiangmiao Pang, and Xudong Xu. Egosim: Egocentric world simulator for em- bodied interaction generation, 2026. 8
2026
-
[17]
Yoon, Mouli Sivapu- rapu, and Jian Zhang
Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapu- rapu, and Jian Zhang. Egodex: Learning dexterous manip- ulation from large-scale egocentric video. InInternational Conference on Learning Representations, 2026
2026
-
[18]
Video prediction policy: A generalist robot policy with predictive visual representations
Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, and Jianyu Chen. Video prediction policy: A generalist robot policy with predictive visual representations. InPro- ceedings of the International Conference on Machine Learn- ing, 2025
2025
-
[19]
EgoMimic: Scaling imitation learning via egocentric video
Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. EgoMimic: Scaling imitation learning via egocentric video. InProceedings of the IEEE International Conference on Robotics and Automation, 2025
2025
-
[20]
Egowam: World action models beyond pixels with in-the-wild egocentric human data, 2026
Baoyu Li, Xinchen Yin, Mengying Lin, Yixin Zhang, and Danfei Xu. Egowam: World action models beyond pixels with in-the-wild egocentric human data, 2026
2026
-
[21]
Egocen- tric world model for photorealistic hand-object interaction synthesis, 2026
Dayou Li, Lulin Liu, Bangya Liu, Shijie Zhou, Jiu Feng, Ziqi Lu, Minghui Zheng, Chenyu You, and Zhiwen Fan. Egocen- tric world model for photorealistic hand-object interaction synthesis, 2026
2026
-
[22]
EgoGen: An egocentric synthetic data generator
Gen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu, Mi- hai Dusmanu, Yan Zhang, Marc Pollefeys, and Siyu Tang. EgoGen: An egocentric synthetic data generator. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
2024
-
[23]
Mask2iv: Interaction-centric video generation via mask tra- jectories
Gen Li, Bo Zhao, Jianfei Yang, and Laura Sevilla-Lara. Mask2iv: Interaction-centric video generation via mask tra- jectories. InProceedings of the AAAI Conference on Artifi- cial Intelligence, pages 6091–6099, 2026
2026
-
[24]
Causal world modeling for robot control, 2026
Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, and Yinghao Xu. Causal world modeling for robot control, 2026
2026
-
[25]
Cameras as relative positional encoding, 2025
Ruilong Li, Brent Yi, Junchen Liu, Hang Gao, Yi Ma, and Angjoo Kanazawa. Cameras as relative positional encoding, 2025
2025
-
[26]
Libero: Benchmarking knowl- edge transfer for lifelong robot learning
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning. InAdvances in Neu- ral Information Processing Systems, 2023
2023
-
[27]
Robotwin: Dual-arm robot benchmark with gen- erative digital twins
Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Mingkun Xu, Lunkai Lin, Zhiqiang Xie, Mingyu Ding, and Ping Luo. Robotwin: Dual-arm robot benchmark with gen- erative digital twins. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2025
2025
-
[28]
Cosmos 3: Omnimodal world models for physical AI, 2026
NVIDIA, Aditi, Niket Agarwal, et al. Cosmos 3: Omnimodal world models for physical AI, 2026
2026
-
[29]
Reconstruct- ing hands in 3d with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstruct- ing hands in 3d with transformers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
2024
-
[30]
Humanego: Zero-shot robot learn- ing from minutes of human egocentric videos, 2024
Kenneth Shaw et al. Humanego: Zero-shot robot learn- ing from minutes of human egocentric videos, 2024. arXiv preprint
2024
-
[31]
Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
2024
-
[32]
Sruthi Sudhakar, Ruoshi Liu, Basile Van Hoorick, Carl V on- drick, and Richard S. Zemel. Controlling the world by sleight of hand. InComputer Vision – ECCV 2024, pages 414–430, 2024
2024
-
[33]
VLA-JEPA: Enhancing vision-language-action model with latent world model, 2026
Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, and Zhibo Chen. VLA-JEPA: Enhancing vision-language-action model with latent world model, 2026
2026
-
[34]
Wan: Open and advanced large-scale video gen- erative models, 2025
Wan Team. Wan: Open and advanced large-scale video gen- erative models, 2025
2025
-
[35]
Qian, Podshara Chanrungmaneekul, and Kaiyu Hang
Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen, Howard H. Qian, Podshara Chanrungmaneekul, and Kaiyu Hang. Egoinfinity: A web-scale 4d hand-object interac- tion data engine for any-view robot retargeting and video- to-action robot learning, 2026
2026
-
[36]
Vggt: Visual geometry grounded trans- former, 2025
Jianyuan Wang et al. Vggt: Visual geometry grounded trans- former, 2025. arXiv preprint
2025
-
[37]
VideoComposer: Compositional video synthesis with motion controllability
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Ji- uniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jin- gren Zhou. VideoComposer: Compositional video synthesis with motion controllability. InAdvances in Neural Informa- tion Processing Systems, 2023
2023
-
[38]
EgoVid-5M: A large-scale video-action dataset for egocentric video generation, 2024
Xiaofeng Wang, Kang Zhao, Feng Liu, Jiayu Wang, Gu- osheng Zhao, Xiaoyi Bao, Zheng Zhu, Yingya Zhang, and Xingang Wang. EgoVid-5M: A large-scale video-action dataset for egocentric video generation, 2024
2024
-
[39]
Mo- tionCtrl: A unified and flexible motion controller for video generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tian- shui Chen, Menghan Xia, Ping Luo, and Ying Shan. Mo- tionCtrl: A unified and flexible motion controller for video generation. InACM SIGGRAPH Conference Papers, 2024
2024
-
[40]
Unleashing large-scale video generative pre-training for visual robot manipulation
Hongtao Wu, Ya Jing, Chi-Lam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video generative pre-training for visual robot manipulation. InInternational Conference on Learning Representations, 2024
2024
-
[41]
iVideoGPT: Interactive VideoGPTs are scalable world models
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long. iVideoGPT: Interactive VideoGPTs are scalable world models. InAdvances in Neu- ral Information Processing Systems, 2024
2024
-
[42]
Learning interactive real-world simulators
Sherry Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. InInternational Conference on Learning Representations, 2024
2024
-
[43]
DragNUW A: Fine-grained control in video generation by integrating text, image, and trajectory
Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Ming Gong, and Nan Duan. DragNUW A: Fine-grained control in video generation by integrating text, image, and trajectory. InInternational Conference on Learning Repre- sentations, 2024
2024
-
[44]
Controllable egocentric video generation via occlusion- aware sparse 3d hand joints, 2026
Chenyangguang Zhang, Botao Ye, Boqi Chen, Alexandros Delitzas, Fangjinhua Wang, Marc Pollefeys, and Xi Wang. Controllable egocentric video generation via occlusion- aware sparse 3d hand joints, 2026. 9
2026
-
[45]
EgoLCD: Egocentric video generation with long con- text diffusion, 2025
Liuzhou Zhang, Jiarui Ye, Yuanlei Wang, Ming Zhong, Mingju Cao, Wanke Xia, Bowen Zeng, Zeyu Zhang, and Hao Tang. EgoLCD: Egocentric video generation with long con- text diffusion, 2025
2025
-
[46]
Imagewam: Do world action models really need video generation, or just image editing?, 2026
Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang, Haitao Lin, Jingbo Zhang, Yao Mu, Xiaokang Yang, Wenjun Zeng, and Xin Jin. Imagewam: Do world action models really need video generation, or just image editing?, 2026
2026
-
[47]
Rynnworld-teleop: An action-conditioned world model for digital teleoperation, 2026
Haoyu Zhao, Xingyue Zhao, Hangyu Li, Biao Gong, Ke- han Li, Siteng Huang, Xin Li, Deli Zhao, and Zhongyu Li. Rynnworld-teleop: An action-conditioned world model for digital teleoperation, 2026
2026
-
[48]
TASTE-Rob: Advancing video gener- ation of task-oriented hand-object interaction for generaliz- able robotic manipulation
Hongxiang Zhao et al. TASTE-Rob: Advancing video gener- ation of task-oriented hand-object interaction for generaliz- able robotic manipulation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025
2025
-
[49]
DINO-WM: World models on pre-trained visual features en- able zero-shot planning
Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto. DINO-WM: World models on pre-trained visual features en- able zero-shot planning. InProceedings of the International Conference on Machine Learning, 2025
2025
-
[50]
RoboDreamer: Learning composi- tional world models for robot imagination, 2024
Siyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, and Chuang Gan. RoboDreamer: Learning composi- tional world models for robot imagination, 2024
2024
-
[51]
Haizhe Zhu et al. Causal forcing: Autoregressive video gen- eration with causal diffusion models, 2026. arXiv preprint. 10 Supplementary Material for EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE Zexuan Yan, Yuzhou Wu, Yue Ma, Zonghang He, Kaibo Yin, Xiaobing Tu Yinggui Wang, Jinkui Ren, Xiantao Zha...
arXiv 2026
-
[52]
contacts, grasps, support, and pushes are credible
-
[53]
object motion is caused by plausible manipulator contact
-
[54]
objects avoid penetration, floating, and violations of grav- ity
-
[55]
Rate from 0–5, respond with only an integer
contact and object dynamics remain coherent over time. Rate from 0–5, respond with only an integer. For the returned scores phys ∈ {0, . . . ,5}, we report Phys.Faith = sphys 5 .(22) 19
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.