Pith. sign in

REVIEW 3 major objections 6 minor 55 references

Geometry-aware egocentric video synthesis can expand scarce robot demos and raise real OOD manipulation success.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 13:31 UTC pith:6WBUGRLB

load-bearing objection Solid geometry-conditioned egocentric generator with a real robot data-mix win; the OOD lift is directionally credible but under-powered and not cleanly attributed to the two modules. the 3 major comments →

arxiv 2607.28243 v1 pith:6WBUGRLB submitted 2026-07-30 cs.CV cs.AI

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

classification cs.CV cs.AI
keywords egocentric video generationworld-action modelsembodied data augmentation3D scene memoryrotary position embeddingrobot manipulationaction-conditioned synthesis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Collecting diverse first-person manipulation data is expensive, so this paper builds a video generator that turns limited real trajectories into controllable synthetic ones. It keeps a fixed 3D snapshot of the first frame while periodically updating a recent scene state, and it encodes end-effector motion as camera-aware 3D rotary coordinates in cross-attention so generated hands and grippers stay on the commanded path. The claim is that these two geometry controls cut scene drift and action misalignment enough for the videos to be useful training fuel. When 400 real trajectories are paired with 400 generated ones under a fixed downstream policy setup, out-of-distribution real-robot success rises from 77% to 84% on single-arm tasks and from 53% to 70% on dual-arm tasks. A sympathetic reader cares because scarce teleop and human data become a bootstrap for broader embodiment, object, and layout coverage without more physical resets.

Core claim

EgoGenesis shows that an autoregressive video prior conditioned with Online Anchored Projective Memory and Action-3D RoPE can synthesize action-aligned egocentric rollouts whose visual and geometric quality is high enough that mixing equal amounts of real and generated trajectories measurably improves downstream world-action model generalization on held-out real-robot layouts and appearances.

What carries the argument

Two coupled conditioners: OAPM (an immutable first-frame 3D scene anchor plus a replace-only recent snapshot read by gated projective cross-attention) and A3D-RoPE (metric, camera-aware 3D rotary phases on skeleton/end-effector patches in skeleton-to-video cross-attention).

Load-bearing premise

Resimulating a real trajectory under edited looks while keeping the same action labels produces synthetic contact dynamics that help, rather than hurt, a real robot policy on new layouts.

What would settle it

Hold the downstream model and schedule fixed, train on 400 real plus 400 EgoGenesis videos versus 800 real only (or versus 400 real plus an ablated generator without OAPM/A3D-RoPE), and check whether OOD success on the eight-task single- and dual-arm suite still rises by the reported margins over 25 trials per task.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Scarce real egocentric demos can be doubled with appearance- and scene-edited resimulations without changing the action label space.
  • Long autoregressive egocentric rollouts need both a frozen first-frame 3D anchor and online recent-state refresh to limit object and background drift.
  • Metric 3D rotary encoding of end-effectors in cross-attention is a practical control path for hands, grippers, and dual-arm skeletons under a shared prior.
  • Downstream bimanual policies stand to gain more from this style of augmentation than single-arm ones when real OOD data are limited.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the dual-slot memory is doing most of the anti-drift work, similar anchored-plus-refresh memories may transfer to third-person robot world models that also suffer long-horizon identity collapse.
  • Cross-embodiment retargeting (hand skeleton to gripper) in unseen rooms suggests a path to train one generator once and serve multiple robot morphologies from human video.
  • Judge-based physical-faithfulness scores may become a cheap filter for which synthetic clips enter policy training, not only a generation leaderboard metric.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. EgoGenesis is an autoregressive egocentric video–action generator built on a pretrained DiT prior. It adds two geometry-aware conditioners: Online Anchored Projective Memory (OAPM), which keeps an immutable first-frame 3D scene anchor and periodically refreshes a recent slot via VGGT-Ω features and gated cross-attention, and Action-3D Rotary Position Embedding (A3D-RoPE), which injects camera-aware metric skeleton/end-effector coordinates into skeleton-to-video cross-attention. Under matched scene and action conditioning on held-out trajectories, the method reports best or near-best scores on fidelity, keypoint error, physical faithfulness, and consistency (Table 1), with component ablations and geometric-drift curves supporting both modules (Tables 2/4, Fig. 5). The central applied claim is that adding 400 EgoGenesis rollouts to 400 real trajectories, with downstream LingBot-VA architecture and schedule fixed, raises OOD real-robot success from 77% to 84% (single-arm) and 53% to 70% (dual-arm) on a Tianji M6 eight-task suite (Table 3; Supp. Table 6).

Significance. If the generation and data-augmentation results hold under tighter controls, the paper is a useful contribution to embodied data engines: it targets a concrete failure mode of egocentric generators (scene drift and weak metric action control) with two implementable modules, shows multi-metric gains over strong generic and egocentric baselines, and links synthetic rollouts to real-robot OOD generalization under a fixed policy recipe. Strengths include complementary-module ablations, long-horizon depth/camera drift measurements, embodiment-balanced training, and a separately initialized downstream WAM that does not inherit generator weights. The work is systems-empirical rather than theoretical; its value hinges on whether the robot lift can be attributed to the proposed geometry mechanisms and complementary variation rather than to evaluation noise or generic dataset enlargement.

major comments (3)
  1. [Section 5.5, Table 3] Section 5.5 and Table 3 (and Supp. Table 6): the headline claim that EgoGenesis-synthesized data “substantially improve downstream WAM generalization” is only partially isolated. The experiment correctly fixes LingBot-VA architecture and schedule and varies data mix (400 real vs 400 synth. vs 400+400), and synthetic rollouts are described as resimulations that preserve the action-label space while editing appearance/scene. There is, however, no control that pairs the same preserved action labels and edit protocol with a weaker generator (e.g., Wan2.2-5B-Control or RynnWorld-TeleOp from Table 1). Generation ablations (Table 2, Fig. 5) support OAPM/A3D-RoPE for PSNR/Kpt.Err/drift, but the robot table cannot yet attribute the OOD lift to those modules rather than to extra trajectories, appearance duplication, or training-set size. A weaker-generator augmentation arm—or an explicit down-weig
  2. [Table 3, Supp. Table 6, Figure 9] Table 3 / Supp. Table 6: statistical support for the OOD lifts is thin relative to how they are stated. Each task uses 25 trials (4-point granularity); suite aggregates are 100 trials with no confidence intervals, no multi-seed fine-tunes, and no hypothesis test. Under binomial sampling noise alone, the dual-arm 53%→70% change is on the order of a few standard errors, and the single-arm 77%→84% change is weaker. Stage-progress histograms in Fig. 9 are helpful but still single-run. Please report CIs or bootstrap intervals, at least two independent fine-tune seeds (or repeated trial blocks), and temper abstract/conclusion language to match the precision of a 25-trial-per-task protocol.
  3. [Section 5.5 (synthetic data pipeline)] Synthetic supervision pipeline (paragraph preceding Table 3; Section 5.5): the weakest load-bearing assumption is that resimulating real trajectories under edited appearance/scene while “preserving the action-label space” yields net-helpful, label-consistent contact dynamics for policy learning. The manuscript does not quantify residual contact/physics error on the synthetic set used for fine-tuning (e.g., Kpt.Err/Phys.Faith on the 400 augmented trajectories, or failure filters), nor how aggressive the appearance/layout edits are. Without that, gains could partly reflect benign visual diversity or, conversely, be limited by systematic contact mismatch. A short audit of synthetic label fidelity and edit distribution would substantially strengthen attribution of the OOD improvement to complementary variation.
minor comments (6)
  1. [Abstract, Figure 1, Conclusion] Notation inconsistency: the method is written as EGOGENESIS / EgoGenesis / \method, and WAM appears as “W AM” with a stray space in several places (abstract, Fig. 1 caption, conclusion). Normalize naming throughout.
  2. [Section 4, OAPM] Eqs. (4)–(6) and (10)–(12): refresh stride s_r is central to OAPM but never given a default value or sensitivity study in the main text; state the operating value used for Tables 1–3.
  3. [Eq. (8), Supplementary A] A3D-RoPE hyperparameters s=4 and κ=10^4 (Eq. 8) and tube radius r_0=0.10 (Supp. A) are fixed without ablation; a one-row sensitivity or justification would help reproducibility.
  4. [Table 1, Supplementary F] Phys.Faith relies on Kimi K2.7 with a 0–5 rubric (Supp. F). Disclose judge variance (repeated queries) or a small human correlation check so the metric is easier to interpret beside PSNR/Kpt.Err.
  5. [Figure 2, Section 2] Figure 2 and related-work framing: “overfitted hand” / gripper morphology failures are clear qualitatively; a quantitative embodiment-confusion rate on baselines would make the motivation sharper.
  6. [Conclusion, Section 5.1, Figure 5] Typos and wording: “spanding”→“expanding” (Conclusion); “momory slots”→“memory slots” (Section 5.1); “Pl ¨ucker” spacing; arXiv date “30 Jul 2026” looks like a metadata error.

Circularity Check

0 steps flagged

No derivation circularity: empirical systems paper with held-out generation metrics and separately initialized real-robot evaluation.

full rationale

EgoGenesis does not present a first-principles derivation whose outputs reduce to its inputs by construction. OAPM and A3D-RoPE are architectural conditioning mechanisms (dual-slot VGGT-Ω scene memory; metric 3D rotary phases in skeleton-to-video cross-attention) trained with a standard flow-matching objective on a source-balanced corpus; claims are empirical comparisons under matched scene/action conditioning (Table 1) and controlled ablations (Table 2, Fig. 5). Downstream WAM results fix LingBot-VA architecture and schedule, independently re-initialize from the official checkpoint (no weight inheritance from the generator), and measure physical task success on ID/OOD real-robot trials (Table 3, Supp. Table 6). Generation quality is scored against held-out real videos (PSNR/SSIM/LPIPS/Kpt.Err), not against quantities fitted from the same targets. Self-citations are ordinary related-work pointers, not load-bearing uniqueness theorems. Dual use of VGGT-Ω for OAPM features and for Depth/Cam-ERR ablations is an evaluation-bias concern, not a circular reduction of a claimed prediction. No self-definitional loop, fitted-input-as-prediction, or ansatz-smuggled-via-self-citation pattern is present.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

Load-bearing structure is engineering assumptions plus pretrained components, not a short axiom list. The central downstream claim rests on pretrained video/VAE/scene encoders, flow-matching AR generation, the dual-slot memory discipline, metric skeleton unprojection validity, and the premise that synthetic observation diversity with preserved actions transfers to a fixed WAM on one robot platform.

free parameters (5)
  • A3D-RoPE scale s and base κ = s=4, κ=10^4
    Rotary angle uses θ=s X_a κ^{-m/M_a} with s=4 and κ=10^4 chosen in the method; these set metric-to-phase mapping sensitivity.
  • OAPM recent-memory refresh stride s_r
    How often M_r is replaced from decoded AR frames is a design hyperparameter controlling staleness vs. drift.
  • Skeleton tube radius r_0 and adaptive edge width = r_0=0.10
    Patch support for A3D-RoPE uses r_te=r_0+0.2||u_tk-u_tj|| with r_0=0.10; controls which latent patches receive metric rotations.
  • Two-stage training schedule (6k+6k steps) and source mixture weights = 6k SFT + 6k AR; 210K clips
    SFT then AR fine-tuning lengths and 210K source-balanced mix (100K EgoDex, 100K AgiBot, etc.) shape the generator; not predicted from theory.
  • Downstream WAM optimization hyperparameters = lr=1e-5, 20000 steps, video CFG=5, action CFG=1
    LingBot-VA lr 1e-5, 20k steps, CFG scales, denoising steps, action channel selection—fixed across data mixes but still free choices affecting reported SR.
axioms (5)
  • domain assumption A frozen pretrained video diffusion/flow prior (Wan-family) plus VAE latents is a suitable backbone for long egocentric manipulation synthesis.
    Section 3–4 instantiate chunkwise DiT flow matching on this prior; all gains are relative to that substrate.
  • domain assumption VGGT-Ω 3D reconstruction features from anchor⊕recent RGB slots are an adequate scene memory for gated cross-attention without overwriting identity.
    Eq. (4)/(10) and OAPM refresh define M_b entirely via this encoder.
  • domain assumption Rendered skeleton/EEF trajectories with camera intrinsics/extrinsics can be unprojected into a stable anchor-frame metric 3D coordinate field on latent patches.
    A3D-RoPE coordinate construction (Eqs. 16–18) assumes valid depths and calibrated cameras.
  • ad hoc to paper Preserving the action-label space while editing appearance/scene yields complementary supervision rather than conflicting visual-action pairs for policy learning.
    Stated in the downstream protocol before Table 3; required for interpreting SR gains as synthetic-data quality.
  • domain assumption Flow-matching / Euler integration of block-causal DiT with KV cache produces committed history suitable both as world-model rollouts and as training videos.
    Preliminaries Eqs. (1)–(3) and AR procedure in supplement.
invented entities (2)
  • Online Anchored Projective Memory (OAPM) no independent evidence
    purpose: Maintain immutable first-frame 3D anchor plus replace-only recent slot for autoregressive egocentric scene consistency.
    Named module; dual-slot projective read/refresh is the paper's scene-conditioning construct, evaluated via ablations not external prior measurement.
  • Action-3D Rotary Position Embedding (A3D-RoPE) no independent evidence
    purpose: Inject camera-aware metric end-effector/skeleton geometry into skeleton-to-video cross-attention via axis-split RoPE phases.
    Named positional mechanism restricted to skeleton-supported patches; evidence is internal ablations vs RoPE/PRoPE and drift plots.

pith-pipeline@v1.2.0-daily-grok45 · 23542 in / 4233 out tokens · 75150 ms · 2026-07-31T13:31:08.859157+00:00 · methodology

0 comments
read the original abstract

Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77\% to 84\% on single-arm tasks and from 53\% to 70\% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.

Figures

Figures reproduced from arXiv: 2607.28243 by Jinghong Liu, Jinkui Ren, Kaibo Yin, Linfeng Zhang, Shijian Wang, Xiantao Zhang, Xiaobing Tu, Yinggui Wang, Yue Ma, Yuzhou Wu, Zexuan Yan, Zonghang He.

Figure 1
Figure 1. Figure 1: EGOGENESIS expands scarce real demonstrations with controllable egocentric videos that improve downstream WAM general￾ization on real robots. Abstract Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present EGOGENESIS, an egocentric world-action sim￾ulator that synthesizes contro… view at source ↗
Figure 2
Figure 2. Figure 2: Motivation from diagnosing failures in existing egocen [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: EGOGENESIS architecture. Text, noisy video, and skeleton embeddings provide the conditioning inputs. Autoregressive DiT generates and appends video frames. OAPM maintains anchored and recent 3D scene slots and refreshes the recent state online. A3D￾RoPE injects metric action geometry into skeleton-to-video cross-attention. outs for prediction and robot control [4, 11, 15, 18, 20, 33, 40–42, 46, 49, 50]. Re… view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison on flattening shorts and assembling a square table; E [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Accumulated depth and camera errors over an 80-frame [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: A3D-RoPE concentrates spatial influence on end [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: OAPM preserves persistent scene content and updates [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗
Figure 9
Figure 9. Figure 9: OOD task progress with and without EGOGENESIS￾generated training data. outs progress, as demonstrated in [PITH_FULL_IMAGE:figures/full_fig_p008_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Composition of the 210K-clip egocentric training cor [PITH_FULL_IMAGE:figures/full_fig_p012_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Detailed execution sequences for four bimanual and four single-arm real-robot tasks. Five checkpoints per row show the pro [PITH_FULL_IMAGE:figures/full_fig_p014_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Tianji M6 real-robot environment from front, side, [PITH_FULL_IMAGE:figures/full_fig_p015_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Additional comparisons on egg transfer and cup-lid removal; E [PITH_FULL_IMAGE:figures/full_fig_p018_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Cross-embodiment simulation in an unseen environ [PITH_FULL_IMAGE:figures/full_fig_p019_14.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

55 extracted references

  1. [1]

    Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems, 2025

    AgiBot-World-Contributors. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems, 2025

  2. [2]

    Black, Dimitrios Tzionas, and Victoria Fern ´andez Abrevaya

    Rick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas, and Victoria Fern ´andez Abrevaya. Interdyn: Con- trollable interactive dynamics with video diffusion models, 2024

  3. [3]

    Masked visual actions for unified world modeling, 2026

    Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, and Jia-Bin Huang. Masked visual actions for unified world modeling, 2026

  4. [4]

    Fan, et al

    Mido Assran, Adrien Bardes, David P. Fan, et al. V-JEPA 2: Self-supervised video models enable understanding, predic- tion and planning, 2025

  5. [5]

    HOT3D: Hand and object tracking in 3d from egocentric multi-view videos

    Prithviraj Banerjee et al. HOT3D: Hand and object tracking in 3d from egocentric multi-view videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025

  6. [6]

    Gen2Act: Human video gener- ation in novel scenarios enables generalizable robot manipu- lation, 2024

    Homanga Bharadhwaj et al. Gen2Act: Human video gener- ation in novel scenarios enables generalizable robot manipu- lation, 2024

  7. [7]

    Sta- ble video diffusion: Scaling latent video diffusion models to large datasets, 2023

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, et al. Sta- ble video diffusion: Scaling latent video diffusion models to large datasets, 2023

  8. [8]

    VideoCrafter2: Overcoming data limitations for high-quality video diffusion models

    Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. VideoCrafter2: Overcoming data limitations for high-quality video diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  9. [9]

    HandsOnWorld: Unconstrained ego- centric video generation with camera-disentangled hand con- trol, 2026

    Yushuo Chen, Xiaoyu Shi, Xintao Wu, Xintao Wang, Pengfei Wan, and Yebin Liu. HandsOnWorld: Unconstrained ego- centric video generation with camera-disentangled hand con- trol, 2026

  10. [10]

    Semantically controllable augmentations for generaliz- able robot learning.The International Journal of Robotics Research, 2024

    Zoey Qiuyu Chen, Zhao Mandi, Homanga Bharadhwaj, Mo- hit Sharma, Shuran Song, Abhishek Gupta, and Vikash Ku- mar. Semantically controllable augmentations for generaliz- able robot learning.The International Journal of Robotics Research, 2024

  11. [11]

    Large video planner enables generalizable robot control, 2025

    Yilun Du et al. Large video planner enables generalizable robot control, 2025. arXiv preprint

  12. [12]

    TokenFlow: Consistent diffusion features for consistent video editing

    Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. TokenFlow: Consistent diffusion features for consistent video editing. InProceedings of the IEEE/CVF International Conference on Computer Vision, 2023

  13. [13]

    Byrne, et al

    Kristen Grauman, Andrew Westbury, Eugene H. Byrne, et al. Ego4D: Around the world in 3,000 hours of egocentric video. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022

  14. [14]

    Ego-Exo4D: Understanding skilled human activity from first- and third-person perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, et al. Ego-Exo4D: Understanding skilled human activity from first- and third-person perspectives. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  15. [15]

    Ctrl-world: A controllable generative world model for robot manipulation, 2025

    Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, and Chelsea Finn. Ctrl-world: A controllable generative world model for robot manipulation, 2025

  16. [16]

    Egosim: Egocentric world simulator for em- bodied interaction generation, 2026

    Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu, Jiafei Cao, Xihui Liu, Ran Yi, Lizhuang Ma, Jiangmiao Pang, and Xudong Xu. Egosim: Egocentric world simulator for em- bodied interaction generation, 2026. 8

  17. [17]

    Yoon, Mouli Sivapu- rapu, and Jian Zhang

    Ryan Hoque, Peide Huang, David J. Yoon, Mouli Sivapu- rapu, and Jian Zhang. Egodex: Learning dexterous manip- ulation from large-scale egocentric video. InInternational Conference on Learning Representations, 2026

  18. [18]

    Video prediction policy: A generalist robot policy with predictive visual representations

    Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, and Jianyu Chen. Video prediction policy: A generalist robot policy with predictive visual representations. InPro- ceedings of the International Conference on Machine Learn- ing, 2025

  19. [19]

    EgoMimic: Scaling imitation learning via egocentric video

    Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu. EgoMimic: Scaling imitation learning via egocentric video. InProceedings of the IEEE International Conference on Robotics and Automation, 2025

  20. [20]

    Egowam: World action models beyond pixels with in-the-wild egocentric human data, 2026

    Baoyu Li, Xinchen Yin, Mengying Lin, Yixin Zhang, and Danfei Xu. Egowam: World action models beyond pixels with in-the-wild egocentric human data, 2026

  21. [21]

    Egocen- tric world model for photorealistic hand-object interaction synthesis, 2026

    Dayou Li, Lulin Liu, Bangya Liu, Shijie Zhou, Jiu Feng, Ziqi Lu, Minghui Zheng, Chenyu You, and Zhiwen Fan. Egocen- tric world model for photorealistic hand-object interaction synthesis, 2026

  22. [22]

    EgoGen: An egocentric synthetic data generator

    Gen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu, Mi- hai Dusmanu, Yan Zhang, Marc Pollefeys, and Siyu Tang. EgoGen: An egocentric synthetic data generator. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  23. [23]

    Mask2iv: Interaction-centric video generation via mask tra- jectories

    Gen Li, Bo Zhao, Jianfei Yang, and Laura Sevilla-Lara. Mask2iv: Interaction-centric video generation via mask tra- jectories. InProceedings of the AAAI Conference on Artifi- cial Intelligence, pages 6091–6099, 2026

  24. [24]

    Causal world modeling for robot control, 2026

    Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, and Yinghao Xu. Causal world modeling for robot control, 2026

  25. [25]

    Cameras as relative positional encoding, 2025

    Ruilong Li, Brent Yi, Junchen Liu, Hang Gao, Yi Ma, and Angjoo Kanazawa. Cameras as relative positional encoding, 2025

  26. [26]

    Libero: Benchmarking knowl- edge transfer for lifelong robot learning

    Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning. InAdvances in Neu- ral Information Processing Systems, 2023

  27. [27]

    Robotwin: Dual-arm robot benchmark with gen- erative digital twins

    Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Mingkun Xu, Lunkai Lin, Zhiqiang Xie, Mingyu Ding, and Ping Luo. Robotwin: Dual-arm robot benchmark with gen- erative digital twins. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2025

  28. [28]

    Cosmos 3: Omnimodal world models for physical AI, 2026

    NVIDIA, Aditi, Niket Agarwal, et al. Cosmos 3: Omnimodal world models for physical AI, 2026

  29. [29]

    Reconstruct- ing hands in 3d with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstruct- ing hands in 3d with transformers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

  30. [30]

    Humanego: Zero-shot robot learn- ing from minutes of human egocentric videos, 2024

    Kenneth Shaw et al. Humanego: Zero-shot robot learn- ing from minutes of human egocentric videos, 2024. arXiv preprint

  31. [31]

    Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

  32. [32]

    Sruthi Sudhakar, Ruoshi Liu, Basile Van Hoorick, Carl V on- drick, and Richard S. Zemel. Controlling the world by sleight of hand. InComputer Vision – ECCV 2024, pages 414–430, 2024

  33. [33]

    VLA-JEPA: Enhancing vision-language-action model with latent world model, 2026

    Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, and Zhibo Chen. VLA-JEPA: Enhancing vision-language-action model with latent world model, 2026

  34. [34]

    Wan: Open and advanced large-scale video gen- erative models, 2025

    Wan Team. Wan: Open and advanced large-scale video gen- erative models, 2025

  35. [35]

    Qian, Podshara Chanrungmaneekul, and Kaiyu Hang

    Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen, Howard H. Qian, Podshara Chanrungmaneekul, and Kaiyu Hang. Egoinfinity: A web-scale 4d hand-object interac- tion data engine for any-view robot retargeting and video- to-action robot learning, 2026

  36. [36]

    Vggt: Visual geometry grounded trans- former, 2025

    Jianyuan Wang et al. Vggt: Visual geometry grounded trans- former, 2025. arXiv preprint

  37. [37]

    VideoComposer: Compositional video synthesis with motion controllability

    Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Ji- uniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jin- gren Zhou. VideoComposer: Compositional video synthesis with motion controllability. InAdvances in Neural Informa- tion Processing Systems, 2023

  38. [38]

    EgoVid-5M: A large-scale video-action dataset for egocentric video generation, 2024

    Xiaofeng Wang, Kang Zhao, Feng Liu, Jiayu Wang, Gu- osheng Zhao, Xiaoyi Bao, Zheng Zhu, Yingya Zhang, and Xingang Wang. EgoVid-5M: A large-scale video-action dataset for egocentric video generation, 2024

  39. [39]

    Mo- tionCtrl: A unified and flexible motion controller for video generation

    Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tian- shui Chen, Menghan Xia, Ping Luo, and Ying Shan. Mo- tionCtrl: A unified and flexible motion controller for video generation. InACM SIGGRAPH Conference Papers, 2024

  40. [40]

    Unleashing large-scale video generative pre-training for visual robot manipulation

    Hongtao Wu, Ya Jing, Chi-Lam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video generative pre-training for visual robot manipulation. InInternational Conference on Learning Representations, 2024

  41. [41]

    iVideoGPT: Interactive VideoGPTs are scalable world models

    Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long. iVideoGPT: Interactive VideoGPTs are scalable world models. InAdvances in Neu- ral Information Processing Systems, 2024

  42. [42]

    Learning interactive real-world simulators

    Sherry Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Leslie Pack Kaelbling, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. InInternational Conference on Learning Representations, 2024

  43. [43]

    DragNUW A: Fine-grained control in video generation by integrating text, image, and trajectory

    Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Ming Gong, and Nan Duan. DragNUW A: Fine-grained control in video generation by integrating text, image, and trajectory. InInternational Conference on Learning Repre- sentations, 2024

  44. [44]

    Controllable egocentric video generation via occlusion- aware sparse 3d hand joints, 2026

    Chenyangguang Zhang, Botao Ye, Boqi Chen, Alexandros Delitzas, Fangjinhua Wang, Marc Pollefeys, and Xi Wang. Controllable egocentric video generation via occlusion- aware sparse 3d hand joints, 2026. 9

  45. [45]

    EgoLCD: Egocentric video generation with long con- text diffusion, 2025

    Liuzhou Zhang, Jiarui Ye, Yuanlei Wang, Ming Zhong, Mingju Cao, Wanke Xia, Bowen Zeng, Zeyu Zhang, and Hao Tang. EgoLCD: Egocentric video generation with long con- text diffusion, 2025

  46. [46]

    Imagewam: Do world action models really need video generation, or just image editing?, 2026

    Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang, Haitao Lin, Jingbo Zhang, Yao Mu, Xiaokang Yang, Wenjun Zeng, and Xin Jin. Imagewam: Do world action models really need video generation, or just image editing?, 2026

  47. [47]

    Rynnworld-teleop: An action-conditioned world model for digital teleoperation, 2026

    Haoyu Zhao, Xingyue Zhao, Hangyu Li, Biao Gong, Ke- han Li, Siteng Huang, Xin Li, Deli Zhao, and Zhongyu Li. Rynnworld-teleop: An action-conditioned world model for digital teleoperation, 2026

  48. [48]

    TASTE-Rob: Advancing video gener- ation of task-oriented hand-object interaction for generaliz- able robotic manipulation

    Hongxiang Zhao et al. TASTE-Rob: Advancing video gener- ation of task-oriented hand-object interaction for generaliz- able robotic manipulation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025

  49. [49]

    DINO-WM: World models on pre-trained visual features en- able zero-shot planning

    Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto. DINO-WM: World models on pre-trained visual features en- able zero-shot planning. InProceedings of the International Conference on Machine Learning, 2025

  50. [50]

    RoboDreamer: Learning composi- tional world models for robot imagination, 2024

    Siyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li, Dit-Yan Yeung, and Chuang Gan. RoboDreamer: Learning composi- tional world models for robot imagination, 2024

  51. [51]

    [TASK PROMPT]

    Haizhe Zhu et al. Causal forcing: Autoregressive video gen- eration with causal diffusion models, 2026. arXiv preprint. 10 Supplementary Material for EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE Zexuan Yan, Yuzhou Wu, Yue Ma, Zonghang He, Kaibo Yin, Xiaobing Tu Yinggui Wang, Jinkui Ren, Xiantao Zha...

  52. [52]

    contacts, grasps, support, and pushes are credible

  53. [53]

    object motion is caused by plausible manipulator contact

  54. [54]

    objects avoid penetration, floating, and violations of grav- ity

  55. [55]

    Rate from 0–5, respond with only an integer

    contact and object dynamics remain coherent over time. Rate from 0–5, respond with only an integer. For the returned scores phys ∈ {0, . . . ,5}, we report Phys.Faith = sphys 5 .(22) 19