Pith. sign in

REVIEW 4 major objections 42 references

Debiasing latent actions from unlabeled video makes robot world models follow commands with far less labeled data.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 04:49 UTC pith:VSSITAYM

load-bearing objection Solid empirical fix for confounded latent actions in LAM-based world models; the efficiency and intervention results are real, even if the causal packaging and SAM3-shared metrics overclaim a bit. the 4 major comments →

arxiv 2607.09185 v1 pith:VSSITAYM submitted 2026-07-10 cs.CV cs.RO

Causally Debiased Latent Action Model for Embodied Action Conditioned World Models

classification cs.CV cs.RO
keywords latent action modelsaction-conditioned world modelscausal debiasingembodied AIrobot action followingvideo world modelsadaptation efficiency
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Action-conditioned world models can simulate how a robot and its scene would evolve under a sequence of controls, but they usually need large amounts of costly action-labeled robot video. Latent action models try to dodge that cost by inventing compact action codes from unlabeled videos alone. This paper argues that the usual reconstruction objective for those codes is the problem: it lets backgrounds, camera-like shifts, and non-interacted objects leak into the latent action, so the world model is not really being conditioned on embodiment dynamics. CD-LAM is a three-stage fine-tuning recipe that re-trains the latent action space with embodiment-weighted reconstruction, action-primitive contrastive learning, and zero-transition calibration, then carries the cleaned latents into the world model and finally maps real robot commands into the same space. On 2B and 14B backbones the method cuts action-following error, raises visual fidelity, and matches a much longer baseline adaptation budget with more than twelve times fewer robot-action updates. A sympathetic reader cares because the bottleneck for controllable simulators is often labeled data, and the paper claims a short, targeted clean-up of the latent action is enough to unlock that controllability.

Core claim

Reconstruction-only latent actions entangle embodiment dynamics with action-irrelevant visual factors, confounding the downstream world model; three short fine-tuning objectives that force embodiment focus, action-aware neighborhoods, and calibrated non-collapse produce debiased latents that measurably improve action following, visual fidelity, and robot-action adaptation efficiency on both 2B and 14B action-conditioned world models.

What carries the argument

CD-LAM: a three-objective LAM fine-tuning loss (embodiment-centric weighted reconstruction, action-centric contrastive learning over coarse verb primitives, and latent-space calibration with free-bit KL plus zero-transition anchoring) applied in a three-stage pipeline that first debiases the latent action, then debiases the world model on those latents, then bridges executable robot actions into the same space.

Load-bearing premise

The method assumes that automatic embodiment masks and coarse caption-verb clusters are faithful enough proxies for true action factors that the three losses remove confounding rather than merely re-weighting toward another correlated visual cue.

What would settle it

Hold the world-model architecture fixed, swap only the latent-action encoder for a reconstruction-only baseline, and check whether mean foreground displacement error under identical robot-action sequences still drops by roughly thirty percent and whether zero-action and camera-shift diagnostics remain low; if the gains vanish or the diagnostics stay high, the causal-debiasing claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 0 minor

Summary. The paper argues that reconstruction-only latent action models (LAMs) encode action-irrelevant confounders (background, non-interacted objects, camera-like factors) into the latent action z_t, which then confounds action-conditioned world models (ACWMs). It proposes CD-LAM, a three-stage fine-tuning pipeline whose Stage-1 LAM objectives—embodiment-centric weighted reconstruction (Eqs. 8–9), action-centric contrastive learning over 12-way caption-verb primitives (Eq. 10), and latent-space calibration via free-bit KL plus zero-transition anchoring (Eqs. 11–12)—produce debiased latents. On 2B and 14B DreamDojo-style ACWMs, the method reports lower FDCE after latent-action and robot-action conditioning, higher PSNR/SSIM, stronger zero-action and target-action interventions, and matching of the DreamDojo 50k-update reference with more than 12× fewer robot-action adaptation steps (final 3k/6k checkpoints). Supporting evidence includes a LAM confounding audit (Table I), multi-stage rollout tables (II–III), data-tier scaling (Table IV), objective ablations (Table V), and qualitative rollouts.

Significance. If the results hold under independent scrutiny, the work is a practically useful contribution to embodied world models: it isolates a concrete failure mode of reconstruction-trained LAMs, supplies diagnostic metrics (zero-transition response, camera-shift response, shortcut leakage, FDCE), and shows that a short, targeted LAM fine-tune can improve controllability and cut robot-action adaptation cost by more than an order of magnitude at both 2B and 14B. The multi-stage evaluation design (latent-only rollouts, robot-action adaptation, zero-action and target-transfer interventions), objective ablations that map each loss to a distinct failure mode, and the promised release of debiased LAMs/ACWMs, protocols, and code are genuine strengths. The efficiency finding—that debiasing the condition is cheaper than unlearning a confounded condition downstream—is of clear engineering value for robot world-model pipelines that rely on unlabeled video pretraining.

major comments (4)
  1. The primary action-following metric and a core training signal share the same tooling. Embodiment-centric reconstruction reweights pixels by SAM3 masks M_t (Eqs. 8–9); FDCE seeds tracks inside SAM3 foreground masks and scores only those tracks (Appendix A, Eq. A.4). Action-centric contrast and the shortcut-leakage diagnostic further share the 12-way caption-verb clusters (Appendix B). A model that concentrates capacity on SAM3-selected regions and verb-cluster neighborhoods can therefore improve headline FDCE and Table I diagnostics without necessarily purifying the causal factor A_t from C_t/V_t. Zero-action residual FDCE and full-frame PSNR provide partially independent evidence, but the claimed 35%/30% FDCE reductions and the “causally debiased” framing rest on a metric aligned with the training signal. Please add at least one mask-independent motion metric (e.g., full-frame or random
  2. Tables II and III (and the efficiency curves in Fig. 8) report point estimates only—no standard errors, bootstrap intervals, or multi-seed variance—despite multi-scale claims and percentage reductions that are central to the abstract. With 300 evaluation clips, seed or clip-level variability is estimable. Without it, it is hard to judge whether the 2B→14B baseline FDCE worsening, the 12× efficiency claim, or the per-action breakdowns in Fig. A.1 are stable. Please report uncertainty for the main FDCE/PSNR numbers and for the step at which CD-LAM crosses the DreamDojo reference.
  3. The causal analysis in §III (Eq. 5, Fig. 1d) and the title/abstract language (“causally debiased,” “confounding path”) go beyond what the experiments strictly establish. The interventions show improved sensitivity to the supplied action, and Table I shows reduced responses to static pairs and synthetic shifts, but there is no identification argument or interventional test that separates removal of C_t/V_t from reweighting toward correlated embodiment appearance. Table V’s footnote already notes that removing zero-transition calibration can lower FDCE while failing the camera-shift diagnostic—evidence that FDCE alone does not certify causal purity. Soften or operationalize the causal claims (e.g., “reduces measured action-irrelevant responses and improves action following under fixed context”) unless additional identification-style evidence is added.
  4. Comparisons are limited to DreamDojo with its original reconstruction-trained LAM (§V-A). The related-work section cites Genie, LAPO/LAPA, AdaWorld, Moto, IGOR, and ConLA, but none appear as empirical baselines for the LAM audit or for Stage-2/3 rollouts. At minimum, a reconstruction-only LAM re-finetuned for the same 1k steps without the three CD-LAM terms (or with only L_emb) should be reported as a compute-matched control beyond the partial ablations in Table V, so that gains are not confounded with extra Stage-1/2 fine-tuning budget alone.

Circularity Check

3 steps flagged

No derivation-level circularity; only mild training–diagnostic alignment on Stage-1 audits, while held-out FDCE/efficiency claims remain empirical.

specific steps
  1. other [Sec. IV-B3 Eq. (12); Table I zero-transition diagnostic; App. A Eq. (A.5)]
    "L_zero = E_ot [ ( [ ∥z0_t∥2 / sg(s_Δ)+ε − m_zero ]_+ )^2 ]. ... Zero-transition response (static pair (o_t, o_t); rel. norm ↓) Median response 0.527 → 0.043"

    L_zero is defined to push the norm of duplicated-frame latents below a margin times ordinary-transition RMS. Table I’s primary zero-transition audit is that same relative norm. Reporting a large drop after training with L_zero is measuring the optimized objective, not an independent causal prediction of purified A_t. Downstream FDCE claims do not inherit this reduction by construction.

  2. other [Sec. IV-B1 Eqs. (8)–(9); App. A FDCE definition; Sec. V-A Metrics]
    "Let M_t ∈ [0,1]^{h×w} be the embodiment–object foreground mask obtained by SAM3... W_t = α_fg M_t + α_bg(1−M_t)... Foreground masks select embodiment and interacted-object regions using SAM3, and point tracks are computed only within valid foreground regions"

    Embodiment-centric reconstruction reweights pixels by SAM3 masks; FDCE seeds and scores tracks inside SAM3 foreground masks. A model that concentrates capacity on SAM3-selected regions can improve FDCE partly by aligning to the same proxy used in training, so the causal reading ‘purified A_t vs C_t/V_t’ is not fully independent of the evaluation tooling. This is metric–training alignment, not a definitional identity of the headline efficiency claim.

  3. other [Sec. IV-B2 Eq. (10); Table I shortcut leakage; App. A Eq. (A.6)]
    "L_ctr = (1/|P|) Σ softplus(−y_ij(τ v_i^⊤ v_j + b)) ... Shortcut leakage 0.151 → 0.014 ... L_shortcut = E[cos | same episode, diff. primitive] − E[cos | diff. episode, same primitive]"

    Action-centric contrast pulls same-primitive pairs together and pushes different-primitive pairs apart using the 12-way caption-verb clusters. Shortcut leakage is exactly the cosine gap between those two pair types. Improving Table I’s shortcut-leakage number is therefore largely the direct effect of L_ctr, not an external test that the latent encodes A_t rather than verb-correlated context. Again, this does not force the held-out robot-action FDCE or adaptation-efficiency results.

full rationale

CD-LAM is an empirical methods paper: causal analysis motivates three fine-tuning losses, which are then evaluated on held-out EgoDex/AgiBot rollouts and action-replacement interventions. There is no uniqueness theorem, no self-citation load-bearing premise, no fitted constant renamed as a prediction of an independent quantity, and no first-principles derivation that reduces to its inputs by construction. The central claims (≈30–35% FDCE drop after robot-action adaptation, PSNR gains, >12× fewer adaptation updates) are measured outcomes on external robot data under fixed observation context, not algebraic identities. Mild circularity risk is confined to Stage-1 LAM audits that re-measure quantities the losses directly optimize (zero-transition norm under L_zero; primitive-neighborhood structure under L_ctr) and to shared SAM3 tooling between embodiment-centric training weights and FDCE foreground selection. Those are validity/interpretation caveats, not forced reductions of the main system-level results. Score 2 reflects that minor diagnostic–objective alignment only.

Axiom & Free-Parameter Ledger

7 free parameters · 5 axioms · 3 invented entities

The work is an empirical ML method paper. Its load-bearing content is not a formal derivation from first principles but a set of modeling choices (what counts as embodiment, how to label primitives, how to weight losses) plus free hyperparameters that shape the debiased latent. The invented pieces are the CD-LAM objective package and the FDCE/audit metrics used to claim success.

free parameters (7)
  • λ_ctr(k) contrastive weight schedule
    Step-dependent weight on the action-centric contrastive term; chosen during fine-tuning and not derived from a principle.
  • λ_cal calibration weight
    Scalar balancing free-bit KL plus zero-transition loss against reconstruction.
  • α_fg / α_bg foreground–background reconstruction weights
    Hand-set spatial weights with α_fg > α_bg that define what the model is told is embodiment-centric.
  • m_zero zero-transition margin
    Margin relative to running RMS of ordinary transition latents that defines the calibrated zero.
  • contrastive temperature τ and bias b
    Learned or tuned parameters of the softplus pairwise loss.
  • free-bits KL floor (per-dimension)
    Capacity-control threshold that prevents collapse while still regularizing; value is a design choice.
  • Stage-1/2/3 step budgets (1k / 2k / 3k–6k) and data-tier hours
    Compute and data quantities that determine the reported efficiency claims; not predicted a priori.
axioms (5)
  • domain assumption Next-frame reconstruction sufficiency alone admits action-irrelevant factors into z_t (Eq. 5 factorization into A_t, C_t, V_t).
    Section III-A treats this as the reason reconstruction-only LAMs confound ACWMs; it is a modeling claim, not a theorem.
  • domain assumption SAM3 masks identify embodiment and interacted-object regions that should dominate the action latent.
    Used both as training weight (Eq. 8–9) and in FDCE evaluation; correctness of the method hinges on this proxy.
  • domain assumption 12-way caption-verb clusters are a valid coarse action-primitive space for contrastive structure without executable actions.
    Appendix B; only 36.6% of clean-ego pairs are labeled, long-tailed, and verb-level only.
  • domain assumption A lightweight MLP bridge g_η can map executable robot actions into the debiased latent space without reintroducing the original confounders.
    Stage 3; required for the robot-action following claims.
  • standard math Standard conditional video / diffusion training losses and latent-action encoder–decoder form (Eqs. 1–4).
    Inherited from prior ACWM/LAM literature; not re-derived.
invented entities (3)
  • CD-LAM three-objective package (L_emb + L_ctr + L_cal) no independent evidence
    purpose: Define a debiased latent action z^CD_t that remains drop-in compatible with existing ACWM conditioning.
    The combination and the staged pipeline are the paper’s proposed method; individual losses have precedents but the package is new here.
  • FDCE (Foreground Displacement Chamfer Error) no independent evidence
    purpose: Quantify action following via Chamfer distance on SAM3-masked CoWTracker displacement tracks.
    Primary downstream metric for the controllability claim; defined in Appendix A and not a community standard.
  • LAM confounding audit suite (zero-transition response, camera-shift response, shortcut leakage) no independent evidence
    purpose: Measure action-irrelevant bias in the encoder before world-model rollouts.
    Used to justify that Stage 1 repairs the confounders identified in Section III.

pith-pipeline@v1.1.0-grok45 · 22630 in / 3886 out tokens · 43637 ms · 2026-07-13T04:49:37.483516+00:00 · methodology

0 comments
read the original abstract

Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect in the real world. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from unlabeled videos, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual factors such as backgrounds and untouched objects. In this work, we identify this action-irrelevant bias as a key obstacle to controllable ACWMs and introduce evaluation metrics to measure latent-action bias, action following, and robustness. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three efficient fine-tuning objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and calibrated non-collapsed latent action representations. Experiments on 2B and 14B ACWM backbones show that CD-LAM substantially improves latent-action controllability, downstream robot-action following, visual fidelity, and adaptation efficiency, requiring only 6k fine-tuning steps and more than 12$\times$ fewer robot-action adaptation updates than the baseline.

Figures

Figures reproduced from arXiv: 2607.09185 by Biwei Huang, Fan Feng, Kun Zhou, Lingjun Mao, Ruobing Han, Shuang Liang, Xinyue Wang, Yuchen Yan, Yufan Wei, Zijun Zhang, Ziming Xu, Ziqiao Xi.

Figure 1
Figure 1. Figure 1: Performance overview and the underlying confounding mechanism. CD-LAM substantially improves action following, visual fidelity, and data efficiency on DreamDojo. (a) CD-LAM lowers embodiment action-following error (FDCE) at both 2B and 14B. (b) CD-LAM also raises PSNR at both scales (gains annotated in dB). (c) CD-LAM uses 3k and 6k robot action adaptation updates for the final 2B and 14B checkpoints, comp… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of CD-LAM. CD-LAM debiases the LAM’s latent action space in three stages while keeping the downstream action conditioning format unchanged. Stage 1 (LAM debiased fine-tuning) debiases the LAM with the three CD-LAM objectives; Stage 2 (ACWM debiased fine-tuning) trains the ACWM on the debiased latent actions; Stage 3 (robot action adaptation) aligns executable robot actions to the same space throug… view at source ↗
Figure 3
Figure 3. Figure 3: Zero robot action inputs still produce motion. Frames are generated by the 2B DreamDojo ACWM after robot action adaptation on the AgiBot dataset [18], with the initial frame fixed and all relative robot action inputs replaced with zero, do(ut = 0). The rollout still produces embodiment motion [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The rollout does not follow the transferred target action. The first row shows the source context, and the second row shows the target video that provides the robot action sequence. The third row is the ACWM rollout conditioned on that target action under the fixed source context. The rollout does not reproduce the target embodiment dynamics, indicating target-action misalignment after robot action adaptat… view at source ↗
Figure 5
Figure 5. Figure 5: Representative rollouts after robot action adaptation at 2B and 14B scale. Rows: ground truth, DreamDojo, and CD-LAM at 2B and 14B; all model rows start from the same initial frame and receive the same robot action sequence. DreamDojo’s arm pose drifts from the ground-truth trajectory and scaling to 14B does not repair the drift, while CD-LAM tracks the commanded motion at both scales; scene appearance sta… view at source ↗
Figure 6
Figure 6. Figure 6: (a), FDCE compares induced foreground displacement rather than raw pixel appearance. FDCE is measured in pixels; we report both the mean, which is sensitive to occasional large failures, and the median, which reflects typical behavior. The full metric definition and reporting conventions are provided in Appendix A. As [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Action-following behavior after robot action adaptation, beyond aggregate scores. (a) CD-LAM reduces FDCE across the eight action categories shown (full breakdown in Fig. A.1) at both 2B and 14B scales, showing that the gain is not concentrated in a single primitive. (b) On the 2B ACWM after robot action adaptation, CD-LAM shifts the rollout distribution toward lower FDCE and higher PSNR. E. Robot Action A… view at source ↗
Figure 8
Figure 8. Figure 8: Robot action adaptation efficiency. Under the aligned protocol, CD￾LAM crosses the DreamDojo reference within 3k–4k updates (more than 12× fewer than the 50k reference), and clearly surpasses it by the 6k final checkpoint. Curves show the 14B model on a monitoring subset; absolute values are not directly comparable with Table III. latent action space. More recent variants explore additively compositional l… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

42 extracted references · 25 linked inside Pith

  1. [1]

    World models,

    D. Ha and J. Schmidhuber, “World models,”arXiv preprint arXiv:1803.10122, 2018

  2. [2]

    Mastering diverse do- mains through world models,

    D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap, “Mastering diverse do- mains through world models,”arXiv preprint arXiv:2301.04104, 2023

  3. [3]

    ACWM-Phys: Investigating generalized physical interaction in action- conditioned video world models,

    H. Xue, Y . Chen, L. Ma, Z. Zhao, L. Moukheiber, Y . Zhu, and Y . Chen, “ACWM-Phys: Investigating generalized physical interaction in action- conditioned video world models,”arXiv preprint arXiv:2605.08567, 2026

  4. [4]

    iVideoGPT: Interactive VideoGPTs are scalable world models,

    J. Wu, S. Yin, N. Feng, X. He, D. Li, J. Hao, and M. Long, “iVideoGPT: Interactive VideoGPTs are scalable world models,” inAdvances in Neu- ral Information Processing Systems (NeurIPS), 2024, arXiv:2405.15223

  5. [5]

    Open X-Embodiment: Robotic learning datasets and RT-X models,

    Open X-Embodiment Collaboration, “Open X-Embodiment: Robotic learning datasets and RT-X models,”arXiv preprint arXiv:2310.08864, 2023

  6. [6]

    DROID: A large-scale in-the-wild robot manipulation dataset,

    A. Khazatsky, K. Pertschet al., “DROID: A large-scale in-the-wild robot manipulation dataset,” inRobotics: Science and Systems (RSS), 2024

  7. [7]

    Ego4D: Around the world in 3,000 hours of egocentric video,

    K. Grauman, A. Westbury, E. Byrneet al., “Ego4D: Around the world in 3,000 hours of egocentric video,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  8. [8]

    Ego-Exo4D: Understand- ing skilled human activity from first- and third-person perspectives,

    K. Grauman, A. Westbury, L. Torresaniet al., “Ego-Exo4D: Understand- ing skilled human activity from first- and third-person perspectives,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  9. [9]

    Scaling egocentric vision: The EPIC-KITCHENS dataset,

    D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kaza- kos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray, “Scaling egocentric vision: The EPIC-KITCHENS dataset,”European Conference on Computer Vision (ECCV), 2018

  10. [10]

    Genie: Generative interactive environ- ments,

    J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y . Aytar, S. Bech- tle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, and T. Rocktäschel, “Genie: Generative interactive environ- ments,” inInternatio...

  11. [11]

    Learning to act without actions,

    D. Schmidt and M. Jiang, “Learning to act without actions,” inInter- national Conference on Learning Representations (ICLR), 2024, lAPO; arXiv:2312.10812

  12. [12]

    Latent action pretraining from videos,

    S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo, “Latent action pretraining from videos,” inInternational Conference on Learning Representations (ICLR), 2025

  13. [13]

    AdaWorld: Learning adaptable world models with latent actions,

    S. Gao, S. Zhou, Y . Du, J. Zhang, and C. Gan, “AdaWorld: Learning adaptable world models with latent actions,” inInternational Conference on Machine Learning (ICML), 2025

  14. [14]

    DreamDojo: A generalist robot world model from large-scale human videos,

    S. Gao, W. Liang, K. Zheng, A. Malik, S. Ye, S. Yu, W.-C. Tseng, Y . Dong, K. Mo, C.-H. Lin, Q. Ma, S. Nah, L. Magne, J. Xiang, Y . Xie, R. Zheng, D. Niu, Y . L. Tan, K. R. Zentner, G. Kurian, S. Indupuru, P. Jannaty, J. Gu, J. Zhang, J. Malik, P. Abbeel, M.-Y . Liu, Y . Zhu, J. Jang, and L. Fan, “DreamDojo: A generalist robot world model from large-scale...

  15. [15]

    What do latent action models actually learn?

    C. Zhang, T. Pearce, P. Zhang, K. Wang, X. Chen, W. Shen, L. Zhao, and J. Bian, “What do latent action models actually learn?” inAdvances in Neural Information Processing Systems (NeurIPS), 2025

  16. [16]

    Latent action learning requires supervision in the presence of distractors,

    A. Nikulin, I. Zisman, D. Tarasov, N. Lyubaykin, A. Polubarov, I. Kise- lev, and V . Kurenkov, “Latent action learning requires supervision in the presence of distractors,” inInternational Conference on Machine Learning (ICML), 2025

  17. [17]

    ConLA: Contrastive latent action learning from human videos for robotic manipulation,

    W. Dai, K. Lan, J. Zhou, B. Zhao, X. Su, J. Tong, W. Guan, and S. Yang, “ConLA: Contrastive latent action learning from human videos for robotic manipulation,”arXiv preprint arXiv:2602.00557, 2026

  18. [18]

    AgiBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems,

    AgiBot-World-Contributors, Q. Bu, J. Cai, L. Chen, X. Cui, Y . Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, S. Jiang, Y . Jiang, C. Jing, H. Li, J. Li, C. Liu, Y . Liu, Y . Lu, J. Luo, P. Luo, Y . Mu, Y . Niu, Y . Pan, J. Pang, Y . Qiao, G. Ren, C. Ruan, J. Shan, Y . Shen, C. Shi, M. Shi, M. Shi, C. Sima, J. Song, H. Wang, W. Wang, D. Wei, C. Xie, G. Xu...

  19. [19]

    SAM 3: Segment anything with concepts,

    N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huanget al., “SAM 3: Segment anything with concepts,”arXiv preprint arXiv:2511.16719, 2025

  20. [20]

    Sigmoid loss for language image pre-training,

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inIEEE/CVF International Conference on Computer Vision (ICCV), 2023

  21. [21]

    EgoDex: Learning dexterous manipulation from large-scale egocentric video,

    R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang, “EgoDex: Learning dexterous manipulation from large-scale egocentric video,” 2025. [Online]. Available: https://arxiv.org/abs/2505.11709

  22. [22]

    CoWTracker: Tracking by warping instead of correlation,

    Z. Lai, E. Insafutdinov, E. Sucar, and A. Vedaldi, “CoWTracker: Tracking by warping instead of correlation,” 2026. [Online]. Available: https://arxiv.org/abs/2602.04877

  23. [23]

    Imitating la- tent policies from observation,

    A. D. Edwards, H. Sahni, Y . Schroecker, and C. L. Isbell, “Imitating la- tent policies from observation,” inInternational Conference on Machine Learning (ICML), 2019, iLPO; arXiv:1805.07914

  24. [24]

    Video PreTraining (VPT): Learning to act by watching unlabeled online videos,

    B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecof- fet, B. Houghton, R. Sampedro, and J. Clune, “Video PreTraining (VPT): Learning to act by watching unlabeled online videos,” in Advances in Neural Information Processing Systems (NeurIPS), 2022, arXiv:2206.11795

  25. [25]

    Moto: Latent motion token as the bridging language for learning robot manipulation from videos,

    Y . Chen, Y . Ge, W. Tang, Y . Li, Y . Ge, M. Ding, Y . Shan, and X. Liu, “Moto: Latent motion token as the bridging language for learning robot manipulation from videos,” inIEEE/CVF International Conference on Computer Vision (ICCV), 2025

  26. [26]

    IGOR: Image-GOal representations are the atomic con- trol units for foundation models in embodied AI,

    X. Chen, J. Guo, T. He, C. Zhang, P. Zhang, D. C. Yang, L. Zhao, and J. Bian, “IGOR: Image-GOal representations are the atomic con- trol units for foundation models in embodied AI,”arXiv preprint arXiv:2411.00785, 2024

  27. [27]

    Learning additively compositional latent actions for embodied AI,

    H. Wei, X. Chen, C. Zhang, T. Pearce, J. Chen, A. Lamb, L. Zhao, and J. Bian, “Learning additively compositional latent actions for embodied AI,”arXiv preprint arXiv:2604.03340, 2026

  28. [28]

    Co- evolving latent action world models,

    Y . Wang, F. Zhang, D.-C. Zhan, L. Zhao, K. Wang, and J. Bian, “Co- evolving latent action world models,”arXiv preprint arXiv:2510.26433, 2025

  29. [29]

    MVP-LAM: Learning action-centric latent action via cross-viewpoint reconstruction,

    J. M. Lee, D. Lee, S. Ju, T. Cho, J. W. Koo, L. Zhao, S. Hong, and J. Lee, “MVP-LAM: Learning action-centric latent action via cross-viewpoint reconstruction,”arXiv preprint arXiv:2602.03668, 2026

  30. [30]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  31. [31]

    Learning interactive real-world simula- tors,

    M. Yang, Y . Du, S. K. S. Ghasemipour, J. Tompson, L. P. Kaelbling, D. Schuurmans, and P. Abbeel, “Learning interactive real-world simula- tors,” inInternational Conference on Learning Representations (ICLR), 2024, uniSim; arXiv:2310.06114

  32. [32]

    Cosmos world foundation model platform for physical AI,

    NVIDIA, “Cosmos world foundation model platform for physical AI,” arXiv preprint arXiv:2501.03575, 2025

  33. [33]

    Towards accurate generative models of video: A new metric and challenges,

    T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Towards accurate generative models of video: A new metric and challenges,”arXiv preprint arXiv:1812.01717, 2018, fVD

  34. [34]

    MotionPro: A precise motion controller for image-to-video generation,

    Z. Zhang, F. Long, Z. Qiu, Y . Pan, W. Liu, T. Yao, and T. Mei, “MotionPro: A precise motion controller for image-to-video generation,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

  35. [35]

    TAP-Vid: A benchmark for tracking any point in a video,

    C. Doersch, A. Gupta, L. Markeeva, A. Recasens, L. Smaira, Y . Aytar, J. a. Carreira, A. Zisserman, and Y . Yang, “TAP-Vid: A benchmark for tracking any point in a video,” inAdvances in Neural Information Processing Systems (NeurIPS), 2022, arXiv:2211.03726

  36. [36]

    TAPIR: Tracking any point with per-frame initialization and temporal refinement,

    C. Doersch, Y . Yang, M. Vecerik, D. Gokay, A. Gupta, Y . Aytar, J. a. Carreira, and A. Zisserman, “TAPIR: Tracking any point with per-frame initialization and temporal refinement,” inIEEE/CVF International Con- ference on Computer Vision (ICCV), 2023, arXiv:2306.08637

  37. [37]

    CoTracker: It is better to track together,

    N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “CoTracker: It is better to track together,” inEuropean Conference on Computer Vision (ECCV), 2024

  38. [38]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Dollár, and R. Gir- shick, “Segment anything,” inIEEE/CVF International Conference on Computer Vision (ICCV), 2023, sAM; arXiv:2304.02643

  39. [39]

    SAM 2: Segment anything in images and videos,

    N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. Rädle, C. Rolland, L. Gustafsonet al., “SAM 2: Segment anything in images and videos,”arXiv preprint arXiv:2408.00714, 2024

  40. [40]

    Invariant risk minimization,

    M. Arjovsky, L. Bottou, I. Gulrajani, and D. Lopez-Paz, “Invariant risk minimization,”arXiv preprint arXiv:1907.02893, 2019

  41. [41]

    Shortcut learning in deep neural networks,

    R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,”Nature Machine Intelligence, vol. 2, pp. 665–673, 2020

  42. [42]

    Causal confusion in imita- tion learning,

    P. de Haan, D. Jayaraman, and S. Levine, “Causal confusion in imita- tion learning,” inAdvances in Neural Information Processing Systems (NeurIPS), 2019. APPENDIXA METRICDETAILS PSNR Reporting.We report PSNR as the visual-fidelity metric, computed on full frames in dB. For images normalized to[0,1], PSNR(x,ˆx) = 10 log10 1 MSE(x,ˆx).(A.1) Fig. 1(b) report...