Pith. sign in

REVIEW 3 major objections 6 minor 35 references

WALA learns executable robot latent actions from both labeled demonstrations and unlabeled videos by predicting future semantic and geometric scene changes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 05:51 UTC pith:LGQXWUBT

load-bearing objection Solid empirical VLA paper that turns unlabeled video into joint latent-action + dynamics supervision; SOTA and low-label claims are real contributions but rest on unreproduced baselines and single-seed rates. the 3 major comments →

arxiv 2607.11397 v1 pith:LGQXWUBT submitted 2026-07-13 cs.RO

WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos

classification cs.RO
keywords latent actionsvision-language-actionaction-free videoworld modelrobot manipulationDINOv3semantic-geometric dynamicspolicy learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Robot policies usually need expensive action-labeled demonstrations, while abundant human and robot videos show rich physical interaction but almost never include robot-executable action labels. WALA argues that those videos still carry usable control signal if you treat the evolution of a scene—what changes between the current frame and sparsely sampled future frames—as a latent action. It pretrains an encoder-decoder that explains those changes in frozen DINOv3 feature space and dense depth rather than by reconstructing pixels, then freezes the encoder as a stable target provider and keeps the decoder as a trainable latent world model while a vision-language backbone is trained. The backbone’s latent actions are pulled toward robot motor commands when labels exist, and toward the same future deltas when they do not. The result is a policy that can absorb dynamics supervision from action-free video at training time and, at deployment, needs only the backbone and action head.

Core claim

Executable latent actions can be learned jointly from action-labeled robot demonstrations and action-free videos by supervising a vision-language backbone with three signals at once: robot action prediction, matching of frozen latent action targets extracted from observed future semantic-geometric deltas, and prediction of those same future deltas through a trainable latent world model. Action-free videos therefore contribute dynamics supervision without robot action annotations, and the policy still deploys without running the world model.

What carries the argument

Semantic-geometric latent action model: an encoder that maps current DINOv3 features, depth, and sparse future feature/depth deltas into latent action tokens, plus a decoder that predicts those future deltas; during policy training the encoder is frozen as a target provider and the decoder acts as a latent world model.

Load-bearing premise

That future differences in frozen image features and depth maps are a good enough stand-in for the actions a robot should take, even when the video is human egocentric footage with no robot action labels.

What would settle it

Fix labeled demonstrations, then add action-free videos whose frames are temporally shuffled so feature and depth deltas no longer match real physical evolution; if RoboCasa and real multi-task success still rise by the same margin as with correctly ordered videos, the claim that future semantic-geometric deltas supply useful dynamics supervision is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • With only 10% of RoboCasa action labels, adding action-free videos can raise success from the low-50s toward the high-60s without extra robot annotations.
  • A single multi-task real-robot policy can nearly match 200-demo performance using 50 demos plus 400 similar-scene human videos per task.
  • Human egocentric videos of an unseen task (e.g., bread pick-and-place) can inject dynamics that support zero-shot robot success on that task.
  • Deployment stays a pure vision-language-action forward pass; latent encoder, depth estimator, and world-model decoder are training-only.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If DINOv3-plus-depth deltas transfer across embodiments, the same pretraining recipe could absorb large web-scale human video corpora without embodiment-specific retargeting.
  • The three-way loss (action, latent match, future prediction) suggests a general template for any VLA: keep a frozen transition encoder as a dynamics teacher even when world-model rollout is never used at test time.
  • Failure modes should concentrate on tasks where semantic change is subtle but contact geometry is critical, or vice versa—ablating depth or DINOv3 alone would map that boundary.
  • Scaling LAM pretraining to larger, more diverse video sets, as the authors flag, is the most direct next test of whether the latent action space saturates or keeps improving.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. WALA proposes a two-stage framework for learning executable latent actions from both action-labeled robot demonstrations and action-free videos. Stage 1 pretrains a semantic-geometric latent action model (LAM) that encodes observed future deltas in frozen DINOv3 feature space and dense depth space (Eqs. 1–6, §III.B) and decodes predicted deltas without pixel reconstruction. Stage 2 freezes the LAM encoder as a target provider, keeps the decoder as a trainable latent world model, and trains a Qwen3-VL-4B vision-language backbone so that its latent actions are jointly supervised by robot action prediction, latent-target matching, and future dynamics prediction (Eqs. 7–9, §III.C), with the action loss masked on action-free data. At inference only the backbone and action head are used. Empirically, the paper reports 90.6%/92.8% on RoboTwin Clean/Random (Table I), a claimed SOTA 75.2% average on RoboCasa-GR1-Tabletop (Table II), ablations isolating each supervision term (Table III), labeled and action-free scaling curves (Fig. 4), and real-robot multi-task gains including a low-label setting (50 demos + 400 human videos ≈ 200 demos) and a zero-shot bread pick-and-place transfer (Table IV, Fig. 8).

Significance. If the results hold under controlled re-evaluation, the work is significant for robot learning: it gives a concrete training-time interface that lets large action-free video corpora supply dynamics supervision for VLAs without world-model cost at deployment. Strengths that should be credited include (i) a clear three-way joint objective that cleanly separates control, latent-target, and dynamics losses with an explicit mask for unlabeled video; (ii) controlled Base-Policy ablations and component ablations (Table III) that isolate LAM pretraining, semantic vs. geometric prediction, and target matching; (iii) data-scaling curves that separately vary labeled demos and action-free videos (Fig. 4); and (iv) real-robot multi-task and low-label/zero-shot human-video experiments with measured latency. These elements go beyond pure representation pretraining and address a practical bottleneck in scaling manipulation policies.

major comments (3)
  1. [Table II, §IV.B] Table II (and the abstract SOTA claim of 75.2% vs. DIAL 70.2%): all success rates are single-point means over 50 episodes/task with no standard errors, no multi-seed runs, and no re-implementation of the strongest baselines under a matched backbone, data split, or evaluation protocol. Baselines are “collected from publicly released reports.” The only fully controlled comparison is Base Policy vs. WALA (same Qwen3-VL-4B, action loss only). Without variance estimates or matched re-runs, the 5.0-point SOTA gap is not yet load-bearing evidence; please report multi-seed means ± stderr (or bootstrap CIs) for WALA and at least re-evaluate the top 1–2 baselines under the same protocol, or qualify the SOTA claim accordingly.
  2. [Fig. 4, Table IV, §IV.C / §IV.F] Fig. 4 (right) and Table IV: the central claim that action-free videos nearly replace robot demos (10% labels + videos → 67.8% vs. Base 100% labels 54.2%; real-robot 50 demos + 400 human videos → 74.2% ≈ 200 demos 75.0%) rests on the same single-run success rates (N=50 sim episodes/task; N=30 real trials/task) without error bars or seeds. These are the paper’s most consequential practical claims. Please add multi-seed or multi-split statistics and, for the real-robot low-label setting, report per-task binomial confidence intervals so that “nearly matches” can be assessed quantitatively rather than by point estimates alone.
  3. [§III.B–C, Eqs. (2)–(7), Table III] §III.B–C and the weakest modeling assumption: latent targets are defined as frozen-encoder outputs of DINOv3 and depth deltas (Eqs. 2, 7). Ablations (Table III) show that semantic+geometric world losses help, but there is no analysis of how well these targets align with robot action manifolds across embodiments (human egocentric vs. robot multi-view), nor of sensitivity to K, τ_k, or depth-estimator noise. A short diagnostic—e.g., correlation of z* with ground-truth robot actions on labeled data, or cross-embodiment retrieval beyond Fig. 5—would make the transfer claim falsifiable rather than only performance-supported.
minor comments (6)
  1. [Title, Abstract] Title and running text inconsistently space the acronym (“W ALA” / “WALA”); standardize to WALA throughout.
  2. [§III.B–C] Eqs. (4)–(6) and (9): λ_cos, λ_grad, λ_dep, λ_align, λ_wm and the sampling schedule (K, τ_k) are free hyperparameters but values and selection procedure are not stated; add a short hyperparameter table or appendix note.
  3. [Fig. 3, §IV.B] Fig. 3 visualizes predicted future DINOv3/depth states but does not quantify prediction error (e.g., feature ℓ1/cosine or depth MAE on held-out transitions); a small quantitative panel would strengthen the qualitative claim.
  4. [Abstract, Table I] Table I Clean: WALA (90.6) is below LingBot-VA (92.9) and Fast-WAM (91.9); the text correctly notes Random is best, but the abstract’s “strong performance on RoboTwin” could briefly acknowledge the Clean ranking to avoid overstatement.
  5. [§IV.F, §V] Real-world zero-shot transfer is reported for a single OOD task (bread pick-and-place, 3/10 → 9/10). The limitations section already flags this; consider moving a one-sentence caveat into the main real-world results paragraph.
  6. [References] References include several 2025–2026 arXiv entries; ensure citation keys and years are consistent with the submitted bibliography style.

Circularity Check

0 steps flagged

No circularity: WALA is an empirical two-stage training pipeline whose success rates are measured on held-out episodes, not quantities forced by definition or self-citation.

full rationale

Walk of the claimed chain shows no reduction of a prediction to its inputs. Stage 1 (Eqs. 1–6, §III.B) pretrains an encoder–decoder on observed future DINOv3 and depth deltas; that is a standard reconstruction/latent-dynamics objective, not a claim that the deltas are derived from first principles. Stage 2 freezes the encoder to supply stop-gradient targets z*_t (Eq. 7), trains the VLA backbone to match those targets plus robot actions and decoder dynamics (Eq. 9), and evaluates success on held-out RoboTwin/RoboCasa/real-robot episodes. Latent targets are produced from external observation transitions, not from the policy’s own outputs, so L_align is not self-definitional. No parameter is fitted to a subset and then reported as an independent prediction of a closely related quantity. No uniqueness theorem or load-bearing premise is imported from overlapping-author citations; DINOv3, Depth Anything, Qwen3-VL, and prior latent-action/WAM works are external methodological choices. Benchmark numbers (Tables I–II, IV; Fig. 4) are empirical success rates against external baselines, not renamings of fitted constants. Concerns about unreproduced baselines or single-seed variance are evaluation-reliability issues, not circularity. steps is empty by design.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The central empirical claim rests on standard deep-learning practice plus a handful of modeling choices (delta prediction in frozen DINOv3 and depth spaces, frozen encoder as target provider, loss masking for unlabeled data) and free loss weights. No new physical entities are postulated; the latent actions are learned representations whose utility is measured by downstream success rates.

free parameters (4)
  • λ_cos, λ_grad, λ_dep (LAM losses)
    Scalar weights balancing L1, cosine, and gradient terms in the pretraining objective (Eqs. 4–6); chosen by authors, not derived.
  • λ_align, λ_wm (policy losses)
    Weights on latent-target matching and world-model losses relative to action loss (Eq. 9); free hyperparameters.
  • K and τ_k (future sampling)
    Number and temporal offsets of sparsely sampled future frames used to form deltas; design choices that affect the latent targets.
  • latent action token dimension / count
    Architecture size of z_t produced by the encoder and consumed by the VLA; selected rather than derived.
axioms (4)
  • domain assumption Frozen DINOv3 features plus dense depth deltas are a sufficient proxy for task-relevant semantic and geometric change.
    Invoked throughout §III.B and the design of L_rgb + L_dep; not proved, only motivated by prior DINO-space work.
  • domain assumption A latent action space learned from observation transitions can be aligned to executable robot actions via joint L_act + L_align + L_wm supervision.
    Core premise of the policy stage (§III.C); success is measured empirically rather than guaranteed.
  • domain assumption Action-free human/robot videos share transferable dynamics with the target robot embodiment when projected into the same latent space.
    Required for the real-world human-video experiments and the claim that unlabeled video helps control.
  • standard math Standard supervised learning and transformer VLA training dynamics hold for the joint objective.
    Background optimization assumptions used throughout training.
invented entities (2)
  • semantic-geometric latent action model (LAM) with DINOv3+depth delta prediction no independent evidence
    purpose: Produces stable latent action targets from unlabeled video and supplies a trainable latent world model decoder.
    New architectural module introduced by the paper; utility is demonstrated only via downstream policy metrics inside this work.
  • executable latent actions jointly supervised by action, target-matching, and dynamics losses no independent evidence
    purpose: Unified interface that lets labeled demos and unlabeled videos train the same VLA backbone.
    The paper’s central representational claim; no external measurement of the latent space outside the reported success rates.

pith-pipeline@v1.1.0-grok45 · 20949 in / 3106 out tokens · 36432 ms · 2026-07-14T05:51:23.653925+00:00 · methodology

0 comments
read the original abstract

Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot videos contain rich physical interactions but often lack executable robot action labels. We present WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos. WALA first pretrains a semantic-geometric latent action model from videos by modeling the evolution between current observations and sparsely sampled future observations. Instead of reconstructing raw pixels, WALA predicts future deltas in the DINOv3 feature space and dense depth space, preserving task-relevant semantic and geometric structure while reducing sensitivity to appearance details. During policy training, the pretrained encoder provides stable latent action targets, and the decoder serves as a trainable latent world model. The latent actions generated by the vision-language backbone are jointly supervised by robot action prediction, latent action target matching, and future dynamics prediction. This enables action-labeled demonstrations to provide executable control supervision, while action-free videos contribute dynamics supervision without requiring robot action annotations. Experiments show that WALA achieves strong performance on RoboTwin, sets a new state-of-the-art result on RoboCasa with 75.2% average success, and improves both policy performance and generalization in real-world manipulation tasks.

Figures

Figures reproduced from arXiv: 2607.11397 by Chaoyue Li, Dongbin Zhao, Haoran Li, Huangrui Li, Jiahao Liu, Jing Li, Linbo Wang, Ning Ma, Shangqing Zhou, Shuai Tian, Xiaotian Liu, Xin Fu, Yixian Li, Yuhang Zheng, Zebin Xing, Zhongpu Xia.

Figure 1
Figure 1. Figure 1: Overview of the semantic-geometric latent action model pretraining. Given the current observation and multiple sparsely sampled future [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Policy training with WALA. The pretrained latent action encoder is frozen and provides stable latent action targets from observed future changes. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Visualization of future semantic and geometric prediction. For each example, WALA observes the current DINOv3 feature visualization and [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Data scaling on RoboCasa-GR1-Tabletop. Left: WALA consistently improves over the Base Policy when both use the same amount of action [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Latent action retrieval. We encode each query transition into latent action tokens and retrieve nearest neighbors from the dataset using latent-action [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: RGB attention visualization of the latent action encoder. The [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Real-world manipulation tasks used in our evaluation. We test a single multi-task policy on Basic Pick-Place, Stack Paper Cups, Insert Flowers, [PITH_FULL_IMAGE:figures/full_fig_p009_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Zero-shot transfer from action-free egocentric human videos to a [PITH_FULL_IMAGE:figures/full_fig_p009_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

35 extracted references · 29 linked inside Pith

  1. [1]

    Siméoni, H

    O. Siméoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H. Jégou, P. Labatut, and P. Bojanowski, “Dinov3,” 2025. [Online]. Available: https...

  2. [2]

    Depth anything v2,

    L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09414

  3. [3]

    Egodex: Learning dexterous manipulation from large-scale egocentric video,

    R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang, “Egodex: Learning dexterous manipulation from large-scale egocentric video,” 2026. [Online]. Available: https://arxiv.org/abs/2505.11709

  4. [4]

    Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation,

    S. Wu, X. Liu, S. Xie, P. Wang, X. Li, B. Yang, Z. Li, K. Zhu, H. Wu, Y . Liu, Z. Long, R. Xu, Y . Wang, C. Liu, D. Wang, Z. Ni, X. Yang, Y . Liu, R. Feng, L. Zhang, D. Huang, C. Jin, A. Yin, X. Wang, Z. Sun, J. Zhao, M. Du, M. Cao, X. Chen, H. Cheng, X. Zhang, Y . Fu, N. Chen, C. Chi, S. Chen, H. Lyu, X. Hao, Y . Wang, B. Lei, D. Liu, X. Yang, Y . Jiao, ...

  5. [5]

    Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation,

    T. Chen, Z. Chen, B. Chen, Z. Cai, Y . Liu, Z. Li, Q. Liang, X. Lin, Y . Ge, Z. Gu, W. Deng, Y . Guo, T. Nian, X. Xie, Q. Chen, K. Su, T. Xu, G. Liu, M. Hu, H. ang Gao, K. Wang, Z. Liang, Y . Qin, X. Yang, P. Luo, and Y . Mu, “Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation,” ...

  6. [6]

    Robocasa: Large-scale simulation of everyday tasks for generalist robots,

    S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y . Zhu, “Robocasa: Large-scale simulation of everyday tasks for generalist robots,” 2024. [Online]. Available: https://arxiv.org/abs/2406.02523

  7. [7]

    Rt-1: Robotics transformer for real-world control at scale,

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J...

  8. [8]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control,

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y . Lu, H. Michalewski, I. Mordatch, K. Pe...

  9. [9]

    Openvla: An open-source vision-language-action model,

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “Openvla: An open-source vision-language-action model,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09246

  10. [10]

    Octo: An open-source generalist robot policy,

    O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” 2024. [Online]. Available: https://arxiv.org/abs/2405.12213

  11. [11]

    π0: A vision-language-action flow model for general robot control,

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky, “π0: A vision-language-action flow model for general robot control,”

  12. [12]

    Available: https://arxiv.org/abs/2410.24164

    [Online]. Available: https://arxiv.org/abs/2410.24164

  13. [13]

    Fast-wam: Do world action models need test-time future imagination?

    T. Yuan, Z. Dong, Y . Liu, and H. Zhao, “Fast-wam: Do world action models need test-time future imagination?” 2026. [Online]. Available: https://arxiv.org/abs/2603.16666

  14. [14]

    Causal world modeling for robot control,

    L. Li, Q. Zhang, Y . Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y . Shen, and Y . Xu, “Causal world modeling for robot control,” 2026. [Online]. Available: https://arxiv.org/abs/2601.21998

  15. [15]

    Lda-1b: Scaling latent dynamics action model via universal embodied data ingestion,

    J. Lyu, K. Liu, X. Zhang, H. Liao, Y . Feng, W. Zhu, T. Shen, J. Chen, J. Zhang, Y . Dong, W. Cui, S. Qi, S. Wang, Y . Zheng, M. Yan, X. Shi, H. Li, D. Zhao, M.-Y . Liu, Z. Zhang, L. Yi, Y . Wang, and H. Wang, “Lda-1b: Scaling latent dynamics action model via universal embodied data ingestion,” 2026. [Online]. Available: https://arxiv.org/abs/2602.12215

  16. [16]

    World action models are zero-shot policies,

    S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xiang, A. Malik, K. Lee, W. Liang, N. Ranawaka, J. Gu, Y . Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y . Xie, J. Wu, Q. Wang, R. Julian, D. Xu, Y . Du, Y . Chebotar, S. Reed, J. Kautz, Y . Zhu, L. J. Fan, and J. Jang, “World action mo...

  17. [17]

    Motus: A unified latent action world model,

    H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y . Feng, C. Xiang, Y . Rong, H. Zhao, H. Liu, Z. Su, L. Ma, H. Su, and J. Zhu, “Motus: A unified latent action world model,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2026, pp. 35 101–35 113

  18. [18]

    Latent action pretraining from videos,

    S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo, “Latent action pretraining from videos,” 2025. [Online]. Available: https://arxiv.org/abs/2410.11758

  19. [19]

    Moto: Latent motion token as the bridging language for learning robot manipulation from videos,

    Y . Chen, Y . Ge, W. Tang, Y . Li, Y . Ge, M. Ding, Y . Shan, and X. Liu, “Moto: Latent motion token as the bridging language for learning robot manipulation from videos,” 2025. [Online]. Available: https://arxiv.org/abs/2412.04445

  20. [20]

    Univla: Learning to act anywhere with task-centric latent actions,

    Q. Bu, Y . Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li, “Univla: Learning to act anywhere with task-centric latent actions,”

  21. [21]

    Available: https://arxiv.org/abs/2505.06111

    [Online]. Available: https://arxiv.org/abs/2505.06111

  22. [22]

    Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling,

    B. Chen, Y . Chen, L. Qiu, J. Bai, Y . Ge, and Y . Ge, “Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling,” 2026. [Online]. Available: https://arxiv.org/abs/2604.19734

  23. [23]

    villa-x: Enhancing latent action modeling in vision-language-action models,

    X. Chen, H. Wei, P. Zhang, C. Zhang, K. Wang, Y . Guo, R. Yang, Y . Wang, X. Xiao, L. Zhao, J. Chen, and J. Bian, “villa-x: Enhancing latent action modeling in vision-language-action models,” 2025. [Online]. Available: https://arxiv.org/abs/2507.23682

  24. [24]

    Qwen3-vl technical report,

    S. Bai, Y . Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y . Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y . Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang,...

  25. [25]

    π 0.5: a vision-language-action model with open-world generalization,

    P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. V...

  26. [26]

    X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model,

    J. Zheng, J. Li, Z. Wang, D. Liu, X. Kang, Y . Feng, Y . Zheng, J. Zou, Y . Chen, J. Zeng, Y .-Q. Zhang, J. Pang, J. Liu, T. Wang, and X. Zhan, “X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model,” 2025. [Online]. Available: https://arxiv.org/abs/2510.10274

  27. [27]

    Starvla-α: Reducing complexity in vision-language-action systems,

    J. Ye, N. Gao, S. Yang, J. Zheng, Z. Wang, Y . Chen, P. Chen, Y . Chen, S. Liu, and J. Jia, “Starvla-α: Reducing complexity in vision-language-action systems,” 2026. [Online]. Available: https://arxiv.org/abs/2604.11757

  28. [28]

    Internvla-a1: Unifying understanding, generation and action for robotic manipulation,

    J. Cai, Z. Cai, J. Cao, Y . Chen, Z. He, L. Jiang, H. Li, H. Li, Y . Li, Y . Liu, Y . Lu, Q. Lv, H. Ma, J. Pang, Y . Qiao, Z. Qiu, Y . Shen, X. Shi, Y . Tian, B. Wang, H. Wang, J. Wang, T. Wang, X. Wei, C. Wu, Y . Xie, B. Xing, Y . Yang, Y . Yang, Q. Yu, F. Yuan, J. Zeng, J. Zhang, S. Zhang, S. Zhang, Z. Zhaxi, B. Zhou, Y . Zhou, Y . Zhou, H. Zhu, Y . Zhu...

  29. [29]

    Starvla: A lego-like codebase for vision-language- action model developing,

    S. Community, “Starvla: A lego-like codebase for vision-language- action model developing,” 2026. [Online]. Available: https://arxiv.org/ abs/2604.05014

  30. [30]

    Gr00t n1: An open foundation model for generalist humanoid robots,

    NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y . Xie, Y . Xu, Z. Xu, S. Ye, Z. ...

  31. [31]

    Abot-m0: Vla foundation model for robotic manipulation with action manifold learning,

    Y . Yang, S. Zeng, T. Lin, X. Chang, D. Qi, J. Xiao, H. Liu, R. Chen, Y . Chen, D. Huo, F. Xiong, X. Wei, Z. Ma, and M. Xu, “Abot-m0: Vla foundation model for robotic manipulation with action manifold learning,” 2026. [Online]. Available: https://arxiv.org/abs/2602.11236

  32. [32]

    Rldx-1 technical report,

    D. Kim, H. Jang, M. Koo, S. Jang, T. Kim, B. Kim, B. Yoon, C. Jang, D. Choi, D. Han, D. Lee, H. Kwon, H. Jeon, J. Kang, J. Bae, J. Lee, J. Lee, J. Won, J. Ahn, J. Park, J. Sung, K. Lee, M. Han, M. Yoon, S. Joo, S. Son, S. Park, S. Cho, S. Moon, S. Kim, Y . Dong, Y . Cho, Y . Kim, C. H. Kim, D. Kim, H. Kim, H. Lee, H. Ahn, H. Ryu, H. Choi, H. Shin, J. Jung...

  33. [33]

    Frameskip: Learning from fewer but more informative frames in vla training,

    B. Yu, S. Lian, X. Lin, Z. Shen, Y . Wei, C. Wu, H. Yuan, H. Liu, B. Wang, C. Huang, and K. Chen, “Frameskip: Learning from fewer but more informative frames in vla training,” 2026. [Online]. Available: https://arxiv.org/abs/2605.13757

  34. [34]

    Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control,

    T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang, “Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control,” 2026. [Online]. Available: https://arxiv.org/abs/2603.10448

  35. [35]

    Dial: Decoupling intent and action via latent world modeling for end-to-end vla,

    Y . Chen, Y . Ge, H. Zhou, M. Ding, Y . Ge, and X. Liu, “Dial: Decoupling intent and action via latent world modeling for end-to-end vla,” 2026. [Online]. Available: https://arxiv.org/abs/2603.29844