REVIEW 3 major objections 6 minor 35 references
WALA learns executable robot latent actions from both labeled demonstrations and unlabeled videos by predicting future semantic and geometric scene changes.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 05:51 UTC pith:LGQXWUBT
load-bearing objection Solid empirical VLA paper that turns unlabeled video into joint latent-action + dynamics supervision; SOTA and low-label claims are real contributions but rest on unreproduced baselines and single-seed rates. the 3 major comments →
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Executable latent actions can be learned jointly from action-labeled robot demonstrations and action-free videos by supervising a vision-language backbone with three signals at once: robot action prediction, matching of frozen latent action targets extracted from observed future semantic-geometric deltas, and prediction of those same future deltas through a trainable latent world model. Action-free videos therefore contribute dynamics supervision without robot action annotations, and the policy still deploys without running the world model.
What carries the argument
Semantic-geometric latent action model: an encoder that maps current DINOv3 features, depth, and sparse future feature/depth deltas into latent action tokens, plus a decoder that predicts those future deltas; during policy training the encoder is frozen as a target provider and the decoder acts as a latent world model.
Load-bearing premise
That future differences in frozen image features and depth maps are a good enough stand-in for the actions a robot should take, even when the video is human egocentric footage with no robot action labels.
What would settle it
Fix labeled demonstrations, then add action-free videos whose frames are temporally shuffled so feature and depth deltas no longer match real physical evolution; if RoboCasa and real multi-task success still rise by the same margin as with correctly ordered videos, the claim that future semantic-geometric deltas supply useful dynamics supervision is falsified.
If this is right
- With only 10% of RoboCasa action labels, adding action-free videos can raise success from the low-50s toward the high-60s without extra robot annotations.
- A single multi-task real-robot policy can nearly match 200-demo performance using 50 demos plus 400 similar-scene human videos per task.
- Human egocentric videos of an unseen task (e.g., bread pick-and-place) can inject dynamics that support zero-shot robot success on that task.
- Deployment stays a pure vision-language-action forward pass; latent encoder, depth estimator, and world-model decoder are training-only.
Where Pith is reading between the lines
- If DINOv3-plus-depth deltas transfer across embodiments, the same pretraining recipe could absorb large web-scale human video corpora without embodiment-specific retargeting.
- The three-way loss (action, latent match, future prediction) suggests a general template for any VLA: keep a frozen transition encoder as a dynamics teacher even when world-model rollout is never used at test time.
- Failure modes should concentrate on tasks where semantic change is subtle but contact geometry is critical, or vice versa—ablating depth or DINOv3 alone would map that boundary.
- Scaling LAM pretraining to larger, more diverse video sets, as the authors flag, is the most direct next test of whether the latent action space saturates or keeps improving.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. WALA proposes a two-stage framework for learning executable latent actions from both action-labeled robot demonstrations and action-free videos. Stage 1 pretrains a semantic-geometric latent action model (LAM) that encodes observed future deltas in frozen DINOv3 feature space and dense depth space (Eqs. 1–6, §III.B) and decodes predicted deltas without pixel reconstruction. Stage 2 freezes the LAM encoder as a target provider, keeps the decoder as a trainable latent world model, and trains a Qwen3-VL-4B vision-language backbone so that its latent actions are jointly supervised by robot action prediction, latent-target matching, and future dynamics prediction (Eqs. 7–9, §III.C), with the action loss masked on action-free data. At inference only the backbone and action head are used. Empirically, the paper reports 90.6%/92.8% on RoboTwin Clean/Random (Table I), a claimed SOTA 75.2% average on RoboCasa-GR1-Tabletop (Table II), ablations isolating each supervision term (Table III), labeled and action-free scaling curves (Fig. 4), and real-robot multi-task gains including a low-label setting (50 demos + 400 human videos ≈ 200 demos) and a zero-shot bread pick-and-place transfer (Table IV, Fig. 8).
Significance. If the results hold under controlled re-evaluation, the work is significant for robot learning: it gives a concrete training-time interface that lets large action-free video corpora supply dynamics supervision for VLAs without world-model cost at deployment. Strengths that should be credited include (i) a clear three-way joint objective that cleanly separates control, latent-target, and dynamics losses with an explicit mask for unlabeled video; (ii) controlled Base-Policy ablations and component ablations (Table III) that isolate LAM pretraining, semantic vs. geometric prediction, and target matching; (iii) data-scaling curves that separately vary labeled demos and action-free videos (Fig. 4); and (iv) real-robot multi-task and low-label/zero-shot human-video experiments with measured latency. These elements go beyond pure representation pretraining and address a practical bottleneck in scaling manipulation policies.
major comments (3)
- [Table II, §IV.B] Table II (and the abstract SOTA claim of 75.2% vs. DIAL 70.2%): all success rates are single-point means over 50 episodes/task with no standard errors, no multi-seed runs, and no re-implementation of the strongest baselines under a matched backbone, data split, or evaluation protocol. Baselines are “collected from publicly released reports.” The only fully controlled comparison is Base Policy vs. WALA (same Qwen3-VL-4B, action loss only). Without variance estimates or matched re-runs, the 5.0-point SOTA gap is not yet load-bearing evidence; please report multi-seed means ± stderr (or bootstrap CIs) for WALA and at least re-evaluate the top 1–2 baselines under the same protocol, or qualify the SOTA claim accordingly.
- [Fig. 4, Table IV, §IV.C / §IV.F] Fig. 4 (right) and Table IV: the central claim that action-free videos nearly replace robot demos (10% labels + videos → 67.8% vs. Base 100% labels 54.2%; real-robot 50 demos + 400 human videos → 74.2% ≈ 200 demos 75.0%) rests on the same single-run success rates (N=50 sim episodes/task; N=30 real trials/task) without error bars or seeds. These are the paper’s most consequential practical claims. Please add multi-seed or multi-split statistics and, for the real-robot low-label setting, report per-task binomial confidence intervals so that “nearly matches” can be assessed quantitatively rather than by point estimates alone.
- [§III.B–C, Eqs. (2)–(7), Table III] §III.B–C and the weakest modeling assumption: latent targets are defined as frozen-encoder outputs of DINOv3 and depth deltas (Eqs. 2, 7). Ablations (Table III) show that semantic+geometric world losses help, but there is no analysis of how well these targets align with robot action manifolds across embodiments (human egocentric vs. robot multi-view), nor of sensitivity to K, τ_k, or depth-estimator noise. A short diagnostic—e.g., correlation of z* with ground-truth robot actions on labeled data, or cross-embodiment retrieval beyond Fig. 5—would make the transfer claim falsifiable rather than only performance-supported.
minor comments (6)
- [Title, Abstract] Title and running text inconsistently space the acronym (“W ALA” / “WALA”); standardize to WALA throughout.
- [§III.B–C] Eqs. (4)–(6) and (9): λ_cos, λ_grad, λ_dep, λ_align, λ_wm and the sampling schedule (K, τ_k) are free hyperparameters but values and selection procedure are not stated; add a short hyperparameter table or appendix note.
- [Fig. 3, §IV.B] Fig. 3 visualizes predicted future DINOv3/depth states but does not quantify prediction error (e.g., feature ℓ1/cosine or depth MAE on held-out transitions); a small quantitative panel would strengthen the qualitative claim.
- [Abstract, Table I] Table I Clean: WALA (90.6) is below LingBot-VA (92.9) and Fast-WAM (91.9); the text correctly notes Random is best, but the abstract’s “strong performance on RoboTwin” could briefly acknowledge the Clean ranking to avoid overstatement.
- [§IV.F, §V] Real-world zero-shot transfer is reported for a single OOD task (bread pick-and-place, 3/10 → 9/10). The limitations section already flags this; consider moving a one-sentence caveat into the main real-world results paragraph.
- [References] References include several 2025–2026 arXiv entries; ensure citation keys and years are consistent with the submitted bibliography style.
Circularity Check
No circularity: WALA is an empirical two-stage training pipeline whose success rates are measured on held-out episodes, not quantities forced by definition or self-citation.
full rationale
Walk of the claimed chain shows no reduction of a prediction to its inputs. Stage 1 (Eqs. 1–6, §III.B) pretrains an encoder–decoder on observed future DINOv3 and depth deltas; that is a standard reconstruction/latent-dynamics objective, not a claim that the deltas are derived from first principles. Stage 2 freezes the encoder to supply stop-gradient targets z*_t (Eq. 7), trains the VLA backbone to match those targets plus robot actions and decoder dynamics (Eq. 9), and evaluates success on held-out RoboTwin/RoboCasa/real-robot episodes. Latent targets are produced from external observation transitions, not from the policy’s own outputs, so L_align is not self-definitional. No parameter is fitted to a subset and then reported as an independent prediction of a closely related quantity. No uniqueness theorem or load-bearing premise is imported from overlapping-author citations; DINOv3, Depth Anything, Qwen3-VL, and prior latent-action/WAM works are external methodological choices. Benchmark numbers (Tables I–II, IV; Fig. 4) are empirical success rates against external baselines, not renamings of fitted constants. Concerns about unreproduced baselines or single-seed variance are evaluation-reliability issues, not circularity. steps is empty by design.
Axiom & Free-Parameter Ledger
free parameters (4)
- λ_cos, λ_grad, λ_dep (LAM losses)
- λ_align, λ_wm (policy losses)
- K and τ_k (future sampling)
- latent action token dimension / count
axioms (4)
- domain assumption Frozen DINOv3 features plus dense depth deltas are a sufficient proxy for task-relevant semantic and geometric change.
- domain assumption A latent action space learned from observation transitions can be aligned to executable robot actions via joint L_act + L_align + L_wm supervision.
- domain assumption Action-free human/robot videos share transferable dynamics with the target robot embodiment when projected into the same latent space.
- standard math Standard supervised learning and transformer VLA training dynamics hold for the joint objective.
invented entities (2)
-
semantic-geometric latent action model (LAM) with DINOv3+depth delta prediction
no independent evidence
-
executable latent actions jointly supervised by action, target-matching, and dynamics losses
no independent evidence
read the original abstract
Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot videos contain rich physical interactions but often lack executable robot action labels. We present WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos. WALA first pretrains a semantic-geometric latent action model from videos by modeling the evolution between current observations and sparsely sampled future observations. Instead of reconstructing raw pixels, WALA predicts future deltas in the DINOv3 feature space and dense depth space, preserving task-relevant semantic and geometric structure while reducing sensitivity to appearance details. During policy training, the pretrained encoder provides stable latent action targets, and the decoder serves as a trainable latent world model. The latent actions generated by the vision-language backbone are jointly supervised by robot action prediction, latent action target matching, and future dynamics prediction. This enables action-labeled demonstrations to provide executable control supervision, while action-free videos contribute dynamics supervision without requiring robot action annotations. Experiments show that WALA achieves strong performance on RoboTwin, sets a new state-of-the-art result on RoboCasa with 75.2% average success, and improves both policy performance and generalization in real-world manipulation tasks.
Figures
Reference graph
Works this paper leans on
-
[1]
O. Siméoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H. Jégou, P. Labatut, and P. Bojanowski, “Dinov3,” 2025. [Online]. Available: https...
Pith/arXiv arXiv 2025
-
[2]
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09414
Pith/arXiv arXiv 2024
-
[3]
Egodex: Learning dexterous manipulation from large-scale egocentric video,
R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang, “Egodex: Learning dexterous manipulation from large-scale egocentric video,” 2026. [Online]. Available: https://arxiv.org/abs/2505.11709
Pith/arXiv arXiv 2026
-
[4]
Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation,
S. Wu, X. Liu, S. Xie, P. Wang, X. Li, B. Yang, Z. Li, K. Zhu, H. Wu, Y . Liu, Z. Long, R. Xu, Y . Wang, C. Liu, D. Wang, Z. Ni, X. Yang, Y . Liu, R. Feng, L. Zhang, D. Huang, C. Jin, A. Yin, X. Wang, Z. Sun, J. Zhao, M. Du, M. Cao, X. Chen, H. Cheng, X. Zhang, Y . Fu, N. Chen, C. Chi, S. Chen, H. Lyu, X. Hao, Y . Wang, B. Lei, D. Liu, X. Yang, Y . Jiao, ...
Pith/arXiv arXiv 2026
-
[5]
T. Chen, Z. Chen, B. Chen, Z. Cai, Y . Liu, Z. Li, Q. Liang, X. Lin, Y . Ge, Z. Gu, W. Deng, Y . Guo, T. Nian, X. Xie, Q. Chen, K. Su, T. Xu, G. Liu, M. Hu, H. ang Gao, K. Wang, Z. Liang, Y . Qin, X. Yang, P. Luo, and Y . Mu, “Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation,” ...
Pith/arXiv arXiv 2025
-
[6]
Robocasa: Large-scale simulation of everyday tasks for generalist robots,
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y . Zhu, “Robocasa: Large-scale simulation of everyday tasks for generalist robots,” 2024. [Online]. Available: https://arxiv.org/abs/2406.02523
Pith/arXiv arXiv 2024
-
[7]
Rt-1: Robotics transformer for real-world control at scale,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J...
Pith/arXiv arXiv 2023
-
[8]
Rt-2: Vision-language-action models transfer web knowledge to robotic control,
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y . Lu, H. Michalewski, I. Mordatch, K. Pe...
Pith/arXiv arXiv 2023
-
[9]
Openvla: An open-source vision-language-action model,
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn, “Openvla: An open-source vision-language-action model,” 2024. [Online]. Available: https://arxiv.org/abs/2406.09246
Pith/arXiv arXiv 2024
-
[10]
Octo: An open-source generalist robot policy,
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” 2024. [Online]. Available: https://arxiv.org/abs/2405.12213
Pith/arXiv arXiv 2024
-
[11]
π0: A vision-language-action flow model for general robot control,
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky, “π0: A vision-language-action flow model for general robot control,”
-
[12]
Available: https://arxiv.org/abs/2410.24164
[Online]. Available: https://arxiv.org/abs/2410.24164
-
[13]
Fast-wam: Do world action models need test-time future imagination?
T. Yuan, Z. Dong, Y . Liu, and H. Zhao, “Fast-wam: Do world action models need test-time future imagination?” 2026. [Online]. Available: https://arxiv.org/abs/2603.16666
Pith/arXiv arXiv 2026
-
[14]
Causal world modeling for robot control,
L. Li, Q. Zhang, Y . Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y . Shen, and Y . Xu, “Causal world modeling for robot control,” 2026. [Online]. Available: https://arxiv.org/abs/2601.21998
Pith/arXiv arXiv 2026
-
[15]
Lda-1b: Scaling latent dynamics action model via universal embodied data ingestion,
J. Lyu, K. Liu, X. Zhang, H. Liao, Y . Feng, W. Zhu, T. Shen, J. Chen, J. Zhang, Y . Dong, W. Cui, S. Qi, S. Wang, Y . Zheng, M. Yan, X. Shi, H. Li, D. Zhao, M.-Y . Liu, Z. Zhang, L. Yi, Y . Wang, and H. Wang, “Lda-1b: Scaling latent dynamics action model via universal embodied data ingestion,” 2026. [Online]. Available: https://arxiv.org/abs/2602.12215
Pith/arXiv arXiv 2026
-
[16]
World action models are zero-shot policies,
S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xiang, A. Malik, K. Lee, W. Liang, N. Ranawaka, J. Gu, Y . Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y . Xie, J. Wu, Q. Wang, R. Julian, D. Xu, Y . Du, Y . Chebotar, S. Reed, J. Kautz, Y . Zhu, L. J. Fan, and J. Jang, “World action mo...
Pith/arXiv arXiv 2026
-
[17]
Motus: A unified latent action world model,
H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y . Feng, C. Xiang, Y . Rong, H. Zhao, H. Liu, Z. Su, L. Ma, H. Su, and J. Zhu, “Motus: A unified latent action world model,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2026, pp. 35 101–35 113
2026
-
[18]
Latent action pretraining from videos,
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin, L. Liden, K. Lee, J. Gao, L. Zettlemoyer, D. Fox, and M. Seo, “Latent action pretraining from videos,” 2025. [Online]. Available: https://arxiv.org/abs/2410.11758
Pith/arXiv arXiv 2025
-
[19]
Moto: Latent motion token as the bridging language for learning robot manipulation from videos,
Y . Chen, Y . Ge, W. Tang, Y . Li, Y . Ge, M. Ding, Y . Shan, and X. Liu, “Moto: Latent motion token as the bridging language for learning robot manipulation from videos,” 2025. [Online]. Available: https://arxiv.org/abs/2412.04445
arXiv 2025
-
[20]
Univla: Learning to act anywhere with task-centric latent actions,
Q. Bu, Y . Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li, “Univla: Learning to act anywhere with task-centric latent actions,”
-
[21]
Available: https://arxiv.org/abs/2505.06111
[Online]. Available: https://arxiv.org/abs/2505.06111
-
[22]
Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling,
B. Chen, Y . Chen, L. Qiu, J. Bai, Y . Ge, and Y . Ge, “Unit: Toward a unified physical language for human-to-humanoid policy learning and world modeling,” 2026. [Online]. Available: https://arxiv.org/abs/2604.19734
Pith/arXiv arXiv 2026
-
[23]
villa-x: Enhancing latent action modeling in vision-language-action models,
X. Chen, H. Wei, P. Zhang, C. Zhang, K. Wang, Y . Guo, R. Yang, Y . Wang, X. Xiao, L. Zhao, J. Chen, and J. Bian, “villa-x: Enhancing latent action modeling in vision-language-action models,” 2025. [Online]. Available: https://arxiv.org/abs/2507.23682
Pith/arXiv arXiv 2025
-
[24]
S. Bai, Y . Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y . Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y . Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang,...
Pith/arXiv arXiv 2025
-
[25]
π 0.5: a vision-language-action model with open-world generalization,
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. V...
Pith/arXiv arXiv 2025
-
[26]
X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model,
J. Zheng, J. Li, Z. Wang, D. Liu, X. Kang, Y . Feng, Y . Zheng, J. Zou, Y . Chen, J. Zeng, Y .-Q. Zhang, J. Pang, J. Liu, T. Wang, and X. Zhan, “X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model,” 2025. [Online]. Available: https://arxiv.org/abs/2510.10274
Pith/arXiv arXiv 2025
-
[27]
Starvla-α: Reducing complexity in vision-language-action systems,
J. Ye, N. Gao, S. Yang, J. Zheng, Z. Wang, Y . Chen, P. Chen, Y . Chen, S. Liu, and J. Jia, “Starvla-α: Reducing complexity in vision-language-action systems,” 2026. [Online]. Available: https://arxiv.org/abs/2604.11757
Pith/arXiv arXiv 2026
-
[28]
Internvla-a1: Unifying understanding, generation and action for robotic manipulation,
J. Cai, Z. Cai, J. Cao, Y . Chen, Z. He, L. Jiang, H. Li, H. Li, Y . Li, Y . Liu, Y . Lu, Q. Lv, H. Ma, J. Pang, Y . Qiao, Z. Qiu, Y . Shen, X. Shi, Y . Tian, B. Wang, H. Wang, J. Wang, T. Wang, X. Wei, C. Wu, Y . Xie, B. Xing, Y . Yang, Y . Yang, Q. Yu, F. Yuan, J. Zeng, J. Zhang, S. Zhang, S. Zhang, Z. Zhaxi, B. Zhou, Y . Zhou, Y . Zhou, H. Zhu, Y . Zhu...
arXiv 2026
-
[29]
Starvla: A lego-like codebase for vision-language- action model developing,
S. Community, “Starvla: A lego-like codebase for vision-language- action model developing,” 2026. [Online]. Available: https://arxiv.org/ abs/2604.05014
Pith/arXiv arXiv 2026
-
[30]
Gr00t n1: An open foundation model for generalist humanoid robots,
NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y . Xie, Y . Xu, Z. Xu, S. Ye, Z. ...
Pith/arXiv arXiv 2025
-
[31]
Abot-m0: Vla foundation model for robotic manipulation with action manifold learning,
Y . Yang, S. Zeng, T. Lin, X. Chang, D. Qi, J. Xiao, H. Liu, R. Chen, Y . Chen, D. Huo, F. Xiong, X. Wei, Z. Ma, and M. Xu, “Abot-m0: Vla foundation model for robotic manipulation with action manifold learning,” 2026. [Online]. Available: https://arxiv.org/abs/2602.11236
Pith/arXiv arXiv 2026
-
[32]
D. Kim, H. Jang, M. Koo, S. Jang, T. Kim, B. Kim, B. Yoon, C. Jang, D. Choi, D. Han, D. Lee, H. Kwon, H. Jeon, J. Kang, J. Bae, J. Lee, J. Lee, J. Won, J. Ahn, J. Park, J. Sung, K. Lee, M. Han, M. Yoon, S. Joo, S. Son, S. Park, S. Cho, S. Moon, S. Kim, Y . Dong, Y . Cho, Y . Kim, C. H. Kim, D. Kim, H. Kim, H. Lee, H. Ahn, H. Ryu, H. Choi, H. Shin, J. Jung...
Pith/arXiv arXiv 2026
-
[33]
Frameskip: Learning from fewer but more informative frames in vla training,
B. Yu, S. Lian, X. Lin, Z. Shen, Y . Wei, C. Wu, H. Yuan, H. Liu, B. Wang, C. Huang, and K. Chen, “Frameskip: Learning from fewer but more informative frames in vla training,” 2026. [Online]. Available: https://arxiv.org/abs/2605.13757
Pith/arXiv arXiv 2026
-
[34]
Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control,
T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang, “Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control,” 2026. [Online]. Available: https://arxiv.org/abs/2603.10448
arXiv 2026
-
[35]
Dial: Decoupling intent and action via latent world modeling for end-to-end vla,
Y . Chen, Y . Ge, H. Zhou, M. Ding, Y . Ge, and X. Liu, “Dial: Decoupling intent and action via latent world modeling for end-to-end vla,” 2026. [Online]. Available: https://arxiv.org/abs/2603.29844
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.