REVIEW 5 major objections 7 minor 97 references
A world-action model that jointly predicts future vision and future touch, then acts from both, leads contact-rich robot manipulation at scale.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 12:37 UTC pith:5BDJHGS4
load-bearing objection Solid systems paper that makes touch a first-class predicted modality in a large WAM; real-robot lead is real but not cleanly attributed to touch. the 5 major comments →
N₀-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
N0-TWAM is, to the authors' knowledge, the first tactile-native world-action model trained at large scale: it jointly predicts future vision and future contact under a shared causal cascade, conditions action on both predicted foresight and observed force-space touch, and is the strongest method overall on the reported contact-rich benchmarks (84.5% UniVTAC, 49.4% NeoSim, 46.3% real-robot macro average), with both tactile roles and data scale load-bearing in ablations.
What carries the argument
Asymmetric Mixture-of-Transformers cascade: three per-modality experts (full-width video; slim tactile and action) share one self-attention under a frame-id causal mask so video and tactile co-generate the future, then action denoises conditioned on those predictions plus a zero-initialized cross-attention from the current NeoForce force reading.
Load-bearing premise
The decisive gains from native dual-pathway touch and large-scale pre-training will hold outside the authors' private multi-robot tactile corpus, sensors, and force encoder once the promised code and checkpoints are public.
What would settle it
Re-run the UniVTAC, NeoSim, and eight-task real-robot suite with released checkpoints against the same baselines after ablating predicted tactile, observed tactile, and pre-training scale; if either pathway or scale no longer moves success, or vision-only world-action models close the gap on contact-critical tasks, the central claim fails.
If this is right
- World-action policies for contact-rich work should treat future touch as a first-class generation target co-denoised with video, not only as a reactive input.
- Isolating tactile capacity in private expert weights while keeping full shared attention preserves the pretrained visual prior better than gating touch out of the visual stream.
- Contact-onset and release events can segment training demos and advance multi-stage execution without relying only on ambiguous mid-task vision or a fixed language prompt.
- Asymmetric slim action/tactile experts plus KV caching of predicted video/tactile make large video backbones usable in real-time closed-loop control.
- Performance on precise tactile and action prediction continues to improve with more multi-embodiment tactile pre-training data.
Where Pith is reading between the lines
- Once public checkpoints exist, the dual-pathway design is a natural drop-in test for whether other video world-action backbones gain the same contact advantage without full re-pretraining.
- The sim gap—where the observed pathway falls back to a from-scratch latent encoder—suggests force-space transfer across sensors may be the bottleneck for sim-to-real tactile WAMs.
- Tactile-punctuated staging could generalize beyond the paper's fixed queues to open-ended VLM planners that only need a reliable contact-event clock.
- If visual perturbation robustness holds, touch-native WAMs may be the default stack wherever lighting, occlusion, or clutter routinely defeat camera-only policies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents N0-TWAM, a 7.2B-parameter world-action model in which touch is a native modality: an asymmetric Mixture-of-Transformers (full-width video expert warm-started from LingBot-VA; slim, from-scratch action and tactile experts) jointly denoises future video and future tactile latents under a shared self-attention with a frame-id causal "predict-then-act" cascade, and the action expert additionally reads observed touch through a zero-initialized cross-attention port (NeoForce force-space encoder on real robots, a from-scratch latent encoder in simulation). A tactile-punctuated scheduler handles long-horizon tasks. The model is pre-trained on the team's private 30k-hour NeoData corpus (six embodiments, 450 tasks, synchronized tactile) and evaluated on UniVTAC (84.5%), the in-house NeoSim suite (49.4%), and an eight-task real-robot suite (46.3% macro average vs 30.0% for π0.5). Simulation ablations remove each tactile pathway and vary pre-training scale; a tactile-realism study shows a NeoForce encoder pretrained on real gel images transfers to a realistic simulated rendering.
Significance. If the results hold, this is a substantial contribution: the first large-scale tactile-native world-action model, with a clean architectural mechanism (modality-private weights plus fully shared attention, and a mask-only predict-then-act cascade) that is well motivated against concurrent gating/masking designs. The empirical package is broad: multi-suite evaluation against strong VLA and WAM baselines, pathway-specific ablations on two sim suites, a data-scaling ablation, qualitative verification that predicted tactile tracks ground truth (Fig. 11), and a thoughtful sim-to-real signal via the gel-rendered + NeoForce study (Table 6). The promised public release of code and checkpoints would add further value. The main caveats are that the headline real-robot gain is not causally attributed to touch, and that the core infrastructure (NeoData, NeoForce, NeoSim) is private and same-team, limiting independent verification at submission time.
major comments (5)
- [§4.3, Fig. 7 (real-robot suite)] The headline real-robot margin (46.3% vs 30.0% for π0.5) is never ablated for touch on real robots. The real-robot baselines differ from N0-TWAM in at least three ways at once: no tactile pathway, no NeoData/NeoForce pre-training (30k hours of private tactile-synchronized data), and a different backbone/post-training recipe. The paper's own ablation (Fig. 9) identifies pre-training scale as the single largest factor on UniVTAC (84.5→65.4 at 20% data), so a large fraction of the 16.3-point real-robot gap is consistent with data scale and curation alone. The task-level narrative (margins largest on contact-critical tasks; π0.5 ahead on Cup Stacking and Bag Packing) is suggestive but post-hoc over 20-trial, operator-judged cells with ~11% binomial SE. Since §4.3 explicitly frames the real suite as 'where native touch matters most,' the attribution claim is load-bearing: please add a real-ro
- [§2.2.2 vs §4.5 (observed pathway mismatch across domains)] The observed tactile pathway is not the same component in the two settings where it is evaluated. On real robots it is the NeoForce force-space encoder (frozen force estimator gθ plus fine-tuned ViT); in simulation it is a lightweight from-scratch latent encoder with 'no force estimator and no pretrained representation.' Consequently the simulation ablations that carry the entire attribution argument (Fig. 9, Tables 7–8) do not ablate the mechanism actually deployed on the real robot, and the real-robot system uses a pathway whose contribution is never isolated anywhere. This gap should be stated plainly in §4.5, and ideally closed by ablating the NeoForce observed pathway on the real suite (or by reporting the gel-rendered + NeoForce ablation of Table 6 with the observed pathway removed).
- [App. A, Tables 7–8 (heterogeneous ablation signature)] The averaged claim that 'both tactile pathways are load-bearing' hides large per-task sign flips that are never discussed. In Table 8, removing observed touch *improves* Cup Stack (12→38) and Cup Handover (14→30) while collapsing Plate Stack (98→15); in Table 7, w/o-predicted improves Insert HDMI (68→73) and Pull-out Key (79→86), and w/o-observed improves Put Bottle in Shelf (87→92) but zeroes Lift Bottle (58→0). At 100 trials/task (SE ~5%) several of these swings exceed noise. The authors should report per-task trial counts and uncertainty, and explain the mechanism by which removing a tactile pathway helps specific tasks — otherwise the averaged ablation overstates the uniformity of the tactile contribution.
- [§4.4, Table 4 (robustness claim)] The conclusion that N0-TWAM is 'the most robust' rests on a macro average of 51.7 vs 50.0 for π0.5 — a 1.7-point margin over cells that appear to be 20 trials each (~11% SE), i.e., well within noise. N0-TWAM is in fact the *worst* of the three methods on the unseen-object axis (65 vs 80 for π0.5 and 75 for LingBot-VA). The visual-perturbation result (45 vs 25/30) is the one axis with a meaningful gap and is consistent with the tactile story, but the aggregate 'most robust' framing should be qualified, or the trial counts increased.
- [§3, §4.2, and Conclusion (verifiability)] Two of the three evaluation suites and much of the training stack are private and same-team: NeoSim and six of eight real-robot tasks come from the team's own NeoReal [56], the observed pathway depends on the unreleased NeoForce encoder, and pre-training data (NeoData) is inaccessible. Code and checkpoints are promised ('will be made publicly available') but not shipped at submission, so the NeoSim numbers, the NeoForce-dependent results, and the data-scaling claims cannot be independently checked. At minimum, the submission should include the NeoSim task definitions and success criteria, per-task real-robot success rubrics, and a concrete release commitment; the journal may wish to condition acceptance on the release.
minor comments (7)
- [§1, ¶3] Typo: 'so prediction occurs happens outside the network that acts.' Also Fig. 1 labels the baseline 'ω0.5' where π0.5 is meant.
- [§1 / Abstract (novelty claim)] 'First tactile world-action model trained at large scale' should be scoped more carefully against the concurrent touch-aware WAMs the paper itself cites ([52] Dream-Tac, [64] VT-WAM, [70] Tactile-WAM, [94] OmniVTA); a sentence distinguishing 'native prediction of touch at scale' from those designs would make the claim defensible.
- [Table 3 (NeoSim)] Unplug & Plug Charger and Place Gears are at 0% for all methods including N0-TWAM, and N0-TWAM loses several individual tasks it wins elsewhere (Cup Stack 12 vs π0.5 22; Cup Unstack 16 vs 31). Please clarify whether the all-zero tasks are solvable/calibrated, and comment on the per-task inconsistency; the in-house suite's difficulty calibration matters for the 49.4 vs 45.8 average.
- [§2.3.2 (real-time claims)] The section argues real-time operation is achievable but reports no measured latency, control rate, or chunk-execution window on the deployment hardware. Since the FP8/NVFP4 serving path and the asynchronous pipeline are presented as deployment-relevant, a small table of measured timings would strengthen the claim.
- [§4.2 (protocol)] The real-robot success criteria are 'judged by the operator' per a 'fixed per-task criterion,' but the criteria are not listed; please include them (e.g., in an appendix). Also state whether operators were blinded to method identity.
- [App. B, Table 10] The dual-arm entries for 'Gel-rendered + NeoForce' are '/' with no explanation; please state why this condition was not run.
- [§4.1 (pre-training)] 'Tens of thousands hours' (§4.1) vs 'over 30,000 hours' (§3.1) should be made consistent, and the fraction of pre-training episodes that actually carry synchronized tactile should be stated quantitatively rather than as 'a large fraction,' since the tactile-foresight claim depends on it.
Circularity Check
No significant circularity: empirical systems paper whose success claims are external trial metrics, not quantities forced by construction from fitted inputs or self-cited uniqueness.
full rationale
N0-TWAM is a robotics/ML systems paper. Its load-bearing claims are architectural (asymmetric MoT cascade jointly predicting video/tactile then action; dual predicted/observed tactile pathways; tactile-event staging) and empirical (success rates on UniVTAC, NeoSim, and a real-robot suite versus VLA/WAM baselines, plus ablations). Success rates are closed-loop behavioral fractions over randomized trials, not algebraic consequences of a fitted constant renamed as a prediction. The flow-matching objective, residual tactile target, and causal mask implement a generative training recipe; they do not define the reported benchmark numbers. Same-team infrastructure (NeoData, NeoForce/N0-Foundation [56], NeoSim/NeoReal) is ordinary stack dependence and descriptive self-citation, not a uniqueness theorem that forbids alternatives or forces the SOTA claim. Warm-start from LingBot-VA supplies a video prior but does not make UniVTAC/real-robot outcomes true by definition. Experimental confounds (sim-only tactile ablations; real-robot baselines differing in data and backbone) are validity concerns, not circular derivation. Score 1 only for mild in-house-stack self-reference that is not load-bearing for the central result.
Axiom & Free-Parameter Ledger
free parameters (5)
- Expert hidden widths (dv=3072, da=dt=1024) and shared attention d=3072 =
3072 / 1024 / 1024
- Modality loss weights λv:λa:λt = 1:1:1 =
1:1:1
- Pretrain schedule (30k steps, batch 512, lr 1e-4, ~2.2 epochs) and UMI mix ~60% =
30k steps; 60% UMI mix in stage 2
- Condition dropout p=0.1 for language and tactile absence =
0.1
- Action chunking / horizons (e.g., train windows, 24-step chunks, 16-step decode) =
33 latent-frame windows; action horizon 24 train / 16 eval
axioms (6)
- domain assumption Joint future video and tactile can be modeled by conditional flow matching with a predict-then-act factorization inside one masked attention pass (Eq. 1–3, Fig. 3).
- ad hoc to paper Isolating modalities in private expert weights while sharing full self-attention preserves visual priors better than gating/masking tactile tokens.
- domain assumption NeoForce force maps are a sufficiently sensor-grounded observed-touch representation for real-robot action conditioning.
- domain assumption Contact onset/release events in tactile (with gripper aperture fallback) validly segment long-horizon tasks into sub-tasks for training and inference staging.
- domain assumption Chunk-anchored delta end-effector actions (with optional absolute pose on alignment tasks) are an adequate unified multi-embodiment action interface.
- domain assumption Large private NeoData demonstrations with synchronized per-finger tactile are i.i.d.-enough for the reported generalization probes.
invented entities (4)
-
N0-TWAM asymmetric Mixture-of-Transformers (video/tactile/action experts + frame-id cascade)
no independent evidence
-
Dual tactile pathways (VAE-latent predicted residual touch + NeoForce/observed force-space cross-attn port)
no independent evidence
-
Tactile-punctuated sub-task scheduler
no independent evidence
-
NeoForce observed-tactile encoder role inside TWAM
no independent evidence
read the original abstract
We present $N_0$-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on contact-rich tasks. We pre-train $N_0$-TWAM at large scale with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments and 450 tasks. We use NeoForce, a unified force-based tactile representation, to form a physically grounded contact signal that conditions action generation. To improve long-horizon and multi-stage manipulation, we introduce tactile contact events for task staging and advance through them during execution. For real-time efficiency, we adopt an asymmetric Mixture-of-Transformers architecture that pairs a full-width expert for video prediction with slim experts for downstream action and tactile prediction. Evaluations on both real and simulated benchmarks justify the capabilities of $N_0$-TWAM across a range of contact-rich tasks, and demonstrate the benefit of data scaling for precise tactile and action prediction. In summary, $N_0$-TWAM endows a world-action model with predictive capabilities to foresee vision, touch and action, building a solid foundation for fine-grained manipulation on open contact-rich tasks. The codebase and model checkpoints will be made publicly available to foster further research and development in tactile-enabled robotic manipulation.
Reference graph
Works this paper leans on
-
[1]
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691 , 2022
Pith/arXiv arXiv 2022
-
[2]
Flash-W AM: Modality-aware distillation for world action models
Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, and Yanzhi Wang. Flash-W AM: Modality-aware distillation for world action models. arXiv preprint arXiv:2606.05254 , 2026
Pith/arXiv arXiv 2026
-
[3]
Diffusion for world modeling: Visual details matter in atari
Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, and François Fleuret. Diffusion for world modeling: Visual details matter in atari. NeurIPS, 2024
2024
-
[4]
V-JEPA 2: Self-supervised video models enable understanding, prediction and planning
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985 , 2025. 20
Pith/arXiv arXiv 2025
-
[5]
Motus: A unified latent action world model
Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu, Zhizhong Su, Lei Ma, Hang Su, and Jun Zhu. Motus: A unified latent action world model. arXiv preprint arXiv:2512.13030 , 2025
Pith/arXiv arXiv 2025
-
[6]
Vla-touch: Enhancing vision-language-action models with dual-level tactile feedback
Jianxin Bi, Kevin Yuchen Ma, Ce Hao, Mike Zheng Shou, and Harold Soh. Vla-touch: Enhancing vision-language-action models with dual-level tactile feedback. arXiv preprint arXiv:2507.17294 , 2025
Pith/arXiv arXiv 2025
-
[7]
Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, ...
2025
-
[8]
π0: A vision-language- action flow model for general robot control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Robert Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, Laura Smith, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Z...
2025
-
[9]
Ryoo, Grecia Salazar, Pannag R
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alexander Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Mall...
2023
-
[10]
Genie: Generative interactive environments
Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal Behbahani, Stephanie Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott Reed, Jingwei Zhang, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Si...
2024
-
[11]
InternVLA-A1: Unifying understanding, generation and action for robotic manipulation
Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, et al. InternVLA-A1: Unifying understanding, generation and action for robotic manipulation. arXiv preprint arXiv:2601.02456 , 2026
arXiv 2026
-
[12]
Xiaomi-Robotics-0: An open-sourced vision-language-action model with real-time execution
Rui Cai, Jun Guo, Xinze He, Piaopiao Jin, Jie Li, Bingxuan Lin, Futeng Liu, Wei Liu, Fei Ma, Kun Ma, et al. Xiaomi-Robotics-0: An open-sourced vision-language-action model with real-time execution. arXiv preprint arXiv:2602.12684 , 2026
arXiv 2026
-
[13]
WorldVLA: Towards autoregressive action world model
Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, Deli Zhao, and Hao Chen. WorldVLA: Towards autoregressive action world model. arXiv preprint arXiv:2506.21539 , 2025
Pith/arXiv arXiv 2025
-
[14]
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, Hanbo Zhang, and Minzhao Zhu. Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation. arXiv preprint arXiv:2410.06158 , 2024
Pith/arXiv arXiv 2024
-
[15]
Baijun Chen, Weijie Wan, Tianxing Chen, Xianda Guo, Congsheng Xu, Yuanyang Qi, Haojie Zhang, Longyan Wu, Tianling Xu, Zixuan Li, et al. UniVTAC: A unified simulation platform for visuo-tactile manipulation data generation, learning, and benchmarking. arXiv preprint arXiv:2602.10093 , 2026. 21
arXiv 2026
-
[16]
Diffusion forcing: Next-token prediction meets full-sequence diffusion
Boyuan Chen, Diego Martí Monsó, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion. NeurIPS, 2024
2024
-
[17]
Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Weijie Wan, Baijun Chen, Haoran Lu, Haowen Yan, Honghao Su, et al. RoboDojo: A unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. arXiv preprint arXiv:2607.04434 , 2026
Pith/arXiv arXiv 2026
-
[18]
Omnivtla: Vision-tactile-language-action models with semantic-aligned tactile sensing
Zhengxue Cheng, Yiqian Zhang, Anni Tang, Keyu Wang, Wenkang Zhang, Haoyu Li, Hengdi Zhang, and Li Song. Omnivtla: Vision-tactile-language-action models with semantic-aligned tactile sensing. RA-L, 2025
2025
-
[19]
Gemini 3.5 flash
Google DeepMind. Gemini 3.5 flash. https://deepmind.google/models/gemini/flash/, 2025
2025
-
[20]
Tenenbaum, Dale Schuurmans, and Pieter Abbeel
Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B. Tenenbaum, Dale Schuurmans, and Pieter Abbeel. Learning universal policies via text-guided video generation. NeurIPS, 2023
2023
-
[21]
Tenenbaum, Leslie Kaelbling, Andy Zeng, and Jonathan Tompson
Yilun Du, Mengjiao Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B. Tenenbaum, Leslie Kaelbling, Andy Zeng, and Jonathan Tompson. Video language planning. ICLR, 2024
2024
-
[22]
LIBERO-plus: In-depth robustness analysis of vision-language-action models
Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, and Xipeng Qiu. LIBERO-plus: In-depth robustness analysis of vision-language-action models. arXiv preprint arXiv:2510.13626 , 2025
Pith/arXiv arXiv 2025
-
[23]
Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine
Dibya Ghosh, Homer Rich Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Quan Vuong, Ted Xiao, Pannag R. Sanketi, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy. RSS, 2024
2024
-
[24]
Prediction with action: Visual policy learning via joint denoising process
Yanjiang Guo, Yucheng Hu, Jianke Zhang, Yen-Jen Wang, Xiaoyu Chen, Chaochao Lu, and Jianyu Chen. Prediction with action: Visual policy learning via joint denoising process. NeurIPS, 2024
2024
-
[25]
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber. Recurrent world models facilitate policy evolution. NeurIPS, 2018
2018
-
[26]
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson. Learning latent dynamics for planning from pixels. ICML, 2019
2019
-
[27]
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. ICLR, 2020
2020
-
[28]
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world models. ICLR, 2021
2021
-
[29]
TLA: Tactile-language-action model for contact-rich manipulation
Peng Hao, Chaofan Zhang, Dingzhe Li, Xiaoge Cao, Xiaoshuai Hao, Shaowei Cui, and Shuo Wang. TLA: Tactile-language-action model for contact-rich manipulation. arXiv preprint arXiv:2503.08548 , 2025
Pith/arXiv arXiv 2025
-
[30]
Foar: Force-aware reactive policy for contact-rich robotic manipulation
Zihao He, Hongjie Fang, Jingjing Chen, Hao-Shu Fang, and Cewu Lu. Foar: Force-aware reactive policy for contact-rich robotic manipulation. RA-L, 2024
2024
-
[31]
Vitacformer: Learning cross-modal representation for visuo-tactile dexterous manipulation
Liang Heng, Haoran Geng, Kaifeng Zhang, Pieter Abbeel, and Jitendra Malik. Vitacformer: Learning cross-modal representation for visuo-tactile dexterous manipulation. arXiv preprint arXiv:2506.15953 , 2025
Pith/arXiv arXiv 2025
-
[32]
Carolina Higuera, Sergio Arnaud, Byron Boots, Mustafa Mukadam, Francois Robert Hogan, and Franziska Meier. Visuo-tactile world models. arXiv preprint arXiv:2602.06001 , 2026
arXiv 2026
-
[33]
Adaptive compliance policy: Learning approximate compliance for diffusion guided control
Yifan Hou, Zeyi Liu, Cheng Chi, Eric Cousineau, Naveen Kuppuswamy, Siyuan Feng, Benjamin Burch- fiel, and Shuran Song. Adaptive compliance policy: Learning approximate compliance for diffusion guided control. arXiv preprint arXiv:2410.09309 , 2024. 22
Pith/arXiv arXiv 2024
-
[34]
GAIA-1: A generative world model for autonomous driving
Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. GAIA-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023
Pith/arXiv arXiv 2023
-
[35]
3D-ViTac: Learning fine-grained manipulation with visuo-tactile sensing
Binghao Huang, Yixuan Wang, Xinyi Yang, Yiyue Luo, and Yunzhu Li. 3D-ViTac: Learning fine-grained manipulation with visuo-tactile sensing. CoRL, 2025
2025
-
[36]
Jialei Huang, Shuo Wang, Fanqi Lin, Yihang Hu, Chuan Wen, and Yang Gao. Tactile-VLA: Un- locking vision-language-action model’s physical knowledge for tactile generalization. arXiv preprint arXiv:2507.09160, 2025
Pith/arXiv arXiv 2025
-
[37]
Self forcing: Bridging the train-test gap in autoregressive video diffusion
Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train-test gap in autoregressive video diffusion. NeurIPS, 2025
2025
-
[38]
Foster, Pannag R
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P. Foster, Pannag R. Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: An open-source vision-language-action model. CoRL, 2025
2025
-
[39]
Cosmos policy: Fine-tuning video models for visuomotor control and planning
Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shu- ran Song, Ming-Yu Liu, Chelsea Finn, and Jinwei Gu. Cosmos policy: Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163 , 2026
Pith/arXiv arXiv 2026
-
[40]
Tenenbaum
Po-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun, and Joshua B. Tenenbaum. Learning to act from actionless videos through dense correspondences. ICLR, 2024
2024
-
[41]
Geonhyup Lee, Yeongjin Lee, Kangmin Kim, Seongju Lee, Sangjun Noh, Seunghyeok Back, and Kyoobin Lee. Manipforce: Force-guided policy learning with frequency-aware representation for contact-rich manipulation. arXiv preprint arXiv:2509.19047 , 2025
arXiv 2025
-
[42]
Cronusvla: Towards efficient and robust manipulation via multi-frame vision-language-action modeling
Hao Li, Shuai Yang, Yilun Chen, Xinyi Chen, Xiaoda Yang, Yang Tian, Hanqing Wang, Tai Wang, Dahua Lin, Feng Zhao, and Jiangmiao Pang. Cronusvla: Towards efficient and robust manipulation via multi-frame vision-language-action modeling. arXiv preprint arXiv:2506.19816 , 2025
arXiv 2025
-
[43]
Efficient-W AM: A 1B-parameter world-action model with low-cost future imagination
Jiajun Li, Tiecheng Guo, Yifan Ye, Rongyu Zhang, Xiaowei Chi, Qianpu Sun, Ying Li, Yunfan Lou, Yan Huang, Zhihe Lu, Meng Guo, and Shanghang Zhang. Efficient-W AM: A 1B-parameter world-action model with low-cost future imagination. arXiv preprint arXiv:2606.10040 , 2026
Pith/arXiv arXiv 2026
-
[44]
Metis: A generalizable and efficient world-action model for autonomous driving and urban navigation
Jingyu Li, Zhe Liu, Dongnan Hu, Junjie Wu, Zipei Ma, Wenxiao Wu, Chao Han, Zhihui Hao, Zhikang Liu, Kun Zhan, Jiankang Deng, Xiatian Zhu, and Li Zhang. Metis: A generalizable and efficient world-action model for autonomous driving and urban navigation. arXiv preprint arXiv:2606.15869 , 2026
arXiv 2026
-
[45]
Causal world modeling for robot control
Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Zhangluyao, Mingrui Yu, Zelin Gao, Nan Xue, Boyu Zhou, Xing Zhu, Mingyu Ding, Yujun Shen, and Yinghao Xu. Causal world modeling for robot control. arXiv preprint arXiv:2601.21998 , 2026
Pith/arXiv arXiv 2026
-
[46]
Shalfun Li, Victor Yao, Charles Yang, Truth Qu, Regis Cheng, Ryan Yu, Howard Lu, Newton Von, Vincent Chen, Yohann Tang, Maeve Zhang, Ellie Ma, Gody Li, Sage Yang, Lorien Shu, J. W. Gao, Ethan Chen, Colin Ye, Yu Sun, Elise Mon, PS Zhang, Neo Li, Lily Li, James Wang, Ping Yang, Chris Pan, Lucy Liang, Hang Su, Roy Gan, Hao Wang, and Qian Wang. Wall-wm: Carvi...
Pith/arXiv arXiv 2026
-
[47]
Unified video action model
Shuang Li, Yihuai Gao, Dorsa Sadigh, and Shuran Song. Unified video action model. RSS, 2025
2025
-
[48]
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. arXiv preprint arXiv:2209.07753, 2022. 23
Pith/arXiv arXiv 2022
-
[49]
Dreamitate: Real-world visuomotor policy learning via video generation
Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, and Carl Vondrick. Dreamitate: Real-world visuomotor policy learning via video generation. CoRL, 2024
2024
-
[50]
Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models
Weixin Liang, Lili Yu, Liang Luo, Srinivasan Iyer, Ning Dong, Chunting Zhou, Gargi Ghosh, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, and Xi Victoria Lin. Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models. TMLR, 2025
2025
-
[51]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. ICLR, 2023
2023
-
[52]
Dream-tac: A unified tactile world action model for contact-rich robot manipulation
Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen, Xiaowei Chi, Yaoxu Lyu, Peidong Jia, Sirui Han, Zhihe Lu, and Shanghang Zhang. Dream-tac: A unified tactile world action model for contact-rich robot manipulation. arXiv preprint arXiv:2606.08737 , 2026
Pith/arXiv arXiv 2026
-
[53]
Being-h0.7: A latent world-action model from egocentric videos
Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, Haiweng Xu, Chaoyi Xu, Ziheng Xi, Yuhui Fu, and Zongqing Lu. Being-h0.7: A latent world-action model from egocentric videos. arXiv preprint arXiv:2605.00078, 2026
Pith/arXiv arXiv 2026
-
[54]
Haoxiang Ma, Junhao Cai, Xiaoxu Xu, Hao Li, Yuyin Yang, Yang Tian, Jiafei Cao, Hongrui Zhu, Zherui Qiu, et al. Internvla-a1.5: Unifying understanding, latent foresight, and action for compositional generalization. arXiv preprint arXiv:2607.04988 , 2026
Pith/arXiv arXiv 2026
-
[55]
Scaling mixture-of-experts video pretraining for embodied intelligence
Shuailei Ma, Jiaqi Liao, Xinyang Wang, Jingjing Wang, Chaoran Feng, Zijing Hu, Chong Bao, Zichen Xi, Yuqi Gan, Weisen Wang, Yanhong Zeng, Qin Zhao, Zifan Shi, Wei Wu, Hao Ouyang, Qiuyu Wang, Shangzhan Zhang, Jiahao Shao, Yipengjing Sun, Liangxiao Hu, Lunke Pan, Nan Xue, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, and Ka Leong Cheng. Scaling mixture-o...
Pith/arXiv arXiv 2026
-
[56]
N0-foundation: Towards the age of tactile intelligence
NeoteAI Team and Fudan TEAI Team. N0-foundation: Towards the age of tactile intelligence. https: //research.neoteai.com/n0-foundation/, 2026
2026
-
[57]
T-Rex: Tactile-reactive dexterous manipulation
Dantong Niu, Zhuoyang Liu, Zekai Wang, ..., Linxi Fan, and Trevor Darrell. T-Rex: Tactile-reactive dexterous manipulation. arXiv preprint arXiv:2606.17055 , 2026
Pith/arXiv arXiv 2026
-
[58]
Cosmos world foundation model platform for physical AI
NVIDIA, Niket Agarwal, Arslan Ali, Maciej Bala, et al. Cosmos world foundation model platform for physical AI. arXiv preprint arXiv:2501.03575 , 2025
Pith/arXiv arXiv 2025
-
[59]
DINOv2: Learn- ing robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, et al. DINOv2: Learn- ing robust visual features without supervision. arXiv preprint arXiv:2304.07193 , 2023
Pith/arXiv arXiv 2023
-
[60]
Perceiver-actor: A multi-task transformer for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. CoRL, 2022
2022
-
[61]
Taxim: An example-based simulation model for GelSight tactile sensors
Zilin Si and Wenzhen Yuan. Taxim: An example-based simulation model for GelSight tactile sensors. RA-L, 2022
2022
-
[62]
Fast-dvla: Accelerating discrete diffusion VLA to real-time performance
Wenxuan Song, Jiayi Chen, Shuai Chen, Jingbo Wang, Pengxiang Ding, Han Zhao, Yikai Qin, Xinhu Zheng, Donglin Wang, Yan Wang, and Haoang Li. Fast-dvla: Accelerating discrete diffusion VLA to real-time performance. arXiv preprint arXiv:2603.25661 , 2026
Pith/arXiv arXiv 2026
-
[63]
Hy- embodied-0.5: Embodied foundation models for real-world agents
Tencent Robotics X, HY Vision Team, Xumin Yu, Zuyan Liu, Ziyi Wang, He Zhang, Yongming Rao, Fangfu Liu, Yani Zhang, Ruowen Zhao, Oran Wang, Yves Liang, Haitao Lin, Minghui Wang, Yubo Dong, Kevin Cheng, Bolin Ni, Rui Huang, Han Hu, Zhengyou Zhang, Linus, and Shunyu Yao. Hy- embodied-0.5: Embodied foundation models for real-world agents. arXiv preprint arXi...
Pith/arXiv arXiv 2026
-
[64]
VT-W AM: Visual-tactile world action model for contact-rich manipulation
Shuai Tian, Yupeng Zheng, Yuhang Zheng, Songen Gu, Yujie Zang, Yuxing Qin, Weize Li, Haoran Li, Wenchao Ding, and Dongbin Zhao. VT-W AM: Visual-tactile world action model for contact-rich manipulation. arXiv preprint arXiv:2607.02503 , 2026. 24
Pith/arXiv arXiv 2026
-
[65]
Wan: Open and advanced large-scale video generative models
Wan Team, Ang Wang, Baole Ai, Bin Wen, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 , 2025
Pith/arXiv arXiv 2025
-
[66]
Repwam: World action modeling with representation visual-action tokenizers
Junke Wang, Qihang Zhang, Shuai Yang, Yiming Luo, Yujun Shen, Zuxuan Wu, Yu-Gang Jiang, and Yinghao Xu. Repwam: World action modeling with representation visual-action tokenizers. arXiv preprint arXiv:2606.13674 , 2026
Pith/arXiv arXiv 2026
-
[67]
Hy-embodied- vlm-1.0: Efficient physical-world agents
Ziyi Wang, Xumin Yu, Yongming Rao, Yonggen Ling, Yunheng Li, Oran Wang, Mingqi Gao, Yuchen Zhou, Yves Liang, Zuyan Liu, Yani Zhang, Rui Huang, Xiaoran Xu, Bowen Yuan, Yifu Yuan, Xu Tan, He Zhang, Yufei Huang, Shenghao Zhang, Hongsheng Wu, Han Hu, and Zhengyou Zhang. Hy-embodied- vlm-1.0: Efficient physical-world agents. arXiv preprint arXiv:2607.12894 , 2026
Pith/arXiv arXiv 2026
-
[68]
Unleashing large-scale video generative pre-training for visual robot manipulation
Hongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen, Jiafeng Xu, Xinghang Li, Minghuan Liu, Hang Li, and Tao Kong. Unleashing large-scale video generative pre-training for visual robot manipulation. ICLR, 2024
2024
-
[69]
iVideoGPT: Interactive VideoGPTs are scalable world models
Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He, Dong Li, Jianye Hao, and Mingsheng Long. iVideoGPT: Interactive VideoGPTs are scalable world models. NeurIPS, 2024
2024
-
[70]
Tactile-W AM: Touch-aware world action model with tactile asymmetric attention
Siyu Wu, Linjing You, Junjie Zhu, Yaozu Liu, Changhao Zhang, Jian Liu, Weiqiang Wang, Qi Li, Jituo Li, and Hengshuang Zhao. Tactile-W AM: Touch-aware world action model with tactile asymmetric attention. arXiv preprint arXiv:2606.26663 , 2026
Pith/arXiv arXiv 2026
-
[71]
Canonical representation and force-based pretraining of 3d tactile for dexterous visuo-tactile policy learning
Tianhao Wu, Jinzhou Li, Jiyao Zhang, Mingdong Wu, and Hao Dong. Canonical representation and force-based pretraining of 3d tactile for dexterous visuo-tactile policy learning. ICRA, 2024
2024
-
[72]
From foundation to application: Improving vla models in practice
Wei Wu, Fangjing Wang, Fan Lu, He Sun, Shi Liu, Yunnan Wang, Yibin Yan, Yong Wang, Shuailei Ma, Xinyang Wang, Yibin Liu, Shuai Yang, Tianxiang Zhou, Kejia Zhang, Lei Zhou, Cheng Su, Nan Xue, Bin Tan, Han Zhang, Youchao Zhang, Fei Liao, Xing Zhu, Yujun Shen, and Kecheng Zheng. From foundation to application: Improving vla models in practice. arXiv preprint...
Pith/arXiv arXiv 2026
-
[73]
Tacdiffusion: Force-domain diffusion policy for precise tactile manip- ulation
Yansong Wu, Zongxie Chen, Fan Wu, Lingyun Chen, Liding Zhang, Zhenshan Bing, Abdalla Swikir, Sami Haddadin, and Alois Knoll. Tacdiffusion: Force-domain diffusion policy for precise tactile manip- ulation. ICRA, 2024
2024
-
[74]
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangj...
Pith/arXiv arXiv 2026
-
[75]
Next forcing: Causal world modeling with multi-chunk prediction
Gangwei Xu, Qihang Zhang, Jiaming Zhou, Xing Zhu, Yujun Shen, Xin Yang, and Yinghao Xu. Next forcing: Causal world modeling with multi-chunk prediction. arXiv preprint arXiv:2606.11187 , 2026
Pith/arXiv arXiv 2026
-
[76]
Seeing touch from motion: A unified modality-aware visuo-tactile policy with tactile motion correlation
Shengqi Xu, Guojin Zhong, Yang Liu, Fanjie Wang, Hu Luo, Hanyu Zhou, Weiyao Zhang, Ziyi Ye, Zuxuan Wu, and Yu-Gang Jiang. Seeing touch from motion: A unified modality-aware visuo-tactile policy with tactile motion correlation. ECCV, 2026
2026
-
[77]
S-V AM: Shortcut video- action model by self-distilling geometric and semantic foresight
Haodong Yan, Zhide Zhong, Jiaguan Zhu, Junjie He, Weilin Yuan, Wenxuan Song, Xin Gong, Yingjie Cai, Guanyi Zhao, Xu Yan, Bingbing Liu, Ying-Cong Chen, and Haoang Li. S-V AM: Shortcut video- action model by self-distilling geometric and semantic foresight. arXiv preprint arXiv:2603.16195 , 2026
arXiv 2026
-
[78]
Moma-force: Visual-force imitation for real-world mobile manipulation
Taozheng Yang, Ya Jing, Hongtao Wu, Jiafeng Xu, Kuankuan Sima, Guangzeng Chen, Qie Sima, and Tao Kong. Moma-force: Visual-force imitation for real-world mobile manipulation. IROS, 2023
2023
-
[79]
GigaWorld-Policy: An efficient action- centered world-action model
Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, et al. GigaWorld-Policy: An efficient action- centered world-action model. arXiv preprint arXiv:2603.17240 , 2026. 25
arXiv 2026
-
[80]
Learning to feel the future: DreamTacVLA for contact-rich manipulation
Guo Ye, Zexi Zhang, Xu Zhao, Shang Wu, Haoran Lu, Shihan Lu, and Han Liu. Learning to feel the future: DreamTacVLA for contact-rich manipulation. arXiv preprint arXiv:2512.23864 , 2025
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.