REVIEW 3 major objections 6 minor 1 cited by
A tactile-aware world model turns real robot failures into imagined video-and-force corrections that raise contact-rich VLA success by 44 percent.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 06:40 UTC pith:VJWSS2BC
load-bearing objection Solid systems paper: real multi-task contact gains from a tactile Recognize–Imagine–Label loop, with the main open bet being whether imagined forces are truly executable recoveries. the 3 major comments →
TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
TACO establishes that localized contact failures in VLA policies can be repaired at scale without repeated human intervention by converting real rollouts into imagined visuo-tactile corrections: failure-adjacent states are recognized from progress estimates, local video-and-force recovery segments are generated by joint denoising, those segments are labeled with corrective actions, and the resulting supervision is fed back through knowledge-insulated, advantage-conditioned post-training that preserves pretrained visual-language priors.
What carries the argument
The tactile-aware world model—a visuo-tactile generator that jointly denoises future video and 12-D force via temporal RoPE alignment, plus a unified progress-action model that both detects progress stalls and labels corrective actions—together with knowledge-insulated tactile adaptation that routes tactile learning only into the action expert.
Load-bearing premise
The imagined video-and-force corrections must be contact-consistent enough that the actions labeled from them actually recover the real robot rather than inventing force patterns the hardware cannot execute.
What would settle it
On held-out real recovery trajectories, measure force prediction error of the imagined segments and whether the labeled corrective actions execute successfully on the robot; systematic force inconsistency at contact transitions plus low real execution success would falsify the claim that offline imagination yields valid recovery supervision.
If this is right
- Contact-rich VLA post-training can run as a closed real-to-imagine-to-real loop without continuous human recovery demos.
- Tactile signals must participate in both imagination and action labeling; vision-only imagination collapses real recovery performance.
- Freezing the pretrained vision-language backbone while training only the action expert is required to keep pre-contact approach and grounding intact.
- Increasing the ratio of imagined corrections to real data continues to raise success rates on the evaluated tasks.
Where Pith is reading between the lines
- The same offline imagination loop could transfer to multi-arm or mobile platforms once a comparable force channel is available.
- If online generation during deployment becomes practical, recovery could close without waiting for the next post-training round.
- Progress-based anchor selection may serve as a cheap failure detector even when full corrective imagination is unavailable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TACO proposes a tactile-aware world-model framework for scalable post-training of Vision-Language-Action (VLA) policies on contact-rich manipulation. From real rollouts, a Recognize–Imagine–Label loop uses a unified progress-action model to locate failure-adjacent states, a visuo-tactile generation model (joint video–force denoising with temporal RoPE and first-frame force anchoring) to synthesize local correction segments, and the same progress-action model to label corrective actions with binary advantage. Post-training combines knowledge-insulated tactile adaptation (stop-gradient on the VLM backbone; tactile/advantage conditioning only in the action expert) with advantage-conditioned flow matching. On six real Franka tasks (40 episodes each, two iterations), TACO reports average success rate rising from 0.38 (base π0.5) to 0.82, a 44% absolute gain and 32% over the same pipeline without knowledge insulation, with supporting ablations on tactile generation/labeling, data scaling, advantage conditioning, and anchor selection, plus OOD and action-distribution analyses.
Significance. Contact-rich recovery remains a clear bottleneck for VLAs, and human corrective intervention does not scale. If the reported gains hold under stronger validation of imagined contact dynamics, TACO offers a practical closed loop that turns real failures into tactile-aware corrective supervision without repeated teleoperation. The combination of joint visuo-tactile imagination, knowledge-insulated adaptation of a strong pretrained VLA (π0.5), and advantage-conditioned offline RL is a coherent systems contribution with real-robot evidence across six tasks, ablations, and OOD probes. Strengths include multi-task physical evaluation, explicit insulation of VLM priors, and component ablations that separate tactile generation, tactile labeling, KI, advantage, and anchor selection. The work is of clear interest to robot learning and embodied foundation-model communities.
major comments (3)
- [§3.1–3.2, Fig. 5, Appendix C.1] §3.1–3.2 and the abstract treat contact-consistent imagined force trajectories as load-bearing for the self-corrector claim, yet validation is only indirect. Appendix C.1 reports force prediction loss F on held-out real segments and Fig. 5 shows SR drops without tactile generation (0.28) or tactile labeling (0.65), but there is no direct comparison of imagined force sequences (or contact-event timing) against real recovery trajectories at failure-adjacent anchors, nor open-loop replay of labeled actions on the Franka. Without such checks, the 44% SR gain in Table 1 could be driven largely by KI, advantage conditioning, and extra data volume rather than faithful tactile self-correction. Please add quantitative force/contact consistency metrics (e.g., force RMSE, contact onset timing) and/or replay success of labeled corrections, or clearly scope the claim to empirical post-training gains.
- [Table 1, §4.2] Table 1 reports point success rates over 40 episodes per task with no standard errors, confidence intervals, or statistical tests across methods or iterations. Given that the headline claims are absolute SR deltas of 44% and 32%, and that several per-task jumps are large (e.g., Move Hanoi Rings 0.08→0.79), the central comparison needs uncertainty quantification (binomial CIs or bootstrap over episodes) and, where possible, a simple significance test between TACO, TACO (w/o KI), and Filtered BC. Without this, it is hard to judge robustness of the ranking, especially for intermediate Iteration-1 results where some tasks are closer.
- [§3.1–3.2, Appendix B.5, C.1, C.4] The Recognize step (§3.2) and UPA training (§3.1) depend on manually annotated task-stage progress labels and a stall criterion (window Δ, threshold ε). Appendix B.5 further notes up to 10 anchors per failed trajectory and binary advantage assignment from the first recognized onset. This makes failure localization and advantage labels partly human-defined rather than fully autonomous. Please report sensitivity of SR to ε/Δ and to the number of anchors, and clarify how much of the Recognize quality (FL in Appendix C.1) is driven by the manual stage taxonomy versus the learned progress head. If manual stage labels are required per task, the scalability claim relative to human intervention should be qualified.
minor comments (6)
- [Figure 5] Fig. 5 mixes a table-like ablation with a scaling plot; axis labels, sample sizes for Val. Loss / VOC / FL, and whether Real SR is after one or two iterations should be stated in the caption.
- [§3.1, §3.3] Notation reuses λ_f for both the force term in L_joint (§3.1) and the force conditioning weight in L_π (§3.3); rename one of them to avoid confusion.
- [Table 1] Completion steps (CS) in Table 1 are averaged only over successful episodes; note this explicitly in the main text so lower CS is not misread as always-faster execution including failures.
- [§2] Related Work cites many concurrent arXiv world-model/VLA papers; a short positioning paragraph against the closest tactile world models [61,62] and knowledge-insulation [30] would help readers separate TACO’s loop from prior components.
- [Appendix B.2] Algorithm 1 in the supplement is useful; consider moving a condensed version or a pointer into the main §3.2 so the iterative real-to-imagine-to-real loop is fully specified without the supplement.
- [Title, Table 1] Typos/formatting: title spacing “forScalable”; Table 1 Iteration-2 TACO (w/o KI) cell “0.65510.52” appears concatenated; fix for camera-ready.
Circularity Check
No circular derivation: TACO is an empirical robotics pipeline whose success claims are measured on physical rollouts, not reduced to fitted identities or self-citation theorems.
full rationale
The paper's load-bearing claim is an empirical success-rate gain (Table 1: base 0.38 → TACO 0.82 avg SR after two iterations; +32% vs TACO w/o KI) measured on 40 real Franka episodes per task with randomized object positions. That outcome is not derived from an equation that reuses its own fit. Progress targets for the unified progress-action model come from manually annotated task-stage labels (Sec. 3.1), not from the quantity later reported as success. Advantage labels are assigned by construction (y=1 for expert demos, successful rollouts, and imagined corrections; y=0 for failed segments after the first progress-stall anchor; Algorithm 1 / Sec. 3.2–3.3)—this is an offline RL training design choice, not a claimed first-principles prediction of advantage. The visuo-tactile generator is trained with joint flow-matching on real video-force data and evaluated via held-out force loss, VOC, FL, and real SR ablations (Fig. 5, App. C.1); none of these equate a fitted parameter to a reported prediction. Knowledge insulation and advantage-conditioned flow cite external work ([30], [31], [32]), not author-only uniqueness theorems. The Recognize–Imagine–Label loop is standard iterative self-improvement (policy → real rollouts → synthetic supervision → policy), which is self-dependent as a training procedure but not circular in the sense that a claimed derivation collapses to its inputs. Weaknesses (e.g., limited direct force-consistency checks of imagined recoveries) are correctness/validation risks, not circularity.
Axiom & Free-Parameter Ledger
free parameters (6)
- progress stall threshold ε and window Δ
- correction horizon T=49
- force loss weight λ_f and advantage/force conditioning weights λ_f, λ_a
- real-to-imagined data ratio (≈1:4–1:8, up to 1:10 in ablations)
- manual task-stage progress labels for UPA training
- CFG null-condition dropout probability 0.1 and positive-advantage inference condition
axioms (5)
- domain assumption Contact-rich VLA failures are primarily localized contact-state errors rather than task-level semantic errors, so local corrective post-training is sufficient.
- domain assumption Vision-only imagined rollouts can be contact-inconsistent; joint video–force generation is necessary for useful corrections.
- domain assumption Stop-gradient knowledge insulation of the VLM backbone preserves pre-contact visual-language priors while allowing the action expert to absorb tactile corrections.
- ad hoc to paper Binary advantage labels (y∈{0,1}) on failed vs corrective segments are an adequate offline RL signal for recovery.
- ad hoc to paper Dense progress from a unified progress-action model reliably localizes failure-adjacent anchors (stall/decrease criterion).
invented entities (3)
-
TACO Recognize–Imagine–Label loop with tactile-aware world model
no independent evidence
-
Visuo-tactile joint denoising generation with temporal RoPE force–video alignment and first-frame force anchor
no independent evidence
-
Unified progress-action model (UPA) for joint progress and corrective action prediction
no independent evidence
read the original abstract
Vision-Language-Action (VLA) models have shown promising generalization in robotic manipulation, but they still struggle with contact-rich tasks, where minor contact perturbations can cause unrecoverable failures that are hard to detect from vision alone. Since these failures are localized rather than task-level semantic errors, tactile-aware corrective post-training offers an efficient way to improve recovery. However, scaling such supervision through human intervention is costly. Recent works have explored world models to synthesize imagined rollouts for policy improvement, but vision-only world models may produce visually plausible yet contact-inconsistent trajectories. We therefore introduce TACO, a tactile-aware world-model-driven framework for scalable VLA post-training in contact-rich manipulation. Given real robot rollouts, TACO follows a Recognize-Imagine-Label loop with a tactile-aware world model: a unified progress-action model recognizes failure-adjacent states using progress estimates, a visuo-tactile generation model imagines local correction segments, and the progress-action model labels them with executable corrective actions. To incorporate tactile corrective supervision into VLA post-training, TACO combines knowledge-insulated tactile adaptation with advantage-conditioned training, enabling the policy to learn from imagined corrections without degrading pretrained visual-language priors. These components enable TACO to convert real-world failures into imagined visuo-tactile corrections for iterative VLA post-training. Experiments on real-world contact-rich manipulation tasks show that TACO achieves 44% absolute success rate improvement over the base policy and 32% over the policy without knowledge-insulated tactile adaptation.
Figures
Forward citations
Cited by 1 Pith paper
-
{\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
Action-conditioned JEPA-style future-visual latent prediction yields dynamics-aware tactile tokens that lift contact-rich VLA success rates from ~30% to ~70% average on four real tasks.
Reference graph
Works this paper leans on
-
[1]
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, et al. Rt-1: Robotics transformer for real-world control at scale.arXiv preprint arXiv:2212.06817, 2022
Pith/arXiv arXiv 2022
-
[2]
Zitkovich, T
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, pages 2165–2183. PMLR, 2023
2023
-
[3]
O’Neill, A
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024
2024
-
[4]
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al. Octo: An open-source generalist robot policy.arXiv preprint arXiv:2405.12213, 2024
Pith/arXiv arXiv 2024
-
[5]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Pith/arXiv arXiv 2024
-
[6]
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al.π 0 : A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024
Pith/arXiv arXiv 2024
-
[7]
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al.π 0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025. 9
Pith/arXiv arXiv 2025
-
[8]
P. Intelligence, B. Ai, A. Amin, R. Aniceto, A. Balakrishna, G. Balke, K. Black, G. Bokinsky, S. Cao, T. Charbonnier, et al.π 0.7 : a steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026
Pith/arXiv arXiv 2026
-
[9]
Q. Li, Y . Liang, Z. Wang, L. Luo, X. Chen, M. Liao, F. Wei, Y . Deng, S. Xu, Y . Zhang, et al. Cogact: A foundational vision-language-action model for synergizing cognition and action in robotic manipulation.arXiv preprint arXiv:2411.19650, 2024
Pith/arXiv arXiv 2024
-
[10]
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. Rdt-1b: a diffu- sion foundation model for bimanual manipulation. InInternational Conference on Learning Representations, volume 2025, pages 29982–30009, 2025
2025
-
[11]
J. Wen, Y . Zhu, M. Zhu, Z. Tang, J. Li, Z. Zhou, X. Liu, C. Shen, Y . Peng, and F. Feng. Diffusionvla: Scaling robot foundation models via unified diffusion and autoregression. In Forty-second International Conference on Machine Learning, 2025
2025
-
[12]
Y . Jia, J. Liu, S. Liu, R. Zhou, W. Yu, Y . Yan, X. Chi, Y . Guo, B. Shi, and S. Zhang. Video2act: A dual-system video diffusion policy with robotic spatio-motional modeling.arXiv preprint arXiv:2512.03044, 2025
arXiv 2025
-
[13]
Z. Liu, J. Liu, H. Chen, J. Yu, Z. Guo, C. Hou, C. Gu, X. Mi, R. Zhang, K. Wu, et al. Last {0}: Latent spatio-temporal chain-of-thought for robotic vision-language-action model.arXiv preprint arXiv:2601.05248, 2026
arXiv 2026
-
[14]
S. Ross, G. Gordon, and D. Bagnell. A reduction of imitation learning and structured predic- tion to no-regret online learning. InProceedings of the fourteenth international conference on artificial intelligence and statistics, pages 627–635. JMLR Workshop and Conference Pro- ceedings, 2011
2011
-
[15]
Hoque, A
R. Hoque, A. Mandlekar, C. Garrett, K. Goldberg, and D. Fox. Intervengen: Interventional data generation for robust and data-efficient robot imitation learning. In2024 IEEE/RSJ In- ternational Conference on Intelligent Robots and Systems (IROS), pages 2840–2846. IEEE, 2024
2024
-
[16]
Korkmaz and E
Y . Korkmaz and E. Bıyık. Mile: Model-based intervention learning. In2025 IEEE Interna- tional Conference on Robotics and Automation (ICRA), pages 15673–15679. IEEE, 2025
2025
-
[17]
X. Xu, Y . Hou, Z. Liu, and S. Song. Compliant residual dagger: Improving real-world contact- rich manipulation with human corrections.Advances in Neural Information Processing Sys- tems, 38:139559–139581, 2026
2026
-
[18]
Y . Wang, R. Syed, F. Wu, M. Zhang, A. Onol, J. Barreiros, H. Nayyeri, T. Dear, H. Zhang, and Y . Li. Interactive world simulator for robot policy training and evaluation.arXiv preprint arXiv:2603.08546, 2026
arXiv 2026
-
[19]
Y . Li, Z. Zhou, Y . Chen, Y . Guo, J. Liu, S. Zhang, J. Chen, and Y . Zhu. Hi-wm: Human-in- the-world-model for scalable robot post-training.arXiv preprint arXiv:2604.21741, 2026
Pith/arXiv arXiv 2026
-
[20]
Q. Xu, J. Liu, R. Zhou, S. Shi, N. Han, Z. Liu, C. Gu, S. Gu, Y . Yue, G. Huang, et al. Twinrl- vla: Digital twin-driven reinforcement learning for real-world robotic manipulation.arXiv preprint arXiv:2602.09023, 2026
Pith/arXiv arXiv 2026
-
[21]
W. Yu, J. Lv, Z. Ying, Y . Jin, C. Wen, and C. Lu. Armada: Autonomous online failure detection and human shared control empower scalable real-world deployment and adaptation.arXiv preprint arXiv:2510.02298, 2025
arXiv 2025
-
[22]
S. Zhou, Y . Du, J. Chen, Y . Li, D.-Y . Yeung, and C. Gan. Robodreamer: Learning composi- tional world models for robot imagination.arXiv preprint arXiv:2404.12377, 2024. 10
Pith/arXiv arXiv 2024
-
[23]
Y . Li, X. Wei, X. Chi, Y . Li, Z. Zhao, H. Wang, N. Ma, M. Lu, and S. Zhang. Manip- dreamer: Boosting robotic manipulation world model with action tree and visual guidance. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), pages 12027–12031. IEEE, 2026
2026
-
[24]
Y . Guo, T. Lee, L. X. Shi, J. Chen, P. Liang, and C. Finn. Vlaw: Iterative co-improvement of vision-language-action policy and world model.arXiv preprint arXiv:2602.12063, 2026
arXiv 2026
-
[25]
F. Zhu, Z. Yan, Z. Hong, Q. Shou, X. Ma, and S. Guo. Wmpo: World model-based policy optimization for vision-language-action models.arXiv preprint arXiv:2511.09515, 2025
arXiv 2025
-
[26]
J. Yang, K. Lin, J. Li, W. Zhang, T. Lin, L. Wu, Z. Su, H. Zhao, Y .-Q. Zhang, L. Chen, et al. Rise: Self-improving robot policy with compositional world model.arXiv preprint arXiv:2602.11075, 2026
Pith/arXiv arXiv 2026
-
[27]
Z. Jiang, S. Zhou, Y . Jiang, Z. Huang, M. Wei, Y . Chen, T. Zhou, Z. Guo, H. Lin, Q. Zhang, et al. Wovr: World models as reliable simulators for post-training vla policies with rl.arXiv preprint arXiv:2602.13977, 2026
Pith/arXiv arXiv 2026
-
[28]
J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y . Fang, F. Hu, S. Huang, K. Kundalia, Y .-C. Lin, et al. Dreamgen: Unlocking generalization in robot learning through video world models. arXiv preprint arXiv:2505.12705, 2025
Pith/arXiv arXiv 2025
-
[29]
A. Yu, Z. Chen, P. Song, Z. Hong, H. Wang, D. Zhang, T. He, Y . Ding, and D. Zhang. Wm- dagger: Enabling efficient data aggregation for imitation learning with world models.arXiv preprint arXiv:2604.11351, 2026
Pith/arXiv arXiv 2026
-
[30]
Driess, J
D. Driess, J. Springenberg, B. Ichter, L. Yu, A. Li-Bell, K. Pertsch, A. Ren, H. Walke, Q. Vuong, L. X. Shi, et al. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better.Advances in Neural Information Processing Systems, 38:102867– 102888, 2026
2026
-
[31]
K. Frans, S. Park, P. Abbeel, and S. Levine. Diffusion guidance is a controllable policy im- provement operator.arXiv preprint arXiv:2505.23458, 2025
Pith/arXiv arXiv 2025
-
[32]
P. Intelligence, A. Amin, R. Aniceto, A. Balakrishna, K. Black, K. Conley, G. Connors, J. Darpinian, K. Dhabalia, J. DiCarlo, et al.π 0.6: a vla that learns from experience.arXiv preprint arXiv:2511.14759, 2025
Pith/arXiv arXiv 2025
-
[33]
Q. Feng, J. Yu, J. Liu, Y . Jia, Z. Wu, H. Chen, Z. Qian, S. Gu, P. Jia, S. Ma, et al. Harmowam: Harmonizing generalizable and precise manipulation via adaptive world action models.arXiv preprint arXiv:2605.10942, 2026
Pith/arXiv arXiv 2026
-
[34]
H. Tan, Y . Feng, X. Mao, S. Huang, G. Liu, Z. Hao, H. Su, and J. Zhu. Anypos: Automated task-agnostic actions for bimanual manipulation.arXiv preprint arXiv:2507.12768, 2025
Pith/arXiv arXiv 2025
-
[35]
W. Mi, Y . Bao, X. Chi, X. Ju, Z. Qin, K. Ge, K. Tang, P. Jia, S. Zhang, and J. Tang. Tc- idm: Grounding video generation for executable zero-shot robot motion.arXiv preprint arXiv:2601.18323, 2026
arXiv 2026
-
[36]
K. Li, Z. Jing, X. Wang, Z. Zhu, Y . Zhou, G. Huang, D. Li, Q. Yang, and H. Huang. Stableidm: Stabilizing inverse dynamics model against manipulator truncation via spatio-temporal refine- ment.arXiv preprint arXiv:2604.17887, 2026
Pith/arXiv arXiv 2026
-
[37]
Y . J. Ma, V . Kumar, A. Zhang, O. Bastani, and D. Jayaraman. Liv: Language-image repre- sentations and rewards for robotic control. InInternational Conference on Machine Learning, pages 23301–23320. PMLR, 2023
2023
-
[38]
T. Lee, A. Wagenmaker, K. Pertsch, P. Liang, S. Levine, and C. Finn. Roboreward: General- purpose vision-language reward models for robotics.arXiv preprint arXiv:2601.00675, 2026. 11
arXiv 2026
-
[39]
J. Lv, H. Li, J. Li, Y . Nie, F. Kong, Y . Wang, X. Wang, Z. Zhu, C. Ni, Q. Deng, et al. Viva: A video-generative value model for robot reinforcement learning.arXiv preprint arXiv:2604.08168, 2026
Pith/arXiv arXiv 2026
-
[40]
H. Tan, S. Chen, Y . Xu, Z. Wang, Y . Ji, C. Chi, Y . Lyu, Z. Zhao, X. Chen, P. Co, et al. Robo- dopamine: General process reward modeling for high-precision robotic manipulation.arXiv preprint arXiv:2512.23703, 2025
arXiv 2025
-
[41]
A. Liang, Y . Korkmaz, J. Zhang, M. Hwang, A. Anwar, S. Kaushik, A. Shah, A. S. Huang, L. Zettlemoyer, D. Fox, et al. Robometer: Scaling general-purpose robotic reward models via trajectory comparisons.arXiv preprint arXiv:2603.02115, 2026
Pith/arXiv arXiv 2026
- [42]
-
[43]
J. Xiao, Y . Yang, X. Chang, R. Chen, F. Xiong, M. Xu, W.-S. Zheng, and Q. Zhang. World- env: Leveraging world model as a virtual environment for vla post-training.arXiv preprint arXiv:2509.24948, 2025
Pith/arXiv arXiv 2025
-
[44]
A. K. Sharma, Y . Sun, N. Lu, Y . Zhang, J. Liu, and S. Yang. World-gymnast: Training robots with reinforcement learning in a world model.arXiv preprint arXiv:2602.02454, 2026
arXiv 2026
-
[45]
X. Liu, Z. Bai, H. Ci, K. Y . Ma, and M. Z. Shou. World-vla-loop: Closed-loop learning of video world model and vla policy.arXiv preprint arXiv:2602.06508, 2026
Pith/arXiv arXiv 2026
-
[46]
G. Team, B. Wang, B. Li, C. Ni, G. Huang, G. Zhao, H. Li, J. Li, J. Lv, J. Liu, et al. Gigabrain- 0.5 m*: a vla that learns from world model-based reinforcement learning.arXiv preprint arXiv:2602.12099, 2026
arXiv 2026
-
[47]
Z. Liu, J. Liu, J. Xu, N. Han, C. Gu, H. Chen, K. Zhou, R. Zhang, K. C. Hsieh, K. Wu, et al. Mla: A multisensory language-action model for multimodal understanding and forecasting in robotic manipulation.arXiv preprint arXiv:2509.26642, 2025
arXiv 2025
-
[48]
J. Yu, H. Liu, Q. Yu, J. Ren, C. Hao, H. Ding, G. Huang, G. Huang, Y . Song, P. Cai, et al. Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation.Ad- vances in Neural Information Processing Systems, 38:93409–93439, 2026
2026
-
[49]
Y . Li, H. Jiang, J. Xia, H. Zhang, J. Du, Y . Zhou, J. Zeng, C. Hao, J. Ren, Q. Yu, et al. Forcevla2: Unleashing hybrid force-position control with force awareness for contact-rich ma- nipulation.arXiv preprint arXiv:2603.15169, 2026
arXiv 2026
-
[50]
J. Huang, S. Wang, F. Lin, Y . Hu, C. Wen, and Y . Gao. Tactile-vla: unlocking vision- language-action model’s physical knowledge for tactile generalization.arXiv preprint arXiv:2507.09160, 2025
Pith/arXiv arXiv 2025
- [51]
-
[52]
C. Morissette, A. Abyaneh, W.-D. Chang, A. Houssaini, D. Meger, H.-C. Lin, J. Tremblay, and G. Dudek. Tactile modality fusion for vision-language-action models.arXiv preprint arXiv:2603.14604, 2026
arXiv 2026
- [53]
- [54]
-
[55]
J. Bi, K. Y . Ma, C. Hao, M. S. Zheng, and H. Soh. Vla-touch: Enhancing vision-language- action model with dual-level tactile feedback.IEEE Robotics and Automation Letters, 2026
2026
-
[56]
S. Yu, K. Lin, A. Xiao, J. Duan, and H. Soh. Octopi: Object property reasoning with large tactile-language models.arXiv preprint arXiv:2405.02794, 2024
Pith/arXiv arXiv 2024
-
[57]
Z. Wang, Y . Wang, M. Ren, P. Li, Y . Liu, Y . Nie, L. Long, Y . Ye, X. Wang, Z. Zhu, et al. Tacmamba: A tactile history compression adapter bridging fast reflexes and slow vla reasoning. arXiv preprint arXiv:2603.01700, 2026
arXiv 2026
-
[58]
B. Huang, Y . Wang, X. Yang, Y . Luo, and Y . Li. 3d-vitac: Learning fine-grained manipulation with visuo-tactile sensing.arXiv preprint arXiv:2410.24091, 2024
Pith/arXiv arXiv 2024
-
[59]
K. Gubernatorov, M. Sannikov, I. Mikhalchuk, E. Kuznetsov, M. Artemov, O. F. Ouwatobi, M. Fernando, A. Asanov, Z. Guo, and D. Tsetserukou. Hapticvla: Contact-rich manipula- tion via vision-language-action model without inference-time tactile sensing.arXiv preprint arXiv:2603.15257, 2026
arXiv 2026
-
[60]
H. Xue, J. Ren, W. Chen, G. Zhang, Y . Fang, G. Gu, H. Xu, and C. Lu. Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation.arXiv preprint arXiv:2503.02881, 2025
Pith/arXiv arXiv 2025
- [61]
-
[62]
C. Higuera, S. Arnaud, B. Boots, M. Mukadam, F. R. Hogan, and F. Meier. Visuo-tactile world models.arXiv preprint arXiv:2602.06001, 2026
arXiv 2026
-
[63]
X. Li, M. Cai, J. Xu, J. Zhu, H. Fan, Y . Shen, G. Ren, and H. Dong. At-vla: Adaptive tactile injection for enhanced feedback reaction in vision-language-action models.arXiv preprint arXiv:2605.07308, 2026
Pith/arXiv arXiv 2026
-
[64]
T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[65]
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haz- iza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023
Pith/arXiv arXiv 2023
-
[66]
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2024
Pith/arXiv arXiv 2024
-
[67]
Q. Bu, J. Cai, L. Chen, X. Cui, Y . Ding, S. Feng, S. Gao, X. He, X. Hu, X. Huang, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669, 2025
Pith/arXiv arXiv 2025
-
[68]
K. Wu, C. Hou, J. Liu, Z. Che, X. Ju, Z. Yang, M. Li, Y . Zhao, Z. Xu, G. Yang, et al. Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation.arXiv preprint arXiv:2412.13877, 2024. 13 TACO: TActile World Model as a Self-COrrector for Scalable VLA Post-Training Supplementary Material We provide additional details, as wel...
Pith/arXiv arXiv 2024
-
[69]
S1: insert the flower into the vase
Insert Flower.The robot picks up a flower from a random location on the table and inserts it into a vase. S1: insert the flower into the vase. S1 is considered successful if the robot securely inserts the flower into the vase without dropping it or knocking over the vase
-
[70]
S1: grasp the eraser; S2: wipe off the star
Wipe Whiteboard.The robot picks up an eraser and wipes off a star drawn at a random location on a whiteboard. S1: grasp the eraser; S2: wipe off the star. S1 is considered successful if the robot securely grasps and lifts the eraser. S2 is considered successful if the robot completely erases the star from the whiteboard
-
[71]
S1: grasp the bottle cap; S2: twist open and lift the cap
Twist Bottle Cap.The robot grasps the cap of a water bottle, applies a twisting motion to loosen it, and then lifts the cap away from the bottle. S1: grasp the bottle cap; S2: twist open and lift the cap. S1 is considered successful if the robot securely grips the cap with stable contact and without displacing the bottle. S2 is considered successful if th...
-
[72]
S1: strike the 1st key; S2: strike the 3rd key; S3: strike the 5th key; S4: strike the 8th key
Play Xylophone.The robot picks up a mallet and sequentially strikes the 1st, 3rd, 5th, and 8th keys of an eight-key xylophone in order. S1: strike the 1st key; S2: strike the 3rd key; S3: strike the 5th key; S4: strike the 8th key. Each stage is considered successful if the robot accurately strikes the designated key with the mallet tip, while following t...
-
[73]
S1: grasp the first slice; S2: insert the first slice into the toaster; S3: grasp the second slice; S4: insert the second slice
T oast Bread.The robot sequentially picks up two slices of bread from a box and inserts them into a toaster. S1: grasp the first slice; S2: insert the first slice into the toaster; S3: grasp the second slice; S4: insert the second slice. Each stage is considered successful if the corresponding action is completed without dropping the bread or misaligning ...
-
[74]
S1: grasp the top ring; S2: place the top ring on the left peg; S3: grasp the second ring; S4: place the second ring on the right peg
Move Hanoi Rings.The robot moves the top ring from the middle peg of a Hanoi tower (where all rings are initially stacked in size order) to the left empty peg, then moves the next ring to the right empty peg. S1: grasp the top ring; S2: place the top ring on the left peg; S3: grasp the second ring; S4: place the second ring on the right peg. S1 is conside...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.