Pith. sign in

REVIEW 3 major objections 6 minor 3 cited by

TouchWorld shows that treating touch as both a predicted contact goal and a fast residual correction signal, on separate timescales from vision-language planning, raises success on long-horizon contact-rich robot tasks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 19:38 UTC pith:COQLFXWB

load-bearing objection Solid multi-rate tactile systems paper with real-robot gains; the hierarchy is useful, but residual refinement may carry more of the lift than the foundation-model framing admits. the 3 major comments →

arxiv 2607.07287 v2 pith:COQLFXWB submitted 2026-07-08 cs.RO

TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation

classification cs.RO
keywords tactile foundation modeldexterous manipulationhierarchical policytactile world modelresidual refinementvision-language-actioncontact-rich manipulationreal-robot benchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Everyday dexterous skills fail when a robot cannot anticipate how contact should evolve and cannot quickly fix slip, misalignment, or force errors. Vision and language give semantic guidance but hide those contact states; most policies still fold touch into one slow monolithic action loop. TouchWorld claims that a multi-rate hierarchy fixes this: a slow planner emits executable subtasks and tactile subgoals, a mid-rate visuo-tactile policy generates nominal action chunks, and a high-rate residual policy corrects those chunks from recent tactile and proprioceptive feedback. On six real-robot tasks spanning watering, clearing, insertion, wiping, and soft-object pulling, this design reaches 65.0% average success clean and 53.7% under human perturbation, beating the strongest baseline by 15.7 and 18.5 points. A sympathetic reader cares because the result says foundation robot policies can keep language-level generalization while becoming robust at contact if touch is given both a predictive and a reactive role.

Core claim

The paper establishes that a predictive-and-reactive tactile foundation model, implemented as a hierarchical multi-timescale policy, improves long-horizon contact-rich dexterous manipulation by separating vision-language subtask planning, tactile world-model subgoal prediction, visuo-tactile goal-conditioned action generation, and high-frequency tactile residual refinement, rather than coupling them in a single monolithic loop.

What carries the argument

TouchWorld hierarchy: High-Level Planning Layer (Subtask Planner plus Tactile World Model) supplies executable subtasks and predicted visual-tactile subgoals; Visuo-Tactile Goal-Conditioned Policy emits nominal action chunks; Tactile-Conditioned Refinement Policy (Tactile Residual Transformer) adds online residuals from high-frequency tactile and proprioceptive history.

Load-bearing premise

The claim assumes that the fixed multi-rate schedule, image-form tactile interface, and residual action subspace used on this hand-and-glove setup will transfer as a general foundation rather than remaining a platform- and task-specific stack.

What would settle it

On the same six tasks with matched data and sensing, a strong monolithic tactile policy that fuses touch at the same rate as vision-language tokens would match or exceed TouchWorld's clean and perturbation success rates, or ablations removing residual refinement and tactile world-model goals would not drop performance as reported.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Contact-rich foundation policies should keep semantic planning slow while dedicating a separate high-rate path to tactile residual correction.
  • Predicted tactile subgoals can condition nominal action chunks without forcing the full model to replan at contact rates.
  • Human bimanual tactile pretraining plus robot fine-tuning can supply contact priors that improve robot tactile subgoal prediction.
  • Perturbation robustness becomes a first-class evaluation axis for tactile VLA systems, not an afterthought.
  • Modular fallbacks (task prompt only, no predicted goals) remain usable when the planner or world model is unavailable.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If residual commit interval and world-model refresh rules were made contact-state adaptive, compute cost could fall during stable contact phases without losing recovery speed.
  • The same predictive-plus-reactive split may help force or torque streams on rigid grippers even without high-resolution tactile images.
  • Uncertainty-aware multi-hypothesis tactile subgoals could address the paper's own limit on longer-horizon contact futures under occlusion or human disturbance.
  • Sensor-layout transfer may reduce to shared image-form tactile goals plus a small residual adapter rather than full policy retraining.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. TouchWorld proposes a hierarchical tactile foundation policy for dexterous manipulation that separates multi-timescale roles: a High-Level Planning Layer (Subtask Planner + Tactile World Model) for executable subtasks and predicted visual-tactile subgoals, a Visuo-Tactile Goal-Conditioned Policy that generates nominal action chunks via flow matching, and a Tactile Residual Transformer that applies high-frequency residual corrections from recent tactile and proprioceptive histories. The system is trained in four stages (planner SFT, human-then-robot world-model adaptation, nominal VLA imitation, residual refinement) and evaluated on six real-robot contact-rich tasks in clean and human-perturbation settings. Table 1 reports 65.0% average clean success and 53.7% under perturbation, outperforming Pi-0.5, FTP-1, and GR00T N1.7 by up to 15.7 and 18.5 percentage points; supporting analyses include stacked ablations (Fig. 5), world-model contact metrics (Table 2), and planner correctness/execution metrics (Table 3).

Significance. If the gains hold under stronger attribution and statistical reporting, the paper would be a useful systems contribution: it shows how to keep VLA-style semantic generalization while giving touch both a predictive subgoal role and a fast residual role, rather than treating tactile as another low-rate token stream. Strengths include a real six-task clean/perturbation benchmark, external baselines (including the prior tactile policy FTP-1), staged training that reuses human EgoTouch data, and separate quantitative checks of the world model and subtask planner. The residual-learning formulation is standard rather than circular, and the hierarchy is a concrete, implementable design for contact-rich dexterous control. The main open question is how much of the lift is the full predictive stack versus residual feedback on this platform.

major comments (3)
  1. §4.3–4.4, Table 1, Fig. 5: The central claim attributes the 15.7/18.5 pp gains over the strongest baseline (FTP-1) to the full predictive-and-reactive hierarchy. Fig. 5 only provides qualitative stacked ablations without numeric per-component success rates, confidence intervals, or matched-compute controls. Because residual correction is already a strong lever in contact-rich control and FTP-1 is a tactile policy, the manuscript needs quantitative ablation numbers (at least average clean vs. perturbation success for each removed component) so readers can test whether residual refinement alone explains most of the lift.
  2. §4.1 and Table 1: Each task reports success over 100 rollouts, but no standard errors, binomial CIs, or significance tests are given. With averages of 65.0% and 53.7%, uncertainty on the order of a few percentage points is material for the claimed margins over FTP-1 (49.3%/35.2%). Adding CIs or bootstrap intervals is load-bearing for the headline comparison.
  3. §5 Limitations and Appendix B: The residual subspace is restricted to 58 of 120 action dimensions, with fixed H=32, W=16, C=4 and an image-form tactile interface tied to the Wuji hand + JQ glove. The foundation-model framing in the abstract and contributions should be tempered, or supported by at least one transfer/sensitivity experiment, so the claim is not overstated relative to a platform-specific multi-rate stack.
minor comments (6)
  1. §2, Eqs. (1)–(4): Define H, W, k, C, and the residual subspace earlier in the main text; several appear only later in §4.1 or Appendix B.
  2. Fig. 5: The stacked bars are hard to read for per-task contribution; a companion numeric table would help.
  3. Table 2: State the pressure threshold τ used for contact/volumetric IoU.
  4. §3.2 / Appendix B: Clarify how EgoTouch pressure maps are converted into the same visual-tactile goal grids as the robot glove (normalization, layout remapping).
  5. Related Work: A short explicit comparison to concurrent reactive tactile / residual VLA lines (e.g., Reactive Diffusion Policy, T-Rex, FTP-1) would sharpen the novelty claim.
  6. Date line says July 10, 2026 while arXiv stamp is 9 Jul 2026; align for consistency.

Circularity Check

0 steps flagged

No circular derivation: claimed gains are empirical real-robot success rates from a staged hierarchical policy, not quantities forced by construction from fitted constants or self-definitional equations.

full rationale

TouchWorld's central claims are (i) a multi-timescale hierarchy (Subtask Planner + Tactile World Model + visuo-tactile goal-conditioned diffusion policy + high-frequency residual refinement) and (ii) measured success rates of 65.0%/53.7% on six real-robot tasks versus external baselines (Pi-0.5, FTP-1, GR00T N1.7). Equations (1)–(4) simply compose the modules; they do not define success in terms of the modules. Residual training (Eq. 9) is ordinary supervised residual learning: the target is the difference between demonstrated high-frequency actions and the frozen nominal VLA chunk, a standard construction that does not force the reported task-success metric. The Tactile World Model is pretrained on external EgoTouch data then fine-tuned and evaluated on held-out contact IoU / temporal accuracy against persistence and nearest-neighbor baselines (Table 2); those metrics are independent of the final success rates. Subtask Planner accuracy (Table 3) is likewise measured against annotated labels. Ablations (Fig. 5) remove components and re-measure success; they do not redefine the metric. Self-citations supply pretraining data or related tactile work and are not load-bearing uniqueness theorems that force the hierarchy or the numbers. No fitted scalar is renamed a “prediction,” no ansatz is smuggled via self-citation into a claimed first-principles result, and the success rates remain externally falsifiable on the robot. The paper is therefore self-contained against its own benchmarks; any debate about attribution of gains to residual versus world-model is a correctness/ablation-strength issue, not circularity.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 2 invented entities

The central claim rests on empirical robotics assumptions and many hand-chosen multi-rate/training choices rather than free physical constants. Invented entities are architectural modules whose value is measured by task success, not by independent physical existence. The ledger is dominated by domain assumptions about contact observability and free scheduling/training parameters that make the hierarchy work on this platform.

free parameters (6)
  • Nominal action horizon H
    Set to 32 in implementation; controls chunk length for the goal-conditioned policy and residual training context.
  • Residual lookahead W and commit interval C
    W=16 and C=4 define how often residual corrections refresh; chosen as deployment hyperparameters, not derived.
  • High-Level Planning Layer update rate (~1 Hz) and world-model refresh rule
    Semantic rate and subgoal reuse-on-stable-phase rule are fixed scheduling choices that shape predictive vs reactive behavior.
  • Residual action subspace dimensionality (58 of 120 dims)
    Only wrist pose blocks and selected hand joints receive residual correction; subspace selection is hand-designed.
  • Training recipe hyperparameters (LoRA ranks, LRs, epochs, residual reg weight 1e-4)
    Stage-wise optimization settings for planner, world model, VLA, and TRT are free engineering choices that affect reported performance.
  • Tactile pressure threshold τ for IoU metrics
    Used when evaluating terminal tactile goals; thresholding choice affects Table 2 contact/volumetric IoU.
axioms (5)
  • domain assumption Vision and language cannot reliably reveal hidden contact states (force, slip, stability), while tactile sensing can.
    Stated in Abstract/Introduction as the motivation for dual predictive/reactive use of touch.
  • domain assumption Separating slow semantic planning, intermediate action chunks, and high-frequency residual correction is better matched to contact-rich control than a monolithic VLA loop.
    Core architectural premise in Introduction and Method; not proved, supported by ablations.
  • ad hoc to paper Image-form tactile representations are a sufficient shared interface between human EgoTouch pretraining and robot glove observations.
    Stage 2 training converts both human pressure maps and robot tactile readings into visual-tactile grids despite layout mismatch.
  • domain assumption Residual targets defined as demonstrated high-frequency actions minus nominal VLA actions train a useful local contact corrector without harming semantic progress.
    Stage 4 loss (Eq. 9) and 'Why residual correction' paragraph.
  • domain assumption Teleoperated demonstrations with manually annotated subtasks are valid supervision for planner, world model, policy, and residual layers.
    Experimental Setup and Training sections; 200 trajectories per task plus 128,866 planner records.
invented entities (2)
  • TouchWorld hierarchical policy (Subtask Planner + Tactile World Model + Visuo-Tactile Goal-Conditioned Policy + Tactile Residual Transformer) no independent evidence
    purpose: Implements predictive tactile subgoals and reactive residual correction at separate timescales for dexterous manipulation.
    The paper's main system object; evaluated only via task success and ablations on the authors' platform.
  • Tactile Residual Transformer (TRT) no independent evidence
    purpose: Predicts residual action windows from nominal lookahead, tactile/proprio histories, and VLA context tokens.
    New residual module name/architecture for the refinement layer; no external validation outside this stack.

pith-pipeline@v1.1.0-grok45 · 21012 in / 3578 out tokens · 34615 ms · 2026-07-10T19:38:30.510079+00:00 · methodology

0 comments
read the original abstract

Dexterous manipulation in everyday environments requires both anticipation and reaction: a robot must predict how contact should evolve while rapidly correcting local errors caused by slip, misalignment, unstable grasping, or force mismatch. Vision and language provide semantic and geometric guidance, but they cannot reliably reveal hidden contact states such as force, slip, and contact stability. Although tactile sensing exposes these physical cues, most existing policies treat touch as a low-frequency observation stream within a monolithic action model, coupling slow task reasoning, action generation, and fast contact feedback in a single loop. We introduce TouchWorld, a predictive-and-reactive tactile foundation model for dexterous manipulation. TouchWorld uses a hierarchical policy that separates vision-language subtask planning, tactile world-model prediction, visuo-tactile goal-conditioned action generation, and high-frequency tactile residual refinement. A High-Level Planning Layer produces executable subtasks and predicts tactile subgoals; a Visuo-Tactile Goal-Conditioned Policy generates nominal action chunks; and a Tactile-Conditioned Refinement Policy performs online residual correction using recent tactile and proprioceptive feedback. By using touch as both a predictive contact reference and a fast feedback signal, TouchWorld preserves the semantic generalization of vision-language-action policies while improving local contact adaptation. Across six long-horizon and contact-rich dexterous manipulation tasks, TouchWorld achieves 65.0% success in the clean setting and 53.7% success under human perturbations, outperforming the strongest baseline by 15.7 and 18.5 percentage points, respectively.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    cs.LG 2026-07 conditional novelty 7.0

    Prediction targets, not fused inputs, decide which physical parameters enter a latent world model; slow ratio-type parameters like drag remain largely unacquired despite high recoverability certificates.

  2. What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    cs.LG 2026-07 conditional novelty 6.0

    Prediction targets, not inputs, decide which physical parameters a latent world model acquires; a certified-recoverable drag parameter stays unlearned under every deterministic prediction objective tested.

  3. ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

    cs.RO 2026-07 conditional novelty 5.0

    An action-conditioned visuo-tactile world model generates synthetic camera-plus-touch rollouts that, mixed with real demonstrations, improve downstream contact-rich manipulation policies.

Reference graph

Works this paper leans on

36 extracted references · 36 canonical work pages · cited by 2 Pith papers · 22 internal anchors

  1. [1]

    Real-Time Execution of Action Chunking Flow Policies

    Kevin Black, Manuel Y. Galliker, and Sergey Levine. Real-time execution of action chunking flow policies, 2025. URLhttps://arxiv.org/abs/2506.07339

  2. [2]

    $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky.π0: A visio...

  3. [3]

    Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

    Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

  4. [4]

    Conla: Contrastive latent action learning from human videos for robotic manipulation, 2026

    Weisheng Dai, Kai Lan, Jianyi Zhou, Bo Zhao, Xiu Su, Junwen Tong, Weili Guan, and Shuo Yang. Conla: Contrastive latent action learning from human videos for robotic manipulation, 2026. URLhttps://arxiv.org/ abs/2602.00557

  5. [5]

    Pressurevision++: Estimating fingertip pressure from diverse rgb images

    Patrick Grady et al. Pressurevision++: Estimating fingertip pressure from diverse rgb images. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  6. [6]

    Tactile-vla: Unlocking vision- language-action model’s physical knowledge for tactile generalization, 2025

    Jialei Huang, Shuo Wang, Fanqi Lin, Yihang Hu, Chuan Wen, and Yang Gao. Tactile-vla: Unlocking vision- language-action model’s physical knowledge for tactile generalization, 2025. URLhttps://arxiv.org/abs/2507. 09160

  7. [7]

    Spatially anchored tactile awareness for robust dexterous manipulation, 2026

    Jialei Huang, Yang Ye, Yuanqing Gong, Xuezhou Zhu, Yang Gao, and Kaifeng Zhang. Spatially anchored tactile awareness for robust dexterous manipulation, 2026. URLhttps://arxiv.org/abs/2510.14647

  8. [8]

    Physical Intelligence, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Kevin Black, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, Jared DiCarlo, Danny Driess, Michael Equi, Adnan Esmail, Yunhao Fang, Chelsea Finn, Catherine Glossop, Thomas Godden, Ivan Goryachev, Lachy Groom, Hunter Hancock, Karol Hausman, Gashon Hussein, Brian Ichter, Szym...

  9. [9]

    Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch...

  10. [10]

    Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, Vedant Choudhary, Foster Collins, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, Maitrayee Dhaka, Jared DiCarlo, Danny Driess, Michael Equi, Adnan Esmail, Yunhao Fang, Chelsea Finn, Catherine...

  11. [11]

    OpenVLA: An Open-Source Vision-Language-Action Model

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024

  12. [12]

    Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

    Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and success, 2025. URLhttps://arxiv.org/abs/2502.19645

  13. [13]

    Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation

    Mike Lambeta, Po-Wei Chou, Stephen Tian, et al. Digit: A novel design for a low-cost compact high-resolution tactile sensor with application to in-hand manipulation. InIEEE Robotics and Automation Letters, 2020

  14. [14]

    HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

    Yi Li, Yuquan Deng, Jesse Zhang, Joel Jang, Marius Memmel, Raymond Yu, Caelan Reed Garrett, Fabio Ramos, Dieter Fox, Anqi Li, Abhishek Gupta, and Ankit Goyal. Hamster: Hierarchical action models for open-world robot manipulation, 2025. URLhttps://arxiv.org/abs/2502.05485

  15. [15]

    Connecting touch and vision via cross-modal prediction

    Yunzhu Li, Jun-Yan Li, Antonio Torralba, et al. Connecting touch and vision via cross-modal prediction. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  16. [16]

    RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

    Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation, 2025. URLhttps://arxiv.org/abs/ 2410.07864

  17. [17]

    T-Rex: Tactile-Reactive Dexterous Manipulation

    Dantong Niu, Zhuoyang Liu, Zekai Wang, Boning Shao, Zhao-Heng Yin, Anirudh Pai, Yuvan Sharma, Stefano Saravalle, Ruijie Zheng, Jing Wang, Ryan Punamiya, Mengda Xu, Yuqi Xie, Yunfan Jiang, Letian Fu, Konstantinos Kallidromitis, Matteo Gioia, Junyi Zhang, Jiaxin Ge, Haiwen Feng, Fabio Galasso, Wei Zhan, David M. Chan, Yutong Bai, Roei Herzig, Jiahui Lei, Li...

  18. [18]

    GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

    NVIDIA, :, Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi "Jim" Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, Guilin Liu, Edith Llontop, Loic Magne, Ajay Mandlekar, Avnish Narayan, Soroush Nasiriany, Scott Reed, Y...

  19. [19]

    Wan: Open and Advanced Large-Scale Video Generative Models

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, T...

  20. [20]

    Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

    Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, Xuhong Huang, Pei Lin, Junyang Lin, Dayiheng Liu, Shuai Bai, Jingren Zhou, Jiazhao Zhang, Haoqi Yuan, Gengze Zhou, Hang Yin, Ye Wang, Yiyang Huang, Zixing Lei, Wujian Peng, Delin Chen, Yingming Zheng, Jingyang Fan, Xianwei Zhuang, Xi...

  21. [21]

    Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation

    Han Xue, Jieji Ren, Wendi Chen, Gu Zhang, Yuan Fang, Guoying Gu, Huazhe Xu, and Cewu Lu. Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation, 2025. URLhttps://arxiv. org/abs/2503.02881

  22. [22]

    Tube Diffusion Policy: Reactive Visual-Tactile Policy Learning for Contact-rich Manipulation

    Teng Xue, Alberto Rigo, Bingjian Huang, Jiayi Shen, Zhengtong Xu, Nick Colonnese, and Amirhossein H Memar. Tube diffusion policy: Reactive visual-tactile policy learning for contact-rich manipulation.arXiv preprint arXiv:2604.23609, 2026

  23. [23]

    Qwen3 Technical Report

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan 14 Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yan...

  24. [24]

    Pressurevision: Estimating hand pressure from a single rgb image

    Patrick Grady Yang, Christian Haase-Schütz, Marcel Leonardi, et al. Pressurevision: Estimating hand pressure from a single rgb image. InEuropean Conference on Computer Vision (ECCV), 2022

  25. [25]

    Ftp-1: A generalist foundation tactile policy across tactile sensors for contact-rich manipulation,

    Chengbo Yuan, Zicheng Zhang, Mingjie Zhou, Wendi Chen, Yi Wang, Zhuoyang Liu, Dantong Niu, Shuo Wang, Hui Zhang, Wenkang Zhang, Yingdong Hu, Yuanqing Gong, Wanli Xing, Chuan Wen, Cewu Lu, Kaifeng Zhang, and Yang Gao. Ftp-1: A generalist foundation tactile policy across tactile sensors for contact-rich manipulation,

  26. [26]

    URLhttps://arxiv.org/abs/2606.13102

  27. [27]

    Vtam: Video-tactile-action models for complex physical interaction beyond vlas, 2026

    Haoran Yuan, Weigang Yi, Zhenyu Zhang, Wendi Chen, Yuchen Mo, Jiashi Yin, Xinzhuo Li, Xiangyu Zeng, Chuan Wen, Cewu Lu, Katherine Driggs-Campbell, and Ismini Lourentzou. Vtam: Video-tactile-action models for complex physical interaction beyond vlas, 2026. URLhttps://arxiv.org/abs/2603.23481

  28. [28]

    Gelsight: High-resolution robot tactile sensors for estimating geometry and force.Sensors, 17(12):2762, 2017

    Wenzhen Yuan, Siyuan Dong, and Edward H Adelson. Gelsight: High-resolution robot tactile sensors for estimating geometry and force.Sensors, 17(12):2762, 2017

  29. [29]

    TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation

    Yujie Zang, Yuhang Zheng, Xian Nie, Yupeng Zheng, Shuai Tian, Songen Gu, Chen Gao, Zining Wang, Shuicheng Yan, and Wenchao Ding. Tacforesight: Force-guided tactile world model for contact-rich manipulation, 2026. URLhttps://arxiv.org/abs/2606.11184

  30. [30]

    TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models

    Zongzheng Zhang, Haobo Xu, Zhuo Yang, Chenghao Yue, Zehao Lin, Huan ang Gao, Ziwei Wang, and Hao Zhao. Ta-vla: Elucidating the design space of torque-aware vision-language-action models, 2025. URLhttps: //arxiv.org/abs/2509.07962

  31. [31]

    Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

    Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware, 2023. URLhttps://arxiv.org/abs/2304.13705

  32. [32]

    Touch begins where vision ends: Generalizable policies for contact-rich manipulation

    Zifan Zhao, Siddhant Haldar, Jinda Cui, Lerrel Pinto, and Raunaq Bhirangi. Touch begins where vision ends: Generalizable policies for contact-rich manipulation, 2025. URLhttps://arxiv.org/abs/2506.13762

  33. [33]

    Egoscale: Scaling dexterous manipulation with diverse egocentric human data, 2026

    Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Castañeda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, and Linxi Fan. Egoscale: Scaling dexterous manipulation with diverse egocentric human data, 2026. URLhttps://arxiv.org/abs/2602.16710

  34. [34]

    Omnivta: Visuo-tactile world modeling for contact-rich robotic manipulation, 2026

    Yuhang Zheng, Songen Gu, Weize Li, Yupeng Zheng, Yujie Zang, Shuai Tian, Xiang Li, Ce Hao, Chen Gao, Si Liu, Haoran Li, Yilun Chen, Shuicheng Yan, and Wenchao Ding. Omnivta: Visuo-tactile world modeling for contact-rich robotic manipulation, 2026. URLhttps://arxiv.org/abs/2603.19201

  35. [35]

    TouchAnything: A Dataset and Framework for Bimanual Tactile Estimation from Egocentric Video

    Jianyi Zhou, Ziteng Gao, Feiyang Hong, Zirui Liu, Guannan Zhang, Weisheng Dai, Ruichen Zhen, Chuqiao Lyu, Haotian Wu, Yinian Mao, Xushi Wang, Yuxiang Jiang, Wenbo Ding, and Shuo Yang. Touchanything: A dataset and framework for bimanual tactile estimation from egocentric video, 2026. URLhttps://arxiv.org/abs/2605.13083

  36. [36]

    Act2goal: From world model to general goal-conditioned policy, 2025

    Pengfei Zhou, Liliang Chen, Shengcong Chen, Di Chen, Wenzhi Zhao, Rongjun Jin, Guanghui Ren, and Jianlan Luo. Act2goal: From world model to general goal-conditioned policy, 2025. URLhttps://arxiv.org/abs/2512.23541. 15 Appendix A Related Work Vision-language-action robot foundation policies.Recent robot foundation policies have rapidly scaled vision- lang...