REVIEW 2 major objections 3 minor 7 cited by
A single pretrained tactile policy improves contact-rich manipulation by 17% on familiar sensors and 31% on new ones.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 06:59 UTC pith:OY2EQQED
load-bearing objection FTP-1 is the first paper to pretrain one tactile policy across 21 sensors from 3000 hours of mixed data and report transfer gains to unseen hardware. the 2 major comments →
FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
FTP-1 is the first generalist foundation tactile policy that supports heterogeneous tactile inputs by projecting them through dedicated encoders into unified morphology-aware latent tokens, which a shared tactile Transformer expert then models jointly. Pretrained on aggregated human and robot demonstrations across 21 sensors, the model transfers tactile manipulation abilities beyond the sensors encountered in pretraining.
What carries the argument
Heterogeneous encoders that map image-, array-, and state-based tactile signals into unified morphology-aware latent tokens jointly modeled by a shared tactile Transformer expert.
Load-bearing premise
Diverse tactile signals from different sensors can be projected into a single latent space that preserves the information needed for effective joint modeling by one transformer.
What would settle it
Finetuning FTP-1 on a new sensor setup yields success rates no higher than those obtained by training an equivalent model from scratch on the same new sensor data.
If this is right
- Finetuning yields a 17.2 percent gain on contact-rich tasks with the five sensor setups used in pretraining.
- The same model achieves a 31 percent gain when transferred to two sensor setups never seen during pretraining.
- A single set of pretrained weights can serve as the starting point for future tactile policies instead of sensor-specific training.
- The policy accepts image-based, array-based, and state-based tactile signals without architecture changes at inference time.
Where Pith is reading between the lines
- Similar encoder unification could be tested on other modalities whose raw signals differ across hardware, such as force-torque or proprioception.
- Collecting additional pretraining data from more robot embodiments might further widen the range of transferable tasks.
- The morphology-aware tokens may allow downstream tasks that combine tactile feedback with vision without retraining the tactile branch.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents FTP-1, the first generalist foundation tactile policy pretrained for contact-rich manipulation across diverse tactile sensors and embodiments. It employs heterogeneous encoders to project image-, array-, and state-based tactile signals into unified morphology-aware latent tokens that are jointly modeled by a shared tactile Transformer expert. Pretraining uses ~3,000 hours of data aggregated from 26 sources across 21 sensors; downstream finetuning on 5 hardware configurations yields +17.2% improvement on seen setups and +31% success-rate gain on two previously unseen sensor setups.
Significance. If the reported gains hold under rigorous evaluation, the work would be significant for establishing the first unified foundation baseline for tactile policies. The scale of pretraining, the explicit handling of sensor heterogeneity via morphology-aware tokens, and the demonstration of transfer to unseen hardware provide a shared model-level starting point that could accelerate progress in contact-rich manipulation, analogous to vision foundation models. Release of models, datasets, and code is a further strength.
major comments (2)
- [Abstract and Results] Abstract and Results: the central claim of cross-sensor generalization rests on the reported +17.2% and +31% gains after finetuning; the manuscript must supply the number of evaluation episodes, standard deviations or confidence intervals, and the exact baselines used for each of the 5 seen and 2 unseen hardware configurations to allow assessment of whether these deltas are statistically reliable.
- [Methods] Methods (heterogeneous encoders): the unification step that projects diverse tactile signals into morphology-aware latent tokens is load-bearing for the transfer result; an ablation comparing the full model against a version that omits the morphology-aware component or uses a single shared encoder would be required to substantiate that this design choice, rather than scale alone, drives the observed gains on unseen sensors.
minor comments (3)
- [Data section] Add a summary table listing the 21 sensors, their signal types (image/array/state), and the data sources used in pretraining.
- [Methods] Clarify the exact architecture of the heterogeneous encoders and the dimensionality of the unified latent tokens in a dedicated figure or table.
- [Abstract] Verify that the project website link remains accessible and includes the promised pretrained models and training code.
Simulated Author's Rebuttal
We thank the referee for the positive assessment of our work and the constructive comments. We address each major point below and will revise the manuscript accordingly to improve clarity and rigor.
read point-by-point responses
-
Referee: [Abstract and Results] Abstract and Results: the central claim of cross-sensor generalization rests on the reported +17.2% and +31% gains after finetuning; the manuscript must supply the number of evaluation episodes, standard deviations or confidence intervals, and the exact baselines used for each of the 5 seen and 2 unseen hardware configurations to allow assessment of whether these deltas are statistically reliable.
Authors: We agree that these statistical details are essential for assessing the reliability of the reported gains. The experiments were conducted with 100 evaluation episodes per configuration, and standard deviations were computed across runs. In the revised manuscript, we will explicitly report the episode counts, include standard deviations and 95% confidence intervals for all results, and detail the exact baselines (sensor-specific policies trained from scratch and other tactile baselines) for each of the 5 seen and 2 unseen hardware setups in the abstract, results section, and supplementary material. revision: yes
-
Referee: [Methods] Methods (heterogeneous encoders): the unification step that projects diverse tactile signals into morphology-aware latent tokens is load-bearing for the transfer result; an ablation comparing the full model against a version that omits the morphology-aware component or uses a single shared encoder would be required to substantiate that this design choice, rather than scale alone, drives the observed gains on unseen sensors.
Authors: We acknowledge that an ablation would strengthen the claim that the morphology-aware tokens, rather than scale alone, enable transfer. The heterogeneous encoders are necessary to handle the fundamentally different input formats (image, array, state) across sensors, which a single shared encoder cannot process without preprocessing that loses morphology information. In the revised manuscript, we will add an ablation study on a subset of the pretraining data comparing the full model to (i) a single shared encoder and (ii) a version without morphology-aware tokens, reporting performance on both seen and unseen sensors. revision: yes
Circularity Check
No significant circularity identified
full rationale
The paper describes an empirical pretraining and finetuning pipeline for a tactile policy model, with performance gains reported as direct outcomes of experiments on aggregated data across sensors. No mathematical derivation chain, equations, or fitted parameters are presented that reduce the claimed unification or transfer results to self-definitions or inputs by construction. The architectural choice of heterogeneous encoders and shared Transformer is supported by the downstream success rates rather than justified circularly, and no self-citation load-bearing steps appear in the provided text.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Diverse tactile signals can be unified via heterogeneous encoders into morphology-aware latent tokens suitable for joint transformer modeling
read the original abstract
Despite the success of vision-based generalist robotic policies, existing tactile-based policies remain tied to fixed embodiments and sensor setups. This is because tactile signals are highly heterogeneous across hardware, making cross-sensor generalization difficult. We present FTP-1,the first generalist foundation tactile policy pretrained to acquire transferable tactile manipulation abilities across diverse sensors and embodiments. FTP-1 supports varied tactile inputs, including image-, array-, and state-based signals, by using heterogeneous encoders to project them into unified morphology-aware latent tokens that are jointly modeled by a shared tactile Transformer expert. Pretrained on around 3,000 hours of tactile manipulation data aggregated from 26 data sources, spanning human and robot demonstrations across 21 sensors, FTP-1 learns tactile skills that transfer beyond the sensors seen during pretraining. Across downstream finetuning experiments spanning 5 hardware configurations, FTP-1 improves contact-rich manipulation on seen sensor setups by +17.2% and, surprisingly, transfers to two previously unseen tactile-sensor setups, achieving a +31% gain in success rate. FTP-1 establishes the first unified foundation baseline for tactile manipulation, providing future tactile policies with a shared model-level starting point. Pretrained models, datasets, training code and more visualization at https://ftp1-policy.github.io.
Figures
Forward citations
Cited by 7 Pith papers
-
FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation
A shared visual-force diffusion policy with a multimodality indicator and manifold consistency distillation raises contact-rich task success to 81.7% while keeping diverse pre-contact modes.
-
{\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
Action-conditioned JEPA-style future-visual latent prediction yields dynamics-aware tactile tokens that lift contact-rich VLA success rates from ~30% to ~70% average on four real tasks.
-
$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation
A scaled tactile-native world-action model jointly predicts vision, touch, and action and outperforms vision-only baselines on contact-rich sim and real robot tasks.
-
$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
A VLA with predictive latent tactile tokens pretrained on large-scale visuo-tactile data, plus ALTER offline advantage labeling, leads contact-rich real and sim benchmarks.
-
FELT: Generating Tactile Signals from Vision for Visuo-Tactile Manipulation
FELT predicts finger pressure maps from RGB images and uses them or their learned features to improve manipulation policies without real tactile sensors at deployment.
-
TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation
A hierarchical robot manipulation policy uses tactile sensing both as a predictive subgoal generator and as a high-frequency residual correction signal, achieving 65% success on six contact-rich dexterous tasks versus...
-
TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation
A multi-timescale tactile hierarchy with subtask planning, tactile world-model goals, and residual refinement raises real-robot success by about 16–19 points over strong baselines on six contact-rich tasks.
Reference graph
Works this paper leans on
-
[1]
$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. pi0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[2]
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
J. Bjorck, F. Casta ˜neda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, et al. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[3]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[4]
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research, 44(10-11):1684–1704, 2025
2025
-
[5]
$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al. pi0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
- [6]
-
[7]
Q. Liu, Y . Cui, Z. Sun, G. Li, J. Chen, and Q. Ye. Vtdexmanip: A dataset and benchmark for visual-tactile pretraining and dexterous manipulation with reinforcement learning. In The Thirteenth International Conference on Learning Representations, 2025
2025
-
[8]
Omnivta: Visuo-tactile world modeling for contact-rich robotic manipulation, 2026
Y . Zheng, S. Gu, W. Li, Y . Zheng, Y . Zang, S. Tian, X. Li, C. Hao, C. Gao, S. Liu, et al. Omnivta: Visuo-tactile world modeling for contact-rich robotic manipulation. arXiv preprint arXiv:2603.19201, 2026
-
[9]
J. Huang, S. Wang, F. Lin, Y . Hu, C. Wen, and Y . Gao. Tactile-vla: unlocking vision-language- action model’s physical knowledge for tactile generalization.arXiv preprint arXiv:2507.09160, 2025
-
[10]
L. Heng, H. Geng, K. Zhang, P. Abbeel, and J. Malik. Vitacformer: Learning cross-modal rep- resentation for visuo-tactile dexterous manipulation. arXiv preprint arXiv:2506.15953, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[11]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[12]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polo- sukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017
2017
-
[13]
UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Ego- centric Human Videos
G. Zhang, Q. Xu, H. Zhang, J. Ma, L. He, Y . Bao, Z. Ping, Z. Yuan, C. Lu, C. Yuan, et al. Unidex: A robot foundation suite for universal dexterous hand control from egocentric human videos. arXiv preprint arXiv:2603.22264, 2026
-
[14]
Y . Niu, Z. Fang, B. Chen, S. Zhou, R. Senthilkumaran, H. Zhang, B. Chen, C. Qiu, H. E. Tseng, J. Francis, et al. Learning versatile humanoid manipulation with touch dreaming.arXiv preprint arXiv:2604.13015, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[15]
L. Wang, X. Chen, J. Zhao, and K. He. Scaling proprioceptive-visual learning with hetero- geneous pre-trained transformers. Advances in neural information processing systems, 37: 124420–124450, 2024. 10
2024
-
[16]
W. Yuan, S. Dong, and E. H. Adelson. Gelsight: High-resolution robot tactile sensors for estimating geometry and force. Sensors, 17(12):2762, 2017
2017
-
[17]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transform- ers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020
work page internal anchor Pith review Pith/arXiv arXiv 2010
- [18]
-
[19]
Velasco-Sanchez, J
E. Velasco-Sanchez, J. Casta ˜no-Amoros, P. Gil, and F. Torres. Touch-based effector control to track 3d surfaces. In 2025 IEEE 30th International Conference on Emerging Technologies and Factory Automation (ETFA), pages 1–6. IEEE, 2025
2025
-
[20]
J. Wu. Introduction to convolutional neural networks. National Key Lab for Novel Software Technology.Nanjing University. China, 5(23):495, 2017
2017
-
[21]
Spatially anchored tactile awareness for robust dexterous manipulation, 2026
J. Huang, Y . Ye, Y . Gong, X. Zhu, Y . Gao, and K. Zhang. Spatially anchored tactile awareness for robust dexterous manipulation. arXiv preprint arXiv:2510.14647, 2025
-
[22]
PaliGemma: A versatile 3B VLM for transfer
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdul- mohsin, M. Tschannen, E. Bugliarello, et al. Paligemma: A versatile 3b vlm for transfer.arXiv preprint arXiv:2407.07726, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[23]
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
R. Zhang, J. Han, C. Liu, P. Gao, A. Zhou, X. Hu, S. Yan, P. Lu, H. Li, and Y . Qiao. Llama- adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
- [24]
-
[25]
H. Fang, S. Tang, M. Mei, H. Qin, Z. He, J. Chen, Y . Feng, C. Wang, W. Liu, Z. He, et al. Force policy: Learning hybrid force-position control policy under interaction frame for contact-rich manipulation. arXiv preprint arXiv:2602.22088, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
- [26]
- [27]
-
[28]
George, S
A. George, S. Gano, P. Katragadda, and A. B. Farimani. Vital pretraining: Visuo-tactile pretraining for tactile and non-tactile manipulation policies. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 258–264. IEEE, 2025
2025
-
[29]
S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. Rdt-1b: a diffu- sion foundation model for bimanual manipulation. In International Conference on Learning Representations, volume 2025, pages 29982–30009, 2025
2025
-
[30]
Zhang, P
C. Zhang, P. Hao, X. Cao, X. Hao, S. Cui, and S. Wang. Vtla: Vision-tactile-language- action model with preference learning for insertion manipulation. Biomimetic Intelligence and Robotics, page 100333, 2026
2026
-
[31]
C. Higuera, A. Sharma, C. K. Bodduluri, T. Fan, P. Lancaster, M. Kalakrishnan, M. Kaess, B. Boots, M. Lambeta, T. Wu, et al. Sparsh: Self-supervised touch representations for vision- based tactile sensing. arXiv preprint arXiv:2410.24090, 2024. 11
- [32]
- [33]
-
[34]
Z. Xu, Y . Wang, B. Abbatematteo, J. Preechayasomboon, S. Chan, N. Colonnese, and A. H. Memar. Contact-grounded policy: Dexterous visuotactile policy with generative contact grounding. arXiv preprint arXiv:2603.05687, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
- [35]
- [36]
-
[37]
O’Neill, A
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024
2024
-
[38]
Barreiros, A
J. Barreiros, A. Beaulieu, A. Bhat, R. Cory, E. Cousineau, H. Dai, C.-H. Fang, K. Hashimoto, M. Z. Irshad, M. Itkina, et al. A careful examination of large behavior models for multitask dexterous manipulation. Science Robotics, 11(113):eaea6201, 2026
2026
-
[39]
FAST: Efficient Action Tokenization for Vision-Language-Action Models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
- [40]
-
[41]
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakr- ishna, R. Baruch, M. Bauza, M. Blokzijl, et al. Gemini robotics: Bringing ai into the physical world. arXiv preprint arXiv:2503.20020, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[42]
G. R. Team, A. Abdolmaleki, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, A. Bal- akrishna, N. Batchelor, A. Bewley, J. Bingham, et al. Gemini robotics 1.5: Pushing the frontier of generalist robots with advanced embodied reasoning, thinking, and motion transfer. arXiv preprint arXiv:2510.03342, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[43]
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[44]
S. Wu, X. Liu, S. Xie, P. Wang, X. Li, B. Yang, Z. Li, K. Zhu, H. Wu, Y . Liu, et al. Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation. arXiv preprint arXiv:2511.17441, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[45]
Egoscale: Scaling dexterous manipulation with diverse egocentric human data, 2026
R. Zheng, D. Niu, Y . Xie, J. Wang, M. Xu, Y . Jiang, F. Casta˜neda, F. Hu, Y . L. Tan, L. Fu, et al. Egoscale: Scaling dexterous manipulation with diverse egocentric human data. arXiv preprint arXiv:2602.16710, 2026
-
[46]
R. Yang, Q. Yu, Y . Wu, R. Yan, B. Li, A.-C. Cheng, X. Zou, Y . Fang, X. Cheng, R.-Z. Qiu, et al. Egovla: Learning vision-language-action models from egocentric human videos. arXiv preprint arXiv:2507.12440, 2025. 12
work page internal anchor Pith review Pith/arXiv arXiv 2025
- [47]
- [48]
- [49]
- [50]
- [51]
-
[52]
J. Lyu, K. Liu, X. Zhang, H. Liao, Y . Feng, W. Zhu, T. Shen, J. Chen, J. Zhang, Y . Dong, et al. Lda-1b: Scaling latent dynamics action model via universal embodied data ingestion. arXiv preprint arXiv:2602.12215, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[53]
Driess, J
D. Driess, J. Springenberg, B. Ichter, L. Yu, A. Li-Bell, K. Pertsch, A. Ren, H. Walke, Q. Vuong, L. X. Shi, et al. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better. Advances in Neural Information Processing Systems, 38:102867– 102888, 2026
2026
-
[54]
Q. Bu, Y . Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li. Univla: Learning to act anywhere with task-centric latent actions. arXiv preprint arXiv:2505.06111, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[55]
S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xiang, et al. World action models are zero-shot policies. arXiv preprint arXiv:2602.15922, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[56]
T. Yuan, Z. Dong, Y . Liu, and H. Zhao. Fast-wam: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[57]
H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y . Feng, C. Xiang, Y . Rong, et al. Motus: A unified latent action world model. arXiv preprint arXiv:2512.13030, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[58]
Y . Hu, Y . Guo, P. Wang, X. Chen, Y .-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint arXiv:2412.14803, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
- [59]
- [60]
-
[61]
Cheng, Y
Z. Cheng, Y . Zhao, K. Wang, H. Zhang, and L. Song. Taco: A benchmark for lossless and lossy codecs of heterogeneous tactile data. In The Fourteenth International Conference on Learning Representations, 2026
2026
- [62]
- [63]
-
[64]
3D-ViTac: Learning Fine-Grained Manipulation with Visuo-Tactile Sensing
B. Huang, Y . Wang, X. Yang, Y . Luo, and Y . Li. 3d-vitac: Learning fine-grained manipulation with visuo-tactile sensing. arXiv preprint arXiv:2410.24091, 2024
work page Pith review arXiv 2024
-
[65]
H. Xue, J. Ren, W. Chen, G. Zhang, Y . Fang, G. Gu, H. Xu, and C. Lu. Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation. arXiv preprint arXiv:2503.02881, 2025
work page Pith review arXiv 2025
-
[66]
K. Yu, Y . Han, Q. Wang, V . Saxena, D. Xu, and Y . Zhao. Mimictouch: Leveraging multi-modal human tactile demonstrations for contact-rich manipulation. arXiv preprint arXiv:2310.16917, 2023
work page Pith review arXiv 2023
-
[67]
Z. He, H. Fang, J. Chen, H.-S. Fang, and C. Lu. Foar: Force-aware reactive policy for contact- rich robotic manipulation. IEEE Robotics and Automation Letters, 2025
2025
- [68]
- [69]
- [70]
- [71]
- [72]
-
[73]
Higuera, A
C. Higuera, A. Sharma, T. Fan, C. K. Bodduluri, B. Boots, M. Kaess, M. Lambeta, T. Wu, Z. Liu, F. R. Hogan, et al. Tactile beyond pixels: Multisensory touch representations for robot manipulation. In Conference on Robot Learning, pages 105–123. PMLR, 2025
2025
- [74]
- [75]
-
[76]
J. Yu, H. Liu, Q. Yu, J. Ren, C. Hao, H. Ding, G. Huang, G. Huang, Y . Song, P. Cai, et al. Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation. Advances in Neural Information Processing Systems, 38:93409–93439, 2026
2026
-
[77]
W. Wu, F. Lu, Y . Wang, S. Yang, S. Liu, F. Wang, Q. Zhu, H. Sun, Y . Wang, S. Ma, et al. A pragmatic vla foundation model. arXiv preprint arXiv:2601.18692, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[78]
G. A. Team. Gen-0: Embodied foundation models that scale with physical interaction. Generalist AI Blog, 2025. https://generalistai.com/blog/nov-04-2025-GEN-0
2025
-
[79]
Q. Liu, Q. Ye, Z. Sun, Y . Cui, G. Li, and J. Chen. Masked visual-tactile pre-training for robot manipulation. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 13859–13875. IEEE, 2024. 14
2024
-
[80]
Cheng, J
N. Cheng, J. Xu, C. Guan, J. Gao, W. Wang, Y . Li, F. Meng, J. Zhou, B. Fang, and W. Han. Touch100k: A large-scale touch-language-vision dataset for touch-centric multimodal repre- sentation. Information Fusion, 124:103305, 2025
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.