REVIEW 4 major objections 3 minor 71 references
FELT predicts per-finger pressure tactile images from a single RGB fisheye frame, and the paper shows that policies using these synthetic tactile signals or the model's latent features outperform vision-only policies on four contact-rich ma
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 09:39 UTC pith:I2J3GYLL
load-bearing objection Solid vision-to-tactile generation paper with honest ablations; occlusion and reproducibility gaps are addressable, not fatal. the 4 major comments →
FELT: Generating Tactile Signals from Vision for Visuo-Tactile Manipulation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's claim is that per-finger pressure tactile images can be predicted from a single RGB frame with enough fidelity to recover much of the benefit of real tactile sensing. FELT implements this with a single feed-forward pass: a frozen pretrained vision transformer extracts features from the fisheye image; two attention-based decoder branches, one per finger panel, each hold a grid of learnable queries matching the physical 12×32 sensor layout; they attend to the visual features and exchange left/right information through gated cross-panel attention; and a shared convolutional readout predicts per-cell contact probability and pressure intensity for the two panels. The generator is cons
What carries the argument
The mechanism that carries the argument is the panel-aware query decoder. The model gives each of the two finger panels its own learnable 12×32 grid of query embeddings, sized to match the physical pressure-sensing pads, and runs each grid through an independent transformer decoder that cross-attends to frozen visual features. A gated cross-panel exchange lets the left and right branches share information while preserving their separate identities, and a panel-aware convolutional readout with learned side embeddings turns the decoded queries into per-cell contact probabilities and pressure intensities. This topology is load-bearing: ablations that remove the dual-branch decoder, the panel ex
Load-bearing premise
The load-bearing premise, stated in the paper's Section 3, is that a single fisheye RGB frame shows the gripper, object, and contact region clearly enough to determine the per-finger pressure distribution; the paper itself acknowledges in Section 7 that heavy occlusion breaks this, and the supplementary material adds that predicting each frame independently can produce temporally unstable tactile signals.
What would settle it
Collect a held-out set of paired fisheye-RGB and real tactile frames in which the gripper or object occludes the contact region — for example during cup nesting or peg insertion — and measure FELT's panel-level contact accuracy and left-right pressure balance on those frames against the real sensor; if accuracy collapses to near chance on occluded frames while a policy trained on FELT features still fails to beat vision-only, the claim that RGB context suffices for tactile synthesis is falsified.
If this is right
- Vision-only demonstration sets can be converted into visuo-tactile ones offline, and in the reported tasks this improves downstream success over vision-only policies.
- A tactile-informed policy can run with only a camera at deployment: generated tactile images arrive in about 20 ms per frame, fast enough for 10 Hz closed-loop control.
- The latent-feature route removes the sensor at both training and deployment, and in these experiments it matches or exceeds real-tactile performance on some tasks with only 60 demonstrations per task.
- Tactile information helps specifically at contact-critical stages — reorientation, insertion, nesting, and wiping — rather than at grasping, which is already visually reliable.
- Respecting the physical sensor topology is part of the method: a single shared decoder or a decoder without left-right information flow produces worse pressure images and lowers insertion and nesting success.
Where Pith is reading between the lines
- The paper's own limitation notes imply a direct extension: heavy occlusion breaks the RGB-to-pressure mapping, so a temporal model or a second camera view could recover the missing contact information; this is the most natural next experiment.
- The frame-independent prediction, which the supplementary material flags, suggests generated tactile sequences can flicker during sustained contact; conditioning on a short history of generated frames is a testable fix that should improve wiping and insertion rollouts.
- Because a generic frozen visual feature does not substitute for the contact-trained latent features in the paper's comparisons, the value lies in the tactile-specific training, not merely in adding a second vision backbone; one could verify by training the same policy with raw visual tokens instead of FELT features.
- A practical extension is to ask how little paired visuo-tactile data is enough: the generator currently needs a large paired corpus, and measuring the augmentation benefit as that corpus shrinks would bound the method's scalability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FELT, a single-frame cross-modal generator that predicts per-finger 12x32 pressure tactile images from a wrist-mounted fisheye RGB camera. The architecture uses a frozen DINOv2 encoder, separate query decoders for the left and right tactile panels, a gated cross-panel exchange, and a convolutional readout head. FELT is trained on 2,700 paired UMI-style visuo-tactile demonstrations and evaluated on an out-of-distribution xArm test set, then deployed in four real-robot contact-rich tasks (tube insertion, cup nesting, eraser wiping, triangle peg insertion). The authors compare vision-only, real-tactile, FELT-generated-tactile, and FELT latent-feature policies. They report that generated tactile images and latent features improve final task success over vision-only baselines, with latent features not requiring real tactile data during policy training or deployment. The paper also includes ablations of the generator architecture and deployment-time ablations.
Significance. If the central claim holds, the paper makes a significant practical contribution: it offers a way to convert large vision-only demonstrations into visuo-tactile ones and to deploy tactile-informed policies without physical tactile sensors at test time. Strengths of the work include real-robot evaluation over four tasks, an out-of-distribution generator test set collected on a different embodiment, careful sensor synchronization, open data/code commitments, a threshold-sensitivity sweep in S4.4, and a zero-tactile control in S7.3 showing that the policy does not simply ignore the tactile channel. The main unresolved issue is whether the single-frame generator produces physically correct tactile signals under the heavy occlusion that characterizes the decisive stages of the evaluated tasks; this directly bears on the paper's core assumption and requires additional analysis before the claims can be accepted.
major comments (4)
- [Section 3 / S6.3 / Section 7] The method assumes that a single fisheye RGB frame contains the gripper and contact region with sufficient visual context to condition tactile synthesis. Yet the evaluation tasks are heavily occluded at the contact-critical stages: S6.3 states that during tube insertion 'the rack is heavily occluded by the tube and gripper,' during cup nesting the target cup is 'heavily occluded by the carried cup and gripper,' and during triangle peg insertion the hole is 'partially occluded by the peg and gripper.' Since FELT is single-frame and has no temporal model (S9 acknowledges frame-to-frame instability), occluded frames provide no direct visual evidence of contact. The paper does not quantify how often or how severely the contact region is occluded in the task episodes, nor does it report tactile prediction accuracy conditioned on occlusion. The Section 7 limitation acknowledges this, but the c
- [Table 2 / S7.1] The central empirical claim rests on success-rate differences from 20 trials per cell, where the paper itself notes binomial standard errors of about 7-11 percentage points. Several headline final-task differences versus Vision Only are 10 points (Tube Insertion 50 vs 40, Cup Nesting 35 vs 25, Eraser Wiping 75 vs 65), which is within one standard error. The supplement computes Wilson intervals but does not use them to support the claim that 'both generated tactile images and latent tactile features improve policy success over vision-only baselines.' Please provide confidence intervals for the relevant pairwise differences, or a pooled analysis across tasks, so that the abstract's language matches the statistical strength of the evidence.
- [Section 4.2 / S6.1] There is a direct inconsistency in the definition of the FELT Features variant. Section 4.2 says the per-panel feature maps z_L, z_R in R^{12x32x256} are concatenated and provided to the policy, with the implication that spatial contact structure is preserved. S6.1 config (4) instead states that these maps are 'spatially mean-pooled, concatenated to 512 dims, and linearly projected to a 768-dim token.' Mean pooling discards the spatial arrangement, so the main text overstates what the implemented latent feature actually does. Please align the main text with the supplement and, ideally, report whether using full feature maps changes downstream performance.
- [Section 5.3 / Table 1 / S3.4] The 'FELT w/o dual-branch decoder' ablation replaces the per-finger branches with a single 24x32 decoder, but it also removes panel exchange, the convolutional readout, and side embeddings, as described in S3.4. The main text attributes the resulting performance drop to 'respecting the physical sensor topology,' but this ablation is confounded: it changes several architectural components simultaneously. Please isolate the dual-branch decoder while keeping the panel exchange and readout head, or soften the attribution accordingly.
minor comments (3)
- [Section 3 / S2.1 / S6.1] The tactile representation is presented inconsistently: the main text uses I_tac in R^{H'xW'x2} with H'=12, W'=32, S2.1 stores a 12x64 image, and S6.1 stacks the panels into a 24x32 grid for the policy. Please state the canonical representation once and use it throughout.
- [Section 7 / S9] The temporal-stability limitation is only in the supplement. Because frame-to-frame inconsistency is a known failure mode for single-frame generation and is directly relevant to downstream closed-loop control, it should be stated in the main limitations section.
- [Section 5.3] The downstream policy encoder is pretrained on the same Zhu et al. [13] dataset used to train the generator. This is not circular, but the main text should state clearly that the FELT Features variant benefits from a tactile-aware pretraining prior shared with the real-tactile baseline, and discuss whether this affects the comparison.
Circularity Check
No significant circularity: FELT's predictions are empirically trained on real paired data and evaluated on independently collected data; no step reduces to its inputs by construction.
full rationale
The paper's central claim is empirical rather than derivational. G_phi is trained by the supervised composite loss in Eq. 1: L(phi) = E[L_contact(c_hat, c) + L_intensity(p_hat, p)], where ground-truth c and p come from real FlexiTac readings. It is then evaluated on a separately collected xArm test set ('We use an 80/20 train/evaluation split... we additionally collect an xArm test set of 30 GELLO demonstrations with ~72K visuo-tactile frames, using objects and task scenarios not seen during generator training'), so its tactile predictions are not fitted on the evaluation frames. The downstream claims are likewise held out: 'For downstream policy learning, we collect 60 GELLO demonstrations per task on the xArm, totaling 240 demonstrations... collected independently from both the generator training data and the xArm generator test set.' The generator is trained on the overlapping-authors dataset Zhu et al. [13], and the policy encoder is also pretrained on it, but this is citation of empirical paired data, not an unverified theorem; the authors also include control configurations (Vision + Zero Tactile, Vision + DINOv2 Features, Vision + FELT Tactile Train) that isolate whether the tactile pathway or a generic visual quantity drives the benefit. Section 7 and S9 explicitly concede the real limitations of occlusion and the continued need for paired visuo-tactile data for generator pretraining; these are correctness risks, not circular reductions. I find no prediction that is equivalent by construction to a fitted parameter, no self-citation chain that forces the central result, and no renamed known result. Score 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- Generator loss weights and contact binarization threshold =
alpha=0.75, gamma=2.0, dice=0.5, logit_reg=0.01, intensity_w=4.0/2.0/0.25, bg=0.05, shape=0.05, eps_c=0.01
- Frame Accuracy evaluation thresholds =
n_min=12 active cells; tau=0.20 intensity
- Policy and evaluation protocol hyperparameters =
60 epochs; lr 3e-4/3e-5/2e-5; chunk 16; execute 6; 20 trials/task
- Sensor synchronization latencies =
GoPro latency 0.315 s; tactile offset 20 ms; max gap 100 ms
axioms (4)
- domain assumption A single RGB frame from the wrist fisheye camera contains sufficient visible context of the gripper, object, and contact region to determine the tactile pressure distribution.
- domain assumption Contact state is determined by the current frame alone; no temporal information is required.
- domain assumption The paired visuo-tactile dataset of Zhu et al. is large and representative enough to train a generator that transfers to xArm GELLO demonstrations.
- domain assumption Pressure contact maps from FlexiTac pads are the relevant tactile signal for downstream policies and are interchangeable with generated maps.
read the original abstract
The sense of touch is central to manipulation, especially when vision is occluded or ambiguous. Although combining vision and touch improves manipulation, learning robust visuo-tactile policies requires substantial tactile data. Such data remains scarcer than visual data, because tactile sensors are fragile, specialized, and hard to standardize. To address this, we present Feature-Extracted Latent Tactile (FELT), a learning-based framework that synthesizes per-finger pressure tactile images from RGB observations, reducing the need for tactile-equipped data collection. FELT uses a large frozen visual encoder and a lightweight query decoder to predict tactile signals in a single feed-forward pass. To respect the physical topology of dual-finger tactile sensors, FELT decodes the left and right tactile sensor panels through separate branches, capturing the asymmetric contact patterns during interactions such as wiping, insertion, and in-hand rotation. At inference time, FELT only requires RGB data, allowing us to augment existing vision-only data with tactile observations, either as generated tactile images or as latent tactile features. Experiments on four contact-rich manipulation tasks demonstrate that both generated tactile images and latent tactile features improve policy success over vision-only baselines, with latent feature requiring no real tactile sensor during policy training or deployment. Supplementary material is available on our anonymous website: https://felt-tactile.github.io/.
Figures
Reference graph
Works this paper leans on
-
[1]
K. Kawaharazuka, T. Matsushima, A. Gambardella, J. Guo, C. Paxton, and A. Zeng. Real-World Robot Applications of Foundation Models: A Review.arXiv preprint arXiv:2402.05741, 2024
Pith/arXiv arXiv 2024
-
[2]
R. Firoozi, J. Tucker, S. Tian, A. Majumdar, J. Sun, W. Liu, Y . Zhu, S. Song, A. Kapoor, K. Hausman, B. Ichter, D. Driess, J. Wu, C. Lu, and M. Schwager. Foundation Models in Robotics: Applications, Challenges, and the Future.arXiv preprint arXiv:2312.07843, 2023
Pith/arXiv arXiv 2023
-
[3]
Y . Hu, Q. Xie, V . Jain, J. Francis, J. Patrikar, N. Keetha, S. Kim, Y . Xie, T. Zhang, S. Zhao, Y . Q. Chong, C. Wang, K. Sycara, M. Johnson-Roberson, D. Batra, X. Wang, S. Scherer, Z. Kira, F. Xia, and Y . Bisk. Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis.arXiv preprint arXiv:2312.08782, 2023
Pith/arXiv arXiv 2023
-
[4]
O. X.-E. Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Her- zog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakrishna, A. Wah...
2024
-
[5]
Huang, Y
B. Huang, Y . Wang, X. Yang, Y . Luo, and Y . Li. 3D-ViTac: Learning Fine-Grained Manipula- tion with Visuo-Tactile Sensing. InConference on Robot Learning (CoRL), 2024
2024
-
[6]
Guzey, Y
I. Guzey, Y . Dai, B. Evans, S. Chintala, and L. Pinto. See to Touch: Learning Tactile Dexterity through Visual Incentives. InIEEE International Conference on Robotics and Automation (ICRA), 2024
2024
-
[7]
Y . Yuan, H. Che, Y . Qin, B. Huang, Z.-H. Yin, K.-W. Lee, Y . Wu, S.-C. Lim, and X. Wang. Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensing. InIEEE International Conference on Robotics and Automation (ICRA), 2024
2024
-
[8]
A. Zorin, Z. Si, M. Park, J. Park, A. Buynitsky, S. Bhadang, T. Park, S. J. Yoon, Y .-L. Park, O. Kroemer, Z. Temel, M. T. Tolley, S. Yi, and X. Wang. TacO: Benchmarking Tactile Sensors for Object Manipulation.arXiv preprint arXiv:2605.21976, 2026
Pith/arXiv arXiv 2026
-
[9]
Z. Chen, Z. Mandi, H. Bharadhwaj, M. Sharma, S. Song, A. Gupta, and V . Kumar. Semanti- cally Controllable Augmentations for Generalizable Robot Learning. InInternational Journal of Robotics Research (IJRR), 2024
2024
-
[10]
T. Yu, T. Xiao, A. Stone, J. Tompson, A. Brohan, S. Wang, J. Singh, C. Tan, D. M, J. Per- alta, B. Ichter, K. Hausman, and F. Xia. Scaling Robot Learning with Semantically Imagined Experience. InRobotics: Science and Systems (RSS), 2023
2023
-
[11]
S. Tian, B. Wulfe, K. Sargent, K. Liu, S. Zakharov, V . Guizilini, and J. Wu. View-Invariant Pol- icy Learning via Zero-Shot Novel View Synthesis. InConference on Robot Learning (CoRL), 2024
2024
-
[12]
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song. Uni- versal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. In Robotics: Science and Systems (RSS), 2024
2024
-
[13]
X. Zhu, B. Huang, and Y . Li. Touch in the wild: Learning fine-grained manipulation with a portable visuo-tactile gripper. InNeural Information Processing Systems (NeurIPS), 2025
2025
-
[14]
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. InInternational Journal of Robotics Research (IJRR), 2025
2025
-
[15]
Bhirangi, T
R. Bhirangi, T. Hellebrekers, C. Majidi, and A. Gupta. ReSkin: Versatile, Replaceable, Lasting Tactile Skins. InConference on Robot Learning (CoRL), 2021
2021
-
[16]
Bhirangi, V
R. Bhirangi, V . Pattabiraman, E. Erciyes, Y . Cao, T. Hellebrekers, and L. Pinto. Anyskin: Plug- and-Play Skin Sensing for Robotic Touch. InIEEE International Conference on Robotics and Automation (ICRA), 2025
2025
-
[17]
V . Pattabiraman, Z. Huang, D. Panozzo, D. Zorin, L. Pinto, and R. Bhirangi. eFlesh: Highly customizable Magnetic Touch Sensing using Cut-Cell Microstructures.arXiv preprint arXiv:2506.09994, 2025
Pith/arXiv arXiv 2025
-
[18]
Donlon, S
E. Donlon, S. Dong, M. Liu, J. Li, E. Adelson, and A. Rodriguez. GelSlim: A High-Resolution, Compact, Robust, and Calibrated Tactile-sensing Finger. InIEEE/RSJ International Confer- ence on Intelligent Robots and Systems (IROS), 2018
2018
-
[19]
A. Sipos, W. van den Bogert, and N. Fazeli. GelSlim 4.0: Focusing on Touch and Repro- ducibility.arXiv preprint arXiv:2409.19770, 2024. 10
Pith/arXiv arXiv 2024
-
[20]
Patel, R
R. Patel, R. Ouyang, B. Romero, and E. Adelson. Digger Finger: GelSight Tactile Sensor for Object Identification Inside Granular Media. InInternational Symposium on Experimental Robotics (ISER), 2021
2021
-
[21]
M. Lambeta, T. Wu, A. Sengul, V . R. Most, N. Black, K. Sawyer, R. Mercado, H. Qi, A. Sohn, B. Taylor, N. Tydingco, G. Kammerer, D. Stroud, J. Khatha, K. Jenkins, K. Most, N. Stein, R. Chavira, T. Craven-Bartle, E. Sanchez, Y . Ding, J. Malik, and R. Calandra. Digitizing Touch with an Artificial Multimodal Fingertip.arXiv preprint arXiv:2411.02479, 2024
Pith/arXiv arXiv 2024
-
[22]
Cheng, K
T. Cheng, K. Chen, L. Chen, L. Zhang, Y . Zhang, Y . Ling, M. Hamad, Z. Bing, F. Wu, K. Sharma, and A. Knoll. TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks. InIEEE International Conference on Robotics and Automation (ICRA), 2026
2026
-
[23]
Tirumala, T
S. Tirumala, T. Weng, D. Seita, O. Kroemer, Z. Temel, and D. Held. Learning to Singulate Layers of Cloth Using Tactile Feedback. InIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022
2022
-
[24]
P. Lin, Y . Huang, W. Li, J. Ma, C. Xiao, and Z. Jiao. PP-Tac: Paper Picking Using Tactile Feedback in Dexterous Robotic Hands. InRobotics: Science and Systems (RSS), 2025
2025
-
[25]
Romero, H.-S
B. Romero, H.-S. Fang, P. Agrawal, and E. Adelson. Eyesight hand: Design of a fully-actuated dexterous robot hand with integrated vision-based tactile sensors and compliant actuation. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024
2024
-
[26]
H. Zhou, H. Lou, W. Liu, E. Zhao, Y . Wang, and D. Seita. The MOTIF Hand: A Robotic Hand for Multimodal Observations with Thermal, Inertial, and Force Sensors. InInternational Symposium on Experimental Robotics (ISER), 2025
2025
-
[27]
Z. Si, T. C. Yu, K. Morozov, J. McCann, and W. Yuan. RobotSweater: Scalable, General- izable, and Customizable Machine-Knitted Tactile Skins for Robots. InIEEE International Conference on Robotics and Automation (ICRA), 2023
2023
-
[28]
B. Huang and Y . Li. Flexitac: A low-cost, open-source, scalable tactile sensing solution for robotic systems.arXiv preprint arXiv:2604.28156, 2026
Pith/arXiv arXiv 2026
-
[29]
Y . Huang, J. Wu, J. Jiang, H. Lin, A. Aierken, Y . Wang, K. Cheng, Z. Jiao, and Y . Zhong. HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Vision.arXiv preprint arxiv:2606.19161, 2026
Pith/arXiv arXiv 2026
-
[30]
Rodriguez, Y
S. Rodriguez, Y . Dou, M. Oller, A. Owens, and N. Fazeli. Cross-Sensor Touch Generation. In Conference on Robot Learning (CoRL), 2025
2025
-
[31]
C. Yuan, Z. Zhang, M. Zhou, W. Chen, Y . Wang, Z. Liu, D. Niu, S. Wang, H. Zhang, W. Zhang, Y . Hu, Y . Gong, W. Xing, C. Wen, C. Lu, K. Zhang, and Y . Gao. FTP-1: A Generalist Foun- dation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation.arXiv preprint arXiv:2606.13102, 2026
Pith/arXiv arXiv 2026
-
[32]
J. Huang, S. Wang, F. Lin, Y . Hu, C. Wen, and Y . Gao. Tactile-VLA: Unlocking Vision- Language-Action Model’s Physical Knowledge for Tactile Generalization.arXiv preprint arXiv:2507.09160, 2025
Pith/arXiv arXiv 2025
-
[33]
J. Bi, K. Y . Ma, C. Hao, M. Z. Shou, and H. Soh. VLA-Touch: Enhancing Vision-Language- Action Models with Dual-Level Tactile Feedback. InIEEE Robotics and Automation Letters (RA-L), 2026
2026
-
[34]
Guzey, B
I. Guzey, B. Evans, S. Chintala, and L. Pinto. Dexterity from Touch: Self-Supervised Pre- Training of Tactile Representations with Robotic Play. InConference on Robot Learning (CoRL), 2023. 11
2023
-
[35]
Z. Zhao, S. Haldar, J. Cui, L. Pinto, and R. Bhirangi. Touch Begins Where Vision Ends: Generalizable Policies for Contact-rich Manipulation.arXiv preprint arXiv:2506.13762, 2025
Pith/arXiv arXiv 2025
-
[36]
M. A. Lee, Y . Zhu, K. Srinivasan, P. Shah, S. Savarese, L. Fei-Fei, A. Garg, and J. Bohg. Mak- ing Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks. InIEEE International Conference on Robotics and Automation (ICRA), 2019
2019
-
[37]
Y . Chen, A. Sipos, M. V . der Merwe, and N. Fazeli. Visuo-Tactile Transformers for Manipula- tion. InConference on Robot Learning (CoRL), 2022
2022
-
[38]
S. Jiang, S. Zhao, Y . Fan, and P. Yin. GelFusion: Enhancing Robotic Manipulation under Visual Constraints via Visuotactile Fusion.arXiv preprint arXiv:2505.07455, 2025
Pith/arXiv arXiv 2025
-
[39]
H. Qi, B. Yi, S. Suresh, M. Lambeta, Y . Ma, R. Calandra, and J. Malik. General In-Hand Object Rotation with Vision and Touch. InConference on Robot Learning (CoRL), 2023
2023
-
[40]
Z.-H. Yin, B. Huang, Y . Qin, Q. Chen, and X. Wang. Rotating without Seeing: Towards In-hand Dexterity through Touch. InRobotics: Science and Systems (RSS), 2023
2023
-
[41]
T. Lin, Y . Zhang, Q. Li, H. Qi, B. Yi, S. Levine, and J. Malik. Learning Visuotactile Skills with Two Multifingered Hands. InIEEE International Conference on Robotics and Automation (ICRA), 2025
2025
-
[42]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Park, I. Radosavovic, K. Wang, A. Zhan, K. Black, C. Chi, K. B. Hatch, S. Lin, J. ...
2024
-
[43]
G. Ji, H. Polavaram, L. Y . Chen, S. Bajamahal, Z. Ma, S. Adebola, C. Xu, and K. Goldberg. OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Pol- icy Learning. InInternational Conference on Machine Learning (ICML), 2026
2026
-
[44]
C. Yuan, S. Joshi, S. Zhu, H. Su, H. Zhao, and Y . Gao. RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation. InIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025
2025
-
[45]
Bharadhwaj, J
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V . Kumar. RoboAgent: Gen- eralization and Efficiency in Robot Manipulation via Semantic Augmentations and Action Chunking. InIEEE International Conference on Robotics and Automation (ICRA), 2024
2024
-
[46]
A. Zhou, M. J. Kim, L. Wang, P. Florence, and C. Finn. NeRF in the Palm of Your Hand: Corrective Augmentation for Robotics via Novel-View Synthesis. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[47]
Zhang, M
X. Zhang, M. Chang, P. Kumar, and S. Gupta. Diffusion Meets DAgger: Supercharging Eye- in-hand Imitation Learning. InRobotics: Science and Systems (RSS), 2024. 12
2024
-
[48]
I.-C. A. Liu, J. Chen, G. Sukhatme, and D. Seita. D-CODA: Diffusion for Coordinated Dual- Arm Data Augmentation. InConference on Robot Learning (CoRL), 2025
2025
-
[49]
L. Y . Chen, C. Xu, K. Dharmarajan, M. Z. Irshad, R. Cheng, K. Keutzer, M. Tomizuka, Q. Vuong, and K. Goldberg. RoVi-Aug: Robot and Viewpoint Augmentation for Cross- Embodiment Robot Learning. InConference on Robot Learning (CoRL), 2024
2024
-
[50]
Y . Guo, L. X. Shi, J. Chen, and C. Finn. Ctrl-World: A Controllable Generative World Model for Robot Manipulation. InInternational Conference on Learning Representations (ICLR), 2026
2026
-
[51]
J. Jang, S. Ye, Z. Lin, J. Xiang, J. Bjorck, Y . Fang, F. Hu, S. Huang, K. Kundalia, Y .-C. Lin, L. Magne, A. Mandlekar, A. Narayan, Y . L. Tan, G. Wang, J. Wang, Q. Wang, Y . Xu, X. Zeng, K. Zheng, R. Zheng, M.-Y . Liu, L. Zettlemoyer, D. Fox, J. Kautz, S. Reed, Y . Zhu, and L. Fan. DreamGen: Unlocking Generalization in Robot Learning through Video World...
Pith/arXiv arXiv 2025
-
[52]
Chen, I.-C
J. Chen, I.-C. A. Liu, G. Sukhatme, and D. Seita. CRAFT: Video Diffusion for Bimanual Robot Data Generation. InIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2026
2026
-
[53]
C. Higuera, S. Arnaud, B. Boots, M. Mukadam, F. R. Hogan, and F. Meier. Visuo-Tactile World Models.arXiv preprint arXiv:2602.06001, 2026
arXiv 2026
-
[54]
Y . Lou, Y . Ye, Y . Fu, J. Cen, X. Chi, Y . Lyu, P. Jia, S. Han, Z. Lu, and S. Zhang. Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation.arXiv preprint arXiv:2606.08737, 2026
Pith/arXiv arXiv 2026
-
[55]
Y . Zang, Y . Zheng, X. Nie, Y . Zheng, S. Tian, S. Gu, C. Gao, Z. Wang, S. Yan, and W. Ding. TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation.arXiv preprint arXiv:2606.11184, 2026
Pith/arXiv arXiv 2026
-
[56]
D. Luo, K. Yu, A.-H. Shahidzadeh, C. Ferm ¨uller, Y . Aloimonos, and R. Gao. ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image. arXiv preprint arXiv:2505.20498, 2025
Pith/arXiv arXiv 2025
-
[57]
Z. Wu, Y . Lin, Y . Zhao, X. Zhang, Z. Chen, N. Lepora, and S. Luo. ViTacGen: Robotic Pushing with Vision-to-Touch Generation. InIEEE Robotics and Automation Letters (RA-L), 2025
2025
-
[58]
K. Lyu, L. Xiao, J. Zeng, D. Wu, L. Shu, and J. Hao. VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios.arXiv preprint arXiv:2607.14728, 2026
Pith/arXiv arXiv 2026
-
[59]
Z. Zhang, A. Desai, J.-C. Hu, Y . Saka, Q. K. Luu, J. Lei, D. Soleymanzadeh, B. Zhang, M. Zheng, and Y . She. TacImag: Touch-Informed Manipulation through Imagined Tactile Representations.arXiv preprint arXiv:2607.01684, 2026
Pith/arXiv arXiv 2026
-
[60]
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haz- iza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski. DINOv2: Learning Robust Visual Features without Sup...
Pith/arXiv arXiv 2024
-
[61]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. InInternational Conference on Learning Representations (ICLR), 2021
2021
-
[62]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polo- sukhin. Attention Is All You Need. InNeural Information Processing Systems (NeurIPS), 2017. 13
2017
-
[63]
T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Doll´ar. Focal Loss for Dense Object Detection. InIEEE/CVF International Conference on Computer Vision (ICCV), 2017
2017
-
[64]
Milletari, N
F. Milletari, N. Navab, and S.-A. Ahmadi. V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation. InInternational Conference on 3D Vision (3DV), 2016
2016
-
[65]
P. J. Huber. Robust Estimation of a Location Parameter.The Annals of Mathematical Statistics, 35(1):73–101, 1964
1964
-
[66]
P. Wu, Y . Shentu, Z. Yi, X. Lin, and P. Abbeel. GELLO: A General, Low-Cost, and Intuitive Teleoperation Framework for Robot Manipulators. InIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024
2024
-
[67]
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image Quality Assessment: From Error Visibility to Structural Similarity.IEEE Transactions on Image Processing, 13(4):600– 612, Apr. 2004
2004
-
[68]
Zhang, P
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[69]
K. He, X. Chen, S. Xie, Y . Li, P. Doll´ar, and R. Girshick. Masked Autoencoders Are Scalable Vision Learners. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022
2022
-
[70]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. InInternational Conference on Machine Learning (ICML), 2021
2021
-
[71]
H. Kress-Gazit, K. Hashimoto, N. Kuppuswamy, P. Shah, P. Horgan, G. Richardson, S. Feng, and B. Burchfiel. Robot Learning as an Empirical Science: Best Practices for Policy Evalua- tion.arXiv preprint arXiv:2409.09491, 2024. 14 Supplementary Material for FELT Generating Tactile Signals from Vision for Visuo-Tactile Manipulation S1 Overview of Supplementar...
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.