Pith. sign in

Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it
abstract

Contact-rich manipulation tasks in unstructured environments often require both haptic and visual feedback. However, it is non-trivial to manually design a robot controller that combines modalities with very different characteristics. While deep reinforcement learning has shown success in learning control policies for high-dimensional inputs, these algorithms are generally intractable to deploy on real robots due to sample complexity. We use self-supervision to learn a compact and multimodal representation of our sensory inputs, which can then be used to improve the sample efficiency of our policy learning. We evaluate our method on a peg insertion task, generalizing over different geometry, configurations, and clearances, while being robust to external perturbations. Results for simulated and real robot experiments are presented.

fields

cs.RO 4

representative citing papers

TactX: Learning Shared Tactile Representations Across Diverse Sensors

cs.RO · 2026-06-30 · unverdicted · novelty 6.0

TactX learns a shared latent representation across three tactile sensor modalities via joint training on paired contacts, enabling zero-shot policy transfer and higher success on pick-and-place, insertion, wiping, and reorientation tasks.

citing papers explorer

Showing 4 of 4 citing papers.

  • TactX: Learning Shared Tactile Representations Across Diverse Sensors cs.RO · 2026-06-30 · unverdicted · none · ref 38 · internal anchor

    TactX learns a shared latent representation across three tactile sensor modalities via joint training on paired contacts, enabling zero-shot policy transfer and higher success on pick-and-place, insertion, wiping, and reorientation tasks.

  • Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation cs.RO · 2026-06-04 · unverdicted · none · ref 32 · internal anchor

    MiTaS fuses multi-resolution tactile data from GelSight and Evetac sensors with vision using modality-specific stems and transformer fusion to condition flow-matching policies, reporting 80% average success on five contact-rich tasks versus 31-54% baselines.

  • Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning cs.RO · 2025-11-18 · conditional · none · ref 4 · internal anchor

    MSDP pre-trains a transformer encoder with masked multisensory autoencoding, then uses an asymmetric actor-critic bridge (cross-attention for critic, pooling for actor) to accelerate and robustify contact-rich RL across simulation and real robots.

  • Grasping Using Tactile Sensing and Deep Calibration cs.RO · 2019-07-23 · unverdicted · none · ref 1 · internal anchor

    A tactile feedback approach for robot grasping evaluated on a real robot, using deep learning to eliminate bias in force-torque sensor data.