Pith. sign in

REVIEW 7 cited by

Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.10191 v2 pith:Y7A353KS submitted 2018-10-24 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords learningcontact-richdifferentinputsmultimodalrealrobotsample
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Contact-rich manipulation tasks in unstructured environments often require both haptic and visual feedback. However, it is non-trivial to manually design a robot controller that combines modalities with very different characteristics. While deep reinforcement learning has shown success in learning control policies for high-dimensional inputs, these algorithms are generally intractable to deploy on real robots due to sample complexity. We use self-supervision to learn a compact and multimodal representation of our sensory inputs, which can then be used to improve the sample efficiency of our policy learning. We evaluate our method on a peg insertion task, generalizing over different geometry, configurations, and clearances, while being robust to external perturbations. Results for simulated and real robot experiments are presented.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A modular backpack-based teleoperation interface enables bimanual mobile manipulation with haptic feedback and active perception across multiple robot platforms.

  2. TactX: Learning Shared Tactile Representations Across Diverse Sensors

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    TactX learns a shared latent representation across three tactile sensor modalities via joint training on paired contacts, enabling zero-shot policy transfer and higher success on pick-and-place, insertion, wiping, and...

  3. Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    MiTaS fuses multi-resolution tactile data from GelSight and Evetac sensors with vision using modality-specific stems and transformer fusion to condition flow-matching policies, reporting 80% average success on five co...

  4. Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning

    cs.RO 2025-11 unverdicted novelty 6.0 of 10

    MSDP pre-trains a transformer encoder with masked multisensory autoencoding, then uses an asymmetric actor-critic bridge (cross-attention for critic, pooling for actor) to accelerate and robustify contact-rich RL acro...

  5. AdaDexGrasp: Adaptive Dexterous Grasping via 3D Visuo-Tactile Representation Fusion

    cs.RO 2026-08 conditional novelty 5.0 of 10

    AdaDexGrasp learns to fuse point clouds with finger-level tactile labels to generate, judge, and correct dexterous grasps, reporting 91%/82%/83% success on seen, unseen-object, and unseen-category sets in simulation.

  6. QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning

    cs.RO 2024-12 conditional novelty 5.0 of 10

    Compressing 10-step action chunks into discrete latent codes lets an 8B multimodal model drive a quadruped at controller frequency and raises average task success by about 65%.

  7. Grasping Using Tactile Sensing and Deep Calibration

    cs.RO 2019-07 unverdicted novelty 3.0 of 10

    A tactile feedback approach for robot grasping evaluated on a real robot, using deep learning to eliminate bias in force-torque sensor data.

Pith tools