Pith. sign in

REVIEW 4 cited by

The Power of the Senses: Generalizable Manipulation from Vision and Touch through Masked Multimodal Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00924 v1 pith:YFO3WVPT submitted 2023-11-02 cs.RO cs.AI

classification cs.ROcs.AI
keywords learningmultimodalsensesmanipulationmaskedrepresentationstouchvision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans rely on the synergy of their senses for most essential tasks. For tasks requiring object manipulation, we seamlessly and effectively exploit the complementarity of our senses of vision and touch. This paper draws inspiration from such capabilities and aims to find a systematic approach to fuse visual and tactile information in a reinforcement learning setting. We propose Masked Multimodal Learning (M3L), which jointly learns a policy and visual-tactile representations based on masked autoencoding. The representations jointly learned from vision and touch improve sample efficiency, and unlock generalization capabilities beyond those achievable through each of the senses separately. Remarkably, representations learned in a multimodal setting also benefit vision-only policies at test time. We evaluate M3L on three simulated environments with both visual and tactile observations: robotic insertion, door opening, and dexterous in-hand manipulation, demonstrating the benefits of learning a multimodal policy. Code and videos of the experiments are available at https://sferrazza.cc/m3l_site.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Feel the Force: Contact-Driven Learning from Humans

    cs.RO 2025-06 conditional novelty 7.0 of 10

    FeelTheForce trains a robot policy on human tactile demonstrations, predicting desired contact forces and using a PD controller to track them on the robot gripper, achieving 77% success across five force-sensitive tasks.

  2. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  3. Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning

    cs.RO 2025-11 unverdicted novelty 6.0 of 10

    MSDP pre-trains a transformer encoder with masked multisensory autoencoding, then uses an asymmetric actor-critic bridge (cross-attention for critic, pooling for actor) to accelerate and robustify contact-rich RL acro...

  4. Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

    cs.RO 2025-01 conditional novelty 6.0 of 10

    FuSe fine-tunes generalist robot policies on touch and audio data using language as a bridge, enabling multimodal prompts and compositional cross-modal tasks.

Pith tools