Pith. sign in

REVIEW 4 cited by

Visuo-Tactile Transformers for Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.00121 v1 pith:EWVK3MRV submitted 2022-09-30 cs.RO cs.LG

classification cs.ROcs.LG
keywords learningrepresentationvisuo-tactileapproachattentiondomainfeedbackmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning representations in the joint domain of vision and touch can improve manipulation dexterity, robustness, and sample-complexity by exploiting mutual information and complementary cues. Here, we present Visuo-Tactile Transformers (VTTs), a novel multimodal representation learning approach suited for model-based reinforcement learning and planning. Our approach extends the Visual Transformer \cite{dosovitskiy2021image} to handle visuo-tactile feedback. Specifically, VTT uses tactile feedback together with self and cross-modal attention to build latent heatmap representations that focus attention on important task features in the visual domain. We demonstrate the efficacy of VTT for representation learning with a comparative evaluation against baselines on four simulated robot tasks and one real world block pushing task. We conduct an ablation study over the components of VTT to highlight the importance of cross-modality in representation learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TACTIC: Tactile and Vision Conditioned Contact-Centric Control for Whole-Arm Manipulation

    cs.RO 2026-07 conditional novelty 6.5 of 10

    A hybrid contact-centric MPC with tactile-vision latents and Jacobian-biased sampling outperforms pure learned and pure kinematic baselines on multi-contact whole-arm tasks in sim and on a manikin/maze robot.

  2. Self-Supervised Multisensory Pretraining for Contact-Rich Robot Reinforcement Learning

    cs.RO 2025-11 unverdicted novelty 6.0 of 10

    MSDP pre-trains a transformer encoder with masked multisensory autoencoding, then uses an asymmetric actor-critic bridge (cross-attention for critic, pooling for actor) to accelerate and robustify contact-rich RL acro...

  3. TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Feeding torque history as a single decoder token and adding torque prediction as an auxiliary objective improves pretrained VLA success rates on contact-rich manipulation, with large gains on button pushing and charge...

  4. Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision

    cs.RO 2025-09 reject novelty 3.0 of 10

    Surformer v2 combines an ImageNet-pretrained CNN for vision with a handcrafted-feature transformer for touch, fuses their outputs via learned weights, and reports 97.4% accuracy with 0.0239 ms inference on Touch and Go.

Pith tools