Pith. sign in

REVIEW 5 cited by

Self-Supervised Visuo-Tactile Pretraining to Locate and Follow Garment Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.13042 v2 pith:2I6UQJHP submitted 2022-09-26 cs.RO

classification cs.RO
keywords featurerepresentationsself-supervisedtactilevisualvisuo-tactilecross-modalgarment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans make extensive use of vision and touch as complementary senses, with vision providing global information about the scene and touch measuring local information during manipulation without suffering from occlusions. While prior work demonstrates the efficacy of tactile sensing for precise manipulation of deformables, they typically rely on supervised, human-labeled datasets. We propose Self-Supervised Visuo-Tactile Pretraining (SSVTP), a framework for learning multi-task visuo-tactile representations in a self-supervised manner through cross-modal supervision. We design a mechanism that enables a robot to autonomously collect precisely spatially-aligned visual and tactile image pairs, then train visual and tactile encoders to embed these pairs into a shared latent space using cross-modal contrastive loss. We apply this latent space to downstream perception and control of deformable garments on flat surfaces, and evaluate the flexibility of the learned representations without fine-tuning on 5 tasks: feature classification, contact localization, anomaly detection, feature search from a visual query (e.g., garment feature localization under occlusion), and edge following along cloth edges. The pretrained representations achieve a 73-100% success rate on these 5 tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. {\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Action-conditioned JEPA-style future-visual latent prediction yields dynamics-aware tactile tokens that lift contact-rich VLA success rates from ~30% to ~70% average on four real tasks.

  2. Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    Splash partitions MLLM parameters into dormant and critical subspaces via significance quantification, updating only the dormant subspace for tactile alignment while preserving general capabilities and achieving SOTA ...

  3. UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    UniTac is the first unified multimodal model for cross-sensor tactile understanding and generation, using dual-level representations, two new understanding tasks, and a two-stage training paradigm with sensor-prior sa...

  4. Heterogeneous Tactile Transformer

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    HTT learns shared representations across heterogeneous tactile sensors using a new paired dataset and pretraining objectives, enabling transfer to unseen sensors and tasks.

  5. Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    The model uses dense visuo-tactile feature interactions and material-diversity pairing on expanded datasets to generate tactile saliency maps for material segmentation, outperforming prior global-alignment methods.

Pith tools