Pith. sign in

REVIEW 19 cited by

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12191 v3 pith:IX2HLOEB submitted 2025-02-15 cs.LG cs.CVcs.RO

classification cs.LGcs.CVcs.RO
keywords sensorstactilemulti-sensorunifiedvisuo-tactilelearningperceptionvarious
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding in completing various tasks. However, the distinct data characteristics of these low-standardized visuo-tactile sensors hinder the establishment of a powerful tactile perception system. We consider that the key to addressing this issue lies in learning unified multi-sensor representations, thereby integrating the sensors and promoting tactile knowledge transfer between them. To achieve unified representation of this nature, we introduce TacQuad, an aligned multi-modal multi-sensor tactile dataset from four different visuo-tactile sensors, which enables the explicit integration of various sensors. Recognizing that humans perceive the physical environment by acquiring diverse tactile information such as texture and pressure changes, we further propose to learn unified multi-sensor representations from both static and dynamic perspectives. By integrating tactile images and videos, we present AnyTouch, a unified static-dynamic multi-sensor representation learning framework with a multi-level structure, aimed at both enhancing comprehensive perceptual abilities and enabling effective cross-sensor transfer. This multi-level architecture captures pixel-level details from tactile data via masked modeling and enhances perception and transferability by learning semantic-level sensor-agnostic features through multi-modal alignment and cross-sensor matching. We provide a comprehensive analysis of multi-sensor transferability, and validate our method on various datasets and in the real-world pouring task. Experimental results show that our method outperforms existing methods, exhibits outstanding static and dynamic perception capabilities across various sensors.

Discussion (0). Sign in to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. You Only Touch Once: 6-DoF Object Pose Estimation from Single Tactile Contact

    cs.RO 2026-06 unverdicted novelty 7.0 of 10

    A tactile system recovers 6-DoF object pose from one contact pair by coarse-to-fine localization of point clouds on a known model followed by normal-aware SVD.

  2. TacVerse: A Multi-Sensor Dataset and Benchmark for Cross-Sensor Vision-Based Tactile Perception

    cs.RO 2026-06 unverdicted novelty 7.0 of 10

    TacVerse is a new multi-sensor tactile dataset with 106,800 images from seven VBTS designs that benchmarks within-sensor performance, zero-shot cross-sensor transfer, and few-shot adaptation on shape, grating, and for...

  3. FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation

    cs.RO 2026-06 unverdicted novelty 7.0 of 10

    FTP-1 is the first foundation tactile policy pretrained on ~3000 hours of data from 26 sources across 21 sensors that improves performance on seen setups by 17.2% and transfers to unseen sensors with 31% success rate gain.

  4. Touch-R1: Reinforcing Touch Reasoning in MLLMs

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Touch-R1 applies GRPO reinforcement learning on a new 1M tactile dataset and benchmark to train a Qwen2.5-VL-7B model that outperforms baselines on tactile perception and visual-tactile conflict tasks.

  5. TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A dynamic-aware tactile encoder plus TouchCoT-10k chain-of-thought data lets a 7B model outperform larger tactile-language baselines on physical-property and real-world reasoning tasks.

  6. UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    UniTac is the first unified multimodal model for cross-sensor tactile understanding and generation, using dual-level representations, two new understanding tasks, and a two-stage training paradigm with sensor-prior sa...

  7. TactX: Learning Shared Tactile Representations Across Diverse Sensors

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    TactX learns a shared latent representation across three tactile sensor modalities via joint training on paired contacts, enabling zero-shot policy transfer and higher success on pick-and-place, insertion, wiping, and...

  8. HT-Bench: Benchmarking and Learning Dexterous Full-Hand Tactile Representations with Egocentric Vision

    cs.RO 2026-06 conditional novelty 6.0 of 10

    HT-Bench is a large egocentric vision-plus-full-hand-tactile benchmark with four evaluation tasks; the proposed HandTouch encoder improves Recall@5 from 74.65% to 85.23%, reduces inpainting RMSE from 0.022 to 0.010, a...

  9. Tac-DINO: Learning Vision-Tactile Features with Patch Alignment

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Tac-DINO constructs a large tactile dataset and Vis-Tac Holographic Matching Benchmark, then proposes Vision-Tactile Patch Alignment (VTPA) methods that outperform non-aligned baselines on local-to-global feature matching.

  10. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

    cs.AI 2026-06 conditional novelty 6.0 of 10

    Large multi-source tactile data plus question-guided Gaussian temporal MoE yields ~7-point gains over VTV-LLM on tactile property and commonsense reasoning tasks, with improved unseen-sensor generalization.

  11. TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    TacForeSight trains a force-conditioned tactile world model to predict latent dynamics and uses those predictions as anticipatory priors inside a visuo-tactile policy for real-time contact-rich manipulation.

  12. RGB-S: Image-Aligned Tactile Saliency for Robust Dexterous Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    RGB-S projects tactile contacts onto images as force-modulated Gaussian saliency maps via kinematics and zero-initialized conditioning, raising real-world occluded dexterous manipulation success by 26.7 percentage poi...

  13. Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    MiTaS fuses multi-resolution tactile data from GelSight and Evetac sensors with vision using modality-specific stems and transformer fusion to condition flow-matching policies, reporting 80% average success on five co...

  14. Tactile Modality Fusion for Vision-Language-Action Models

    cs.RO 2026-03 conditional novelty 6.0 of 10

    A FiLM-based tactile fusion method that conditions VLA visual features on frozen pretrained touch embeddings improves real-robot insertion success, speed, and force control relative to vision-only and concatenation baselines.

  15. Imagining the Sense of Touch: Touch-Informed Manipulation via Imagined Tactile Representations

    cs.RO 2026-07 unverdicted novelty 5.0 of 10

    TacImag framework trains on paired visuotactile data to predict tactile observations from vision, improving performance on six simulated and four real-world manipulation tasks.

  16. Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    A visuo-tactile policy learning method that exploits tactile motion correlation for contact state distinction and Mixture-of-Transformers for cross-modal fusion.

  17. TacCoRL: Integrating Tactile Feedback into VLA via Simulation

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    TacCoRL integrates tactile feedback into VLA policies via real-aligned simulation co-training and RL, raising average success from 50% to 72.5% on four bimanual contact-rich tasks with direct real-robot transfer.

  18. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    TouchThinker introduces a 1M-scale multi-source tactile dataset and action-aware modeling to scale commonsense reasoning from tactile observations, reporting competitive performance on new and existing benchmarks.

  19. Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

    cs.RO 2026-05 unverdicted novelty 4.0 of 10

    A survey proposing a hierarchical taxonomy for multimodal tactile fusion datasets and methods across perception, generation, and interaction in embodied intelligence.

Pith tools