REVIEW 18 cited by
Touch and Go: Learning from Human-Collected Vision and Touch
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The ability to associate touch with sight is essential for tasks that require physically interacting with objects in the world. We propose a dataset with paired visual and tactile data called Touch and Go, in which human data collectors probe objects in natural environments using tactile sensors, while simultaneously recording egocentric video. In contrast to previous efforts, which have largely been confined to lab settings or simulated environments, our dataset spans a large number of "in the wild" objects and scenes. To demonstrate our dataset's effectiveness, we successfully apply it to a variety of tasks: 1) self-supervised visuo-tactile feature learning, 2) tactile-driven image stylization, i.e., making the visual appearance of an object more consistent with a given tactile signal, and 3) predicting future frames of a tactile signal from visuo-tactile inputs.
Forward citations
Cited by 18 Pith papers
-
HapTile: A Haptic-Informed Vision-Tactile-Language-Action Dataset for Contact-Rich Imitation Learning
HapTile introduces a visuotactile dataset with haptic-informed teleoperation for language-conditioned contact-rich manipulation tasks and provides baseline policy benchmarks.
-
Touch-R1: Reinforcing Touch Reasoning in MLLMs
Touch-R1 applies GRPO reinforcement learning on a new 1M tactile dataset and benchmark to train a Qwen2.5-VL-7B model that outperforms baselines on tactile perception and visual-tactile conflict tasks.
-
TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance
TouchGuide improves contact-rich robot manipulation by steering diffusion or flow-matching visuomotor policies with tactile feasibility scores from a contrastively trained Contact Physical Model.
-
New York Smells: A Large Multimodal Dataset for Olfaction
New York Smells is an in-the-wild dataset of 7,000 co-captured image–e-nose smell pairs covering 3,500 objects, and contrastive vision-smell training on it yields olfactory representations that outperform hand-crafted...
-
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Janus decouples visual encoding into task-specific pathways inside a single autoregressive transformer to unify multimodal understanding and generation while outperforming earlier unified models.
-
{\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
Action-conditioned JEPA-style future-visual latent prediction yields dynamics-aware tactile tokens that lift contact-rich VLA success rates from ~30% to ~70% average on four real tasks.
-
TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios
A dynamic-aware tactile encoder plus TouchCoT-10k chain-of-thought data lets a 7B model outperform larger tactile-language baselines on physical-property and real-world reasoning tasks.
-
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
UniTac is the first unified multimodal model for cross-sensor tactile understanding and generation, using dual-level representations, two new understanding tasks, and a two-stage training paradigm with sensor-prior sa...
-
Heterogeneous Tactile Transformer
HTT learns shared representations across heterogeneous tactile sensors using a new paired dataset and pretraining objectives, enabling transfer to unseen sensors and tasks.
-
Tac-DINO: Learning Vision-Tactile Features with Patch Alignment
Tac-DINO constructs a large tactile dataset and Vis-Tac Holographic Matching Benchmark, then proposes Vision-Tactile Patch Alignment (VTPA) methods that outperform non-aligned baselines on local-to-global feature matching.
-
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Large multi-source tactile data plus question-guided Gaussian temporal MoE yields ~7-point gains over VTV-LLM on tactile property and commonsense reasoning tasks, with improved unseen-sensor generalization.
-
HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals
HapticCap is the first large human-annotated vibration-caption dataset, and a contrastive retrieval model using T5 and AST achieves the best caption-matching performance among the tested baselines.
-
VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios
VQ-Touch applies VQGAN with deformable convolutions and discrete diffusion to generate tactile images across sensors with few-shot mixed training.
-
TacCoRL: Integrating Tactile Feedback into VLA via Simulation
TacCoRL integrates tactile feedback into VLA policies via real-aligned simulation co-training and RL, raising average success from 50% to 72.5% on four bimanual contact-rich tasks with direct real-robot transfer.
-
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
TouchThinker introduces a 1M-scale multi-source tactile dataset and action-aware modeling to scale commonsense reasoning from tactile observations, reporting competitive performance on new and existing benchmarks.
-
Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
A survey proposing a hierarchical taxonomy for multimodal tactile fusion datasets and methods across perception, generation, and interaction in embodied intelligence.
-
Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
Surformer v1 is a cross-modal transformer for tactile-visual surface classification that claims 99.4% accuracy at 0.77 ms inference, but the submitted full text belongs to an unrelated statistics paper.
-
Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision
Surformer v2 combines an ImageNet-pretrained CNN for vision with a handcrafted-feature transformer for touch, fuses their outputs via learned weights, and reports 97.4% accuracy with 0.0239 ms inference on Touch and Go.
Discussion (0). Sign in to comment.