Pith. sign in

REVIEW 9 cited by

Touch and Go: Learning from Human-Collected Vision and Touch

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12498 v2 pith:WQ55W4A7 submitted 2022-11-22 cs.CV

classification cs.CV
keywords tactiletouchdatasetobjectsdataenvironmentslearningsignal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to associate touch with sight is essential for tasks that require physically interacting with objects in the world. We propose a dataset with paired visual and tactile data called Touch and Go, in which human data collectors probe objects in natural environments using tactile sensors, while simultaneously recording egocentric video. In contrast to previous efforts, which have largely been confined to lab settings or simulated environments, our dataset spans a large number of "in the wild" objects and scenes. To demonstrate our dataset's effectiveness, we successfully apply it to a variety of tasks: 1) self-supervised visuo-tactile feature learning, 2) tactile-driven image stylization, i.e., making the visual appearance of an object more consistent with a given tactile signal, and 3) predicting future frames of a tactile signal from visuo-tactile inputs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. New York Smells: A Large Multimodal Dataset for Olfaction

    cs.CV 2025-11 conditional novelty 7.0 of 10

    New York Smells is an in-the-wild dataset of 7,000 co-captured image–e-nose smell pairs covering 3,500 objects, and contrastive vision-smell training on it yields olfactory representations that outperform hand-crafted...

  2. {\tau}: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Action-conditioned JEPA-style future-visual latent prediction yields dynamics-aware tactile tokens that lift contact-rich VLA success rates from ~30% to ~70% average on four real tasks.

  3. TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A dynamic-aware tactile encoder plus TouchCoT-10k chain-of-thought data lets a 7B model outperform larger tactile-language baselines on physical-property and real-world reasoning tasks.

  4. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    TouchThinker introduces a 1M-scale multi-source tactile dataset and action-aware modeling to scale commonsense reasoning from tactile observations, reporting competitive performance on new and existing benchmarks.

  5. HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals

    cs.CL 2025-07 conditional novelty 6.0 of 10

    HapticCap is the first large human-annotated vibration-caption dataset, and a contrastive retrieval model using T5 and AST achieves the best caption-matching performance among the tested baselines.

  6. CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new multimodal dataset of simulated operating-room team dialogues shows that existing LLMs and fine-tuned models reach only about 51% macro F1 on team reflection behavior classification, indicating substantial room ...

  7. VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios

    cs.CV 2026-07 reject novelty 5.0 of 10

    VQ-Touch applies VQGAN with deformable convolutions and discrete diffusion to generate tactile images across sensors with few-shot mixed training.

  8. Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Surformer v1 is a cross-modal transformer for tactile-visual surface classification that claims 99.4% accuracy at 0.77 ms inference, but the submitted full text belongs to an unrelated statistics paper.

  9. Surformer v2: A Multimodal Classifier for Surface Understanding from Touch and Vision

    cs.RO 2025-09 reject novelty 3.0 of 10

    Surformer v2 combines an ImageNet-pretrained CNN for vision with a handcrafted-feature transformer for touch, fuses their outputs via learned weights, and reports 97.4% accuracy with 0.0239 ms inference on Touch and Go.

Pith tools