Pith. sign in

REVIEW 11 cited by

Touch begins where vision ends: Generalizable policies for contact-rich manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.13762 v1 pith:DWAGU6BY submitted 2025-06-16 cs.RO cs.CV

Touch begins where vision ends: Generalizable policies for contact-rich manipulation

classification cs.RO cs.CV
keywords vitalcontact-richmanipulationpolicieslearninglocalrobusttasks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Data-driven approaches struggle with precise manipulation; imitation learning requires many hard-to-obtain demonstrations, while reinforcement learning yields brittle, non-generalizable policies. We introduce VisuoTactile Local (ViTaL) policy learning, a framework that solves fine-grained manipulation tasks by decomposing them into two phases: a reaching phase, where a vision-language model (VLM) enables scene-level reasoning to localize the object of interest, and a local interaction phase, where a reusable, scene-agnostic ViTaL policy performs contact-rich manipulation using egocentric vision and tactile sensing. This approach is motivated by the observation that while scene context varies, the low-level interaction remains consistent across task instances. By training local policies once in a canonical setting, they can generalize via a localize-then-execute strategy. ViTaL achieves around 90% success on contact-rich tasks in unseen environments and is robust to distractors. ViTaL's effectiveness stems from three key insights: (1) foundation models for segmentation enable training robust visual encoders via behavior cloning; (2) these encoders improve the generalizability of policies learned using residual RL; and (3) tactile sensing significantly boosts performance in contact-rich tasks. Ablation studies validate each of these insights, and we demonstrate that ViTaL integrates well with high-level VLMs, enabling robust, reusable low-level skills. Results and videos are available at https://vitalprecise.github.io.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Feeling the Unexpected: ResTacVLA for Contact-Rich Manipulation via Residual Tactile Representation

    cs.RO 2026-07 conditional novelty 6.5

    Residual tactile representations plus surprise-aware gating let VLA policies master contact-rich robot tasks that vision-only models fail.

  2. $N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    A scaled tactile-native world-action model jointly predicts vision, touch, and action and outperforms vision-only baselines on contact-rich sim and real robot tasks.

  3. FELT: Generating Tactile Signals from Vision for Visuo-Tactile Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    FELT predicts finger pressure maps from RGB images and uses them or their learned features to improve manipulation policies without real tactile sensors at deployment.

  4. TactiDex: A Real-World Tactile-Guided Benchmark for Human-Like Dexterous Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    A tactile-rich HOI dataset plus a tri-component force reward improves contact fidelity and success of human-to-robot dexterous transfer over kinematic imitation alone.

  5. TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    A hierarchical robot manipulation policy uses tactile sensing both as a predictive subgoal generator and as a high-frequency residual correction signal, achieving 65% success on six contact-rich dexterous tasks versus...

  6. Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation

    cs.RO 2026-06 unverdicted novelty 6.0

    MiTaS fuses multi-resolution tactile data from GelSight and Evetac sensors with vision using modality-specific stems and transformer fusion to condition flow-matching policies, reporting 80% average success on five co...

  7. From Reach to Insert: Tactile-Augmented Precision Assembly under Sub-Millimeter Tolerances

    cs.RO 2026-05 unverdicted novelty 6.0

    A two-stage IL-RL method with tactile group sampling and a tactile critic achieves 67% success at 0.05 mm clearance while cutting max force by 60% and torque by 44%.

  8. FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception

    cs.RO 2026-04 conditional novelty 6.0

    FingerViP equips each finger with a miniature camera and trains a multi-view diffusion policy that achieves 80.8% success on real-world dexterous tasks previously limited by wrist-camera occlusion.

  9. TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation

    cs.RO 2026-07 conditional novelty 5.5

    A multi-timescale tactile hierarchy with subtask planning, tactile world-model goals, and residual refinement raises real-robot success by about 16–19 points over strong baselines on six contact-rich tasks.

  10. Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation

    cs.RO 2026-06 unverdicted novelty 5.0

    A visuo-tactile policy learning method that exploits tactile motion correlation for contact state distinction and Mixture-of-Transformers for cross-modal fusion.

  11. TacCoRL: Integrating Tactile Feedback into VLA via Simulation

    cs.RO 2026-06 unverdicted novelty 5.0

    TacCoRL integrates tactile feedback into VLA policies via real-aligned simulation co-training and RL, raising average success from 50% to 72.5% on four bimanual contact-rich tasks with direct real-robot transfer.