Pith. sign in

REVIEW 18 cited by

ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.06156 v2 pith:7AXBNEUM submitted 2025-04-08 cs.RO

ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface

classification cs.RO
keywords manipulationtaskstactilelearningvitamincontact-richdatagripper
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Tactile information plays a crucial role for humans and robots to interact effectively with their environment, particularly for tasks requiring the understanding of contact properties. Solving such dexterous manipulation tasks often relies on imitation learning from demonstration datasets, which are typically collected via teleoperation systems and often demand substantial time and effort. To address these challenges, we present ViTaMIn, an embodiment-free manipulation interface that seamlessly integrates visual and tactile sensing into a hand-held gripper, enabling data collection without the need for teleoperation. Our design employs a compliant Fin Ray gripper with tactile sensing, allowing operators to perceive force feedback during manipulation for more intuitive operation. Additionally, we propose a multimodal representation learning strategy to obtain pre-trained tactile representations, improving data efficiency and policy robustness. Experiments on seven contact-rich manipulation tasks demonstrate that ViTaMIn significantly outperforms baseline methods, demonstrating its effectiveness for complex manipulation tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design

    cs.RO 2026-07 conditional novelty 7.0

    A single diffusion transformer trains on tokenized robot bodies and motions to generate and optimize robot designs for unseen rewards and trajectories, outpacing evolutionary search in speed and often in reward.

  2. FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors for Contact-Rich Manipulation

    cs.RO 2026-06 unverdicted novelty 7.0

    FTP-1 is the first foundation tactile policy pretrained on ~3000 hours of data from 26 sources across 21 sensors that improves performance on seen setups by 17.2% and transfers to unseen sensors with 31% success rate gain.

  3. TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

    cs.RO 2026-01 unverdicted novelty 7.0

    TouchGuide improves contact-rich robot manipulation by steering diffusion or flow-matching visuomotor policies with tactile feasibility scores from a contrastively trained Contact Physical Model.

  4. S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

    cs.RO 2026-07 conditional novelty 6.0

    S2A2 adds microphone-array spatial audio and spectrograms to imitation-learning policies, substantially improving success on manipulation tasks where vision alone cannot identify the target or destination.

  5. $N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

    cs.RO 2026-07 conditional novelty 6.0

    A VLA with predictive latent tactile tokens pretrained on large-scale visuo-tactile data, plus ALTER offline advantage labeling, leads contact-rich real and sim benchmarks.

  6. Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

    cs.RO 2026-07 conditional novelty 6.0

    Pre-training a VLA model on 100k hours of auto-labeled UMI trajectories, then post-training on robot data, yields SOTA simulated manipulation and data-efficient fine-tuning.

  7. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  8. VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity

    cs.RO 2026-06 unverdicted novelty 6.0

    VibeAct bridges real vibro-acoustic sensing and sim-based RL via a shared contact/slip representation, outperforming proprioception baselines on contact-rich dexterous tasks with successful real-world transfer.

  9. RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning

    cs.RO 2026-06 unverdicted novelty 6.0

    A wearable interface with a shared dexterous hand module enables retargeting-free teleoperation and matched data collection, yielding policies with 88.75% average success across eight real-robot tasks that generalize ...

  10. OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform

    cs.RO 2026-06 conditional novelty 6.0

    OpenEAI-Platform delivers an open-source low-cost robotic arm and VLA model that outperforms commercial arms and matches large pretrained baselines on four real-world manipulation tasks using limited open data.

  11. HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations

    cs.RO 2026-03 unverdicted novelty 6.0

    HoMMI learns whole-body mobile manipulation policies from robot-free human demonstrations by augmenting UMI with egocentric sensing and bridging the embodiment gap through an agnostic visual representation, relaxed he...

  12. ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

    cs.RO 2026-07 conditional novelty 5.0

    An action-conditioned visuo-tactile world model generates synthetic camera-plus-touch rollouts that, mixed with real demonstrations, improve downstream contact-rich manipulation policies.

  13. Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation

    cs.RO 2026-06 unverdicted novelty 5.0

    A visuo-tactile policy learning method that exploits tactile motion correlation for contact state distinction and Mixture-of-Transformers for cross-modal fusion.

  14. OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction

    cs.RO 2026-04 unverdicted novelty 5.0

    OmniUMI introduces a multimodal handheld interface that synchronously records RGB, depth, trajectory, tactile, internal grasp force, and external wrench data for training diffusion policies on contact-rich robot manipulation.

  15. Causal World Modeling for Robot Control

    cs.CV 2026-01 unverdicted novelty 5.0

    LingBot-VA combines video world modeling with policy learning via Mixture-of-Transformers, closed-loop rollouts, and asynchronous inference to improve robot manipulation in simulation and real settings.

  16. NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation

    cs.RO 2026-06 unverdicted novelty 4.0

    A multimodal transformer fuses RGB-D vision and proprioception to predict binary contact states, supporting RL agents for in-hand reorientation that generalize to novel objects in simulation and on a real robot.

  17. Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

    cs.AI 2025-09 conditional novelty 4.0

    Agent spatial intelligence is organized into six neuroscience-inspired modules, and the field is reviewed through that lens without any experimental validation.

  18. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.