Pith. sign in

REVIEW 11 cited by

ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19464 v2 pith:QPCYXW4R submitted 2024-06-27 cs.RO cs.AIcs.CVcs.SDeess.AS

ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

classification cs.RO cs.AIcs.CVcs.SDeess.AS
keywords robotmanipulationdemonstrationsin-the-wildlearningaudiodatainformation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Audio signals provide rich information for the robot interaction and object properties through contact. This information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio data in robot manipulation has been constrained to teleoperated demonstrations collected by either attaching a microphone to the robot or object, which significantly limits its usage in robot learning pipelines. In this work, we introduce ManiWAV: an 'ear-in-hand' data collection device to collect in-the-wild human demonstrations with synchronous audio and visual feedback, and a corresponding policy interface to learn robot manipulation policy directly from the demonstrations. We demonstrate the capabilities of our system through four contact-rich manipulation tasks that require either passively sensing the contact events and modes, or actively sensing the object surface materials and states. In addition, we show that our system can generalize to unseen in-the-wild environments by learning from diverse in-the-wild human demonstrations.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A-SLIP: Acoustic Sensing for Continuous In-hand Slip Estimation

    cs.RO 2026-04 unverdicted novelty 7.0

    A four-microphone acoustic system with a CNN achieves 14.1-degree mean directional error for continuous in-hand slip estimation and outperforms single-channel baselines.

  2. Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction

    cs.RO 2025-03 unverdicted novelty 7.0

    Introduces the Kaiwu multimodal dataset and framework with 11,664 synchronized assembling demonstrations including hand motions, pressures, sounds, multi-view videos, motion capture, eye gaze, and EMG signals with tim...

  3. Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force

    cs.RO 2026-06 conditional novelty 6.0

    MuSe adapts a vision-action policy to force-torque via multi-stage fusion, multisensory future prediction, and experience replay, improving both new contact-rich tasks and original vision tasks.

  4. Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force

    cs.RO 2026-06 unverdicted novelty 6.0

    MuSe adapts vision-only visuomotor policies to force-torque sensing with multi-stage fusion, multisensory future prediction, and experience replay, showing strong performance on contact-rich tasks while preserving ori...

  5. VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity

    cs.RO 2026-06 unverdicted novelty 6.0

    VibeAct bridges real vibro-acoustic sensing and sim-based RL via a shared contact/slip representation, outperforming proprioception baselines on contact-rich dexterous tasks with successful real-world transfer.

  6. RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning

    cs.RO 2026-06 unverdicted novelty 6.0

    A wearable interface with a shared dexterous hand module enables retargeting-free teleoperation and matched data collection, yielding policies with 88.75% average success across eight real-robot tasks that generalize ...

  7. OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform

    cs.RO 2026-06 conditional novelty 6.0

    OpenEAI-Platform delivers an open-source low-cost robotic arm and VLA model that outperforms commercial arms and matches large pretrained baselines on four real-world manipulation tasks using limited open data.

  8. Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force

    cs.RO 2026-06 unverdicted novelty 5.0

    MuSe adapts vision-only pretrained visuomotor policies to force-torque sensing via multi-stage fusion, multisensory future prediction, and experience replay, achieving strong contact-rich performance while preserving ...

  9. World Models for Robotic Manipulation: A Survey

    cs.RO 2026-05 accept novelty 5.0

    Survey organizing world models for robotic manipulation into representation families, a functional taxonomy, and infrastructure roles across pretraining, post-training, and inference, while reviewing 34 datasets and e...

  10. You're Pushing My Buttons: Instrumented Learning of Gentle Button Presses

    cs.RO 2026-04 unverdicted novelty 5.0

    Training-time instrumentation with audio and privileged button-state signals produces contact policies that match success rates but apply lower forces using only vision and audio at inference.

  11. Causal World Modeling for Robot Control

    cs.CV 2026-01 unverdicted novelty 5.0

    LingBot-VA combines video world modeling with policy learning via Mixture-of-Transformers, closed-loop rollouts, and asynchronous inference to improve robot manipulation in simulation and real settings.