REVIEW 11 cited by
ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data
read the original abstract
Audio signals provide rich information for the robot interaction and object properties through contact. This information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio data in robot manipulation has been constrained to teleoperated demonstrations collected by either attaching a microphone to the robot or object, which significantly limits its usage in robot learning pipelines. In this work, we introduce ManiWAV: an 'ear-in-hand' data collection device to collect in-the-wild human demonstrations with synchronous audio and visual feedback, and a corresponding policy interface to learn robot manipulation policy directly from the demonstrations. We demonstrate the capabilities of our system through four contact-rich manipulation tasks that require either passively sensing the contact events and modes, or actively sensing the object surface materials and states. In addition, we show that our system can generalize to unseen in-the-wild environments by learning from diverse in-the-wild human demonstrations.
Forward citations
Cited by 11 Pith papers
-
A-SLIP: Acoustic Sensing for Continuous In-hand Slip Estimation
A four-microphone acoustic system with a CNN achieves 14.1-degree mean directional error for continuous in-hand slip estimation and outperforms single-channel baselines.
-
Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction
Introduces the Kaiwu multimodal dataset and framework with 11,664 synchronized assembling demonstrations including hand motions, pressures, sounds, multi-view videos, motion capture, eye gaze, and EMG signals with tim...
-
Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force
MuSe adapts a vision-action policy to force-torque via multi-stage fusion, multisensory future prediction, and experience replay, improving both new contact-rich tasks and original vision tasks.
-
Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force
MuSe adapts vision-only visuomotor policies to force-torque sensing with multi-stage fusion, multisensory future prediction, and experience replay, showing strong performance on contact-rich tasks while preserving ori...
-
VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity
VibeAct bridges real vibro-acoustic sensing and sim-based RL via a shared contact/slip representation, outperforming proprioception baselines on contact-rich dexterous tasks with successful real-world transfer.
-
RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning
A wearable interface with a shared dexterous hand module enables retargeting-free teleoperation and matched data collection, yielding policies with 88.75% average success across eight real-robot tasks that generalize ...
-
OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform
OpenEAI-Platform delivers an open-source low-cost robotic arm and VLA model that outperforms commercial arms and matches large pretrained baselines on four real-world manipulation tasks using limited open data.
-
Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force
MuSe adapts vision-only pretrained visuomotor policies to force-torque sensing via multi-stage fusion, multisensory future prediction, and experience replay, achieving strong contact-rich performance while preserving ...
-
World Models for Robotic Manipulation: A Survey
Survey organizing world models for robotic manipulation into representation families, a functional taxonomy, and infrastructure roles across pretraining, post-training, and inference, while reviewing 34 datasets and e...
-
You're Pushing My Buttons: Instrumented Learning of Gentle Button Presses
Training-time instrumentation with audio and privileged button-state signals produces contact policies that match success rates but apply lower forces using only vision and audio at inference.
-
Causal World Modeling for Robot Control
LingBot-VA combines video world modeling with policy learning via Mixture-of-Transformers, closed-loop rollouts, and asynchronous inference to improve robot manipulation in simulation and real settings.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.