Pith. sign in

REVIEW 7 cited by

ARCap: Collecting High-quality Human Demonstrations for Robot Learning with Augmented Reality Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08464 v1 pith:VRHOSXSQ submitted 2024-10-11 cs.RO cs.AI

classification cs.ROcs.AI
keywords arcapdatarobotcollectiondemonstrationsfeedbackmanipulationaugmented
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent progress in imitation learning from human demonstrations has shown promising results in teaching robots manipulation skills. To further scale up training datasets, recent works start to use portable data collection devices without the need for physical robot hardware. However, due to the absence of on-robot feedback during data collection, the data quality depends heavily on user expertise, and many devices are limited to specific robot embodiments. We propose ARCap, a portable data collection system that provides visual feedback through augmented reality (AR) and haptic warnings to guide users in collecting high-quality demonstrations. Through extensive user studies, we show that ARCap enables novice users to collect robot-executable data that matches robot kinematics and avoids collisions with the scenes. With data collected from ARCap, robots can perform challenging tasks, such as manipulation in cluttered environments and long-horizon cross-embodiment manipulation. ARCap is fully open-source and easy to calibrate; all components are built from off-the-shelf products. More details and results can be found on our website: https://stanford-tml.github.io/ARCap

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Robot-free HiFi-UMI demonstrations can replace teleoperated real-robot data in post-training: three policy backbones matched in-domain teleoperation within 3.1 percentage points, including 85% success on a precision i...

  2. Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Dexplore learns dexterous robotic hand control from human MoCap demonstrations by treating them as soft, adaptively shrinking spatial references, then distills the policy into a vision-based controller.

  3. Learning human-to-robot handovers through 3D scene reconstruction

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A handover policy trained only on images rendered from a sparse-view Gaussian Splatting scene can deploy on a real robot without real-robot training data.

  4. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  5. TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A type-guided teleoperation system that selects predefined dexterous hand poses with a language model outperforms retargeting-based teleoperation on nine real-world tasks and improves imitation learning success.

  6. Where Do Humans Look When Demonstrating to Robots? Human Gaze Behavior in Pick-and-Place Tasks Across Demonstration Devices

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Demonstration devices that emulate a robot's body or camera view shift human gaze away from task objects toward the end-effector, and gaze from natural wearable recording improves imitation learning robustness.

  7. TacPrint: A Wearable Fingertip Tactile Sensor for Human-to-Robot Contact Reproduction

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A low-cost wearable fingertip sensor estimates dense contact-depth maps from 24 capacitive channels and uses them to substantially improve robot grasping and wiping in human-to-robot replay.

Pith tools