Pith. sign in

REVIEW 2 cited by

VIEW: Visual Imitation Learning with Waypoints

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.17906 v3 pith:4IGZ4FZV submitted 2024-04-27 cs.RO

classification cs.RO
keywords viewvideolearningtasksvisualdemonstrationsefficiencyimitation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Robots can use Visual Imitation Learning (VIL) to learn manipulation tasks from video demonstrations. However, translating visual observations into actionable robot policies is challenging due to the high-dimensional nature of video data. This challenge is further exacerbated by the morphological differences between humans and robots, especially when the video demonstrations feature humans performing tasks. To address these problems we introduce Visual Imitation lEarning with Waypoints (VIEW), an algorithm that significantly enhances the sample efficiency of human-to-robot VIL. VIEW achieves this efficiency using a multi-pronged approach: extracting a condensed prior trajectory that captures the demonstrator's intent, employing an agent-agnostic reward function for feedback on the robot's actions, and utilizing an exploration algorithm that efficiently samples around waypoints in the extracted trajectory. VIEW also segments the human trajectory into grasp and task phases to further accelerate learning efficiency. Through comprehensive simulations and real-world experiments, VIEW demonstrates improved performance compared to current state-of-the-art VIL methods. VIEW enables robots to learn manipulation tasks involving multiple objects from arbitrarily long video demonstrations. Additionally, it can learn standard manipulation tasks such as pushing or moving objects from a single video demonstration in under 30 minutes, with fewer than 20 real-world rollouts. Code and videos here: https://collab.me.vt.edu/view/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robotic Visual Instruction

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Hand-drawn symbolic sketches (arrows and circles) can serve as precise, silent robot instructions, and a vision-language pipeline can execute them on unseen tasks.

  2. Efficient Sensorimotor Learning for Open-world Robot Manipulation

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A PhD dissertation argues that object, spatial, and behavioral regularities, extracted with foundation models, enable data-efficient, generalizable robot manipulation, and presents seven systems and a benchmark built ...

Pith tools