Pith. sign in

REVIEW 7 cited by

R+X: Retrieval and Execution from Everyday Human Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12957 v2 pith:4JXPGGJZ submitted 2024-07-17 cs.RO cs.LG

R+X: Retrieval and Execution from Everyday Human Videos

classification cs.RO cs.LG
keywords videoseverydayhumanskillsbehaviourexecutionin-contextlanguage
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present R+X, a framework which enables robots to learn skills from long, unlabelled, first-person videos of humans performing everyday tasks. Given a language command from a human, R+X first retrieves short video clips containing relevant behaviour, and then executes the skill by conditioning an in-context imitation learning method (KAT) on this behaviour. By leveraging a Vision Language Model (VLM) for retrieval, R+X does not require any manual annotation of the videos, and by leveraging in-context learning for execution, robots can perform commanded skills immediately, without requiring a period of training on the retrieved videos. Experiments studying a range of everyday household tasks show that R+X succeeds at translating unlabelled human videos into robust robot skills, and that R+X outperforms several recent alternative methods. Videos and code are available at https://www.robot-learning.uk/r-plus-x.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

    cs.RO 2026-05 unverdicted novelty 6.0

    A dual-contrastive disentanglement method factorizes videos into independent task and embodiment latents, then uses a parameter-efficient adapter on a frozen video diffusion model to synthesize robot executions from s...

  2. A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics

    cs.RO 2026-04 unverdicted novelty 6.0

    A two-level hierarchical vector quantization tokenizer that clusters actions spatially and temporally achieves new state-of-the-art results in in-context imitation learning for robotics.

  3. WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations

    cs.RO 2026-04 unverdicted novelty 6.0

    WARPED synthesizes realistic wrist-view observations from monocular egocentric human videos via foundation models, hand-object tracking, retargeting, and Gaussian Splatting to train visuomotor policies that match tele...

  4. Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

    cs.RO 2024-09 unverdicted novelty 6.0

    Gen2Act enables generalizable robot manipulation for unseen objects and novel motions by using zero-shot human video generation from web data to condition a policy trained on an order of magnitude less robot interaction data.

  5. A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics

    cs.RO 2026-04 unverdicted novelty 5.0

    A two-level vector quantization tokenizer that jointly reconstructs robot actions and timestamps to improve in-context imitation learning performance on manipulation benchmarks.

  6. A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics

    cs.RO 2026-04 unverdicted novelty 5.0

    A two-level hierarchical vector-quantization tokenizer that recovers actions and timestamps sets a claimed SOTA for in-context robot imitation learning.

  7. VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation

    cs.RO 2025-09 unverdicted novelty 5.0

    VLBiMan framework enables generalizable bimanual manipulation from single human demonstrations via vision-language anchored task decomposition and adaptation without retraining.