Pith. sign in

REVIEW 2 cited by

SDS -- See it, Do it, Sorted: Quadruped Skill Synthesis from Single Video Demonstration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.11571 v2 pith:I4HS3SXI submitted 2024-10-15 cs.RO

classification cs.RO
keywords engineeringinputlocomotionquadrupedsrewardsingleskillsorted
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Imagine a robot learning locomotion skills from any single video, without labels or reward engineering. We introduce SDS ("See it. Do it. Sorted."), an automated pipeline for skill acquisition from unstructured demonstrations. Using GPT-4o, SDS applies novel prompting techniques, in the form of spatio-temporal grid-based visual encoding ($G_{v}$) and structured input decomposition (SUS). These produce executable reward functions (RF) from the raw input videos. The RFs are used to train PPO policies and are optimized through closed-loop evolution, using training footage and performance metrics as self-supervised signals. SDS allows quadrupeds (e.g. Unitree Go1) to learn four gaits -- trot, bound, pace, and hop -- achieving 100% gait matching fidelity, Dynamic Time Warping (DTW) distance in the order of $10^{-6}$, and stable locomotion with zero failures, both in simulation and the real world. SDS generalizes to morphologically different quadrupeds (e.g. ANYmal) and outperforms prior work in data efficiency, training time and engineering effort. Further materials and the code are open-source under: https://rpl-cs-ucl.github.io/SDSweb/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. APEX: Action Priors Enable Efficient Exploration for Robust Motion Tracking on Legged Robots

    cs.RO 2025-05 conditional novelty 5.0 of 10

    APEX trains gait-tracking policies with decaying action priors and separate style and task critics, achieving reference-free deployment, faster convergence, and reward-robustness over DeepMimic.

  2. MA-ROESL: Motion-aware Rapid Reward Optimization for Efficient Robot Skill Learning from Single Videos

    cs.RO 2025-05 conditional novelty 5.0 of 10

    MA-ROESL cuts training time for learning quadruped gaits from single videos by about 68% using motion-aware frame selection and offline-to-online RL.

Pith tools