Pith. sign in

REVIEW 3 major objections 5 minor 102 references

Reinforcement Learning from Wild Animal Videos

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A video classifier trained on 8,791 wild-animal clips supplies the reward that teaches a quadruped robot to walk, jump, and stand still, without reference trajectories or per-skill reward functions; the policy transfers to a real robot.

desk verdict A genuine existence proof that wild-animal video classifiers can reward quadruped locomotion, but the reward grounding on robot renders is unvalidated and the walking/running distinction is weak. read the letter →

arxiv 2412.04273 v1 pith:S6FNZAKX submitted 2024-12-05 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords reinforcementlearningleggedlocomotionvideo-basedrewardcross-embodimenttransferactionrecognitionquadrupedrobotsim-to-realanimalvideos
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a legged robot can learn locomotion skills by watching internet footage of wild animals, and answers yes. The authors train a video classifier on 8,791 clips from a large animal-behavior dataset to recognize four actions — keeping still, walking, running, jumping — and then use the classifier's score on third-person videos of a simulated quadruped as the only task reward in reinforcement learning. Physical plausibility comes entirely from embodiment-agnostic constraints, not from skill-specific reward design. The resulting multi-skill policy transfers directly to a real Solo-12 robot, which trots forward on command, jumps in place, and stands still. The intended conclusion is that raw, cross-embodiment video data can substitute for hand-designed rewards and reference motions in locomotion learning.

What carries the argument

The central object is the classifier-as-reward loop: a frozen video classifier $f_\theta$ (a Uniformer-S transformer fine-tuned from Kinetics-400 pretrained weights, with random-convolution augmentations and weight-averaged model soups) turns pixels into a reward $R(s_t,a_t) = \alpha_y f_\theta(x_{\mathrm{robot}}, y) + \beta_y$ on frames captured every five simulation steps, where $y$ is the commanded skill and $\alpha_y, \beta_y$ are per-skill normalization terms. The second load-bearing piece is constrained reinforcement learning: Proximal Policy Optimization combined with Constraints as Terminations (CaT), where every constraint is independent of the skill command so that behavioral differences arise only from the video reward. A symmetry loss on the policy is added because the camera sees the robot from one side only. The argument therefore rests on a division of labor: the video classifier supplies what each skill means, and the embodiment-agnostic constraints supply what is physically possible for the robot.

What would settle it

Feed the trained classifier robot clips with the body masked or blurred while the background and camera motion are preserved: if the 'Walking' or 'Running' probability stays high, the reward is tracking visual confounds rather than the skill. A direct correlation measurement between the classifier score and ground-truth forward velocity or foot air time across policy rollouts would likewise reveal how much of the reward is skill-grounded, and the two-leg walking-in-place failure reported in the paper is a concrete case where the score stayed high while the skill was absent.

Watch

Extended reading notes

Core claim

The paper establishes that a video action classifier trained purely on animals in natural habitats can act as a valid scalar reward for robotic locomotion, bridging what the authors call the extreme gap in domain and embodiment between animals and robots. A Uniformer video model is fine-tuned on a single-label subset of the Animal Kingdom dataset covering keeping still, walking, running, and jumping, then frozen. During policy training in a physics simulator, a third-person camera records 8-frame clips of the robot and the classifier's probability for the commanded skill is used as the reward, with zero reward on non-rendering steps and per-skill reward normalization. Task-agnostic constraints (joint limits, torques, foot air time, base orientation) enforce physical plausibility and sim-to-real transferability, and a symmetry loss compensates for the single camera view. The authors report that distinct, recognizable behaviors emerge for each skill command and that the policy deploys on a physical Solo-12 robot, with walking and running appearing as trotting gaits and jumping as rhythmic in-place pumping with flight phases; they also concede that running never produces a proper flying phase and that the behaviors lag the state of the art in learning-based locomotion.

Load-bearing premise

The pipeline assumes that the video classifier, trained only on wild-animal clips, scores the simulated and real robot's movements for the right reason — recognizing the skill in the motion — rather than because of background, camera angle, body pose, or incidental visual artifacts.

Editorial extensions

If this is right

  • A video classifier trained on animal footage can reward a quadruped policy, so the same recipe should extend to other skills and other legged embodiments whenever the classifier can recognize the skill across morphologies.
  • The no-curating ablation shows the reward signal must be tailored to single labels: training the classifier on the full multi-label dataset with binary cross-entropy prevents walking and running from emerging at all.
  • The learned running is only a faster trot without an aerial phase, and walking and running look similar on the real robot; the paper attributes this gap to conventional video classification and standard on-policy RL rather than to the overall approach.
  • Sim-to-real transfer succeeds directly, without domain randomization, a result the authors attribute to the task-agnostic constraints — especially the foot air-time constraint — rather than to the reward function itself.
  • Because rewards fire every five steps from one camera, the symmetry loss and a camera angle that shows the full body are necessary for walking and running to emerge; extreme camera angles degrade or destroy individual skills.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the method turns reward engineering into a recognition problem, so the quality ceiling is set by how well the frozen classifier generalizes out of distribution; improvements in video understanding should translate directly into better locomotion skills without changing the RL loop.
  • Editorial inference: the two-leg walking-in-place failure reported for the walking skill shows the reward can be gamed by visual artifacts; a direct test would measure how strongly the classifier score correlates with forward velocity and foot contact across rollouts, and a fix would add multi-view or motion-focused scoring.
  • Editorial inference: the same scheme could point at other video corpora — human sports footage, finer-grained animal behavior labels, or egocentric clips — to produce rewards for skills that are hard to specify by hand, whenever the visual gap between source and robot is comparable to the one tested here.
  • Editorial inference: the sensitivity to dataset size (skills degrade when training drops below 50% of the 8,791 videos) suggests the reward's robustness scales with video diversity rather than with the RL algorithm, making internet-scale data the main lever for future work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces RLWAV, a pipeline in which a video action classifier trained on naturally occurring animal videos from the Animal Kingdom dataset is used as the scalar reward for a constrained reinforcement-learning policy that controls a Solo-12 quadruped in IsaacGym, after which the policy is transferred to the physical robot. The authors report simulation metrics across four skills, ablations of both the classifier and the policy-learning choices, and qualitative real-robot demonstrations. The central claim is that distinct locomotion skills can be acquired from wild animal videos without reference trajectories or skill-specific reward functions.

Significance. If the claim is upheld, the work is a useful step toward using internet-scale video corpora as reward sources for legged locomotion, complementing prior work on human-video-based manipulation rewards. Strengths include the use of a public dataset, multi-seed reporting, systematic ablations (classifier curation, model soup, camera position, constraint removal), and an honest discussion of failure cases and limitations. The main risks are that the learned reward is not demonstrated to be grounded in the intended animal actions when applied to synthetic robot renders, and that the real-robot evidence is qualitative; these gaps directly affect the strength of the central claim. There is no circularity concern because the reward is explicitly a learned classifier score and the evaluation metrics are independent, but the out-of-distribution validity of that classifier score is the load-bearing issue.

major comments (3)
  1. [§3.2, Eq. (4); §4.2, Fig. 4] The reward used for RL is the classifier score fθ(x_robot, y) on synthetic robot videos, but the classifier is only trained on, and as far as reported only evaluated on, Animal Kingdom natural videos. The paper provides no accuracy, confusion matrix, or per-class reliability of this classifier on robot-domain inputs. This matters because the only validation of the reward is through downstream policy behavior, and the paper's own Figure 4 documents a failure mode in which a walking-in-place policy 'fool[s] the video reward function.' That example shows that high classifier reward can be obtained without performing the commanded skill, so the central claim that the learned behaviors are grounded in animal motion is conditional on an unvalidated out-of-distribution generalization step. I recommend adding a direct evaluation of the video classifier on robot renders (e.g., per-skill classification accuracy or reward-versus-behavior correlation) and, if possible, a reward-hacking analysis or a calibration against the quantitative skill metrics.
  2. [§4.3] The real-robot experiments are qualitative: no measured forward speeds, jump heights, contact timings, or success rates are reported, and no video classifier scores or manual style ratings are taken on the real robot. As a result, the claim of 'successful transfer' and 'directly deploy' in §4.3 is not quantitatively supported. I ask for quantitative real-robot measurements (e.g., forward velocity, vertical displacement, and optionally stance/duty-factor estimates) and, if style is to be assessed, a small multi-rater study with reported inter-rater agreement.
  3. [§4.2, Table 1; §4.3] The distinction between walking and running is not established. In Table 1, walking velx = 23.0 ± 15.4 cm/s and running velx = 35.1 ± 12.9 cm/s overlap by more than one standard deviation, and the text concedes that the running policy does not produce flying phases and that on the real robot the two skills 'appear alike.' Since one of the four considered action classes is 'Running,' the abstract-level claim that the robot acquires skills corresponding to the action classes needs either a measurable gait criterion (e.g., duty factor, Froude number, flight-phase duration) or a softened claim that only three distinct behaviors are demonstrated.
minor comments (5)
  1. [§1, §3.2] There are several typos and minor awkward phrasings: 'activites' in §1, 'fishs' in §3.2, 'adaptatively optimized' in Appendix A.2, and the consistent spacing artifacts in 'RLW A V' throughout the captions. These should be corrected in a revision.
  2. [Figure 4 and Table 2] The alternative camera positions (Camera 1 through Camera 4) are described only as 'more extreme' and shown in a schematic; the exact placement relative to the robot, field of view, and distance are not specified, making the camera ablation difficult to reproduce. Please provide numerical camera parameters or a precise description in the appendix.
  3. [Appendix A.2, Eq. (6)] The adaptive reward normalization coefficients α_y and β_y are said to be 'adaptatively optimized' based on reward statistics across actors, but the precise update rule, initialization, and effect on the learned reward scale are not described. Since these coefficients are per-skill free parameters, a brief specification would improve reproducibility.
  4. [§4.1] The style scores are presented as single numbers without stating how many raters evaluated the rollout videos, whether the ratings were blind to the skill command, or whether any inter-rater agreement was computed. This makes the style column difficult to interpret as evidence of skill quality.
  5. [§3.3, Appendix A.2] The claim 'without skill-specific reward functions' should be read carefully: the foot air-time constraint is applied uniformly across skills but is specifically designed to enable walking and running (Appendix B shows its removal destroys those skills). This is not a fatal issue, but the wording in §1 and §5 should acknowledge that the method still relies on locomotion-specific inductive biases.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the video-classifier reward is an explicit learned input, and skill acquisition is evaluated with independent physical metrics and ablations.

full rationale

The load-bearing claim is that a video classifier trained on Animal Kingdom clips can serve as a reward for training a quadruped policy that exhibits distinct locomotion skills without reference trajectories or skill-specific rewards. The reward is explicitly defined in Section 3.3, Eq. (4), as R(s_t,a_t)=f_theta(x_robot,y), i.e., the classifier score on synthetic robot renders. This is a learned input, not a hidden derivation: the paper never claims the classifier score itself proves skill acquisition. Instead, Section 4.1 defines independent evaluation metrics (|velxy| for keeping still, velx for walking/running, Delta z for jumping, plus manual style ratings), and all reported results in Tables 1 and 2 use those quantities. Crucially, Figure 4 documents a failure case where the robot 'fool[s] the video reward function,' showing the authors explicitly treat high classifier score as insufficient evidence of the commanded skill. This makes the central validation non-circular: the policy is not evaluated by the same objective it optimizes. The skill command y is a one-hot input to the policy, and the constraints are task-agnostic, with differences across skills attributed solely to the learned reward. The self-citations ([20] CaT, [43] SoloParkour, [90] video-conditioned policies) are algorithm/tool citations and related-work context; they are not invoked to justify the video-to-robot grounding claim. In particular, CaT is used to enforce constraints but is not the source of the skill definitions. The paper's principal weakness, namely that the video classifier is never directly evaluated on robot renders and may be exploited via domain shift, is a soundness and generalization concern, not a circularity. The derivation chain is therefore self-contained relative to its stated inputs, and no circular step can be exhibited from the paper's own equations.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claim rests on transfer assumptions rather than mathematical derivation. The main free parameters are design choices in the reward pipeline: camera placement, image update interval, air-time threshold, and adaptive reward normalization. The key axioms are that Animal Kingdom labels are correct, that a Kinetics-pretrained Uniformer generalizes from wild animals to robot renders, that IsaacGym is a faithful enough surrogate for Solo-12, and that manual style ratings capture skill identity. No new physical entities are introduced.

free parameters (6)
  • Reward normalization coefficients alpha_y and beta_y per skill = not reported
    Equation 6; adaptively optimized based on reward statistics across actors; handles different reward scales and is fit during RL.
  • Third-person camera position = nominal base camera; Camera 1-4 ablated in Table 2
    Section 3.3 and Figure 4; camera view strongly affects which skills emerge and their quality.
  • Image generation interval = every 5 RL steps (8 in the 'update 8' ablation)
    Section 3.3; the ablation shows that an 8-step interval degrades jumping, so this choice affects the central result.
  • Foot air-time constraint threshold t_des_air_time = not reported
    Appendix A.2, Table 3; removing this constraint destroys walking and running (Appendix B), so the threshold is load-bearing.
  • Constraint thresholds (torque, joint limits, orientation, contact force) = not reported
    Appendix A.2, Table 3; standard values inherited from prior work are not given in the text.
  • Video classifier training hyperparameters = lr 3e-5, batch 64, 8 frames, stride 4, stochastic depth in {0.3, 0.4}, weight decay in {0.05, 0.01, 0.1}, epochs in…
    Appendix A.1; these are chosen by the authors and affect reward quality, though they are not the central claim.
assumptions (6)
  • domain assumption Animal Kingdom multi-label annotations are reliable enough to train a four-class single-label classifier.
    Section 3.2: 8,791 videos are curated by labels; no label verification or inter-annotator agreement is reported.
  • domain assumption A Uniformer pretrained on Kinetics-400 and fine-tuned on animals will generalize zero-shot to simulated robot videos.
    Sections 3.2 and 4.2; this is the load-bearing transfer assumption, mitigated only by augmentations and model soup, never tested on robot videos.
  • domain assumption IsaacGym accurately models Solo-12 dynamics and contact so that policies transfer without domain randomization.
    Sections 3.3 and 4.3; sim-to-real transfer is direct, and the paper notes foot slippage as a sim-to-real artifact.
  • domain assumption Manual style ratings from 0 to 1 are a valid measure of whether a motion is identifiable as the target skill.
    Section 4.1; ratings are used as an evaluation metric, but no inter-rater reliability or variance is reported.
  • domain assumption The same fixed constraints (air time, symmetry, torques, orientation) are sufficient physical grounding for all skills.
    Section 3.3; the air-time constraint is shown in Appendix B to be critical for walking and running, so the 'no skill-specific rewards' claim still relies on skill-specific physics priors.
  • ad hoc to paper A single third-person camera capturing 8-frame 128x128 videos in simulation provides sufficient information for skill identification.
    Section 3.3; camera position is ablated in Table 2, and different positions strongly affect which skills emerge, indicating the reward is view-sensitive.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reinforcement Learning from Wild Animal Videos." pith.science (2026). https://pith.science/paper/S6FNZAKX

@misc{pith2026241204273,
  author       = {Pith},
  title        = {Pith review of: Reinforcement Learning from Wild Animal Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S6FNZAKX}},
  note         = {Machine review of arXiv:2412.04273}
}
read the original abstract

We propose to learn legged robot locomotion skills by watching thousands of wild animal videos from the internet, such as those featured in nature documentaries. Indeed, such videos offer a rich and diverse collection of plausible motion examples, which could inform how robots should move. To achieve this, we introduce Reinforcement Learning from Wild Animal Videos (RLWAV), a method to ground these motions into physical robots. We first train a video classifier on a large-scale animal video dataset to recognize actions from RGB clips of animals in their natural habitats. We then train a multi-skill policy to control a robot in a physics simulator, using the classification score of a third-person camera capturing videos of the robot's movements as a reward for reinforcement learning. Finally, we directly transfer the learned policy to a real quadruped Solo. Remarkably, despite the extreme gap in both domain and embodiment between animals in the wild and robots, our approach enables the policy to learn diverse skills such as walking, jumping, and keeping still, without relying on reference trajectories nor skill-specific rewards.

Figures

Figures reproduced from arXiv: 2412.04273 by the authors.

Figure 1
Figure 1. RLWAV trains a video classifier on 8,791 wild animal videos in natural environments to [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. (Left) We train a video classifier to recognize actions from the Animal Kingdom [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Impact of animal video dataset size (% of total videos) on skill emergence. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: (Top) Example of failure case for the ”Walking” skill, where the robot performs walking [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Rollout examples for the 4 skills considered on the real Solo-12. [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Rollout examples in simulation for the 4 skills considered. [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Additional illustrations of videos from the Animal Kingdom dataset [ [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

102 extracted references · 33 canonical work pages

  1. [1]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language super- vision. In International conference on machine learning, pages 8748–8763. PMLR, 2021

  2. [2]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image syn- thesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022

  3. [3]

    Y . Wang, K. Li, Y . Li, Y . He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y . Liu, Z. Wang, et al. Internvideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191, 2022

  4. [4]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  5. [5]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 10

  6. [6]

    Molenberghs, R

    P. Molenberghs, R. Cunnington, and J. B. Mattingley. Is the mirror neuron system involved in imitation? a short review and meta-analysis. Neuroscience & Biobehavioral Reviews, 33 (7):975–980, 2009. ISSN 0149-7634. doi:https://doi.org/10.1016/j.neubiorev.2009.03.010

  7. [7]

    L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg. Concept2robot: Learning manip- ulation concepts from instructions and human demonstrations. The International Journal of Robotics Research, 40(12-14):1419–1434, 2021

  8. [8]

    S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak. Affordances from human videos as a versatile representation for robotics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13778–13790, 2023

Show all 102 references
  1. [9]

    Bharadhwaj, D

    H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation. arXiv preprint arXiv:2409.16283, 2024

  2. [10]

    L. Feng, Y . Zhao, Y . Sun, W. Zhao, and J. Tang. Action recognition using a spatial-temporal network for wild felines. Animals, 11(2):485, 2021

  3. [11]

    X. L. Ng, K. E. Ong, Q. Zheng, Y . Ni, S. Y . Yeo, and J. Liu. Animal kingdom: A large and diverse dataset for animal behavior understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19023–19034, 2022

  4. [12]

    J. Chen, M. Hu, D. J. Coker, M. L. Berumen, B. Costelloe, S. Beery, A. Rohrbach, and M. Elhoseiny. Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, ...

  5. [13]

    X. Peng, E. Coumans, T. Zhang, T. Lee, J. Tan, and S. Levine. Learning agile robotic loco- motion skills by imitating animals. arXiv preprint arXiv:2004.00784, 2020

  6. [14]

    Bohez, S

    S. Bohez, S. Tunyasuvunakool, P. Brakel, F. Sadeghi, L. Hasenclever, Y . Tassa, E. Parisotto, J. Humplik, T. Haarnoja, R. Hafner, et al. Imitate and repurpose: Learning reusable robot movement skills from human and animal behaviors. arXiv preprint arXiv:2203.17138, 2022

  7. [15]

    L. Han, Q. Zhu, J. Sheng, C. Zhang, T. Li, Y . Zhang, H. Zhang, Y . Liu, C. Zhou, R. Zhao, et al. Lifelike agility and play in quadrupedal robots using reinforcement learning and generative pre-trained models. Nature Machine Intelligence, 6(7):787–798, 2024

  8. [16]

    Makoviychuk, L

    V . Makoviychuk, L. Wawrzyniak, Y . Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State. Isaac gym: High performance gpu-based physics simulation for robot learning, 2021

  9. [17]

    Rudin, D

    N. Rudin, D. Hoeller, P. Reist, and M. Hutter. Learning to walk in minutes using massively parallel deep reinforcement learning. In Conference on Robot Learning. PMLR, 2021

  10. [18]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimiza- tion algorithms. arXiv preprint arXiv:1707.06347, 2017

  11. [19]

    Y . Kim, H. Oh, J. Lee, J. Choi, G. Ji, M. Jung, D. Youm, and J. Hwangbo. Not only re- wards but also constraints: Applications on legged robot locomotion. IEEE Transactions on Robotics, 2024

  12. [20]

    Chane-Sane, P.-A

    E. Chane-Sane, P.-A. Leziart, T. Flayols, O. Stasse, P. Sou `eres, and N. Mansard. Cat: Con- straints as terminations for legged locomotion reinforcement learning. In IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems (IROS), 2024

  13. [21]

    Todorov, T

    E. Todorov, T. Erez, and Y . Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pages 5026–

  14. [22]

    C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem. Brax - a differentiable physics engine for large scale rigid body simulation, 2021

  15. [23]

    X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel. Sim-to-real transfer of robotic control with dynamics randomization. In 2018 IEEE international conference on robotics and automation (ICRA), pages 3803–3810. IEEE, 2018

  16. [24]

    J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning quadrupedal locomo- tion over challenging terrain. Science robotics, 5(47):eabc5986, 2020

  17. [25]

    Z. Fu, A. Kumar, J. Malik, and D. Pathak. Minimizing energy consumption leads to the emergence of gaits in legged robots. In Conference on Robot Learning (CoRL), 2021

  18. [26]

    Z. Li, X. B. Peng, P. Abbeel, S. Levine, G. Berseth, and K. Sreenath. Robust and ver- satile bipedal jumping control through multi-task reinforcement learning. arXiv preprint arXiv:2302.09450, 1, 2023

  19. [27]

    Aractingi, P.-A

    M. Aractingi, P.-A. L ´eziart, T. Flayols, J. Perez, T. Silander, and P. Sou`eres. Controlling the solo12 quadruped robot with deep reinforcement learning. Scientific Reports, 13(1):11945, 2023

  20. [28]

    Bellegarda, Y

    G. Bellegarda, Y . Chen, Z. Liu, and Q. Nguyen. Robust high-speed running for quadruped robots via deep reinforcement learning. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022

  21. [29]

    G. B. Margolis, G. Yang, K. Paigwar, T. Chen, and P. Agrawal. Rapid locomotion via rein- forcement learning. The International Journal of Robotics Research, 43(4):572–587, 2024

  22. [30]

    T. He, C. Zhang, W. Xiao, G. He, C. Liu, and G. Shi. Agile but safe: Learning collision-free high-speed legged locomotion. arXiv preprint arXiv:2401.17583, 2024

  23. [31]

    Yang and J

    I. Yang and J. Hwangbo. Learning rapid turning, aerial reorientation, and balancing using manipulator as a tail. arXiv preprint arXiv:2407.10420, 2024

  24. [32]

    Bellegarda, C

    G. Bellegarda, C. Nguyen, and Q. Nguyen. Robust quadruped jumping via deep reinforce- ment learning. arXiv preprint arXiv:2011.07089, 2020

  25. [33]

    G. B. Margolis, T. Chen, K. Paigwar, X. Fu, D. Kim, S. Kim, and P. Agrawal. Learning to jump from pixels. arXiv preprint arXiv:2110.15344, 2021

  26. [34]

    Smith, J

    L. Smith, J. C. Kew, T. Li, L. Luu, X. B. Peng, S. Ha, J. Tan, and S. Levine. Learning and adapting agile locomotion skills by transferring experience.Proceedings of Robotics: Science and Systems, 2023

  27. [35]

    Zhang, J

    C. Zhang, J. Sheng, T. Li, H. Zhang, C. Zhou, Q. Zhu, R. Zhao, Y . Zhang, and L. Han. Learn- ing highly dynamic behaviors for quadrupedal robots. arXiv preprint arXiv:2402.13473 , 2024

  28. [36]

    T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning robust perceptive locomotion for quadrupedal robots in the wild. Science Robotics, 7(62):eabk2822, 2022

  29. [37]

    Hoeller, N

    D. Hoeller, N. Rudin, C. Choy, A. Anandkumar, and M. Hutter. Neural scene representation for locomotion on structured terrain.IEEE Robotics and Automation Letters, 7(4):8667–8674, 2022

  30. [38]

    Agarwal, A

    A. Agarwal, A. Kumar, J. Malik, and D. Pathak. Legged locomotion in challenging terrains using egocentric vision. In Conference on robot learning, pages 403–415. PMLR, 2023. 12

  31. [39]

    R. Yang, G. Yang, and X. Wang. Neural volumetric memory for visual locomotion control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1430–1440, 2023

  32. [40]

    Zhuang, Z

    Z. Zhuang, Z. Fu, J. Wang, C. Atkeson, S. Schwertfeger, C. Finn, and H. Zhao. Robot parkour learning. In Conference on Robot Learning (CoRL), 2023

  33. [41]

    Hoeller, N

    D. Hoeller, N. Rudin, D. Sako, and M. Hutter. Anymal parkour: Learning agile navigation for quadrupedal robots. Science Robotics, 2024

  34. [42]

    Caluwaerts, A

    K. Caluwaerts, A. Iscen, J. C. Kew, W. Yu, T. Zhang, D. Freeman, K.-H. Lee, L. Lee, S. Sal- iceti, V . Zhuang, et al. Barkour: Benchmarking animal-level agility with quadruped robots. arXiv preprint arXiv:2305.14654, 2023

  35. [43]

    Chane-Sane, J

    E. Chane-Sane, J. Amigo, T. Flayols, L. Righetti, and N. Mansard. Soloparkour: Constrained reinforcement learning for visual locomotion from privileged experience. In Conference on Robot Learning (CoRL), 2024

  36. [44]

    S. Luo, S. Li, R. Yu, Z. Wang, J. Wu, and Q. Zhu. Pie: Parkour with implicit-explicit learning framework for legged robots. IEEE Robotics and Automation Letters, 2024

  37. [45]

    J. Lee, L. Schroth, V . Klemm, M. Bjelonic, A. Reske, and M. Hutter. Evaluation of constrained reinforcement learning algorithms for legged locomotion. arXiv preprint arXiv:2309.15430, 2023

  38. [46]

    X. B. Peng, P. Abbeel, S. Levine, and M. Van de Panne. Deepmimic: Example-guided deep reinforcement learning of physics-based character skills. ACM Transactions On Graphics (TOG), 37(4):1–14, 2018

  39. [47]

    C. Li, M. Vlastelica, S. Blaes, J. Frey, F. Grimminger, and G. Martius. Learning agile skills via adversarial imitation of rough partial demonstrations. In Conference on Robot Learning, pages 342–352. PMLR, 2023

  40. [48]

    T. Li, H. Jung, M. Gombolay, Y . K. Cho, and S. Ha. Crossloco: Human motion driven control of legged robots via guided unsupervised reinforcement learning. arXiv preprint arXiv:2309.17046, 2023

  41. [49]

    T. Li, J. Won, A. Clegg, J. Kim, A. Rai, and S. Ha. Ace: Adversarial correspondence em- bedding for cross morphology motion retargeting from human to nonhuman characters. In SIGGRAPH Asia 2023 Conference Papers, pages 1–11, 2023

  42. [50]

    X. B. Peng, Z. Ma, P. Abbeel, S. Levine, and A. Kanazawa. Amp: Adversarial motion priors for stylized physics-based character control. ACM Transactions on Graphics (ToG), 40(4): 1–20, 2021

  43. [51]

    Escontrela, X

    A. Escontrela, X. B. Peng, W. Yu, T. Zhang, A. Iscen, K. Goldberg, and P. Abbeel. Adversarial motion priors make good substitutes for complex reward functions. In 2022 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems (IROS), pages 25–32. IEEE, 2022

  44. [52]

    R. Yang, Z. Chen, J. Ma, C. Zheng, Y . Chen, Q. Nguyen, and X. Wang. Generalized animal imitator: Agile locomotion with versatile motion prior. arXiv preprint arXiv:2310.01408 , 2023

  45. [53]

    Q. Yao, J. Wang, S. Yang, C. Wang, H. Zhang, Q. Zhang, and D. Wang. Imitation and adaptation based on consistency: A quadruped robot imitates animals from videos using deep reinforcement learning. In 2022 IEEE International Conference on Robotics and Biomimetics (ROBIO), pages...

  46. [54]

    J. Z. Zhang, S. Yang, G. Yang, A. L. Bishop, S. Gurumurthy, D. Ramanan, and Z. Manch- ester. Slomo: A general system for legged robot motion imitation from casual videos. IEEE Robotics and Automation Letters, 2023

  47. [55]

    Huang, I

    W. Huang, I. Mordatch, and D. Pathak. One policy to control them all: Shared modular policies for agent-agnostic control. In International Conference on Machine Learning, pages 4455–4464. PMLR, 2020

  48. [56]

    Salhotra, I

    G. Salhotra, I. Liu, C. Arthur, and G. Sukhatme. Learning robot manipulation from cross- morphology demonstration. arXiv preprint arXiv:2304.03833, 2023

  49. [57]

    G. Feng, H. Zhang, Z. Li, X. B. Peng, B. Basireddy, L. Yue, Z. Song, L. Yang, Y . Liu, K. Sreenath, et al. Genloco: Generalized locomotion controllers for quadrupedal robots. In Conference on Robot Learning, pages 1893–1903. PMLR, 2023

  50. [58]

    Devin, A

    C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine. Learning modular neural network policies for multi-task and multi-robot transfer. In 2017 IEEE international conference on robotics and automation (ICRA), pages 2169–2176. IEEE, 2017

  51. [59]

    Padalkar, A

    A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023

  52. [60]

    J. Yang, D. Sadigh, and C. Finn. Polybot: Training one policy across robots while embracing variability. arXiv preprint arXiv:2307.03719, 2023

  53. [61]

    D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine. Gnm: A general navigation model to drive any robot. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 7226–7233. IEEE, 2023

  54. [62]

    Shafiee, G

    M. Shafiee, G. Bellegarda, and A. Ijspeert. Manyquadrupeds: Learning a single locomotion policy for diverse quadruped robots. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 3471–3477. IEEE, 2024

  55. [63]

    J. Yang, C. Glossop, A. Bhorkar, D. Shah, Q. Vuong, C. Finn, D. Sadigh, and S. Levine. Pushing the limits of cross-embodiment learning for manipulation and navigation. arXiv preprint arXiv:2402.19432, 2024

  56. [64]

    Qin, Y .-H

    Y . Qin, Y .-H. Wu, S. Liu, H. Jiang, R. Yang, Y . Fu, and X. Wang. Dexmv: Imitation learning for dexterous manipulation from human videos. InEuropean Conference on Computer Vision, pages 570–587. Springer, 2022

  57. [65]

    Mandikal and K

    P. Mandikal and K. Grauman. Dexvip: Learning dexterous grasping with human hand pose priors from video. In Conference on Robot Learning, pages 651–661. PMLR, 2022

  58. [66]

    K. Shaw, S. Bahl, and D. Pathak. Videodex: Learning dexterity from internet videos. In Conference on Robot Learning, pages 654–665. PMLR, 2023

  59. [67]

    Z. Chen, S. Chen, A. Etienne, I. Laptev, and C. Schmid. ViViDex: Learning vision-based dexterous manipulation from human videos. arXiv:2404.15709, 2024

  60. [68]

    K. Shaw, S. Bahl, A. Sivakumar, A. Kannan, and D. Pathak. Learning dexterity from human hand motion in internet videos. The International Journal of Robotics Research, 43(4):513– 532, 2024

  61. [69]

    X. B. Peng, A. Kanazawa, J. Malik, P. Abbeel, and S. Levine. Sfv: Reinforcement learning of physical skills from videos. ACM Transactions On Graphics (TOG), 37(6):1–14, 2018. 14

  62. [70]

    Xiong, Q

    H. Xiong, Q. Li, Y .-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg. Learning by watching: Physical imitation of manipulation skills from human videos. In2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7827–7834. IEEE, 2021

  63. [71]

    S. Bahl, A. Gupta, and D. Pathak. Human-to-robot imitation in the wild. In RSS, 2022

  64. [72]

    Heppert, M

    N. Heppert, M. Argus, T. Welschehold, T. Brox, and A. Valada. Ditto: Demonstration im- itation by trajectory transformation. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024

  65. [73]

    Smith, N

    L. Smith, N. Dhawan, M. Zhang, P. Abbeel, and S. Levine. Avid: Learning multi-stage tasks via pixel-level translation of human videos. arXiv preprint arXiv:1912.04443, 2019

  66. [74]

    Zakka, A

    K. Zakka, A. Zeng, P. Florence, J. Tompson, J. Bohg, and D. Dwibedi. Xirl: Cross- embodiment inverse reinforcement learning. Conference on Robot Learning (CoRL), 2021

  67. [75]

    M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song. Xskill: Cross embodiment skill discovery. In Conference on Robot Learning, pages 3536–3555. PMLR, 2023

  68. [76]

    C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y . Zhu, and A. Anandkumar. Mimicplay: Long-horizon imitation learning by watching human play. arXiv preprint arXiv:2302.12422, 2023

  69. [77]

    S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta. R3m: A universal visual represen- tation for robot manipulation. arXiv preprint arXiv:2203.12601, 2022

  70. [78]

    T. Xiao, I. Radosavovic, T. Darrell, and J. Malik. Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173, 2022

  71. [79]

    Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training. arXiv preprint arXiv:2210.00030, 2022

  72. [80]

    Majumdar, K

    A. Majumdar, K. Yadav, S. Arnaud, J. Ma, C. Chen, S. Silwal, A. Jain, V .-P. Berges, T. Wu, J. Vakil, et al. Where are we in the search for an artificial visual cortex for embodied intelli- gence? Advances in Neural Information Processing Systems, 36:655–677, 2023

  73. [81]

    Y . Seo, K. Lee, S. L. James, and P. Abbeel. Reinforcement learning with action-free pre- training from videos. In International Conference on Machine Learning, 2022

  74. [82]

    Radosavovic, T

    I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell. Real-world robot learning with masked visual pre-training. In Conference on Robot Learning, pages 416–426. PMLR, 2023

  75. [83]

    Y . J. Ma, V . Kumar, A. Zhang, O. Bastani, and D. Jayaraman. Liv: Language-image represen- tations and rewards for robotic control. In International Conference on Machine Learning , pages 23301–23320. PMLR, 2023

  76. [84]

    Mendonca, S

    R. Mendonca, S. Bahl, and D. Pathak. Structured world models from human videos. arXiv preprint arXiv:2308.10901, 2023

  77. [85]

    Y . Ze, Y . Liu, R. Shi, J. Qin, Z. Yuan, J. Wang, and H. Xu. H-index: Visual reinforcement learning with hand-informed representations for dexterous manipulation. Advances in Neural Information Processing Systems, 36, 2024

  78. [86]

    Schmeckpeper, O

    K. Schmeckpeper, O. Rybkin, K. Daniilidis, S. Levine, and C. Finn. Reinforcement learning with videos: Combining offline observations with interaction. arXiv preprint arXiv:2011.06507, 2020. 15

  79. [87]

    L. Fan, G. Wang, Y . Jiang, A. Mandlekar, Y . Yang, H. Zhu, A. Tang, D.-A. Huang, Y . Zhu, and A. Anandkumar. Minedojo: Building open-ended embodied agents with internet-scale knowledge. Advances in Neural Information Processing Systems, 35:18343–18362, 2022

  80. [88]

    Alakuijala, G

    M. Alakuijala, G. Dulac-Arnold, J. Mairal, J. Ponce, and C. Schmid. Learning reward func- tions for robotic manipulation by observing humans. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 5006–5012. IEEE, 2023

  81. [89]

    A. S. Chen, S. Nair, and C. Finn. Learning generalizable robotic reward functions from” in-the-wild” human videos. arXiv preprint arXiv:2103.16817, 2021

  82. [90]

    Chane-Sane, C

    E. Chane-Sane, C. Schmid, and I. Laptev. Learning video-conditioned policies for unseen manipulation tasks. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 909–916. IEEE, 2023

  83. [91]

    K. Li, Y . Wang, P. Gao, G. Song, Y . Liu, H. Li, and Y . Qiao. Uniformer: Unified transformer for efficient spatiotemporal representation learning. arXiv preprint arXiv:2201.04676, 2022

  84. [92]

    K. Lee, K. Lee, J. Shin, and H. Lee. Network randomization: A simple technique for gener- alization in deep reinforcement learning. arXiv preprint arXiv:1910.05396, 2019

  85. [93]

    Wortsman, G

    M. Wortsman, G. Ilharco, S. Y . Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y . Carmon, S. Kornblith, et al. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In Inter- national con...

  86. [94]

    A. Rame, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, P. Gallinari, and M. Cord. Diverse weight averaging for out-of-distribution generalization. Advances in Neural Information Pro- cessing Systems, 2022

  87. [95]

    W. Yu, G. Turk, and C. K. Liu. Learning symmetric and low-energy locomotion. ACM Transactions on Graphics (TOG), 37(4):1–12, 2018

  88. [96]

    Abdolhosseini, H

    F. Abdolhosseini, H. Y . Ling, Z. Xie, X. B. Peng, and M. Van de Panne. On learning sym- metric locomotion. In Proceedings of the 12th ACM SIGGRAPH Conference on Motion, Interaction and Games, pages 1–10, 2019

  89. [97]

    Grimminger, A

    F. Grimminger, A. Meduri, M. Khadiv, J. Viereck, M. W ¨uthrich, M. Naveau, V . Berenz, S. Heim, F. Widmaier, T. Flayols, et al. An open torque-controlled modular robot architecture for legged locomotion research. IEEE Robotics and Automation Letters , 5(2):3650–3657, 2020

  90. [98]

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017

  91. [99]

    Loshchilov, F

    I. Loshchilov, F. Hutter, et al. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101, 5, 2017

  92. [100]

    H. Zhang. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017

  93. [101]

    Huang, R

    S. Huang, R. F. J. Dossa, C. Ye, J. Braga, D. Chakraborty, K. Mehta, and J. G. Ara´ujo. Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms. Journal of Machine Learning Research, 2022. 16 A Implementation Details A.1 Animal Video Classif...

  94. [5033]

    doi:10.1109/IROS.2012.6386109

    IEEE, 2012. doi:10.1109/IROS.2012.6386109. 11

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.