REVIEW 3 major objections 7 minor 94 references
Bicycle Acrobatics with Reinforcement Learning
T0 review · 3 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that reinforcement learning lets a bicycle robot learn and chain jumps, flips, kips, and wheelies, validated on hardware in repeated long repertoires.
desk verdict Genuine hardware advance in bicycle acrobatics with RL, but the robustness evidence needs tighter reporting before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two mechanisms carry the argument. First, a set of specialist RL policies, each trained with one of five formulations ranging from sparse waypoint following to dense motion imitation; the waypoint-following policy sees a robot-centric heightmap extracted from a global terrain map, so jumping emerges from reaching goals rather than being scripted. Second, an orchestrator: a finite state machine built on sequential composition, in which each stunt is a node and a transition fires only when the robot's current state lies in the next policy's region of attraction. The orchestrator is what turns isolated stunts into long repertoires, because it waits for safe, state-dependent conditions instead of switching on a timer or button alone.
What would settle it
Run the waypoint-following jump experiment with the motion-capture feed disabled and the offline heightmap withheld, so the robot must localize itself and perceive the table from onboard sensors only; if the robot still completes the single- and multi-table jumps, the paper's environmental-support premise is wrong, and if it cannot, the autonomy claim is bounded to instrumented environments.
Extended reading notes
Core claim
The paper's central claim is that reinforcement learning can give a bicycle robot a repertoire of dynamic acrobatic stunts that transfer from simulation to a physical platform, and that these stunts can be chained into extended, repeatable performances. On the Ultra Mobility Vehicle (UMV), a 23.5 kg single-track robot, individual policies trained with waypoint following, pose reaching, twist tracking, guided tracking, and motion imitation produce table jumps up to 1 m, front flips, kip-ups and kip-downs, wheelies, bunny hops, three-point turns, and lateral jumps. A state-driven orchestrator sequences these policies, and the paper reports more than 15 consecutive autonomous waypoint-following jumps, more than 20 consecutive kip-jump-flip-kip-down repertoires, and more than 10 consecutive wheelie-lateral-jump repertoires on hardware. The authors read these results as evidence that RL can give wheeled robots a level of agility previously associated mainly with legged platforms.
Load-bearing premise
The autonomy results are anchored to the laboratory: the robot localizes through an external motion-capture system and reads terrain from an offline map generated before each experiment, because it has no onboard camera or other exteroceptive sensor; if that instrumentation is removed, the waypoint-following jumps and multi-table adaptations would not transfer.
Editorial extensions
If this is right
- Wheeled robots can cross obstacles that previously required legged platforms, since the same policy clears 75 cm and 1 m tables and composes jumps with flips and kips.
- One jump policy generalizes beyond its training distribution: it handles two-table configurations it never saw and adapts to table heights up to 1 m without retraining.
- New stunts can be added to a repertoire by training a specialist policy and defining its entry and exit conditions, without retraining existing behaviors.
- RL discovers energy- and hardware-friendly strategies, such as the snake-like climb before takeoff and the brief wheelie before descent, that a designer would be unlikely to specify by hand.
- Repeated long trials (20 or more kip-jump-flip cycles) suggest the policies recover from landing impacts and maintain robustness over continuous operation, not just in single demonstrations.
Reading between the lines
- Removing the motion-capture and offline-map instrumentation would likely collapse the waypoint-following autonomy; the paper's results therefore do not yet imply outdoor or unstructured-environment operation until onboard perception is added.
- The orchestrator's hand-designed state boundaries and fixed transition timings (for example, 0.8 seconds after a trigger) may limit scalability; learning transition conditions from data is a natural extension.
- Because the five RL formulations are morphology-agnostic, the same pipeline could be retrained for scooters, motorcycles, or cargo bikes, with the main effort going into new reference trajectories and reward shaping.
- The low success rates on multi-table jumps hint that purely reactive, memoryless policies are near the limit of what this architecture can do; the paper's proposed high-level motion generator plus low-level tracker is a testable remedy.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a reinforcement-learning framework for a custom bicycle robot (UMV) and demonstrates a diverse set of acrobatic stunts, including table jumps, flips, kips, wheelies, lateral jumps, bunny hops, and three-point turns. The authors train specialist policies with five different RL formulations (waypoint following, pose reaching, SE2 twist tracking, guided tracking, and motion imitation) and introduce an FSM-based orchestrator that transitions between policies using state-dependent triggers to execute long-horizon repertoires. They report hardware validation with repeated runs, including more than 15 consecutive autonomous jumps, more than 20 consecutive kip-jump-flip-kip-down repertoires, and more than 10 consecutive wheelie-lateral-jump repertoires, alongside simulation ablations in the supplementary material. The central claim is that RL can give bicycle robots a level of agility previously associated with legged platforms.
Significance. If the reported results hold, this is a notable empirical advance: it demonstrates a large repertoire of dynamic, long-horizon acrobatic behaviors on a real underactuated, non-holonomic platform, with repeated hardware trials, MoCap traces, ablations, and careful disclosure of the state-estimation infrastructure (S8). The paper also provides useful evidence on curriculum design, reward engineering, and orchestrator-based policy composition. However, the quantitative evidence for the central 'robust execution' claim is currently unauditable: Table 1 success rates lack denominators and trial definitions, and the Discussion's 'without failures' statement is inconsistent with several non-unit rates. The S5 bailout disclosure further clouds the 'autonomous consecutive repertoires' headline. The absence of released code, data, or policy weights is an additional reproducibility gap.
major comments (3)
- [Table 1 and Discussion] The paper's robustness claim rests on success rates that are not auditable. Table 1 reports success rates such as 0.67 (row 3), 0.50 (row 5), and 0.97 (row 21) with no trial counts, denominators, confidence intervals, or a definition of what counts as a trial. The Discussion states 'We ran these stunts several times on several robots without failures', which is inconsistent with the non-unit success rates unless the rates were computed on a different condition (e.g., simulation or non-consecutive attempts). Please report the number of attempts and successes for every Table 1 entry, state explicitly whether each rate is hardware or simulation, and reconcile the 'without failures' claim with the reported rates.
- [Section S5 and Table 1 rows 22–25] The emergency 'bail out' disclosure directly affects the counted successes behind the abstract's 'more than 10 consecutive autonomous and steerable repertoires'. Section S5 states that if the robot is noticed to be about to fail or fall, the orchestrator transitions to a stabilizing policy before proceeding with the stunt. If any of the consecutive trials counted toward the headline included such a bailout, the claim overstates what the policies alone achieved. Please state whether bailouts occurred in the counted trials, and if so, report the number of trials completed without any human-triggered intervention separately from those with bailouts.
- [Data, code, and materials availability] The availability statement says 'Data and figure generation code for this study will be available if approved by our internal review process', meaning hardware logs, policy weights, and code are not released. Since the central claims are empirical and hardware-specific, the absence of released trial logs makes it impossible for a reader to verify the success counts or the sim-to-real details. Please either release the data/code or provide a detailed raw-data appendix of all hardware trials (including failures) in the supplementary material.
minor comments (7)
- [Figure 4(B)] The radar chart is described as normalized, but the normalization method is not defined; please specify how each metric is scaled.
- [Results section] The text contains missing references/placeholders, e.g., 'All the aforementioned experiments ... are shown in .' and 'show multiple trials.' These need to be filled or the text removed.
- [Section S8] The statement that the robot 'also supports a fully onboard estimator for outdoor operation [1]' should be reconciled with the statement that 'the current platform lacks onboard exteroceptive sensors'; clarify whether the outdoor estimator assumes additional sensors not present in the current platform.
- [Section S4 and Figure S2] The ablation success rates are simulation results; please label them as simulation in the main text and figure captions to avoid confusion with the hardware success rates in Table 1.
- [Table S1(A)] The noise magnitudes for 'Bike Global Position' (N(0,0.1)) are given without units; add units (e.g., meters).
- [Orchestrator section] The phrase 'We designed the orchestrator as an Finite State Machine' should read 'a Finite State Machine'.
- [References [14] and [15]] References [14] and [15] are identical; please correct or replace the duplicate.
Circularity Check
No circular derivation: the hardware stunts are independent external benchmarks, and no reported success measure is reconstructed from a fitted parameter or from a self-citation.
full rationale
The paper's derivation chain is empirical rather than definitional. Each RL policy is trained in simulation from stated rewards, constraints, curricula, and domain randomization, then deployed on the UMV, and the claimed stunts are judged by whether the physical robot clears the table, completes the flip, or finishes the repertoire. These outcomes are external events; they are not quantities that the training objectives or reward weights reproduce by construction. The citations to same-group prior work (UMV system design [1], IMI [65], LineRides [68]) supply hardware and training methodology, but the hardware demonstrations are separately shown in the figures and movies, so the central agility claim does not reduce to those citations. The orchestrator is an FSM based on the external sequential-composition framework [74], with state-dependent transition conditions that are not fitted to the reported success counts. The S5 disclosure of a human-triggered emergency 'bail out' policy for the wheelie/lateral-jump repertoires is a genuine evidence-quality concern: it may mean that some counted 'autonomous' successes included operator intervention, and the Discussion's 'without failures' statement is hard to reconcile with the 0.50-0.97 rates in Table 1. However, these are ambiguity and auditability problems about what was measured, not circularity: no equation in the paper is shown to equal its own input, and no fitted parameter is renamed as a prediction. The paper is therefore self-contained against external benchmarks, with no demonstrated circular step.
Assumptions & free parameters
free parameters (8)
- Reward weights for each training type =
Various; e.g., distance-to-goal weight 6, imitation orientation weight 20 (Tables S3-S7)
- Action scales, stiffness, and damping per task =
e.g., head action scale 0.1 nominal vs 1.5 in pose reaching; head stiffness 120-200; damping 5-10 (Table S2)
- PPO entropy coefficients per task =
0.006, 0.0035, 0.0035, 0.002, 0.0025 (Table S12)
- Curriculum iteration thresholds =
e.g., 5000 iterations for waypoint-following termination curricula; varied per task (Table S10)
- Wheel contact stiffness and damping =
Not reported numerically; 'sampled many values and picked the ones that matched' drop-test data (Section S6)
- PD gains and torque constants =
Torque constants experimentally identified; PD gains tuned by sinusoidal identification (Table S2, Section S7)
- Termination thresholds =
e.g., wheel contact force > 350 N, rear wheel velocity > 77 rad/s (Table S9)
- Domain randomization ranges =
e.g., friction U[0.6,1.6], added mass scale U[0.9,1.1], heightmap drift U[-0.1,0.1] (Table S8)
assumptions (5)
- domain assumption The Isaac Lab dynamics and contact models, after manual wheel-contact tuning from drop tests, are sufficiently accurate for zero-shot policy transfer to hardware.
- domain assumption Each policy in the orchestrator has a funnel or region of attraction that overlaps the entry set of the next policy, as required by sequential composition.
- domain assumption External motion capture and an offline terrain map provide accurate localization and heightmap inputs during hardware experiments.
- domain assumption The reference trajectories used for motion imitation become dynamically feasible for UMV after iterative refinement via IMI.
- ad hoc to paper Hand-designed reward functions and curriculum schedules are sufficient for PPO to discover the intended acrobatic behaviors.
Cite this review
Pith. "Pith review of Bicycle Acrobatics with Reinforcement Learning." pith.science (2026). https://pith.science/paper/GGTAT3WT
@misc{pith2026260800880,
author = {Pith},
title = {Pith review of: Bicycle Acrobatics with Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/GGTAT3WT}},
note = {Machine review of arXiv:2608.00880}
}
read the original abstract
Bicycle robots are fast and energy efficient, but their simple mechanical design and their underactuated and non-holonomic dynamics make highly agile maneuvers difficult to achieve. Here, we use Reinforcement Learning (RL) to enable a bicycle robot to learn and compose a diverse repertoire of dynamic acrobatic stunts. Using different RL formulations such as waypoint following, pose reaching, twist tracking, guided tracking, and motion imitation, the robot acquires autonomous single and multi-table forward and lateral jumps, steerable jumps, front flips, kip-ups, kip-downs, driving, wheelies, bunny hops, and three-point turns. To coordinate these behaviors, we introduce an orchestrator that transitions between policies using state-dependent triggers, enabling robust long-horizon acrobatic stunts. We validate the approach on the Ultra Mobility Vehicle (UMV), a custom bicycle robot, in simulation and hardware. The robot repeatedly traverses tables up to 1 m high, performs more than 15 consecutive autonomous jumps while following waypoints, handles previously unseen multi-table configurations, executes continuous repertoires of kipups, jumps, flips, kip-downs, over more than 20 consecutive trials, and performs more than 10 consecutive autonomous and steerable repertoires of wheelies, lateral jumps, and single-wheel jump downs. These results demonstrate that RL can endow bicycle robots with levels of agility previously associated primarily with legged platforms while preserving the speed and efficiency of wheeled locomotion, establishing a foundation for bicycle acrobatics.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
B. Bokser,et al., System Design of the Ultra Mobility Vehicle: A Driving, Balancing, and Jumping Bicycle Robot.arXiv preprint arXiv:2602.22118(2026)
arXiv 2026
-
[2]
Dynamics, Introducing Handle, Y outube video, https://www.youtube
B. Dynamics, Introducing Handle, Y outube video, https://www.youtube. com/watch?v=-7xvqQeoA8c (2017), accessed: 15 September 2025
2017
-
[3]
Kashiri,et al., Centauro: A hybrid locomotion and high power resilient manipulation platform.IEEE Robotics and Automation Letters 4(2), 1595–1602 (2019)
N. Kashiri,et al., Centauro: A hybrid locomotion and high power resilient manipulation platform.IEEE Robotics and Automation Letters 4(2), 1595–1602 (2019)
2019
-
[4]
Klemm,et al., Ascento: A two-wheeled jumping robot, in2019 International conference on robotics and automation (ICRA)(IEEE) (2019), pp
V. Klemm,et al., Ascento: A two-wheeled jumping robot, in2019 International conference on robotics and automation (ICRA)(IEEE) (2019), pp. 7515–7521
2019
-
[5]
Bjelonic, P
M. Bjelonic, P . K. Sankar, C. D. Bellicoso, H. Vallery, M. Hutter, Rolling in the deep–hybrid locomotion for wheeled-legged robots using online Research Article - Preprint Robotics and AI Institute 14 kipup trigger Kip Up 0.8 sec Kip DownStart/End kipdown trigger0.8 sec 0.8 sec & drive trigger 0.8 sec & kipdown trigger 0.8 sec & drive trigger Fig. 8.Auto...
2020
-
[6]
Bjelonic, V
M. Bjelonic, V. Klemm, J. Lee, M. Hutter, A survey of wheeled-legged robots, inClimbing and walking robots conference(Springer) (2022), pp. 83–94
2022
-
[7]
D. Liu, F . Y ang, X. Liao, X. Lyu, Diablo: A 6-dof wheeled bipedal robot composed entirely of direct-drive joints, in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)(IEEE) (2024), pp. 3605–3612
2024
-
[8]
Unitree Robotics, Unitree B2-W (2025), https://www.unitree.com/b2-w
2025
Show all 94 references
-
[9]
D. E. Jones, From the archives: The stability of the bicycle.Physics Today59(9), 51–56 (2006)
2006
-
[10]
Astrom, R
K. Astrom, R. Klein, A. Lennartsson, Bicycle dynamics and control: adapted bicycles for education and research.IEEE Control Systems Magazine25(4), 26–47 (2005),
2005
-
[11]
J. P . Meijaard, J. M. Papadopoulos, A. Ruina, A. L. Schwab, Linearized dynamics equations for the balance and steer of a bicycle: a benchmark and review.Proceedings of the Royal society A: mathematical, physical and engineering sciences463(2084), 1955–1982 (2007)
2007
-
[12]
J. D. G. Kooijman, J. P . Meijaard, J. M. Papadopoulos, A. Ruina, A. L. Schwab, A Bicycle Can Be Self-Stable Without Gyroscopic or Caster Effects.Science332(6027), 339–342 (2011),
2011
-
[13]
L. Cui,et al., Nonlinear balance control of an unmanned bicycle: De- sign and experiments, in2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)(IEEE) (2020), pp. 7279–7284
2020
-
[15]
Xiong, R
J. Xiong, R. Yu, C. Liu, Steering control and stability analysis for an autonomous bicycle: part II experiments and a modified linear control law.Nonlinear Dynamics112(5), 3107–3132 (2024)
2024
-
[16]
B. Wang, F . Jing, Y . Deng, Z. Chen, B. Liang, Bayesian Optimization- Based Ideal Landing Planning for Ramp Jump of Single-Track Two- Wheeled Robots, in2024 IEEE International Conference on Robotics and Biomimetics (ROBIO)(IEEE) (2024), pp. 396–401
2024
-
[17]
B. Wang, F . Jing, Y . Deng, Z. Chen, B. Liang, Attitude Control During the Flight Phase of Ramp Jump for Single-Track Two-Wheeled Robots, in2024 IEEE International Conference on Robotics and Biomimetics (ROBIO)(IEEE) (2024), pp. 1629–1634
2024
-
[18]
Randløv, P
J. Randløv, P . Alstrøm, Learning to Drive a Bicycle Using Reinforcement Learning and Shaping., inICML, vol. 98 (1998), pp. 463–471
1998
-
[19]
X. Zhu,et al., Deep Reinforcement Learning-Based Control of Bicycle Robots on Rough Terrain, in2023 9th International Conference on Control, Automation and Robotics (ICCAR)(IEEE) (2023), pp. 103– 108
2023
-
[20]
Zheng,et al., Reinforcement learning-based control of single-track two-wheeled robots in narrow terrain.Actuators12(3), 109 (2023)
Q. Zheng,et al., Reinforcement learning-based control of single-track two-wheeled robots in narrow terrain.Actuators12(3), 109 (2023)
2023
-
[21]
Baltes, G
J. Baltes, G. Christmann, S. Saeedvand, A deep reinforcement learning algorithm to control a two-wheeled scooter with a humanoid robot. Engineering Applications of Artificial Intelligence126, 106941 (2023)
2023
-
[22]
Zheng, D
Q. Zheng, D. Wang, Z. Chen, Y . Sun, B. Liang, Continuous reinforce- ment learning based ramp jump control for single-track two-wheeled robots.Transactions of the Institute of Measurement and Control44(4), 892–904 (2022)
2022
-
[23]
Z. Yuan, L. Y e, H. Liu, Z. Chen, B. Liang, Dynamic Jumps of an Agile Bicycle Through Reinforcement Learning, in2024 IEEE 3rd Industrial Electronics Society Annual On-Line Conference (ONCON) (IEEE) (2024), pp. 1–5
2024
-
[24]
J. Tan, Y . Gu, C. K. Liu, G. Turk, Learning bicycle stunts.ACM Transac- tions on Graphics (TOG)33(4), 1–12 (2014)
2014
-
[25]
M. H. Raibert,Legged robots that balance(MIT Press) (1986)
1986
-
[26]
Kuindersma,et al., Optimization-based locomotion planning, esti- mation, and control design for the atlas humanoid robot.Autonomous Robots40(3), 429–455 (2016),
S. Kuindersma,et al., Optimization-based locomotion planning, esti- mation, and control design for the atlas humanoid robot.Autonomous Robots40(3), 429–455 (2016),
2016
-
[27]
Di Carlo, P
J. Di Carlo, P . M. Wensing, B. Katz, G. Bledt, S. Kim, Dynamic Locomo- tion in the MIT Cheetah 3 Through Convex Model-Predictive Control, in2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)(2018), pp. 1–9,
2018
-
[28]
C. D. Bellicoso, F . Jenelten, C. Gehring, M. Hutter, Dynamic Locomotion Through Online Nonlinear Motion Optimization for Quadrupedal Robots. IEEE Robotics and Automation Letters3(3), 2261–2268 (2018),
2018
-
[29]
Fahmi, C
S. Fahmi, C. Mastalli, M. Focchi, C. Semini, Passive Whole-Body Control for Quadruped Robots: Experimental Validation Over Challeng- ing Terrain.IEEE Robotics and Automation Letters4(3), 2553–2560 (2019),
2019
-
[30]
Rudin, D
N. Rudin, D. Hoeller, P . Reist, M. Hutter, Learning to Walk in Minutes Us- ing Massively Parallel Deep Reinforcement Learning, inProceedings of the 5th Conference on Robot Learning, A. Faust, D. Hsu, G. Neumann, Eds. (PMLR), vol. 164 ofProceedings of Machine Learning Research...
2022
-
[31]
Ha,et al., Learning-based legged locomotion: State of the art and future perspectives.Int
S. Ha,et al., Learning-based legged locomotion: State of the art and future perspectives.Int. J. Robot. Res.44(8), 1396–1427 (2025),
2025
-
[32]
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, M. Hutter, Learning quadrupedal locomotion over challenging terrain.Science robotics 5(47), eabc5986 (2020)
2020
-
[33]
T. Miki,et al., Learning robust perceptive locomotion for quadrupedal Research Article - Preprint Robotics and AI Institute 15 robots in the wild.Science Robotics7(62), eabk2822 (2022), , https: //www.science.org/doi/abs/10.1126/scirobotics.abk2822
2022 doi
-
[34]
Hoeller, N
D. Hoeller, N. Rudin, D. Sako, M. Hutter, Anymal parkour: Learn- ing agile navigation for quadrupedal robots.Science Robotics9(88), eadi7566 (2024)
2024
-
[35]
Cheng, K
X. Cheng, K. Shi, A. Agarwal, D. Pathak, Extreme parkour with legged robots, in2024 IEEE International Conference on Robotics and Au- tomation (ICRA)(IEEE) (2024), pp. 11443–11450
2024
-
[36]
Zhuang,et al., Robot Parkour Learning, inConference on Robot Learning(PMLR) (2023), pp
Z. Zhuang,et al., Robot Parkour Learning, inConference on Robot Learning(PMLR) (2023), pp. 73–92
2023
-
[37]
Li,et al., Robust and Versatile Bipedal Jumping Control through Re- inforcement Learning, inRobotics science and systems(RSS) (2023)
Z. Li,et al., Robust and Versatile Bipedal Jumping Control through Re- inforcement Learning, inRobotics science and systems(RSS) (2023)
2023
-
[38]
Lee,et al., Learning robust autonomous navigation and locomotion for wheeled-legged robots.Science Robotics9(89), eadi9641 (2024)
J. Lee,et al., Learning robust autonomous navigation and locomotion for wheeled-legged robots.Science Robotics9(89), eadi9641 (2024)
2024
-
[39]
Kim,et al., High-speed control and navigation for quadrupedal robots on complex and discrete terrain.Science Robotics10(102), eads6192 (2025)
H. Kim,et al., High-speed control and navigation for quadrupedal robots on complex and discrete terrain.Science Robotics10(102), eads6192 (2025)
2025
-
[40]
He,et al., Attention-based map encoding for learning generalized legged locomotion.Science Robotics10(105), eadv3604 (2025)
J. He,et al., Attention-based map encoding for learning generalized legged locomotion.Science Robotics10(105), eadv3604 (2025)
2025
-
[41]
Bussola, M
R. Bussola, M. Focchi, G. Turrisi, C. Semini, L. Palopoli, Guided rein- forcement learning for omnidirectional 3d jumping in quadruped robots. arXiv preprint arXiv:2507.16481(2025)
2025 arXiv
-
[42]
X. B. Peng, P . Abbeel, S. Levine, M. Van de Panne, Deepmimic: Example-guided deep reinforcement learning of physics-based charac- ter skills.ACM Transactions On Graphics (TOG)37(4), 1–14 (2018)
2018
-
[43]
X. B. Peng,et al., Learning agile robotic locomotion skills by imitating animals, inRobotics: Science and Systems (RSS)(2020)
2020
-
[44]
X. B. Peng, Z. Ma, P . Abbeel, S. Levine, A. Kanazawa, Amp: Adver- sarial motion priors for stylized physics-based character control.ACM Transactions on Graphics (ToG)40(4), 1–20 (2021)
2021
-
[45]
X. B. Peng, Y . Guo, L. Halper, S. Levine, S. Fidler, Ase: Large-scale reusable adversarial skill embeddings for physically simulated charac- ters.ACM Transactions On Graphics (TOG)41(4), 1–17 (2022)
2022
-
[46]
Tessler, Y
C. Tessler, Y . Guo, O. Nabati, G. Chechik, X. B. Peng, Maskedmimic: Unified physics-based character control through masked motion in- painting.ACM Transactions on Graphics (TOG)43(6), 1–21 (2024)
2024
-
[47]
G. B. Margolis, M. Wang, N. Fey, P . Agrawal, SoftMimic: Learn- ing Compliant Whole-body Control from Examples.arXiv preprint arXiv:2510.17792(2025)
2025
-
[48]
Kang,et al., Learning Steerable Imitation Controllers from Unstruc- tured Animal Motions.arXiv preprint arXiv:2507.00677(2025)
D. Kang,et al., Learning Steerable Imitation Controllers from Unstruc- tured Animal Motions.arXiv preprint arXiv:2507.00677(2025)
2025
-
[49]
X. B. Peng, A. Kanazawa, J. Malik, P . Abbeel, S. Levine, Sfv: Rein- forcement learning of physical skills from videos.ACM Transactions On Graphics (TOG)37(6), 1–14 (2018)
2018
-
[50]
Miller, S
A. Miller, S. Fahmi, M. Chignoli, S. Kim, Reinforcement learning for legged robots: Motion imitation from model-based optimal control. arXiv preprint arXiv:2305.10989(2023)
2023 arXiv
-
[51]
D. Kang, J. Cheng, M. Zamora, F . Zargarbashi, S. Coros, Rl+ model- based control: Using on-demand optimal control to learn versatile legged locomotion.IEEE Robotics and Automation Letters8(10), 6619–6626 (2023)
2023
-
[52]
Y oum,et al., Imitating and finetuning model predictive control for robust and symmetric quadrupedal locomotion.IEEE Robotics and Automation Letters8(11), 7799–7806 (2023)
D. Y oum,et al., Imitating and finetuning model predictive control for robust and symmetric quadrupedal locomotion.IEEE Robotics and Automation Letters8(11), 7799–7806 (2023)
2023
-
[53]
Fuchioka, Z
Y . Fuchioka, Z. Xie, M. Van de Panne, Opt-mimic: Imitation of opti- mized trajectories for dynamic quadruped behaviors.arXiv preprint arXiv:2210.01247(2022)
2022 arXiv
-
[54]
Liu,et al., Opt2skill: Imitating dynamically-feasible whole-body trajectories for versatile humanoid loco-manipulation.arXiv preprint arXiv:2409.20514(2024)
F . Liu,et al., Opt2skill: Imitating dynamically-feasible whole-body trajectories for versatile humanoid loco-manipulation.arXiv preprint arXiv:2409.20514(2024)
2024
-
[55]
Y ang,et al., OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction, inarxiv preprint arXiv:2509.26633(2026), pp
L. Y ang,et al., OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction, inarxiv preprint arXiv:2509.26633(2026), pp. 1–12
2026 arXiv
-
[56]
Mahmood, N
N. Mahmood, N. Ghorbani, N. F . Troje, G. Pons-Moll, M. J. Black, AMASS: Archive of motion capture as surface shapes, inProceedings of the IEEE/CVF international conference on computer vision(2019), pp. 5442–5451
2019
-
[57]
F . G. Harvey, M. Yurick, D. Nowrouzezahrai, C. Pal, Robust motion in-betweening.ACM Transactions on Graphics (TOG)39(4), 60–1 (2020)
2020
-
[58]
He,et al., OmniH2O: Universal and Dexterous Human-to- Humanoid Whole-Body Teleoperation and Learning.arXiv preprint arXiv:2406.08858(2024)
T. He,et al., OmniH2O: Universal and Dexterous Human-to- Humanoid Whole-Body Teleoperation and Learning.arXiv preprint arXiv:2406.08858(2024)
2024 arXiv
-
[59]
He,et al., Asap: Aligning simulation and real-world physics for learn- ing agile humanoid whole-body skills.arXiv preprint arXiv:2502.01143 (2025)
T. He,et al., Asap: Aligning simulation and real-world physics for learn- ing agile humanoid whole-body skills.arXiv preprint arXiv:2502.01143 (2025)
2025 arXiv
-
[60]
Chen,et al., GMT: General Motion Tracking for Humanoid Whole- Body Control.arXiv preprint arXiv:2506.14770(2025)
Z. Chen,et al., GMT: General Motion Tracking for Humanoid Whole- Body Control.arXiv preprint arXiv:2506.14770(2025)
2025 arXiv
-
[61]
Wu,et al., Perceptive humanoid parkour: Chaining dynamic human skills via motion matching.arXiv preprint arXiv:2602.15827(2026)
Z. Wu,et al., Perceptive humanoid parkour: Chaining dynamic human skills via motion matching.arXiv preprint arXiv:2602.15827(2026)
2026 arXiv
-
[62]
Sleiman, ZEST: Zero-shot Embodied Skill Transfer for Athletic Robot Control.arXiv preprint arXiv:2602.00401(2026)
J.-P . Sleiman, ZEST: Zero-shot Embodied Skill Transfer for Athletic Robot Control.arXiv preprint arXiv:2602.00401(2026)
2026
-
[63]
D. S. Brown, W. Goo, S. Niekum, Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations, inInternational Conference on Machine Learning (ICML)(2019), pp. 783–792
2019
-
[64]
Y .-H. Wu, N. Charoenphakdee, H. Bao, V. Tangkaratt, M. Sugiyama, Imitation learning from imperfect demonstration, inInternational Con- ference on Machine Learning (ICML)(2019), pp. 6818–6827
2019
-
[65]
J. Kim, S. Fahmi, S. Rho, S. Ha, G. Nelson, Flip Stunts on Bicycle Robots using Iterative Motion Imitation, inProc. IEEE Int. Conf. Robot. Automat. (ICRA)(Vienna, Austria) (2026), pp. 1–8
2026
-
[66]
Li,et al., Learning agile skills via adversarial imitation of rough partial demonstrations, inConference on Robot Learning(PMLR) (2023), pp
C. Li,et al., Learning agile skills via adversarial imitation of rough partial demonstrations, inConference on Robot Learning(PMLR) (2023), pp. 342–352
2023
-
[67]
Zargarbashi, J
F . Zargarbashi, J. Cheng, D. Kang, R. Sumner, S. Coros, RobotKeyfram- ing: Learning Locomotion with High-Level Objectives via Mixture of Dense and Sparse Rewards, inConference on Robot Learning(PMLR) (2025), pp. 916–932
2025
-
[68]
Rho,et al., LineRides: Line-Guided Reinforcement Learning for Bicycle Robot Stunts.IEEE Robotics and Automation Letters(2026),
S. Rho,et al., LineRides: Line-Guided Reinforcement Learning for Bicycle Robot Stunts.IEEE Robotics and Automation Letters(2026),
2026
-
[69]
Y . Ma, F . Farshidian, M. Hutter, Learning Arm-Assisted Fall Damage Reduction and Recovery for Legged Mobile Manipulators, in2023 IEEE International Conference on Robotics and Automation (ICRA)(2023), pp. 12149–12155,
2023
-
[70]
Huang,et al., Learning Humanoid Standing-up Control across Di- verse Postures.arXiv preprint arXiv:2502.08378(2025)
T. Huang,et al., Learning Humanoid Standing-up Control across Di- verse Postures.arXiv preprint arXiv:2502.08378(2025)
2025 arXiv
-
[71]
Strauch,et al., Robot Crash Course: Learning Soft and Stylized Falling.arXiv preprint arXiv:2511.10635(2025)
P . Strauch,et al., Robot Crash Course: Learning Soft and Stylized Falling.arXiv preprint arXiv:2511.10635(2025)
2025
-
[72]
MacAskill, Red Bull, Danny MacAskill’s Imaginate (2013), https: //www.youtube.com/watch?v=Sv3xVOs7_No, accessed: 2026-05-18
D. MacAskill, Red Bull, Danny MacAskill’s Imaginate (2013), https: //www.youtube.com/watch?v=Sv3xVOs7_No, accessed: 2026-05-18
2013
-
[73]
MacAskill, Redd Bull Bike, Danny MacAskill’s Gymnasium (2020), https://www.youtube.com/watch?v=fAEBNEscL0c, accessed: 2026-05- 18
D. MacAskill, Redd Bull Bike, Danny MacAskill’s Gymnasium (2020), https://www.youtube.com/watch?v=fAEBNEscL0c, accessed: 2026-05- 18
2020
-
[74]
R. R. Burridge, A. A. Rizzi, D. E. Koditschek, Sequential composition of dynamically dexterous robot behaviors.The International Journal of Robotics Research18(6), 534–555 (1999),
1999
-
[75]
Rudin, D
N. Rudin, D. Hoeller, M. Bjelonic, M. Hutter, Advanced Skills by Learn- ing Locomotion and Local Navigation End-to-End, inProc. IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS)(2022), pp. 2497–2503,
2022
-
[76]
Z. Xie, H. Y . Ling, N. H. Kim, M. Van De Panne, Allsteps: curriculum- driven learning of stepping stone skills, inComputer Graphics Forum (Wiley Online Library), vol. 39 (2020), pp. 213–224
2020
-
[77]
Portelas, C
R. Portelas, C. Colas, L. Weng, K. Hofmann, P .-Y . Oudeyer, Automatic curriculum learning for deep RL: A short survey, inInternational Joint Conference on Artificial Intelligence (IJCAI)(2020)
2020
-
[78]
Hwangbo,et al., Learning agile and dynamic motor skills for legged robots.Science Robotics4(26), eaau5872 (2019)
J. Hwangbo,et al., Learning agile and dynamic motor skills for legged robots.Science Robotics4(26), eaau5872 (2019)
2019
-
[79]
E. Chane-Sane,et al., CaT: Constraints as Terminations for Legged Locomotion Reinforcement Learning, inIEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS)(2024)
2024
-
[80]
Schulman, F
J. Schulman, F . Wolski, P . Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347(2017)
2017 arXiv
-
[81]
Mittal,et al., Isaac Lab: A GPU-Accelerated Simulation Frame- work for Multi-Modal Robot Learning.arXiv preprint arXiv:2511.04831 (2025)
M. Mittal,et al., Isaac Lab: A GPU-Accelerated Simulation Frame- work for Multi-Modal Robot Learning.arXiv preprint arXiv:2511.04831 (2025). Research Article - Preprint Robotics and AI Institute 16
2025 arXiv
-
[82]
Pinto, M
L. Pinto, M. Andrychowicz, P . Welinder, W. Zaremba, P . Abbeel, Asymmetric actor critic for image-based robot learning.arXiv preprint arXiv:1710.06542(2017)
2017 arXiv
-
[83]
Abdolhosseini, H
F . Abdolhosseini, H. Y . Ling, Z. Xie, X. B. Peng, M. Van de Panne, On learning symmetric locomotion, inProceedings of the 12th ACM SIGGRAPH Conference on Motion, Interaction and Games(2019), pp. 1–10
2019
-
[84]
J. Tobin,et al., Domain randomization for transferring deep neural net- works from simulation to the real world, in2017 IEEE/RSJ international conference on intelligent robots and systems (IROS)(IEEE) (2017), pp. 23–30
2017
-
[85]
X. B. Peng, M. Andrychowicz, W. Zaremba, P . Abbeel, Sim-to-real transfer of robotic control with dynamics randomization, in2018 IEEE international conference on robotics and automation (ICRA)(IEEE) (2018), pp. 3803–3810
2018
-
[86]
Tan,et al., Sim-to-real: Learning agile locomotion for quadruped robots.arXiv preprint arXiv:1804.10332(2018)
J. Tan,et al., Sim-to-real: Learning agile locomotion for quadruped robots.arXiv preprint arXiv:1804.10332(2018)
2018 arXiv
-
[87]
Choi,et al., Learning quadrupedal locomotion on deformable terrain
S. Choi,et al., Learning quadrupedal locomotion on deformable terrain. Science Robotics8(74), eade2256 (2023),
2023
-
[88]
Todorov, T
E. Todorov, T. Erez, Y . Tassa, Mujoco: A physics engine for model- based control, in2012 IEEE/RSJ international conference on intelligent robots and systems(IEEE) (2012), pp. 5026–5033
2012
-
[89]
Open Neural Network Exchange (ONNX), https://github.com/onnx/ onnx
-
[90]
Bellicoso, A
D. Bellicoso, A. Wollschläger, Exploy: EXport and dePLOY Rein- forcement Learning policies (2026), https://github.com/rai-opensource/ exploy
2026
-
[91]
UP Xtreme i12, https://up-board.org/up-xtreme-i12/?ADL-P-Board
-
[92]
Specialized Bicycle Components, 2022 Hotwalk Carbon, https://www.specialized.com/us/en/hotwalk-carbon/p/200221? color=322068-200221&searchText=94021-0005 (2022), accessed: 2026-06-11
2022
-
[93]
Dellaert,et al., borglab/gtsam: Release 4.2.Zenodo(2023)
F . Dellaert,et al., borglab/gtsam: Release 4.2.Zenodo(2023)
2023
-
[94]
Goldfain,et al., Autorally: An Open Platform for Aggressive Au- tonomous Driving.IEEE Control Systems Magazine39(1), 26–55 (2019)
B. Goldfain,et al., Autorally: An Open Platform for Aggressive Au- tonomous Driving.IEEE Control Systems Magazine39(1), 26–55 (2019)
2019
-
[95]
bail out
C. Forster, L. Carlone, F . Dellaert, D. Scaramuzza,IMU preintegration on Manifold for Efficient Visual-Inertial Maximum-a-Posteriori Estima- tion(2015). AcknowledgmentsWe thank the entire hardware, simulation, machine learning, and controls team of UMV including A. Pre- ston,...
2015
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.