Pith. sign in

REVIEW 3 cited by

Bridging the Sim-to-Real Gap for Athletic Loco-Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.10894 v1 pith:KL6UWGNC submitted 2025-02-15 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords rewardsathleticrobotbehaviorsexplorationguidehackingleverages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Achieving athletic loco-manipulation on robots requires moving beyond traditional tracking rewards - which simply guide the robot along a reference trajectory - to task rewards that drive truly dynamic, goal-oriented behaviors. Commands such as "throw the ball as far as you can" or "lift the weight as quickly as possible" compel the robot to exhibit the agility and power inherent in athletic performance. However, training solely with task rewards introduces two major challenges: these rewards are prone to exploitation (reward hacking), and the exploration process can lack sufficient direction. To address these issues, we propose a two-stage training pipeline. First, we introduce the Unsupervised Actuator Net (UAN), which leverages real-world data to bridge the sim-to-real gap for complex actuation mechanisms without requiring access to torque sensing. UAN mitigates reward hacking by ensuring that the learned behaviors remain robust and transferable. Second, we use a pre-training and fine-tuning strategy that leverages reference trajectories as initial hints to guide exploration. With these innovations, our robot athlete learns to lift, throw, and drag with remarkable fidelity from simulation to reality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning

    cs.RO 2025-05 conditional novelty 7.0 of 10

    SPI-Active identifies legged-robot physical parameters via massive parallel sampling and uses Fisher-information-optimal command sequences to collect informative real-world data, improving sim-to-real transfer on quad...

  2. NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force Perception

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A Transformer-based actuator model learns torque surrogates and external forces on low-cost servo arms from pose trajectories and motor telemetry, improving force estimation and behavior-cloning control.

  3. Thor: Towards Human-Level Whole-Body Reactions for Intense Contact-Rich Environments

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A decoupled whole-body RL policy with a force-based lean reward enables a Unitree G1 humanoid to pull with up to 167.7 N, beating prior controllers by 69–75%.

Pith tools