Sparse-reward RL policies for whole-body loco-manipulation, bootstrapped with SMPC-generated offline demonstrations, surpass their SMPC teacher in simulated task completion time and transfer to real Spot and G1 hardware.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
Sparse-reward RL policies for whole-body loco-manipulation, bootstrapped with SMPC-generated offline demonstrations, surpass their SMPC teacher in simulated task completion time and transfer to real Spot and G1 hardware.