REVIEW 3 major objections 26 references
Hard humanoid motions stay unsolved under the usual recipe because the recipe itself limits what the controller can learn.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 12:45 UTC pith:SSY6JWMR
load-bearing objection Useful systems paper: residual long-tail failures as recipe/capability mismatch, with a compact dynamic/balance expert pipeline and better high-coverage metrics—solid sim evidence, incomplete residual isolation. the 3 major comments →
Athena-WBC: Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In strong humanoid whole-body control baselines, residual feasible training clips remain unsolved even under targeted training because of a capability bottleneck: a mismatch between motion demands and the effective control regime induced by the default acquisition recipe. Capability-aligned experts—dynamic experts that remove effort and temporal-control penalties while retaining physical constraints and balance experts that use a gravity curriculum—plus motion-routed distillation and RL fine-tuning recover more of that long tail and improve held-out tracking than reallocating data alone under the same recipe.
What carries the argument
Capability-aligned policy experts: dynamic experts train with tracking plus physical-constraint rewards only, plus Grad-CAPS auxiliary smoothness on the policy mean; balance experts train under a gravity-scale curriculum; both residual-set teachers are then motion-routed for DAgger distillation into one student that is RL-finetuned under deployable observations.
Load-bearing premise
That residual failures after data-only interventions mainly reflect a recipe-induced capability mismatch rather than imperfect retargeted references, near-limit physical infeasibility, embodiment limits, or under-optimized training budgets.
What would settle it
Train the same residual clips to saturation under the default recipe with matched budget and cleaner references; if those clips then succeed at high rate without capability changes, the bottleneck claim fails. Conversely, if dynamic and balance recipe changes still do not recover a large share of the residual set, the expert design is insufficient.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Athena-WBC argues that residual training-set failures in strong humanoid whole-body motion trackers are not only data-allocation problems but capability bottlenecks induced by the default acquisition recipe. On a full-size 80 kg humanoid, the authors reimplement a SONIC-style baseline, mine residual clips from a general privileged teacher, and train two capability-aligned experts on the same residual set: a dynamic expert that keeps tracking and physical-constraint rewards while removing effort/temporal penalties and regularizing with Grad-CAPS, and a balance expert trained with a gravity curriculum. Teachers are motion-routed by rollout success, distilled with DAgger into one deployable student, and RL-finetuned. Relative to SONIC-Base, the final policy improves held-out AMASS/Omni SR, TIS, MPJPE, and MPJPE-W, with ablations on no-smoothness, CAPS vs Grad-CAPS, single- vs multi-teacher distillation, and teacher observations, plus new diagnostics (STC/TIS, MPJPE-W).
Significance. If the residual-failure diagnosis holds, the paper usefully reframes long-tail humanoid WBC: changing reward/curriculum capability can matter more than only resampling or partitioning motions, and a compact two-expert bank can recover complementary high-dynamic and balance regimes before distillation. The empirical package is a real strength for a systems paper—multi-seed tables, smoothness-placement ablations with spectra/jitter, qualitative case studies, and threshold-/salience-aware metrics that are well motivated in the high-coverage regime. The contribution is incremental relative to SONIC/OmniH2O/expert-distillation recipes, but the capability-alignment framing and evaluation tools are of clear interest to the humanoid control community.
major comments (3)
- The central claim that residual failures are primarily recipe-induced capability bottlenecks (Secs. 1, 3.2–3.3; Eqs. 6–8) is only partially isolated from reference quality and physical feasibility. Rgen is defined by cSR < 0.8 on the general teacher; Fig. 2 and the Limitations section already admit data artifacts, near-limit motions, and incomplete teacher acquisition, yet Table 4 reports recovery on automatically constructed Ddynamic/Dbalance manifests (App. A.4) rather than on a verified recoverable residual subset with per-clip feasibility/reference filters. Please quantify what fraction of Rgen is recovered by no-smoothness/experts versus remains unsolved for reference/embodiment reasons, and report long-tail metrics on that filtered residual set.
- The abstract and introduction claim residual clips remain unsolved 'even under targeted training,' but the main experimental tables do not include a controlled data-only intervention baseline on the same residual set (e.g., oversampling or training only on Rgen under the SONIC-Base recipe with matched budget). Without that comparison, Table 4’s expert gains show that less conservative recipes track hard clips better, but do not fully establish that exposure alone is insufficient. Add this ablation, or narrow the claim to the evidence actually reported.
- All quantitative results are simulation-only on a proprietary, unreleased platform (Sec. 5.1, Limitations). The paper correctly scopes claims to recipe-level comparison on one embodiment, but for a journal contribution in humanoid WBC the deployability claim (action-rate recovery, RL fine-tuning as deployment stage) needs at least limited real-robot quantitative tracking/smoothness results, or a clearer demotion of hardware-readiness claims until such evidence exists.
Circularity Check
No circular derivation: empirical systems claims measured by independent rollouts, not tautologies of fitted inputs.
full rationale
Athena-WBC is a methods/systems paper whose load-bearing claims are empirical: residual training-set failures under a SONIC-recipe baseline, recovery after changing acquisition recipes (remove effort/temporal reward terms; gravity curriculum), and held-out gains after routed DAgger distillation plus RL fine-tuning. Residual sets R(π0)/Rgen (Eqs. 6–8) are defined by measured rollout success of a trained general teacher, not by the expert objectives or final student metrics; experts are then trained with altered rewards/curricula and evaluated with fixed thresholds, STC/TIS, MPJPE, and MPJPE-W on training and held-out splits. That experimental dependence on a chosen ρ_fail and on the authors’ reimplementation is ordinary methodology, not a reduction of the claimed result to its inputs by construction. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from overlapping authors, and no ansatz smuggled in as a first-principles derivation. Concerns about residual-set confounds (reference quality, physical saturation) affect causal interpretation of the capability-bottleneck claim, not circularity of the derivation chain. Score 0 with empty steps is therefore the correct outcome.
Axiom & Free-Parameter Ledger
free parameters (6)
- residual failure threshold ρ_fail
- Grad-CAPS / CAPS regularization weight λ_reg
- gravity curriculum range α_e ∈ [α_min, 1]
- adaptive sampling mixture ρ, EMA β_ema, violation bonus c_viol, difficulty cap D_max
- default success thresholds δz=0.20 m, δyaw=0.50 rad, δkb=0.50 m
- training budget 40k iterations / 16384 envs and RL fine-tuning schedule
axioms (6)
- ad hoc to paper A residual clip that remains unsolved after data-only interventions but becomes solvable after changing reward/curriculum is a capability bottleneck relative to embodiment, simulator, protocol, and budget.
- domain assumption Physical-constraint penalties encode feasibility, while effort and temporal-control penalties encode conservative control preferences that can suppress feasible high-dynamic actions.
- domain assumption Reduced-gravity continuation improves early survivability of balance-critical motions without changing the final evaluation dynamics.
- domain assumption Privileged teacher actions provide useful supervision for a deployable student via DAgger, and post-distillation RL fine-tuning improves closed-loop generalization.
- domain assumption Simulator rollouts under domain randomization with fixed success predicates are a valid proxy for recoverable tracking performance on the target embodiment.
- standard math PPO optimization, GAE, and standard continuous-control RL math are valid for the reported training procedure.
invented entities (4)
-
capability bottleneck (recipe-induced effective control regime)
no independent evidence
-
Athena-WBC capability-aligned expert bank (dynamic + balance teachers with motion routing)
no independent evidence
-
Success–Tolerance Curve (STC) and Threshold-Integrated Success (TIS)
no independent evidence
-
Motion-Salience Weighted MPJPE (MPJPE-W)
no independent evidence
read the original abstract
Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficult motions are sampled more often, isolated into smaller subsets, or assigned to specialized experts. We show that this view is incomplete. In strong whole-body-control baselines, a residual set of feasible training clips remains unsolved even under targeted training, especially for high-dynamic transitions and balance-critical motions. These failures arise not only from insufficient exposure, but from a mismatch between the motion demands and the effective capability induced by the default training recipe. We propose Athena-WBC, a compact teacher-student pipeline with capability-aligned policy experts for long-tail humanoid whole-body control. Dynamic experts use a tracking-focused, constraint-aware objective that removes conservative effort and temporal-control penalties while preserving physical feasibility constraints; balance experts use a gravity curriculum to improve early-training survivability. The resulting privileged teachers are motion-routed for DAgger distillation and then compressed into a single controller with deployable observations followed by RL fine-tuning. Experiments on a full-size humanoid show improved recovery of training-set long-tail motions and better held-out tracking than a strong SONIC-recipe baseline, using only a small number of experts.
Figures
Reference graph
Works this paper leans on
-
[1]
SONIC: Supersizing motion tracking for natural humanoid whole-body control, 2025
Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Fernando Castañeda, Sirui Chen, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, Jinhyung Park, David Sami, Zi Wang, Xingye Da, Runyu Ding, Cyrus Hogg, Lina Song, Edy Lim, Eugene Jeong, Tairan He, Haoru Xue, Wenli Xiao, Simon Yuen, Jan Kautz, Yan Chang, Umar Iqbal, Linxi Jim Fan, and Yuke Zhu. SONIC: Supersi...
Pith/arXiv arXiv 2025
-
[2]
Perpetual humanoid control for real-time simulated avatars
Zhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani, and Weipeng Xu. Perpetual humanoid control for real-time simulated avatars. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10861–10870, 2023. doi: 10.1109/ICCV51070.2023.01000
-
[3]
Kitani, Changliu Liu, and Guanya Shi
Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris M. Kitani, Changliu Liu, and Guanya Shi. OmniH2O: Universal and dexterous human-to-humanoid whole- body teleoperation and learning. In Pulkit Agrawal, Oliver Kroemer, and Wolfram Burgard, editors, Proceedings of The 8th Conference on Robot Learning, volume 270 ofProceedings of ...
2025
-
[4]
Chao Yang, Yingkai Sun, Peng Ye, Xin Chen, Chong Yu, and Tao Chen. EGM: Efficiently learning general motion tracking policy for high dynamic humanoid whole-body control, 2025. URL https://arxiv.org/abs/2512.19043
arXiv 2025
-
[5]
Zhenguo Sun, Bo-Sheng Huang, Yibo Peng, Xukun Li, Jingyu Ma, Yu Sun, Zhe Li, Haojun Jiang, Biao Gao, Zhenshan Bing, Xinlong Wang, and Alois Knoll. MOSAIC: Bridging the Sim-to-Real gap in generalist humanoid motion tracking and teleoperation with rapid residual adaptation, 2026. URLhttps://arxiv.org/abs/2602.08594. 21
arXiv 2026
-
[6]
Zhenguo Sun, Yibo Peng, Yuan Meng, Xukun Li, Bo-Sheng Huang, Zhenshan Bing, Xinlong Wang, and Alois Knoll. RobotDancing: Residual-action reinforcement learning enables robust long-horizon humanoid motion tracking, 2025. URLhttps://arxiv.org/abs/2509.20717
arXiv 2025
-
[7]
M3imic: Learning a versatile whole-body controller for multimodal motion mimicking, 2026
Zuxing Lu, Ziang Zheng, Yao Lyu, Jingyu Liu, Feihong Zhang, Song Lu, Xin Yuan, Changyin Sun, Xingxing Zuo, and Shengbo Eben Li. M3imic: Learning a versatile whole-body controller for multimodal motion mimicking, 2026. URLhttps://arxiv.org/abs/2606.04829
Pith/arXiv arXiv 2026
-
[8]
Stubborn: A streamlined and unified reinforcement learning framework for robust motion tracking and fall recovery for humanoids,
Xiao Ren, Yuhui Yang, Zongbiao Weng, Zhijie Liu, and He Kong. Stubborn: A streamlined and unified reinforcement learning framework for robust motion tracking and fall recovery for humanoids,
-
[9]
URLhttps://arxiv.org/abs/2606.12814
-
[10]
Agility meets stability: Versatile humanoid control with heterogeneous data, 2025
Yixuan Pan, Ruoyi Qiao, Li Chen, Kashyap Chitta, Liang Pan, Haoguang Mai, Qingwen Bu, Hao Zhao, Cunyuan Zheng, Ping Luo, and Hongyang Li. Agility meets stability: Versatile humanoid control with heterogeneous data, 2025. URLhttps://arxiv.org/abs/2511.17373
arXiv 2025
-
[11]
From experts to a generalist: Toward general whole-body control for humanoid robots,
Yuxuan Wang, Ming Yang, Weishuai Zeng, Yu Zhang, Xinrun Xu, Haobin Jiang, Ziluo Ding, and Zongqing Lu. From experts to a generalist: Toward general whole-body control for humanoid robots,
-
[12]
URLhttps://arxiv.org/abs/2506.12779
-
[13]
Humanoid-GPT: Scaling data and structure for zero-shot motion tracking, 2026
Zekun Qi, Xuchuan Chen, Dairu Liu, Chenghuai Lin, Yunrui Lian, Sikai Liang, Zhikai Zhang, Yu Guan, Jilong Wang, Wenyao Zhang, Xinqiang Yu, He Wang, and Li Yi. Humanoid-GPT: Scaling data and structure for zero-shot motion tracking, 2026. URL https://arxiv.org/abs/2606. 03985
2026
-
[14]
A reduction of imitation learning and structured prediction to no-regret online learning
Stephane Ross, Geoffrey Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Geoffrey Gordon, David Dunson, and Miroslav Dudík, editors,Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, volume 15 ofProceedings of Machine Learning Research, pag...
2011
-
[15]
GMT: General motion tracking for humanoid whole-body control, 2025
Zixuan Chen, Mazeyu Ji, Xuxin Cheng, Xuanbin Peng, Xue Bin Peng, and Xiaolong Wang. GMT: General motion tracking for humanoid whole-body control, 2025. URL https://arxiv.org/abs/ 2506.14770
Pith/arXiv arXiv 2025
-
[16]
Estrada, Francesco Iacobelli, Twan Koolen, Alexander Lambert, Erica Lin, M
Jean Pierre Sleiman, He Li, Alphonsus Adu-Bredu, Robin Deits, Arun Kumar, Kevin Bergamin, Mohak Bhardwaj, Scott Biddlestone, Nicola Burger, Matthew A. Estrada, Francesco Iacobelli, Twan Koolen, Alexander Lambert, Erica Lin, M. Eva Mungai, Zach Nobles, Shane Rozen-Levy, Yuyao Shi, Jiashun Wang, Jakob Welner, Fangzhou Yu, Mike Zhang, Alfred Rizzi, Jessica H...
arXiv 2026
-
[17]
OmniTrack: General motion tracking via physics-consistent reference, 2026
Yuhan Li, Peiyuan Zhi, Yunshen Wang, Tengyu Liu, Sixu Yan, Wenyu Liu, Xinggang Wang, Baoxiong Jia, and Siyuan Huang. OmniTrack: General motion tracking via physics-consistent reference, 2026. URLhttps://arxiv.org/abs/2602.23832
arXiv 2026
-
[18]
Regularizing action policies for smooth control with reinforcement learning
Siddharth Mysore, Bassel Mabsout, Renato Mancuso, and Kate Saenko. Regularizing action policies for smooth control with reinforcement learning. In2021 IEEE International Conference on Robotics and Automation (ICRA), pages 1810–1816, 2021. doi: 10.1109/ICRA48506.2021.9561138
-
[19]
Lee, Hoang-Giang Cao, Cong-Tinh Dao, Yu-Cheng Chen, and I-Chen Wu
I. Lee, Hoang-Giang Cao, Cong-Tinh Dao, Yu-Cheng Chen, and I-Chen Wu. Gradient-based regularization for action smoothness in robotic control with reinforcement learning. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 603–610, 2024. doi: 10.1109/IROS58592.2024.10801464
-
[20]
Regularization matters in policy optimization – an empirical study on continuous control
Zhuang Liu, Xuanlin Li, Bingyi Kang, and Trevor Darrell. Regularization matters in policy optimization – an empirical study on continuous control. InInternational Conference on Learning Representations, 2021. URLhttps://openreview.net/forum?id=yr1mzrH3IC. 22
2021
-
[21]
AWAC: Accelerating online reinforcement learning with offline datasets, 2020
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine. AWAC: Accelerating online reinforcement learning with offline datasets, 2020. URLhttps://arxiv.org/abs/2006.09359
Pith/arXiv arXiv 2020
-
[22]
Residual force control for agile human behavior imitation and extended motion synthesis
Ye Yuan and Kris Kitani. Residual force control for agile human behavior imitation and extended motion synthesis. InAdvances in Neural Information Processing Sys- tems, volume 33, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/ f76a89f0cb91bc419542ce9fa43902dc-Abstract.html
2020
-
[23]
Learning motion skills with adaptive assistive curriculum force in humanoid robots, 2025
Zhanxiang Cao, Yang Zhang, Buqing Nie, Huangxuan Lin, Haoyang Li, and Yue Gao. Learning motion skills with adaptive assistive curriculum force in humanoid robots, 2025. URL https: //arxiv.org/abs/2506.23125
Pith/arXiv arXiv 2025
-
[24]
Nikita Rudin, Junzhe He, Joshua Aurand, and Marco Hutter. Parkour in the wild: Learning a general and extensible agile locomotion policy using multi-expert distillation and RL fine-tuning. The International Journal of Robotics Research, 2026. doi: 10.1177/02783649261455067. URL https://doi.org/10.1177/02783649261455067. OnlineFirst
-
[25]
Troje, Gerard Pons-Moll, and Michael J
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 5441–5450, 2019. doi: 10.1109/ICCV .2019.00554
doi:10.1109/iccv 2019
-
[26]
BEAT: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis
Haiyang Liu, Zihao Zhu, Naoya Iwamoto, Yichen Peng, Zhengqing Li, You Zhou, Elif Bozkurt, and Bo Zheng. BEAT: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis. InComputer Vision – ECCV 2022, pages 612–630, 2022. doi: 10.1007/ 978-3-031-20071-7_36. 23 A Appendix A.1 Training Hyperparameters Table 7 lists the PP...
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.