REVIEW 3 major objections 5 minor 131 references
GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read GenTrack claims that online co-training of a text-to-motion generator and a humanoid tracker improves both generator executability and zero-shot tracking coverage beyond what static replay or one-way alignment achieves.
desk verdict A well-controlled co-training study whose own small margins and missing error bars undercut the 'markedly broader coverage' abstract claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an online sample–execute–update loop with a lagged judge. Each round, the generator samples $K$ motions per prompt, the tracker from the preceding round (kept frozen during generator optimization) executes them, and the execution score $S_{\mathrm{exec}} = (1-c) + [e_j]^2 + [e_t/0.5]^2 + 0.5[e_d/0.5]^2 + 2\mathbb{I}_{\mathrm{fall}}$, with $[x]_2 = \min(x,2)$, turns falls, incomplete rollouts, joint error ($e_j$), root-trajectory error ($e_t$), and root-displacement error ($e_d$) into a group-relative reward: rewards are normalized within each same-prompt group and applied by FlowGRPO with four clipped policy-ratio updates. Drift is constrained by a KL penalty to the frozen initial generator and periodic supervised flow-matching rehearsal on the original text–motion pairs. The same structurally valid generated references are accumulated and mixed in equal share with public references to update the tracker, so the reward model and the tracker co-evolve; the current trainee never gates its own feedback.
What would settle it
Replace the frozen SONIC evaluation with a different closed-loop controller (for example the ProtoMotions tracker) or tighten the fall criterion to 0.10 m pelvis-height deviation and re-run the matched GenTrack versus static-replay comparison; if the co-trained generator's success advantage shrinks or inverts, the claim that the loop yields genuinely more executable motion is falsified.
Extended reading notes
Core claim
GenTrack couples a pretrained text-to-motion generator with a pretrained humanoid tracker for the Unitree G1 in an online, alternating loop. Each round, the generator samples robot-space references from training prompts; a tracker frozen from the previous round executes them in closed loop, producing an execution score $S_{\mathrm{exec}} = (1-c) + [e_j]^2 + [e_t/0.5]^2 + 0.5[e_d/0.5]^2 + 2\mathbb{I}_{\mathrm{fall}}$ with $[x]_2=\min(x,2)$, where $c$ is completion fraction, $e_j$ the maximum wrapped joint error, $e_t$ mean root-trajectory error, $e_d$ root-displacement error, and $\mathbb{I}_{\mathrm{fall}}$ a fall indicator. Rewards are normalized within each same-prompt group and applied through FlowGRPO, while a frozen-generator KL penalty and supervised flow-matching rehearsal on the original text–motion pairs limit drift. The structurally valid on-policy generations are accumulated and mixed in equal share with public references to update the tracker. On the SONIC backbone, the loop raises frozen-SONIC generator success from 92.58% to 94.43%, lowers key-body error from 0.410 m to 0.325 m, and improves fall-only zero-shot tracking success on LAFAN1/AMASS-test/Wild-G1 from 85.0/79.0/47.2 to 90.0/79.7/48.0; matched static-replay and one-way filtering controls do not reproduce the gains. The paper concludes that joint online post-training narrows the executability gap between retargeted references and robot-native motion.
Load-bearing premise
The claims rest on a single simulated evaluator—closed-loop rollouts of a frozen SONIC policy with a 0.25 m pelvis-height fall threshold—being a faithful stand-in for real robot executability.
Editorial extensions
If this is right
- Generator executability improves alongside semantics: the SONIC branch raises frozen-SONIC success from 92.58% to 94.43% while preserving or improving TMR-G1 retrieval and FID.
- Zero-shot tracker coverage extends beyond the static reference pool: fall-only success on the three frozen splits rises on the SONIC backbone, with the largest relative gain on the most out-of-distribution split.
- The recipe transfers across backbone designs: both ProtoMotions (AMP/PPO) and SONIC benefit, so the effect is not tied to one tracker training paradigm.
- Equal-budget offline controls—static replay, final-generator replay, reward-weighted SFT, and DPO—fall short of the online loop, indicating that temporal co-adaptation is the active mechanism.
- Without new robot data, the loop converts language-conditioned generation and closed-loop execution into a self-sustaining source of tracker supervision.
Reading between the lines
- If the loop's gains hold across morphologies, the same alternating recipe could be applied to other robot platforms and to other generative backbones (latent diffusion or autoregressive motion models), which the paper does not test.
- The fall-only success criterion tolerates up to 0.25 m reference-relative pelvis-height deviation; the method's practical value on real hardware would be better probed with a stricter, task- or contact-based success metric that the paper leaves to future work.
- Removing the KL and rehearsal anchors raises raw executability but degrades TMR retrieval and FID in the reported ablations, which suggests the approach implicitly treats preserved text–motion alignment as part of "executability"; a deployment-focused variant might choose a different trade-off.
- The paper's own ablation shows that using the current trainee as sole judge collapses cross-split success, implying a testable extension: varying the lag between tracker updates and reward scoring to find how much judge staleness is needed for stable co-adaptation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes GenTrack, an online co-training framework that alternately updates a pretrained text-to-motion generator and a pretrained humanoid tracker. The generator is aligned with execution-grounded, group-relative rewards computed from a frozen lagged tracker, while the tracker is trained on a mixture of a fixed public reference pool and newly generated references that pass only structural validity checks. Anchoring via a frozen-generator KL penalty and supervised rehearsal is used to limit drift. The method is evaluated on simulated Unitree G1 with ProtoMotions and SONIC backbones, reporting generator executability via frozen-SONIC rollouts and zero-shot tracker coverage on three held-out splits, alongside matched static-replay, one-way, and objective ablations.
Significance. If the reported gains are robust, the contribution is meaningful: it offers a way to extend zero-shot humanoid tracking coverage without additional embodied data collection, while using closed-loop execution feedback to make a text-to-motion generator produce more robot-compatible references. The experimental design is careful in several respects: matched budgets and initializations across arms, separation of the reward judge from the current trainee, no success-gating of generated references admitted to tracker training, and multiple one-way/offline controls. The method is also presented with a clear statement that evaluation is simulation-only. However, the central claims of markedly broader coverage and general robot-native executability currently rest on small margins, single deterministic rollouts, and a single simulated executor, so the evidence is not yet conclusive.
major comments (3)
- [§4, Table 1; §C.1] The tracker-side success gains are small and are reported without uncertainty quantification. Section C.1 disables observation corruption, reset perturbations, and startup randomization, so each number is a single deterministic rollout per test clip. Under the binary fall-only criterion, which tolerates up to 0.25 m reference-relative pelvis-height deviation, a +0.7 or +0.8 percentage-point difference on AMASS-test or Wild-G1 could be a handful of clips, and per-clip counts are not reported. The abstract's phrase 'markedly broader zero-shot coverage' is not supported by margins of this size without per-split counts, bootstrap confidence intervals, or repeated-seed statistics. Please provide such uncertainty quantification or substantially temper the coverage claim.
- [§3, Eq. (3)–(5); §4, Table 2; §E] Generator executability is measured only with a frozen SONIC policy in IsaacLab, and the SONIC branch's generator reward is produced by the same frozen SONIC policy family. Although the ProtoMotions branch provides partial independence, all generator rows are evaluated by the same SONIC executor, so the results could reflect evaluator-specific adaptation rather than a general robot-native executability gain. The paper honestly limits the evaluation to simulation in Section E, but the abstract and conclusion generalize to 'robot-native' motion and 'zero-shot humanoid tracking' beyond that scope. Please add a second independent executor, a cross-controller evaluation, or real-hardware spot checks, or explicitly restrict the executability claims to the single simulated SONIC evaluator.
- [§4, Table 1 (ProtoMotions rows)] The ProtoMotions branch does not consistently improve over matched controls: GenTrack scores 75.0 on LAFAN1 versus 77.5 for G0 replay and equals the 75.0 baseline, while AMASS-test and Wild-G1 improve by only +2.2 and +0.5 percentage points relative to G0 replay. The statement that online co-training 'consistently produces ... trackers with markedly broader zero-shot coverage' across both backbones is therefore an overstatement. The evidence supports at most a split-dependent benefit for SONIC and a small Wild-G1 improvement for ProtoMotions; please revise the framing or provide a principled aggregate analysis that supports the broader claim.
minor comments (5)
- [Table 1] Any2Track's reported 100.0% on LAFAN1 against 5.1% and 10.4% on AMASS-test and Wild-G1 is surprising; please clarify how the external tracker's released success flags and trajectory conversion affect this row.
- [§C.1, Table 4] The metric table lists many auxiliary diagnostics (completion, unexpected fall, foot skate, penetration, kinematic diagnostics, judge calibration) that are not reported in the main text; please state explicitly which of these are available in the supplement and which are omitted.
- [§4, Tables 2 and 5] FID and R-Precision are reported to three decimals from deterministic single evaluations; such precision overstates the reliability of the estimates. Please include uncertainty bounds or state that these are point estimates from a single evaluation.
- [§4, Table 2] The private 1,024-prompt generator suite is described only in the supplement; please include at least its construction, prompt coverage, and exclusion rules in the main text so the reader can assess the generator claims without accessing the supplement.
- [§3, 'Execution reward'] The term 'lagged' is used in the main text before it is defined; consider clarifying at first use that it refers to the tracker from the preceding round, frozen during the generator phase.
Circularity Check
No significant circularity; central claims rest on a frozen external SONIC evaluator and held-out tracker splits, with only a minor non-load-bearing self-citation to RLPF.
-
other
[Section C.1 (Generator metrics) and Table 4 caption]
"Following the high-level/low-level evaluation decomposition of RLPF (Yue et al. 2025), generator quality is measured by semantic generation metrics and by official frozen-SONIC execution in IsaacLab. ... Low-level generator metrics follow the RLPF structure."
The execution-reward/evaluation decomposition is inherited from RLPF, a prior paper sharing an author with this submission, so the measure of 'physical fidelity' is not fully external to the authors' own lineage. This is not a load-bearing reduction, however: the reward is redefined in Eq. 4, the generator comparisons all share one frozen official SONIC executor that is not the training judge, and the core ablation compares GenTrack against matched static-replay and one-way controls. The citation is thus a provenance note rather than a mathematical equivalence between input and output.
full rationale
The main derivation chain is self-contained. The generator is optimized with a group-relative reward computed from a lagged tracker that is frozen during each generator phase, and the current trainee has zero reward weight (Section 3, Eqs. 3-4), so the generator is not its own judge. Generated references are admitted to tracker training by structural validity only, not by the tracker's own rollout success (Section 3, 'Generated-reference curriculum'), and tracker updates mix equal numbers of public and generated transitions. Tracker-side evaluation uses frozen held-out splits (LAFAN1-G1, AMASS-test-G1, Wild-G1-clean) with matched reference-only continuation, G0 replay, and final-generator replay controls, so the coverage gains are not forced by construction. Generator-side evaluation uses a single frozen official SONIC IsaacLab policy shared by every generator row, with TMR-G1 semantic metrics held out from the reward; the improved success of GenTrack(SONIC) is therefore checked against an external executor rather than the in-branch lagged model. The only mild issue is that the reward/evaluation structure follows RLPF, a paper sharing an author, but because Eq. 4 is defined in-paper and the key comparisons use external and matched controls, this citation is not load-bearing. Score 2 reflects that minor self-citation; no prediction reduces to its input by construction.
Assumptions & free parameters
free parameters (4)
- Execution score weights in Eq. (4) =
coefficients: 1 (completion/joint), 4 (root error), 2 (displacement), 2 (fall), saturation cap 2, threshold 0.5
- KL anchor weight lambda_KL =
0.02
- GT anchor weight lambda_GT =
1.0 with update every two GRPO iterations
- Policy ratio replay updates per group =
4
assumptions (6)
- domain assumption MuJoCo simulation of Unitree G1 with the frozen SONIC policy faithfully represents robot executability.
- domain assumption The fall-only success criterion (max reference-relative pelvis height deviation <= 0.25 m) is a sufficient measure of tracking success.
- domain assumption GMD retargeting of the internal human corpus yields valid, representative G1 robot references.
- domain assumption Wild-G1-clean is an out-of-distribution test set.
- domain assumption TMR-G1 latent-space metrics measure semantic preservation.
- domain assumption The lagged tracker's execution score is a reliable reward without degenerate bias beyond the retention anchors.
Cite this review
Pith. "Pith review of GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking." pith.science (2026). https://pith.science/paper/FUB6EBHK
@misc{pith2026260801410,
author = {Pith},
title = {Pith review of: GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking},
year = {2026},
howpublished = {\url{https://pith.science/paper/FUB6EBHK}},
note = {Machine review of arXiv:2608.01410}
}
read the original abstract
General-purpose humanoid trackers can execute diverse references, but their zero-shot coverage depends on large embodied corpora that are costly to extend. Text-to-motion generators offer scalable supervision, yet models trained on human motion or retargeted data inherit a gap between kinematic plausibility and robot executability. Existing one-way pipelines fix either the generated corpus or the reward tracker. We introduce GenTrack, an online generator--tracker framework that alternates execution-grounded, group-relative generator alignment with tracker training on newly generated references; anchoring and rehearsal constrain drift. On Unitree G1, we evaluate GenTrack with ProtoMotions and SONIC backbones across three zero-shot tracking splits including public AMASS and LAFAN benchmarks, and a private out-of-distribution test set of 1,024 prompt-motion pairs in the wild. The online co-training strategy consistently produces generators that output more robot-executable motions with strong semantic alignment, and trackers with markedly broader zero-shot coverage and improved tracking accuracy, especially on out-of-distribution references. These results demonstrate that joint online post-training effectively narrows the executability gap between retargeted references and robot-native motion, advancing zero-shot humanoid control without additional data collection and beyond the limitations of a static reference pool.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[2]
Bao, J.; Yang, H.; Xin, Y.; Liu, J.; Xu, Y.; Liang, H.; Han, P.; Ma, X.; Wang, D.; and Zhao, B. 2026. PhyGile : Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking. arXiv preprint arXiv:2603.19305
work page Pith review arXiv 2026
-
[3]
Barquero, G.; Escalera, S.; and Palmero, C. 2024. Seamless Human Motion Composition with Blended Positional Encodings. In CVPR, 457--469
2024
-
[4]
Cao, J.; Chen, Y.; and Tomizuka, M. 2026. CLAW : Composable Language-Annotated Whole-Body Motion Generation. arXiv preprint arXiv:2604.11251
arXiv 2026
-
[5]
Chen, M.; Wang, K.; Zhang, B.; Ma, X.; Yang, Z.; Ren, Y.; Huang, Q.; Zhu, Z.; Wang, Y.; and Su, Z. 2026. HoloMotion -1 Technical Report. arXiv preprint arXiv:2605.15336
arXiv 2026
-
[6]
Chen, W.; Xiao, H.; Zhang, E.; Hu, L.; Wang, L.; Liu, M.; and Chen, C. 2024. SATO : Stable Text-to-Motion Framework. arXiv preprint arXiv:2405.01461
work page Pith review arXiv 2024
-
[7]
Chen, X.; Jiang, B.; Liu, W.; Huang, Z.; Fu, B.; Chen, T.; and Yu, G. 2023. Executing Your Commands via Motion Diffusion in Latent Space. In CVPR, 18000--18010
2023
-
[9]
Cho, H.; Kim, S.-H.; Kang, J.; and Koo, D. 2026. SafeFlow : Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating. arXiv preprint arXiv:2603.23983
arXiv 2026
-
[10]
Dai, W.; Chen, L.-H.; Wang, J.; Liu, J.; Dai, B.; and Tang, Y. 2024. MotionLCM : Real-Time Controllable Motion Generation via Latent Consistency Model. In ECCV
2024
Show all 131 references
-
[12]
Fu, Z.; Zhao, Q.; Wu, Q.; Wetzstein, G.; and Finn, C. 2024. HumanPlus : Humanoid Shadowing and Imitation from Humans. arXiv preprint arXiv:2406.10454
2024 arXiv
-
[13]
Gao, L.; Yang, F.; Chen, J.; Liu, L.; Zheng, Y.; Cai, Y.; and Li, Z. 2026. QuadFM : Foundational Text-Driven Quadruped Motion Dataset for Generation and Control. arXiv preprint arXiv:2603.24021
2026
-
[14]
G.; Wang, S.; and Cheng, L
Guo, C.; Mu, Y.; Javed, M. G.; Wang, S.; and Cheng, L. 2024. MoMask : Generative Masked Modeling of 3D Human Motions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1900--1910
2024
-
[15]
Guo, C.; Zou, S.; Zuo, X.; Wang, S.; Ji, W.; Li, X.; and Cheng, L. 2022. Generating Diverse and Natural 3D Human Motions from Text. In CVPR
2022
-
[16]
Han, J.; Xie, W.; Zheng, J.; Shi, J.; Zhang, W.; Xiao, T.; and Bai, C. 2025. KungfuBot2 : Learning Versatile Motion Skills for Humanoid Whole-Body Control. arXiv preprint arXiv:2509.16638
2025
-
[17]
Harithas, S.; Kwak, S.; Katara, P.; Deolasee, S.; Kalaria, D.; Sridhar, S.; Vemprala, S.; Kapoor, A.; and Huang, J. C.-K. 2026. DreamControl-v2 : Simpler and Scalable Autonomous Humanoid Skills via Trainable Guided Diffusion Priors. arXiv preprint arXiv:2604.00202
2026
-
[18]
G.; Yurick, M.; Nowrouzezahrai, D.; and Pal, C
Harvey, F. G.; Yurick, M.; Nowrouzezahrai, D.; and Pal, C. 2020. Robust Motion In-Betweening. ACM Transactions on Graphics
2020
-
[19]
He, T.; Gao, J.; Xiao, W.; Zhang, Y.; Wang, Z.; Wang, J.; Luo, Z.; He, G.; Sobanbabu, N.; Pan, C.; Yi, Z.; Qu, G.; Kitani, K.; Hodgins, J.; Fan, L.; Zhu, Y.; Liu, C.; and Shi, G. 2025. ASAP : Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Ski...
2025 arXiv
-
[20]
He, T.; Luo, Z.; He, X.; Xiao, W.; Zhang, C.; Zhang, W.; Kitani, K.; Liu, C.; and Shi, G. 2024 a . OmniH2O : Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning. arXiv preprint arXiv:2406.08858
2024 arXiv
-
[22]
Huang, T.; Yuan, F.; Gu, J.; Fang, S.; Zhang, X.; Wang, Y.; Gao, W.; and Zhang, S. 2026. Human2Humanoid : Physics-Aware Cross-Morphology Motion Retargeting for Humanoid Robots. arXiv preprint arXiv:2606.03476
2026 arXiv
-
[23]
Jiang, B.; Chen, X.; Liu, W.; Yu, J.; Yu, G.; and Chen, T. 2023. MotionGPT : Human Motion as a Foreign Language. In NeurIPS
2023
-
[24]
Li, B.; Zhang, R.; Liang, H.; Zhang, J.; Zhang, J.; Chen, X.; and Wang, J. 2026 a . MIND : Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control. arXiv preprint arXiv:2605.26006
2026 arXiv
-
[25]
Li, P.; Zhuang, Z.; Gao, Y.; Dong, Y.; Li, S.; Jiang, C.; Dou, S.; Xi, Z.; Zhou, E.; Huang, J.; Li, H.; Gong, J.; Ma, X.; Gui, T.; Wu, Z.; Zhang, Q.; Huang, X.; Jiang, Y.-G.; and Qiu, X. 2026 b . FRoM-W1 : Towards General Humanoid Whole-Body Control with Language Instructions....
2026
-
[26]
Li, Y.; Luo, Z.; Zhang, T.; Dai, C.; Kanervisto, A.; Tirinzoni, A.; Weng, H.; Kitani, K.; Guzek, M.; Touati, A.; Lazaric, A.; Pirotta, M.; and Shi, G. 2026 c . BFM-Zero : A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning. I...
2026
-
[27]
Li, Y.; Zhi, P.; Wang, Y.; Liu, T.; Yan, S.; Liu, W.; Wang, X.; Jia, B.; and Huang, S. 2026 d . OmniTrack : General Motion Tracking via Physics-Consistent Reference. arXiv preprint arXiv:2602.23832
2026
-
[28]
Li, Z.; Chi, C.; Wei, Y.; Zhu, B.; Peng, Y.; Huang, T.; Wang, P.; Wang, Z.; Zhang, S.; and Xu, C. 2026 e . From Language to Locomotion: Retargeting-Free Humanoid Control via Motion Latent Guidance. ICLR
2026
-
[29]
E.; Huang, X.; Gao, Y.; Tevet, G.; Sreenath, K.; and Liu, C
Liao, Q.; Truong, T. E.; Huang, X.; Gao, Y.; Tevet, G.; Sreenath, K.; and Liu, C. K. 2025. BeyondMimic : From Motion Tracking to Versatile Humanoid Control via Guided Diffusion. arXiv preprint arXiv:2508.08241
2025 arXiv
-
[30]
Liu, J.; Liu, G.; Liang, J.; Li, Y.; Liu, J.; Wang, X.; Wan, P.; Zhang, D.; and Ouyang, W. 2025 a . Flow-GRPO : Training Flow Matching Models via Online RL. arXiv preprint arXiv:2505.05470
2025 arXiv
-
[32]
Lu, S.; Chen, L.-H.; Zeng, A.; Lin, J.; Zhang, R.; Zhang, L.; and Shum, H.-Y. 2024. HumanTOMATO : Text-Aligned Whole-Body Motion Generation. In ICML
2024
-
[35]
F.; Pons-Moll, G.; and Black, M
Mahmood, N.; Ghorbani, N.; Troje, N. F.; Pons-Moll, G.; and Black, M. J. 2019. AMASS : Archive of Motion Capture as Surface Shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision
2019
-
[36]
NVLabs . 2025. ProtoMotions : GPU-Accelerated Simulation and Learning for Humanoids. https://github.com/NVlabs/ProtoMotions
2025
-
[37]
Pinyoanuntapong, E.; Wang, P.; Lee, M.; and Chen, C. 2024. MMM : Generative Masked Motion Model. In CVPR
2024
-
[38]
Plappert, M.; Mandery, C.; and Asfour, T. 2016. The KIT Motion-Language Dataset. In Big Data
2016
-
[39]
Qi, Z.; Chen, X.; Liu, D.; Lin, C.; Lian, Y.; Liang, S.; Zhang, Z.; Guan, Y.; Wang, J.; Zhang, W.; Yu, X.; Wang, H.; and Yi, L. 2026. Humanoid-GPT : Scaling Data and Structure for Zero-Shot Motion Tracking. arXiv preprint arXiv:2606.03985
2026 arXiv
-
[40]
Rempe, D.; Petrovich, M.; Yuan, Y.; Zhang, H.; Peng, X. B.; Jiang, Y.; Wang, T.; Iqbal, U.; Minor, D.; de Ruyter, M.; Li, J.; Tessler, C.; Lim, E.; Jeong, E.; Wu, S.; Hassani, E.; Huang, M.; Yu, J.-B.; Chung, C.; Song, L.; Dionne, O.; Kautz, J.; Yuen, S.; and Fidler, S. 2026. ...
2026
-
[41]
Serifi, A.; Grandia, R.; Knoop, E.; Gross, M.; and B \"a cher, M. 2024. Robot Motion Diffusion Model : Motion Generation for Robotic Characters. In SIGGRAPH Asia 2024 Conference Papers, 1--9
2024
-
[42]
Tao, Z.; Su, Z.; Liu, P.; Sun, J.; Que, W.; Ma, J.; Yu, J.; Cao, J.; Sun, P.; Liang, H.; Han, G.; Zhao, W.; Xu, Z.; Tang, J.; Zhang, Q.; and Guo, Y. 2026. Heracles : Bridging Precise Tracking and Generative Synthesis for General Humanoid Control. arXiv preprint arXiv:2603.27756
2026
-
[43]
B.; Bermano, A
Tevet, G.; Raab, S.; Cohan, S.; Reda, D.; Luo, Z.; Peng, X. B.; Bermano, A. H.; and van de Panne, M. 2025. CLoSD : Closing the Loop between Simulation and Diffusion for Multi-Task Character Control. In ICLR
2025
-
[44]
Tevet, G.; Raab, S.; Gordon, B.; Shafir, Y.; Cohen-Or, D.; and Bermano, A. H. 2023. Human Motion Diffusion Model. In ICLR
2023
-
[46]
Xie, W.; Zheng, J.; Han, J.; Shi, J.; Zhang, W.; Bai, C.; and Li, X. 2026. TextOp : Real-Time Interactive Text-Driven Humanoid Robot Motion Generation and Control. arXiv preprint arXiv:2602.07439
2026
-
[47]
Xu, M.; Shi, Y.; Yin, K.; and Peng, X. B. 2025. PARC : Physics-based Augmentation with Reinforcement Learning for Character Controllers. In SIGGRAPH Conference Papers
2025
-
[48]
Xu, W.; Wu, Q.; Zhang, J.; Tan, J.; Li, Y.; Fang, Y.; Xiong, J.; Wu, K.; Ou, R.; and Xu, R. 2026. Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control. In CVPR, 16398--16407
2026
-
[49]
Yin, K.; Zeng, W.; Fan, K.; Dai, M.; Wang, Z.; Zhang, Q.; Tian, Z.; Wang, J.; Pang, J.; and Zhang, W. 2025. UniTracker : Learning Universal Whole-Body Motion Tracker for Humanoid Robots. arXiv preprint arXiv:2507.07356
2025
-
[50]
Yuan, X.; Li, Z.; Lyu, B.; Zuo, K.; Lu, Y.; Li, G.; and Yang, J. 2026. RoboForge : Physically Optimized Text-Guided Whole-Body Locomotion for Humanoids. arXiv preprint arXiv:2603.17927
2026
-
[51]
Yuan, Y.; Song, J.; Iqbal, U.; Vahdat, A.; and Kautz, J. 2023. PhysDiff : Physics-Guided Human Motion Diffusion Model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 16010--16021
2023
-
[54]
Zhang, J.; Liang, H.; Zhang, R.; Li, B.; Zhang, J.; Chen, X.; Wang, J.; Xu, L.; and Yu, J. 2026 a . SCRIPT : Scalable Diffusion Policy with Multi-Stage Training for Language-Driven Physics-Based Humanoid Control. arXiv preprint arXiv:2605.22894
2026 arXiv
-
[55]
Zhang, J.; Zhang, Y.; Cun, X.; Zhang, Y.; Zhao, H.; Lu, H.; Shen, X.; and Shan, Y. 2023 a . Generating Human Motion from Textual Descriptions with Discrete Representations. In CVPR, 14730--14740
2023
-
[56]
Zhang, M.; Cai, Z.; Pan, L.; Hong, F.; Guo, X.; Yang, L.; and Liu, Z. 2024. MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6): 4115--4128
2024
-
[57]
Zhang, M.; Guo, X.; Pan, L.; Cai, Z.; Hong, F.; Li, H.; Yang, L.; and Liu, Z. 2023 b . ReMoDiffuse : Retrieval-Augmented Motion Diffusion Model. In ICCV
2023
-
[63]
Zhou, Y.; Barnes, C.; Lu, J.; Yang, J.; and Li, H. 2019. On the Continuity of Rotation Representations in Neural Networks. In CVPR
2019
-
[64]
Zhuang, Z.; Wang, T.; Zou, B.; Luo, X.; Ma, J.; Zhou, H.; Liu, J.; and Wang, D. 2025. Humanoid-R0 : Bridging Text-to-Motion Generation and Physical Deployment via RL. OpenReview, withdrawn ICLR 2026 submission
2025
-
[65]
CVPR , year=
On the Continuity of Rotation Representations in Neural Networks , author=. CVPR , year=
-
[66]
Plappert, Matthias and Mandery, Christian and Asfour, Tamim , booktitle=. The
-
[67]
ACM Transactions on Graphics , year=
Robust Motion In-Betweening , author=. ACM Transactions on Graphics , year=
-
[68]
CVPR , year=
Generating Diverse and Natural 3D Human Motions from Text , author=. CVPR , year=
-
[69]
ICLR , year=
Human Motion Diffusion Model , author=. ICLR , year=
-
[70]
IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=
MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence , volume=. 2024 , doi=
2024
-
[71]
CVPR , pages=
Executing Your Commands via Motion Diffusion in Latent Space , author=. CVPR , pages=
-
[72]
CVPR , pages=
Generating Human Motion from Textual Descriptions with Discrete Representations , author=. CVPR , pages=
-
[73]
Jiang, Biao and Chen, Xin and Liu, Wen and Yu, Jingyi and Yu, Gang and Chen, Tao , booktitle=
-
[74]
Guo, Chuan and Mu, Yuxuan and Javed, Muhammad Gohar and Wang, Sen and Cheng, Li , booktitle=
-
[75]
Rempe, Davis and Petrovich, Mathis and Yuan, Ye and Zhang, Haotian and Peng, Xue Bin and Jiang, Yifeng and Wang, Tingwu and Iqbal, Umar and Minor, David and de Ruyter, Michael and Li, Jiefeng and Tessler, Chen and Lim, Edy and Jeong, Eugene and Wu, Sam and Hassani, Ehsan and H...
-
[76]
arXiv preprint arXiv:2503.09015 , year=
Natural Humanoid Robot Locomotion with Generative Motion Prior , author=. arXiv preprint arXiv:2503.09015 , year=
-
[77]
Yuan, Ye and Song, Jiaming and Iqbal, Umar and Vahdat, Arash and Kautz, Jan , booktitle=
-
[78]
Peng, Xue Bin and Abbeel, Pieter and Levine, Sergey and van de Panne, Michiel , journal=
-
[79]
2022 , doi=
Peng, Xue Bin and Guo, Yunrong and Halper, Lina and Levine, Sergey and Fidler, Sanja , journal=. 2022 , doi=
2022
-
[80]
ICCV , year=
Perpetual Humanoid Control for Real-time Simulated Avatars , author=. ICCV , year=
-
[81]
Fu, Zipeng and Zhao, Qingqing and Wu, Qi and Wetzstein, Gordon and Finn, Chelsea , journal=
-
[82]
He, Tairan and Luo, Zhengyi and He, Xialin and Xiao, Wenli and Zhang, Chong and Zhang, Weinan and Kitani, Kris and Liu, Changliu and Shi, Guanya , journal=
-
[83]
He, Tairan and Gao, Jiawei and Xiao, Wenli and Zhang, Yuanhang and Wang, Zi and Wang, Jiashun and Luo, Zhengyi and He, Guanqi and Sobanbabu, Nikhil and Pan, Chaoyi and Yi, Zeji and Qu, Guannan and Kitani, Kris and Hodgins, Jessica and Fan, Linxi and Zhu, Yuke and Liu, Changliu...
-
[84]
and van de Panne, Michiel , booktitle=
Tevet, Guy and Raab, Sigal and Cohan, Setareh and Reda, Daniele and Luo, Zhengyi and Peng, Xue Bin and Bermano, Amit H. and van de Panne, Michiel , booktitle=
-
[85]
SIGGRAPH Asia 2024 Conference Papers , pages=
Serifi, Agon and Grandia, Ruben and Knoop, Espen and Gross, Markus and B. SIGGRAPH Asia 2024 Conference Papers , pages=. 2024 , doi=
2024
-
[86]
arXiv preprint arXiv:2506.12769 , year=
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control , author=. arXiv preprint arXiv:2506.12769 , year=
-
[87]
Liu, Jie and Liu, Gongye and Liang, Jiajun and Li, Yangguang and Liu, Jiaheng and Wang, Xintao and Wan, Pengfei and Zhang, Di and Ouyang, Wanli , journal=
-
[88]
CVPR , pages=
Diffusion Model Alignment Using Direct Preference Optimization , author=. CVPR , pages=
-
[89]
arXiv preprint arXiv:2404.09445 , year=
Exploring Text-to-Motion Generation with Human Preference , author=. arXiv preprint arXiv:2404.09445 , year=
-
[90]
Chen, Wenshuo and Xiao, Hongru and Zhang, Erhang and Hu, Lijie and Wang, Lei and Liu, Mengyuan and Chen, Chen , journal=
-
[91]
Pinyoanuntapong, Ekkasit and Wang, Pu and Lee, Minwoo and Chen, Chen , booktitle=
-
[92]
Dai, Wenxun and Chen, Ling-Hao and Wang, Jingbo and Liu, Jinpeng and Dai, Bo and Tang, Yansong , booktitle=
-
[93]
Lu, Shunlin and Chen, Ling-Hao and Zeng, Ailing and Lin, Jing and Zhang, Ruimao and Zhang, Lei and Shum, Heung-Yeung , booktitle=
-
[94]
CVPR , pages=
Seamless Human Motion Composition with Blended Positional Encodings , author=. CVPR , pages=
-
[95]
Zhang, Mingyuan and Guo, Xinying and Pan, Liang and Cai, Zhongang and Hong, Fangzhou and Li, Huirong and Yang, Lei and Liu, Ziwei , booktitle=
-
[96]
He, Tairan and Xiao, Wenli and Lin, Toru and Luo, Zhengyi and Xu, Zhenjia and Jiang, Zhenyu and Jan Kautz and Liu, Changliu and Shi, Guanya and Wang, Xiaolong and Fan, Linxi and Zhu, Yuke , journal=
-
[97]
Han, Jinrui and Xie, Weiji and Zheng, Jiakun and Shi, Jiyuan and Zhang, Weinan and Xiao, Ting and Bai, Chenjia , journal=
-
[98]
Yin, Kangning and Zeng, Weishuai and Fan, Ke and Dai, Minyue and Wang, Zirui and Zhang, Qiang and Tian, Zheng and Wang, Jingbo and Pang, Jiangmiao and Zhang, Weinan , journal=
-
[99]
Li, Yitang and Luo, Zhengyi and Zhang, Tonghe and Dai, Cunxi and Kanervisto, Anssi and Tirinzoni, Andrea and Weng, Haoyang and Kitani, Kris and Guzek, Mateusz and Touati, Ahmed and Lazaric, Alessandro and Pirotta, Matteo and Shi, Guanya , booktitle=
-
[100]
Qi, Zekun and Chen, Xuchuan and Liu, Dairu and Lin, Chenghuai and Lian, Yunrui and Liang, Sikai and Zhang, Zhikai and Guan, Yu and Wang, Jilong and Zhang, Wenyao and Yu, Xinqiang and Wang, He and Yi, Li , journal=
-
[101]
arXiv preprint arXiv:2511.07820 , year=
Luo, Zhengyi and Yuan, Ye and Wang, Tingwu and Li, Chenran and Casta. arXiv preprint arXiv:2511.07820 , year=
-
[102]
arXiv preprint arXiv:2601.23080 , year=
Robust and Generalized Humanoid Motion Tracking , author=. arXiv preprint arXiv:2601.23080 , year=
-
[103]
arXiv preprint arXiv:2602.11929 , year=
General Humanoid Whole-Body Control via Pretraining and Fast Adaptation , author=. arXiv preprint arXiv:2602.11929 , year=
-
[104]
Li, Yuhan and Zhi, Peiyuan and Wang, Yunshen and Liu, Tengyu and Yan, Sixu and Liu, Wenyu and Wang, Xinggang and Jia, Baoxiong and Huang, Siyuan , journal=
-
[105]
Tao, Zelin and Su, Zeran and Liu, Peiran and Sun, Jingkai and Que, Wenqiang and Ma, Jiahao and Yu, Jialin and Cao, Jiahang and Sun, Pihai and Liang, Hao and Han, Gang and Zhao, Wen and Xu, Zhiyuan and Tang, Jian and Zhang, Qiang and Guo, Yijie , journal=
-
[106]
arXiv preprint arXiv:2509.13833 , year=
Track Any Motions under Any Disturbances , author=. arXiv preprint arXiv:2509.13833 , year=
-
[107]
Xie, Weiji and Han, Jinrui and Zheng, Jiakun and Li, Huanyu and Liu, Xinzhe and Shi, Jiyuan and Zhang, Weinan and Bai, Chenjia and Li, Xuelong , journal=
-
[108]
Ji, Mazeyu and Peng, Xuanbin and Liu, Fangchen and Li, Jialong and Yang, Ge and Cheng, Xuxin and Wang, Xiaolong , journal=
-
[109]
2024 , doi=
Tessler, Chen and Guo, Yunrong and Nabati, Ofir and Chechik, Gal and Peng, Xue Bin , journal=. 2024 , doi=
2024
-
[110]
and Huang, Xiaoyu and Gao, Yuman and Tevet, Guy and Sreenath, Koushil and Liu, C
Liao, Qiayuan and Truong, Takara E. and Huang, Xiaoyu and Gao, Yuman and Tevet, Guy and Sreenath, Koushil and Liu, C. Karen , journal=
-
[111]
Sun, Zhenguo and Peng, Yibo and Meng, Yuan and Li, Xukun and Huang, Bo-Sheng and Bing, Zhenshan and Wang, Xinlong and Knoll, Alois , journal=
-
[112]
arXiv preprint arXiv:2403.04436 , year=
Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation , author=. arXiv preprint arXiv:2403.04436 , year=
-
[113]
Xu, Sirui and Ling, Hung Yu and Wang, Yu-Xiong and Gui, Liang-Yan , journal=
-
[114]
Weng, Haoyang and Li, Yitang and Sobanbabu, Nikhil and Wang, Zihan and Luo, Zhengyi and He, Tairan and Ramanan, Deva and Shi, Guanya , journal=
-
[115]
Chen, Maiyue and Wang, Kaihui and Zhang, Bo and Ma, Xihan and Yang, Zhiyuan and Ren, Yi and Huang, Qijun and Zhu, Zihao and Wang, Yucheng and Su, Zhizhong , journal=
-
[116]
Xie, Weiji and Zheng, Jiakun and Han, Jinrui and Shi, Jiyuan and Zhang, Weinan and Bai, Chenjia and Li, Xuelong , journal=
-
[117]
Yuan, Xichen and Li, Zhe and Lyu, Bofan and Zuo, Kuangji and Lu, Yanshuo and Li, Gen and Yang, Jianfei , journal=
-
[118]
arXiv preprint arXiv:2603.13228 , year=
Zhang, Yangsong and Muraleedharan, Anujith and Akizhanov, Rikhat and Butt, Abdul Ahad and Varol, G. arXiv preprint arXiv:2603.13228 , year=
-
[119]
CVPR , pages=
Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control , author=. CVPR , pages=
-
[120]
Li, Peng and Zhuang, Zihan and Gao, Yangfan and Dong, Yi and Li, Sixian and Jiang, Changhao and Dou, Shihan and Xi, Zhiheng and Zhou, Enyu and Huang, Jixuan and Li, Hui and Gong, Jingjing and Ma, Xingjun and Gui, Tao and Wu, Zuxuan and Zhang, Qi and Huang, Xuanjing and Jiang, ...
-
[121]
arXiv preprint arXiv:2511.22963 , year=
Commanding Humanoid by Free-Form Language: A Large Language Action Model with Unified Motion Vocabulary , author=. arXiv preprint arXiv:2511.22963 , year=
-
[122]
Harithas, Sudarshan and Kwak, Sangkyung and Katara, Pushkal and Deolasee, Srujan and Kalaria, Dvij and Sridhar, Srinath and Vemprala, Sai and Kapoor, Ashish and Huang, Jonathan Chung-Kuan , journal=
-
[123]
2025 , url=
Zhuang, Zifeng and Wang, Ting and Zou, Binghong and Luo, Xinxin and Ma, Jianfei and Zhou, Huaicheng and Liu, Jinxin and Wang, Donglin , howpublished=. 2025 , url=
2025
-
[124]
arXiv preprint arXiv:2603.09956 , year=
Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization , author=. arXiv preprint arXiv:2603.09956 , year=
-
[125]
Huang, Tianchen and Yuan, Feiyang and Gu, Junchi and Fang, Shurui and Zhang, Xiaohu and Wang, Yu and Gao, Wei and Zhang, Shiwu , journal=
-
[126]
doi:10.1145/3721238.3730616 , year=
Xu, Michael and Shi, Yi and Yin, KangKang and Peng, Xue Bin , booktitle=. doi:10.1145/3721238.3730616 , year=
-
[127]
Gao, Li and Yang, Fuzhi and Chen, Jianhui and Liu, Liu and Zheng, Yao and Cai, Yang and Li, Ziqiao , journal=
-
[128]
arXiv preprint arXiv:2606.26855 , year=
Debbad, Pranav and Thiagarajan, Kanish and Dh. arXiv preprint arXiv:2606.26855 , year=
-
[129]
arXiv preprint arXiv:2604.17335 , year=
Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking , author=. arXiv preprint arXiv:2604.17335 , year=
-
[130]
Cho, Hanbyel and Kim, Sang-Hun and Kang, Jeonguk and Koo, Donghan , journal=
-
[131]
Li, Bin and Zhang, Ruichi and Liang, Han and Zhang, Jingyan and Zhang, Juze and Chen, Xin and Wang, Jingya , journal=
-
[132]
Zhang, Jingyan and Liang, Han and Zhang, Ruichi and Li, Bin and Zhang, Juze and Chen, Xin and Wang, Jingya and Xu, Lan and Yu, Jingyi , journal=
-
[133]
Cao, Jianuo and Chen, Yuxin and Tomizuka, Masayoshi , journal=
-
[134]
Bao, Jiacheng and Yang, Haoran and Xin, Yucheng and Liu, Junhong and Xu, Yuecheng and Liang, Han and Han, Pengfei and Ma, Xiaoguang and Wang, Dong and Zhao, Bin , journal=
-
[135]
ICLR , year=
From Language to Locomotion: Retargeting-Free Humanoid Control via Motion Latent Guidance , author=. ICLR , year=
-
[136]
Peng, Xue Bin and Ma, Ze and Abbeel, Pieter and Levine, Sergey and Kanazawa, Angjoo , booktitle=
-
[137]
ICLR , year=
Universal Humanoid Motion Representations for Physics-Based Control , author=. ICLR , year=
-
[138]
arXiv preprint arXiv:2510.02252 , year=
Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking , author=. arXiv preprint arXiv:2510.02252 , year=
-
[139]
arXiv preprint arXiv:2509.15443 , year=
A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting , author=. arXiv preprint arXiv:2509.15443 , year=
-
[140]
arXiv preprint arXiv:2603.22201 , year=
Make Tracking Easy: Neural Motion Retargeting for Humanoid Whole-Body Control , author=. arXiv preprint arXiv:2603.22201 , year=
-
[141]
arXiv preprint arXiv:1707.06347 , year=
Proximal Policy Optimization Algorithms , author=. arXiv preprint arXiv:1707.06347 , year=
-
[142]
Robotics: Science and Systems , year=
RMA: Rapid Motor Adaptation for Legged Robots , author=. Robotics: Science and Systems , year=
-
[143]
arXiv preprint arXiv:2408.14472 , year=
Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning , author=. arXiv preprint arXiv:2408.14472 , year=
-
[144]
arXiv preprint arXiv:2205.02824 , year=
Rapid Locomotion via Reinforcement Learning , author=. arXiv preprint arXiv:2205.02824 , year=
-
[145]
and Pons-Moll, Gerard and Black, Michael J
Mahmood, Naureen and Ghorbani, Nima and Troje, Nikolaus F. and Pons-Moll, Gerard and Black, Michael J. , booktitle=
-
[146]
Al-Hafez, Firas and Zhao, Guoping and Peters, Jan and Tateo, Davide , booktitle=
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.