Pith. sign in

REVIEW 4 major objections 6 minor 38 references

A fixed head camera is enough for whole-body humanoid dodgeball when safety training matches what the robot can see.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 01:51 UTC pith:I2LLAR2V

load-bearing objection Clean systems result: barrier strength has to match what the policy can see, with a real G1 deploy that mostly backs the fixed-camera Link-CBF choice. the 4 major comments →

arxiv 2607.28623 v1 pith:I2LLAR2V submitted 2026-07-30 cs.RO cs.AI

PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball

classification cs.RO cs.AI
keywords humanoid robotscontrol barrier functionsreinforcement learningperception-aware controlwhole-body evasiondodgeballsim-to-real
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Humanoid robots that must dodge fast-moving objects face a tight coupling problem: safety methods that need exact object state do not match what a head-mounted camera actually provides. This paper argues that the right amount of barrier-based safety structure depends on perceptual observability. With accurate ball state, a strong joint-space control-barrier filter is best. With only fixed-camera, ball-masked depth, a lighter per-link barrier used only during training works better and can be internalized by the policy. On a controlled any-link contact benchmark, that fixed-camera policy comes within a few points of a privileged-state oracle. The same policy is deployed zero-shot on hardware from onboard depth and proprioception alone, succeeding on 95% of throws and dodging different balls via semantic segmentation. A sympathetic reader cares because the result reframes whole-body reactive safety as a co-design of sensing and barrier strength rather than as an independent filter stacked on a vision stack.

Core claim

Usable barrier structure for whole-body humanoid dodgeball is limited by what the policy can observe. Joint-space CBF is strongest when accurate ball states are available at runtime, but degrades under fixed-camera observations when used only as training guidance; it recovers with better tracking or a privileged filter. A lightweight Link-CBF reward that extends clearance to every body link is the best deployable choice under fixed onboard depth, matching a privileged oracle within a few points in simulation and transferring zero-shot to 95% hardware success without runtime ball state.

What carries the argument

PAC-MAN: perception-aware CBF-RL that pairs deployment-realistic segmentation-masked depth with two barrier levels—Link-CBF (per-link clearance reward during training) and Joint-CBF (joint-space projection usable as training guidance or privileged runtime filter)—plus an adversarial motion prior that shapes crouches, leans, and sidesteps without defining the safe set.

Load-bearing premise

The claim rests on a controlled frontal, on-target throw setup and perception noise model being representative enough that sim rankings and a short hand-thrown hardware test generalize to real dodging conditions.

What would settle it

Rerun the same any-link benchmark and hardware protocol with substantially faster, off-axis, or multi-ball throws, or with perception failures outside the training dropout model; if fixed-camera Link-CBF then falls far behind oracle or Joint-CBF-with-filter and hardware success drops well below the reported 95%, the observability-matched design claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Deployable whole-body evasion can rely on fixed-camera masked depth and training-time Link-CBF without a runtime ball-state estimator or CBF filter.
  • Stronger joint-space barriers should be reserved for settings with accurate online ball state or active tracking that keeps the threat observable.
  • Any-link contact, not pelvis clearance alone, is the right success metric for humanoid dodge tasks.
  • Semantic segmentation lets one trained policy dodge different ball types without retuning the controller.
  • Closing the loop with a tracker-aimed gimbal or online ball estimator is the natural next step to unlock Joint-CBF on hardware.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same observability-matched barrier idea likely applies to other brief hazards—flying debris, human limbs in close work, or ball sports—where privileged filters look strong in sim but starve under egocentric vision.
  • If active gaze can be driven from the existing segmentation track rather than oracle aim, hardware may recover much of the Joint-CBF gain without full state estimation.
  • Sparse temporal depth stacks that encode looming may be doing as much work as the barrier terms; ablating stack timing against barrier level would separate those contributions.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. PAC-MAN couples training-time control-barrier guidance with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The policy observes only proprioception and segmentation-masked head-camera depth; Link-CBF (per-link clearance reward) and Joint-CBF (joint-space projection, optionally kept as a privileged runtime filter) supply whole-body safety structure during training, with an AMP motion prior regularizing evasive style. On a statue-calibrated any-link contact benchmark (reset and walk-back deployment regimes), fixed-camera Link-CBF reaches ~90%/89% success, within a few points of strong privileged oracles, while Joint-CBF without a filter degrades under fixed-camera observations and recovers with an oracle gimbal or runtime filter. The authors deploy fixed-camera Link-CBF zero-shot on a Unitree G1 and report 19/20 (95%) successful hand throws from onboard depth and proprioception alone, with semantic segmentation enabling different ball types.

Significance. The paper cleanly frames reactive whole-body safety as a perception–barrier co-design problem rather than as independent modules, and it backs that framing with a structured perception × safety-mode ablation under a stricter any-link contact criterion. Zero-shot hardware transfer of a vision-only Link-CBF policy on the G1, plus a reusable masked-depth interface that retargets across ball appearances, is a concrete systems contribution. If the observability–barrier ranking holds beyond the controlled frontal regime, the work gives a practical design rule for when to internalize lightweight barriers versus when to keep privileged CBF filters—useful for humanoid reaction under partial exteroception. Planned release of the benchmark and training pipeline would further strengthen impact.

major comments (4)
  1. [§IV-A, Table II] Table II, State-oracle row: Link-CBF drops reset success from 98% (no barrier) to 88%, while the same Link-CBF is best under fixed camera. This anomaly is load-bearing for the central claim that “usable barrier structure depends on perceptual observability,” yet it is not analyzed. Please explain (e.g., reward interference with r_core, over-conservative clearance, optimization conflict with AMP) and, if needed, retune or ablate weights so the co-design narrative is not undercut by a large privileged-state regression.
  2. [§IV-A, Table II] Table II reports a single matched 20k-iteration checkpoint with no multi-seed means/std or confidence intervals. Several deployable contrasts are small (fixed-camera no-barrier 86% → Link-CBF 90%/89%; Joint-CBF deployment 76%). Without seed-level uncertainty, the ranking that justifies deploying Link-CBF over no-barrier or bare Joint-CBF is not yet statistically secured. Please add multi-seed evaluation (or equivalent bootstrap over seeded throws) for the main cells.
  3. [§II-B, §IV-A/B, Abstract] §II-B and §IV-B: the adequacy claim (fixed camera “alone is adequate”; 95% real-world success) rests on a narrow threat model—frontal ±25° cone, ~0.6 s flight, statue-calibrated on-target throws, recovery between throws—and a 20-throw hand-thrown hardware test without reported ball speed, aim dispersion, or comparison to the simulator launcher. Please characterize hardware throws (speed/range/aim) against the sim distribution, and either broaden the sim threat set (off-axis, shorter flight, multi-ball, or no full recovery) or explicitly scope the claim to this controlled loop so the abstract does not over-generalize.
  4. [§IV-B, Abstract] §IV-B perception stack (EfficientTAM, distance-aware mask-size checks, optional looming gate, all-far on track loss) is conservative by design. If hard or partial observations are systematically zeroed, the 19/20 figure may partly reflect filtering rather than policy robustness under imperfect perception—the property claimed in the abstract. Please report track-accept/reject rates, how often the policy received all-far frames during the 20 throws, and success conditioned on accepted vs. marginal tracks; ablate the looming/mask gates if feasible.
minor comments (6)
  1. [§IV-A/B, Table II–III] Table II caption refers to “Fig. II (top)” and the hardware section to “Fig. III (top)” for plots that appear to be embedded above the tables; renumber figures consistently and give each plot a proper figure environment.
  2. [§III-B] Eq. (4)–(5) vs. Joint-CBF text: Link-CBF uses linear clearance h_i = ∥p_b−p_i∥−(ρ_b+ρ_i), while Joint-CBF uses squared distance h_i = ∥p_b−p_i∥²−D_i². A brief note on why the two barrier forms differ would help readers compare r_clear and the projection constraint.
  3. [Abstract, §I] Abstract and §I say the policy “comes within a few points of a privileged state oracle,” but oracle +filter reaches 98–99% while deployed Link-CBF is ~90%; “few points” fits oracle Link-CBF/no-barrier better than the oracle ceiling. Soften or qualify the wording.
  4. [§IV-A] Comparison to SMP [15] (§IV-A) is only under oracle observations and an omnidirectional codebase setting; one sentence on whether fixed-camera PAC-MAN was evaluated in that 360° protocol would clarify the baseline claim.
  5. [Abstract, §II-B, §III-C] Typos/spacing: “succeeds on95%of throws,” “within±25,” “roughly0.6s,” “1-DoF,” “20k-iteration” — add spaces between numbers and units/words throughout.
  6. [Table I, §III-A] Table I lists fall termination weight −200 and many regularizers; briefly state whether these weights were tuned jointly with λ_cbf or held fixed across the no-barrier / Link / Joint sweep so ablations remain comparable.

Circularity Check

0 steps flagged

No significant circularity: empirical contact/fall benchmarks are independent of the CBF training rewards and self-citations supply methods only.

full rationale

PAC-MAN is an empirical robotics paper whose central claims are measured success rates under an any-link contact criterion, not quantities derived from first principles. Training injects Link-CBF / Joint-CBF structure via rewards (Eqs. 2–7, Table I) and an AMP style term, but evaluation is external: simulator terminations on ball–link contact or fall, a frozen-statue on-target floor (~4% survival), seeded single-throw and deployment-loop regimes (Table II), and 19/20 physical throws on the Unitree G1 with onboard masked depth (Table III). Success is therefore not the training objective by construction. Self-citations (notably CBF-RL [11] and the Ames CBF line [10]) supply the safety-filtering training technique and barrier formalism; they do not underwrite the reported dodge percentages, the perception–barrier ranking, or the hardware transfer result. There is no fitted universal constant re-presented as a prediction, no uniqueness theorem imported to forbid alternatives, and no renaming of a known empirical law as a derived organization. The paper is self-contained against its stated external benchmarks. Concerns about threat-regime narrowness or small deployable deltas are generalization/correctness issues, not circularity.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

Load-bearing content is methodological and empirical rather than axiomatic physics. The claim rests on standard CBF invariance ideas, PPO/AMP learning assumptions, a hand-specified throw and perception model, and many reward weights chosen for training stability. No new physical entity is postulated; Link-CBF and Joint-CBF are engineered barrier constructions. Free parameters are the usual RL reward and CBF hyperparameters plus hardware perception gates.

free parameters (6)
  • Link-CBF reward weight and barrier params (λ, α=1, clip c=2) = weight 0.27; α=1; c=2
    Chosen training weights/hyperparameters that shape how strongly per-link barrier violations affect the policy; not derived from first principles.
  • Joint-CBF correction/buffer weights and keep-out buffer = correction 0.1; buffer 1.0; d_safe=0.1
    Hand-set costs on filter intervention and skimming (λ_corr, λ_buf, d_safe, δ schedule) that define the stronger training signal.
  • Core task and style reward weights (r_core, AMP, station/posture terms) = r_core 1.0; AMP 0.5; plus listed posture terms
    Full stack in Table I is designer-chosen to balance evasion, stillness, style, and regularizers; different weights could change the learned reflex.
  • Throw distribution and reaction window = ~0.6 s flight; ±25°; 80/20 throw/stand mix
    Frontal cone ±25°, 2–3 m, ~0.6 s flight, mixed descending/low-arc threats calibrated so a statue is almost always hit; defines the benchmark difficulty.
  • Masked depth resolution and temporal stack = 16×9; lags 0,3,8,18
    16×9 pooled ball-only depth with lags [0,3,8,18] at 50 Hz is an observation design choice that limits what the policy can infer.
  • Hardware perception acceptance gates
    Low-percentile pooling, mask-size checks, optional looming gate, and all-far fallback are tuned filters that define when the policy sees a ball on the real robot.
axioms (5)
  • domain assumption Control barrier functions encode forward-invariant safe sets via conditions of the form ḣ + αh ≥ 0 (or the joint-space projection used in Joint-CBF).
    Invoked throughout §III-B from Ames et al. CBF theory; safety guidance is only as meaningful as this model of clearance dynamics.
  • domain assumption A policy trained with privileged barrier signals and randomized masked depth can internalize evasive behavior that transfers to real onboard segmentation-masked depth without runtime privileged state.
    Core sim-to-real premise of §III-A/C and §IV-B; supports zero-shot deployment claims.
  • domain assumption Any-link geometric contact plus fall thresholds are the right success metric for whole-body dodgeball safety.
    §II-C defines success strictly by contact with any link or loss of balance; rankings could shift under softer metrics.
  • domain assumption PPO with an AMP discriminator on retargeted human dodge clips yields dynamically feasible, stylistically appropriate whole-body actions when combined with task/CBF rewards.
    §III-D; standard in motion-imitation RL but still an unproved modeling bet for this task.
  • standard math Standard RL/optimization mathematics (policy gradients, closed-form half-space projection of joint velocity) holds in the training stack.
    Used for PPO updates and Eq. (7) barrier projection.
invented entities (3)
  • PAC-MAN framework (perception-aware pairing of Link-CBF and Joint-CBF with masked-depth observations) no independent evidence
    purpose: Organize training-time barrier guidance levels against deployable onboard sensing for humanoid dodgeball.
    Named system composition of existing CBF-RL, AMP, and depth masking ideas rather than a new physical object; no independent existence outside this engineering stack.
  • Link-CBF per-link clearance reward over the whole-body keep-out set C = ∩_i {h_i ≥ 0} no independent evidence
    purpose: Extend pelvis-centered evasion to every body link as lightweight training guidance without runtime enforcement.
    Engineered reward construction (§III-B); validated only via the paper’s own ablations and hardware trial.
  • Any-link contact dodgeball benchmark (reset and deployment-loop regimes with statue-calibrated throws) no independent evidence
    purpose: Provide a controlled evaluation that penalizes limb grazes and includes walk-back recovery between throws.
    Evaluation protocol introduced here; promised for release but not an externally standardized benchmark yet.

pith-pipeline@v1.2.0-daily-grok45 · 17331 in / 4231 out tokens · 81697 ms · 2026-07-31T01:51:10.093782+00:00 · methodology

0 comments
read the original abstract

We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted camera, while training-time CBF guidance represents clearance to every body link, and an adversarial motion prior regularizes the resulting evasive reflexes. We evaluate on a controlled any-link contact benchmark with seeded throws in two regimes: single throws and a deployment loop in which the robot walks back to its station and recovers between throws. On this benchmark, the policy comes within a few points of a privileged state oracle: a fixed onboard camera alone is adequate for evasion. We find that usable barrier structure depends on perceptual observability: Joint-CBF gives the best performance with accurate ball states, degrades under fixed-camera observations when used only as training guidance, and recovers with a ball-tracking gimbal or privileged runtime filter. We therefore deploy a lightweight Link-CBF policy zero-shot on the Unitree G1 in the real world, where it tolerates imperfect perception, succeeds on 95% of throws, and uses semantic segmentation to dodge different balls.

Figures

Figures reproduced from arXiv: 2607.28623 by Aaron D. Ames, Junheng Li, Lizhi Yang.

Figure 1
Figure 1. Figure 1: Perception-aware dodgeball couples safety and sensing: a humanoid [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Pipeline overview. The policy maps temporally stacked ball-only depth and proprioception to [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Whole-body barrier geometry across representative motion-prior [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Emergent whole-body evasion modes in simulation (left) and hardware (right). Crouching, leaning, and sidestepping coordinate the legs, torso, and [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

38 extracted references · 8 linked inside Pith

  1. [1]

    Humanoid locomotion and manipulation: Current progress and challenges in control, planning, and learning,

    Z. Gu, J. Li, W. Shen, W. Yu, Z. Xie, S. McCrory, X. Cheng, A. Shamsah, R. Griffin, C. K. Liu,et al., “Humanoid locomotion and manipulation: Current progress and challenges in control, planning, and learning,”IEEE/ASME Transactions on Mechatronics, vol. 31, no. 2, pp. 2300–2330, 2026

  2. [2]

    Humanoid parkour learning,

    Z. Zhuang, S. Yao, and H. Zhao, “Humanoid parkour learning,” in Proceedings of The 8th Conference on Robot Learning(P. Agrawal, O. Kroemer, and W. Burgard, eds.), vol. 270 ofProceedings of Machine Learning Research, pp. 1975–1991, PMLR, 06–09 Nov 2025

  3. [3]

    Gmt: General motion tracking for humanoid whole-body control,

    Z. Chen, M. Ji, X. Cheng, X. Peng, X. B. Peng, and X. Wang, “Gmt: General motion tracking for humanoid whole-body control,”arXiv preprint arXiv:2506.14770, 2025

  4. [4]

    Sonic: Supersizing motion tracking for natural humanoid whole-body control,

    Z. Luo, Y . Yuan, T. Wang, C. Li, F. Casta ˜neda, S. Chen, Z.-A. Cao, J. Li, D. Minor, Q. Ben,et al., “Sonic: Supersizing motion tracking for natural humanoid whole-body control,”arXiv preprint arXiv:2511.07820, 2025

  5. [5]

    Beyondmimic: From motion tracking to versatile humanoid control via guided diffusion,

    Q. Liao, T. E. Truong, X. Huang, Y . Gao, G. Tevet, K. Sreenath, and C. K. Liu, “Beyondmimic: From motion tracking to versatile humanoid control via guided diffusion,”arXiv preprint arXiv:2508.08241, 2025

  6. [6]

    Visualmimic: Visual humanoid loco-manipulation via motion tracking and generation,

    S. Yin, Y . Ze, H.-X. Yu, C. K. Liu, and J. Wu, “Visualmimic: Visual humanoid loco-manipulation via motion tracking and generation,” arXiv preprint arXiv:2509.20322, 2025

  7. [7]

    CReF: Cross-modal and recurrent fusion for depth-conditioned humanoid locomotion,

    Y . Hao, R. Yu, S. Luo, G. Zhang, J. Wu, and Q. Zhu, “CReF: Cross-modal and recurrent fusion for depth-conditioned humanoid locomotion,”arXiv preprint arXiv:2603.29452, 2026

  8. [8]

    Agile but safe: Learning collision-free high-speed legged locomotion,

    T. He, C. Zhang, W. Xiao, G. He, C. Liu, and G. Shi, “Agile but safe: Learning collision-free high-speed legged locomotion,” 2024

  9. [9]

    Egocentric tactile and proximity sensors as observation priors for humanoid collision avoidance,

    C. Kohlbrenner, N. Pudasaini, W. Xie, N. Sivagnanadasan, N. Cor- rell, and A. Roncone, “Egocentric tactile and proximity sensors as observation priors for humanoid collision avoidance,”arXiv preprint arXiv:2604.25554, 2026

  10. [10]

    Control barrier function based quadratic programs for safety critical systems,

    A. D. Ames, X. Xu, J. W. Grizzle, and P. Tabuada, “Control barrier function based quadratic programs for safety critical systems,”IEEE Transactions on Automatic Control, vol. 62, no. 8, pp. 3861–3876, 2016

  11. [11]

    CBF-RL: Safety filtering reinforcement learning in training with control barrier func- tions,

    L. Yang, B. Werner, M. de Sa, and A. D. Ames, “CBF-RL: Safety filtering reinforcement learning in training with control barrier func- tions,”2026 IEEE International Conference on Robotics and Automa- tion (ICRA), 2026

  12. [12]

    Deepmimic: Example-guided deep reinforcement learning of physics-based char- acter skills,

    X. B. Peng, P. Abbeel, S. Levine, and M. Van de Panne, “Deepmimic: Example-guided deep reinforcement learning of physics-based char- acter skills,”ACM Transactions On Graphics (TOG), vol. 37, no. 4, pp. 1–14, 2018

  13. [13]

    Amp: Adversarial motion priors for stylized physics-based character con- trol,

    X. B. Peng, Z. Ma, P. Abbeel, S. Levine, and A. Kanazawa, “Amp: Adversarial motion priors for stylized physics-based character con- trol,”ACM Transactions on Graphics (ToG), vol. 40, no. 4, pp. 1–20, 2021

  14. [14]

    Ase: Large- scale reusable adversarial skill embeddings for physically simulated characters,

    X. B. Peng, Y . Guo, L. Halper, S. Levine, and S. Fidler, “Ase: Large- scale reusable adversarial skill embeddings for physically simulated characters,”ACM Transactions On Graphics (TOG), vol. 41, no. 4, pp. 1–17, 2022

  15. [15]

    Smp: Reusable score-matching motion priors for physics-based character control,

    Y . Mu, Z. Zhang, Y . Shi, D. Yang, M. Matsumoto, K. Imamura, G. Tevet, C. Guo, M. Taylor, C. Shu, P. Xi, and X. B. Peng, “Smp: Reusable score-matching motion priors for physics-based character control,”ACM Transactions on Graphics (Proceedings of SIGGRAPH 2026), 2026

  16. [16]

    Reinforcement learning for robust parameterized locomotion control of bipedal robots,

    Z. Li, X. Cheng, X. B. Peng, P. Abbeel, S. Levine, G. Berseth, and K. Sreenath, “Reinforcement learning for robust parameterized locomotion control of bipedal robots,” in2021 IEEE International Conference on Robotics and Automation (ICRA), pp. 2811–2817, IEEE, 2021

  17. [17]

    Sim-to-real learning for humanoid box loco-manipulation,

    J. Dao, H. Duan, and A. Fern, “Sim-to-real learning for humanoid box loco-manipulation,” in2024 IEEE International Conference on Robotics and Automation (ICRA), pp. 16930–16936, IEEE, 2024

  18. [18]

    Walk the PLANC: Physics-guided rl for agile humanoid locomotion on constrained footholds,

    M. Dai, W. D. Compton, J. Li, L. Yang, and A. D. Ames, “Walk the PLANC: Physics-guided rl for agile humanoid locomotion on constrained footholds,”arXiv preprint arXiv:2601.06286, 2026

  19. [19]

    Amo: Adaptive motion optimization for hyper-dexterous humanoid whole- body control,

    J. Li, X. Cheng, T. Huang, S. Yang, R.-Z. Qiu, and X. Wang, “Amo: Adaptive motion optimization for hyper-dexterous humanoid whole- body control,” inProceedings of Robotics: Science and Systems, (LosAngeles, CA, USA), June 2025

  20. [20]

    Opt2skill: Imitating dynamically-feasible whole-body trajectories for versatile humanoid loco-manipulation,

    F. Liu, Z. Gu, Y . Cai, Z. Zhou, H. Jung, J. Jang, S. Zhao, S. Ha, Y . Chen, D. Xu,et al., “Opt2skill: Imitating dynamically-feasible whole-body trajectories for versatile humanoid loco-manipulation,” IEEE Robotics and Automation Letters, 2025

  21. [21]

    Retargeting matters: General motion retargeting for humanoid motion tracking,

    J. P. Araujo, Y . Ze, P. Xu, J. Wu, and C. K. Liu, “Retargeting matters: General motion retargeting for humanoid motion tracking,” arXiv preprint arXiv:2510.02252, 2025

  22. [22]

    Omniretarget: Interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction,

    L. Yang, X. Huang, Z. Wu, A. Kanazawa, P. Abbeel, C. Sferrazza, C. K. Liu, R. Duan, and G. Shi, “Omniretarget: Interaction-preserving data generation for humanoid whole-body loco-manipulation and scene interaction,”2026 IEEE International Conference on Robotics and Automation (ICRA), 2026

  23. [23]

    Extreme parkour with legged robots,

    X. Cheng, K. Shi, A. Agarwal, and D. Pathak, “Extreme parkour with legged robots,” in2024 IEEE International Conference on Robotics and Automation (ICRA), pp. 11443–11450, IEEE, 2024

  24. [24]

    Anymal parkour: Learning agile navigation for quadrupedal robots,

    D. Hoeller, N. Rudin, D. Sako, and M. Hutter, “Anymal parkour: Learning agile navigation for quadrupedal robots,”Science Robotics, vol. 9, no. 88, p. eadi7566, 2024

  25. [25]

    Learning humanoid locomotion with perceptive internal model,

    J. Long, J. Ren, M. Shi, Z. Wang, T. Huang, P. Luo, and J. Pang, “Learning humanoid locomotion with perceptive internal model,” pp. 9997–10003, 2025

  26. [26]

    Perceptive humanoid parkour: Chaining dynamic human skills via motion matching,

    Z. Wu, X. Huang, L. Yang, Y . Zhang, X. Chen, P. Abbeel, R. Duan, A. Kanazawa, C. Sferrazza, G. Shi,et al., “Perceptive humanoid parkour: Chaining dynamic human skills via motion matching,”arXiv preprint arXiv:2602.15827, 2026

  27. [27]

    Deep whole-body parkour,

    Z. Zhuang, S. Zhu, M. Zhao, and H. Zhao, “Deep whole-body parkour,”arXiv preprint arXiv:2601.07701, 2026

  28. [28]

    BAT: Balancing agility and stability via online policy switching for long-horizon whole-body humanoid control,

    D. Baek, S.-H. Kim, and S. Ha, “BAT: Balancing agility and stability via online policy switching for long-horizon whole-body humanoid control,”arXiv preprint arXiv:2604.01064, 2026

  29. [29]

    TAGA: Terrain-aware active gaze learning for general- izable agile humanoid locomotion,

    P. Li, H. Li, M. Fan, F. Xu, S. Liao, Y . Ma, Z. Zeng, Z. Wang, Y . Jin, Y . Cao,et al., “TAGA: Terrain-aware active gaze learning for general- izable agile humanoid locomotion,”arXiv preprint arXiv:2606.05880, 2026

  30. [30]

    Hitter: A humanoid table tennis robot via hierarchical planning and learning,

    Z. Su, B. Zhang, N. Rahmanian, Y . Gao, Q. Liao, C. Regan, K. Sreenath, and S. S. Sastry, “Hitter: A humanoid table tennis robot via hierarchical planning and learning,”2026 IEEE International Conference on Robotics and Automation (ICRA), 2026

  31. [31]

    Capture point: A step toward humanoid push recovery,

    J. Pratt, J. Carff, S. Drakunov, and A. Goswami, “Capture point: A step toward humanoid push recovery,” in2006 6th IEEE-RAS international conference on humanoid robots, pp. 200–207, Ieee, 2006

  32. [32]

    Dynamic balance force control for compliant humanoid robots,

    B. J. Stephens and C. G. Atkeson, “Dynamic balance force control for compliant humanoid robots,” in2010 IEEE/RSJ international conference on intelligent robots and systems, pp. 1248–1255, IEEE, 2010

  33. [33]

    A collision-free mpc for whole-body dynamic locomotion and manipu- lation,

    J.-R. Chiu, J.-P. Sleiman, M. Mittal, F. Farshidian, and M. Hutter, “A collision-free mpc for whole-body dynamic locomotion and manipu- lation,” in2022 international conference on robotics and automation (ICRA), pp. 4686–4693, IEEE, 2022

  34. [34]

    Curobo: Parallelized collision-free robot motion generation,

    B. Sundaralingam, S. K. S. Hari, A. Fishman, C. Garrett, K. Van Wyk, V . Blukis, A. Millane, H. Oleynikova, A. Handa, F. Ramos,et al., “Curobo: Parallelized collision-free robot motion generation,” in2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 8112–8119, IEEE, 2023

  35. [35]

    BONES-SEED: Skeletal everyday embodiment dataset

    Bones Studio, “BONES-SEED: Skeletal everyday embodiment dataset.”https://bones.studio/datasets/seed, 2026

  36. [36]

    Rsl-rl: A learning library for robotics research,

    C. Schwarke, M. Mittal, N. Rudin, D. Hoeller, and M. Hutter, “Rsl-rl: A learning library for robotics research,”arXiv preprint arXiv:2509.10771, 2025

  37. [37]

    mjlab: A lightweight framework for gpu-accelerated robot learning,

    K. Zakka, Q. Liao, B. Yi, L. L. Lay, K. Sreenath, and P. Abbeel, “mjlab: A lightweight framework for gpu-accelerated robot learning,” arXiv preprint arXiv:2601.22074, 2026

  38. [38]

    AMP mjlab: G1 AMP motion control on mjlab + rsl rl

    ccrpRepo, “AMP mjlab: G1 AMP motion control on mjlab + rsl rl.” https://github.com/ccrpRepo/AMP_mjlab, 2025