{"id":"1cb0671d-28fd-4a14-bc14-e3e3c29cf9c7","arxiv_id":"2607.04837","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Capability-aligned dynamic and balance experts recover residual humanoid whole-body tracking failures better than data reallocation alone, then distill into one stronger deployable controller.","lead":"Athena-WBC shows that leftover hard humanoid motions often fail because the training recipe is too conservative or unstable, not only because the data is rare. It trains a few capability-aligned experts, distills them into one deployable controller, and improves long-tail and held-out tracking on a full-size robot.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Residual-set construction does not cleanly isolate recipe capability from reference/feasibility confounds.","rationale":"The reader already identified the load-bearing soft spot: residual failures after data-only interventions are assumed to be recipe-induced capability mismatches rather than reference quality, near-limit infeasibility, embodiment limits, or under-optimized budgets. That assumption is necessary for the paper’s strongest claim (capability bottleneck + capability-aligned experts as the right fix). The manuscript is careful in places—separating rphys from reffort/rtemp, using no-smoothness as a diagnostic, and acknowledging residual artifacts—but the operational residual definition and main recovery tables do not enforce a clean recoverable-feasible subset. This is not an internal contradiction; it is an attribution gap that keeps the contribution systems-level and conditional. Real-robot evidence and public artifacts would help, but the single most decisive check is residual-set purification: if gains persist on verified recoverable clips, the claim strengthens; if not, the paper still shows useful expert recipes for hard motions, but not that long-tail residual coverage is mainly a capability-recipe problem. Verdict remains CONDITIONAL; no upgrade or reject is warranted from this stress pass alone.","tokens_in":21914,"tokens_out":688,"duration_ms":6175,"concrete_test":"On the fixed Rgen mined by Eq. 8, label each clip as (i) recoverable under privileged no-smoothness/dynamic/balance teachers with cSR≥0.8, (ii) reference/retarget artifact, or (iii) near physical saturation (e.g., torque/joint-limit saturation or contact inconsistency under privileged rollouts). Recompute Table 4 SR/MPJPE only on class (i). If most of the reported recovery disappears or class (i) is a small minority of Rgen, the capability-bottleneck premise weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that residual training-set failures after data-only interventions are primarily capability bottlenecks (recipe mismatch), not insufficient exposure, and that capability-aligned experts recover them. That claim rests on R(π0)/Rgen (Eqs. 6–8, Sec. 3.2–3.3, 4.1): clips with cSR < 0.8 after general-teacher training (and, in the narrative, after targeted retraining) are treated as feasible motions whose acquisition recipe is wrong. Sec. 1 and Fig. 2 already note residual concentration in high-dynamic/balance regimes and that some failures are data artifacts or physical saturation; Limitations further admits imperfect retargeting, near-limit motions, and incomplete teacher acquisition. Yet the residual set is defined by a single success threshold on the general teacher, shared by both experts without a published per-clip feasibility/reference-quality filter, and Table 4’s long-tail recovery is reported on automatically constructed Ddynamic/Dbalance manifests (App. A.4) rather than on a verified recoverable residual subset. If a large fraction of Rgen is reference-infeasible or embodiment-saturated, the no-smoothness and expert gains show that less conservative recipes track hard clips better, but do not establish that the default recipe was the primary bottleneck for feasible residuals.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"Athena-WBC argues that residual training-set failures in strong humanoid whole-body motion trackers are not only data-allocation problems but capability bottlenecks induced by the default acquisition recipe. On a full-size 80 kg humanoid, the authors reimplement a SONIC-style baseline, mine residual clips from a general privileged teacher, and train two capability-aligned experts on the same residual set: a dynamic expert that keeps tracking and physical-constraint rewards while removing effort/temporal penalties and regularizing with Grad-CAPS, and a balance expert trained with a gravity curriculum. Teachers are motion-routed by rollout success, distilled with DAgger into one deployable student, and RL-finetuned. Relative to SONIC-Base, the final policy improves held-out AMASS/Omni SR, TIS, MPJPE, and MPJPE-W, with ablations on no-smoothness, CAPS vs Grad-CAPS, single- vs multi-teacher distillation, and teacher observations, plus new diagnostics (STC/TIS, MPJPE-W).","tokens_in":22353,"tokens_out":800,"duration_ms":6326,"significance":"If the residual-failure diagnosis holds, the paper usefully reframes long-tail humanoid WBC: changing reward/curriculum capability can matter more than only resampling or partitioning motions, and a compact two-expert bank can recover complementary high-dynamic and balance regimes before distillation. The empirical package is a real strength for a systems paper—multi-seed tables, smoothness-placement ablations with spectra/jitter, qualitative case studies, and threshold-/salience-aware metrics that are well motivated in the high-coverage regime. The contribution is incremental relative to SONIC/OmniH2O/expert-distillation recipes, but the capability-alignment framing and evaluation tools are of clear interest to the humanoid control community.","major_comments":[{"comment":"The central claim that residual failures are primarily recipe-induced capability bottlenecks (Secs. 1, 3.2–3.3; Eqs. 6–8) is only partially isolated from reference quality and physical feasibility. Rgen is defined by cSR < 0.8 on the general teacher; Fig. 2 and the Limitations section already admit data artifacts, near-limit motions, and incomplete teacher acquisition, yet Table 4 reports recovery on automatically constructed Ddynamic/Dbalance manifests (App. A.4) rather than on a verified recoverable residual subset with per-clip feasibility/reference filters. Please quantify what fraction of Rgen is recovered by no-smoothness/experts versus remains unsolved for reference/embodiment reasons, and report long-tail metrics on that filtered residual set.","section":null},{"comment":"The abstract and introduction claim residual clips remain unsolved 'even under targeted training,' but the main experimental tables do not include a controlled data-only intervention baseline on the same residual set (e.g., oversampling or training only on Rgen under the SONIC-Base recipe with matched budget). Without that comparison, Table 4’s expert gains show that less conservative recipes track hard clips better, but do not fully establish that exposure alone is insufficient. Add this ablation, or narrow the claim to the evidence actually reported.","section":null},{"comment":"All quantitative results are simulation-only on a proprietary, unreleased platform (Sec. 5.1, Limitations). The paper correctly scopes claims to recipe-level comparison on one embodiment, but for a journal contribution in humanoid WBC the deployability claim (action-rate recovery, RL fine-tuning as deployment stage) needs at least limited real-robot quantitative tracking/smoothness results, or a clearer demotion of hardware-readiness claims until such evidence exists.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: under a strong SONIC-style recipe, some training clips stay unsolved even with targeted exposure, and the authors treat that as a capability bottleneck rather than pure data imbalance. Athena-WBC then trains two experts on the residual set—dynamic experts that drop effort/temporal reward penalties and put smoothness in Grad-CAPS, balance experts with a gravity curriculum—routes them by rollout success, DAgger-distills into one student, and RL-finetunes.\n\nWhat is actually new is the framing plus the compact two-expert bank, not the individual pieces. Privileged teachers, DAgger, adaptive sampling, multi-expert distillation, and post-distillation RL are all familiar (SONIC, OmniH2O, GMT/EGM, Parkour in the Wild, CAPS/Grad-CAPS). The contribution is showing that changing the acquisition recipe on residual high-dynamic and balance-critical clips recovers more of the training long tail than data reallocation alone, and packaging that into a small teacher bank that still improves held-out AMASS and hard Omni tracking. The evaluation package is better than average for this subfield: multi-seed tables, no-smoothness vs CAPS/Grad-CAPS ablations, single- vs multi-teacher, action spectra/jerk, and the STC/TIS and MPJPE-W diagnostics. Those metrics are genuinely useful when SR is already high.\n\nThe soft spot the stress-test flags is real but not fatal. Residual sets are defined by cSR < 0.8 on the general teacher; the paper itself notes data artifacts and physical saturation, and Limitations admits imperfect retargeting and incomplete teacher acquisition. So some of Rgen may not be cleanly \"feasible but recipe-mismatched.\" The no-smoothness and expert gains still show that less conservative recipes track hard clips better; they just do not fully prove the bottleneck story for every residual clip. Other limits are ordinary for this stage: sim-only on a proprietary 80 kg platform, incomplete coverage, heavy pipeline with many free knobs, and no systematic real-robot numbers yet.\n\nMath and citations look fine for a systems paper; free parameters are listed and ablated where it matters. This is for people training large humanoid trackers who care about residual training-set failures and deployable smoothness. I would send it to peer review. Engage if you work on humanoid WBC scaling or evaluation; the metrics alone are worth a look.","headline":"Useful systems paper: residual long-tail failures as recipe/capability mismatch, with a compact dynamic/balance expert pipeline and better high-coverage metrics—solid sim evidence, incomplete residual isolation.","tokens_in":23032,"tokens_out":607,"would_cite":true,"duration_ms":5322,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Hard humanoid motions stay unsolved under the usual recipe because the recipe itself limits what the controller can learn.","keywords":["humanoid whole-body control","motion tracking","long-tail learning","capability bottleneck","teacher-student distillation","DAgger","gravity curriculum","policy regularization"],"falsifier":"Train the same residual clips to saturation under the default recipe with matched budget and cleaner references; if those clips then succeed at high rate without capability changes, the bottleneck claim fails. Conversely, if dynamic and balance recipe changes still do not recover a large share of the residual set, the expert design is insufficient.","tokens_in":22755,"feed_emoji":"🤖","tokens_out":682,"duration_ms":5253,"temperature":0.7,"pith_summary":"Strong humanoid whole-body trackers still leave a residual set of feasible training motions unsolved, especially high-dynamic transitions and balance-critical poses. The paper argues this is not only a data-allocation problem: even when those clips are oversampled or trained in isolation, the default reward recipe—effort and smoothness penalties under full gravity—can suppress the aggressive yet feasible actions or early survivability the motions need. Athena-WBC trains a small bank of capability-aligned privileged experts: dynamic experts drop conservative effort and temporal-control rewards while keeping physical limits and imposing smoothness via an auxiliary policy loss; balance experts start under reduced gravity so early rollouts survive long enough to learn. Those teachers are then routed per motion, distilled into one deployable student, and fine-tuned with RL. On a full-size humanoid, the pipeline recovers more of the training long tail and improves held-out tracking versus a strong SONIC-recipe baseline, with only a few experts.","feed_headline":"Usual training recipe, not just rare data, blocks hard humanoid moves","feed_subtitle":"Capability-aligned dynamic and balance experts recover residual motions a strong baseline still misses","key_machinery":"Capability-aligned policy experts: dynamic experts train with tracking plus physical-constraint rewards only, plus Grad-CAPS auxiliary smoothness on the policy mean; balance experts train under a gravity-scale curriculum; both residual-set teachers are then motion-routed for DAgger distillation into one student that is RL-finetuned under deployable observations.","core_discovery":"In strong humanoid whole-body control baselines, residual feasible training clips remain unsolved even under targeted training because of a capability bottleneck: a mismatch between motion demands and the effective control regime induced by the default acquisition recipe. Capability-aligned experts—dynamic experts that remove effort and temporal-control penalties while retaining physical constraints and balance experts that use a gravity curriculum—plus motion-routed distillation and RL fine-tuning recover more of that long tail and improve held-out tracking than reallocating data alone under the same recipe.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Default recipe not rare data blocks residual humanoid motions","Capability bottleneck not data scarcity leaves hard moves unsolved","Aligned dynamic balance experts recover long-tail humanoid control","Training recipe mismatch not exposure fails residual humanoid clips","Capability-aligned experts solve recipe-limited humanoid WBC failures"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That residual failures after data-only interventions mainly reflect a recipe-induced capability mismatch rather than imperfect retargeted references, near-limit physical infeasibility, embodiment limits, or under-optimized training budgets.","fun_headline_variants_meta":{"raw":{"variants":["Default recipe not rare data blocks residual humanoid motions","Capability bottleneck not data scarcity leaves hard moves unsolved","Aligned dynamic balance experts recover long-tail humanoid control","Training recipe mismatch not exposure fails residual humanoid clips","Capability-aligned experts solve recipe-limited humanoid WBC failures"]},"model":"grok-4.5","effort":"low","cost_usd":0.00552,"raw_usage":{"total_tokens":1492,"prompt_tokens":810,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":55200000,"prompt_tokens_details":{"text_tokens":810,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":620,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":810,"tokens_out":62,"duration_ms":4527,"temperature":1.0,"reasoning_tokens":620,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T12:45:05.004836+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train the same residual clips to saturation under the default recipe with matched budget and cleaner references; if those clips then succeed at high rate without capability changes, the bottleneck claim fails. Conversely, if dynamic and balance recipe changes still do not recover a large share of the residual set, the expert design is insufficient.","supporting_citations":[],"review_version":1}