{"id":"ad54f7cc-ef29-4f2f-bc65-8877bb843748","arxiv_id":"2608.01410","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Co-training a text-conditioned motion generator with a humanoid tracker, using execution feedback as reward, improves both generated-motion executability and zero-shot tracking coverage in simulation.","lead":"This paper describes a training loop that alternates between a text-to-motion generator and a humanoid tracking policy, so each improves the other on the fly. It reports small but consistent gains in simulated zero-shot tracking and in how well generated motions can be executed by a fixed robot controller.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Tracker-side gains are within measurement noise: single deterministic rollouts, no repeated seeds, and 0.7–5.0 pp differences cannot support the 'markedly broader coverage' claim.","rationale":"I read the paper as a careful, well-controlled empirical study. The matched budgets, one-way controls, separate ProtoMotions and SONIC backbones, and explicit acknowledgement of simulation-only evaluation are real strengths. The load-bearing problem is that the central comparative claim rests on very small differences. The tracker coverage improvements are 5.0, 0.7 and 0.8 percentage points across the three splits; the generator improvements above G0 are roughly 1–2 percentage points. There are no error bars, no repeated seeds, and the protocol intentionally makes rollouts deterministic, so the reported values are point estimates from single trajectories. The reader's weakest assumption focused on the lenient 0.25 m fall threshold and the absence of real-hardware validation; I partially agree, but I think the more immediate issue is that even under the same simulated evaluator the margins are near the resolution of the experiment. The generator-side claim is additionally fragile because the frozen-SONIC evaluator belongs to the same policy family that provides the training reward in the SONIC branch, and the ProtoMotions branch shows only a marginal generator gain with a slightly worse E_joint. A multi-seed rerun or per-clip bootstrap would settle whether the reported gaps are real. This does not change the reader's CONDITIONAL verdict: the concern is already within the stated conditions, but it sharpens the reason for those conditions and the exact check needed.","tokens_in":20342,"tokens_out":9016,"duration_ms":121521,"concrete_test":"Re-run the SONIC post-training arm and its matched G0-replay control with at least 5 independent training seeds under the same Section C.1 protocol, and report per-seed fall-only SR with the paired difference and a 95% confidence interval for each split. If the GenTrack minus G0-replay difference overlaps zero on AMASS-test-G1 or Wild-G1-clean, the claimed mutual-adaptation advantage over static replay is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is comparative: online co-training must beat matched static replay and one-way alignment. The reported tracker-side margins are very small — SONIC fall-only SR goes from 85.0/79.0/47.2 to 90.0/79.7/48.0, i.e. +5.0, +0.7 and +0.8 percentage points. Table 1 gives no per-clip counts, no confidence intervals, and no repeated-seed statistics. Section C.1 deliberately disables observation corruption, reset perturbations, and startup randomization, so each reported number is a single deterministic trajectory per test clip. With a binary fall-only criterion, a 5.0 pp gain could be a handful of clips on a small split, and the AMASS/Wild gains are below one percentage point. The same issue affects the generator table: GenTrack(SONIC) improves frozen-SONIC success by 1.85 pp over G0, while GenTrack(ProtoMotions) improves it by only 0.97 pp and its E_joint slightly worsens. Because the evaluator for generator executability is the same SONIC policy family used to train the SONIC branch, and no cross-controller or real-hardware check is reported, the evidence that this is a general robot-native executability gain rather than a small, evaluator-specific or noise-level effect is not yet established. The paper is honest about simulation-only scope (Section E), but the central 'markedly broader coverage' wording in the abstract is not supported by margins of this size without uncertainty quantification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes GenTrack, an online co-training framework that alternately updates a pretrained text-to-motion generator and a pretrained humanoid tracker. The generator is aligned with execution-grounded, group-relative rewards computed from a frozen lagged tracker, while the tracker is trained on a mixture of a fixed public reference pool and newly generated references that pass only structural validity checks. Anchoring via a frozen-generator KL penalty and supervised rehearsal is used to limit drift. The method is evaluated on simulated Unitree G1 with ProtoMotions and SONIC backbones, reporting generator executability via frozen-SONIC rollouts and zero-shot tracker coverage on three held-out splits, alongside matched static-replay, one-way, and objective ablations.","tokens_in":20772,"tokens_out":5836,"duration_ms":66633,"significance":"If the reported gains are robust, the contribution is meaningful: it offers a way to extend zero-shot humanoid tracking coverage without additional embodied data collection, while using closed-loop execution feedback to make a text-to-motion generator produce more robot-compatible references. The experimental design is careful in several respects: matched budgets and initializations across arms, separation of the reward judge from the current trainee, no success-gating of generated references admitted to tracker training, and multiple one-way/offline controls. The method is also presented with a clear statement that evaluation is simulation-only. However, the central claims of markedly broader coverage and general robot-native executability currently rest on small margins, single deterministic rollouts, and a single simulated executor, so the evidence is not yet conclusive.","major_comments":[{"comment":"The tracker-side success gains are small and are reported without uncertainty quantification. Section C.1 disables observation corruption, reset perturbations, and startup randomization, so each number is a single deterministic rollout per test clip. Under the binary fall-only criterion, which tolerates up to 0.25 m reference-relative pelvis-height deviation, a +0.7 or +0.8 percentage-point difference on AMASS-test or Wild-G1 could be a handful of clips, and per-clip counts are not reported. The abstract's phrase 'markedly broader zero-shot coverage' is not supported by margins of this size without per-split counts, bootstrap confidence intervals, or repeated-seed statistics. Please provide such uncertainty quantification or substantially temper the coverage claim.","section":"§4, Table 1; §C.1"},{"comment":"Generator executability is measured only with a frozen SONIC policy in IsaacLab, and the SONIC branch's generator reward is produced by the same frozen SONIC policy family. Although the ProtoMotions branch provides partial independence, all generator rows are evaluated by the same SONIC executor, so the results could reflect evaluator-specific adaptation rather than a general robot-native executability gain. The paper honestly limits the evaluation to simulation in Section E, but the abstract and conclusion generalize to 'robot-native' motion and 'zero-shot humanoid tracking' beyond that scope. Please add a second independent executor, a cross-controller evaluation, or real-hardware spot checks, or explicitly restrict the executability claims to the single simulated SONIC evaluator.","section":"§3, Eq. (3)–(5); §4, Table 2; §E"},{"comment":"The ProtoMotions branch does not consistently improve over matched controls: GenTrack scores 75.0 on LAFAN1 versus 77.5 for G0 replay and equals the 75.0 baseline, while AMASS-test and Wild-G1 improve by only +2.2 and +0.5 percentage points relative to G0 replay. The statement that online co-training 'consistently produces ... trackers with markedly broader zero-shot coverage' across both backbones is therefore an overstatement. The evidence supports at most a split-dependent benefit for SONIC and a small Wild-G1 improvement for ProtoMotions; please revise the framing or provide a principled aggregate analysis that supports the broader claim.","section":"§4, Table 1 (ProtoMotions rows)"}],"minor_comments":[{"comment":"Any2Track's reported 100.0% on LAFAN1 against 5.1% and 10.4% on AMASS-test and Wild-G1 is surprising; please clarify how the external tracker's released success flags and trajectory conversion affect this row.","section":"Table 1"},{"comment":"The metric table lists many auxiliary diagnostics (completion, unexpected fall, foot skate, penetration, kinematic diagnostics, judge calibration) that are not reported in the main text; please state explicitly which of these are available in the supplement and which are omitted.","section":"§C.1, Table 4"},{"comment":"FID and R-Precision are reported to three decimals from deterministic single evaluations; such precision overstates the reliability of the estimates. Please include uncertainty bounds or state that these are point estimates from a single evaluation.","section":"§4, Tables 2 and 5"},{"comment":"The private 1,024-prompt generator suite is described only in the supplement; please include at least its construction, prompt coverage, and exclusion rules in the main text so the reader can assess the generator claims without accessing the supplement.","section":"§4, Table 2"},{"comment":"The term 'lagged' is used in the main text before it is defined; consider clarifying at first use that it refers to the tracker from the preceding round, frozen during the generator phase.","section":"§3, 'Execution reward'"}],"recommendation":"major_revision","confidential_remarks":"The private 1,024-prompt OOD suite and several supplement details are not fully specified in the submitted text; if reviewers cannot access the supplement, those claims should be treated as unverified. The related-work discussion is extensive and appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serious, carefully controlled study of an idea that's been floating around — closing the loop between a text-to-motion generator and a humanoid tracker — and it does the comparison better than most prior work. The headline claim, though, is stronger than the numbers support. The tracker-side gains are a few percentage points at best, with no error bars, no repeated seeds, and deterministic rollouts, so 'markedly broader coverage' is not yet established.\n\nWhat's genuinely good: the paper isn't just proposing another co-training loop. It tests the loop against matched one-way and offline controls — reference-only continuation, frozen-generator replay, final-generator replay, tracker-filtered SFT, and a frozen-strong-tracker reward — and the controls are well constructed. The lagged-tracker design (the current trainee never scores its own generated references) is a sensible answer to a real failure mode, and the ablation in Table 5 backs it up: using the current judge hurts, removing the KL/rehearsal anchors degrades semantic metrics. The paper also owns its scope, explicitly limiting claims to simulation and reporting TMR/div/std alongside physical metrics. That's honest.\n\nThe soft spots are real but not disqualifying. Table 1 reports single deterministic rollouts; Section C.1 deliberately disables observation corruption, reset perturbations, and startup randomization. With margins like +5.0, +0.7, and +0.8 pp on SONIC, and +1.4 pp on only one ProtoMotions split, the comparative claim needs seed variance or confidence intervals before I'd call it reliable. The generator gain is also modest (92.58 to 94.43 with SONIC, 93.55 with ProtoMotions) and measured only against a frozen SONIC evaluator — the same policy family used to train the SONIC branch. That's a small risk of evaluator-specific overfitting, and a cross-controller or hardware check would be the natural fix. The private 1,024-prompt test suite and lack of released code/data don't help, though the paper does describe the protocol carefully.\n\nNet: this is a well-designed empirical study with an overreaching abstract. The core idea is plausible, the controls are informative, and the honest limitations section suggests the authors know where the evidence ends. I'd send it to peer review, with a clear request for uncertainty quantification and a softened central claim. I wouldn't cite it yet, but I'd want to see the revision and any released artifacts.","headline":"A well-controlled co-training study whose own small margins and missing error bars undercut the 'markedly broader coverage' abstract claim.","tokens_in":21183,"tokens_out":3590,"would_cite":false,"duration_ms":34778,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GenTrack claims that online co-training of a text-to-motion generator and a humanoid tracker improves both generator executability and zero-shot tracking coverage beyond what static replay or one-way alignment achieves.","keywords":["humanoid motion tracking","zero-shot tracking","text-to-motion generation","online co-training","execution-grounded reward","FlowGRPO","robot-native motion","Unitree G1"],"falsifier":"Replace the frozen SONIC evaluation with a different closed-loop controller (for example the ProtoMotions tracker) or tighten the fall criterion to 0.10 m pelvis-height deviation and re-run the matched GenTrack versus static-replay comparison; if the co-trained generator's success advantage shrinks or inverts, the claim that the loop yields genuinely more executable motion is falsified.","tokens_in":20140,"feed_emoji":"🤖","tokens_out":13998,"duration_ms":124154,"temperature":0.7,"pith_summary":"The paper tries to establish that the two bottlenecks in scaling zero-shot humanoid tracking—the cost of collecting embodied robot data and the residual gap between retargeted human motion and robot-executable motion—can be attacked together. Its proposal, GenTrack, is an online loop in which a pretrained text-to-motion generator and a humanoid tracker improve each other: execution outcomes by the tracker teach the generator which motions are robot-compatible, while newly generated references expand the tracker's training pool. On two different tracker backbones, the loop improves measured success on held-out motion splits and makes generated motions more executable in closed-loop simulation, without collecting new robot demonstrations. The authors argue that joint online post-training is what produces the gains, since equal-budget static replay and one-way alignment do not.","feed_headline":"Co-trained generator and tracker push zero-shot success to 90%","feed_subtitle":"A generator-tracker loop improves executability and broadens zero-shot coverage without new robot data.","key_machinery":"The load-bearing mechanism is an online sample–execute–update loop with a lagged judge. Each round, the generator samples $K$ motions per prompt, the tracker from the preceding round (kept frozen during generator optimization) executes them, and the execution score $S_{\\mathrm{exec}} = (1-c) + [e_j]^2 + [e_t/0.5]^2 + 0.5[e_d/0.5]^2 + 2\\mathbb{I}_{\\mathrm{fall}}$, with $[x]_2 = \\min(x,2)$, turns falls, incomplete rollouts, joint error ($e_j$), root-trajectory error ($e_t$), and root-displacement error ($e_d$) into a group-relative reward: rewards are normalized within each same-prompt group and applied by FlowGRPO with four clipped policy-ratio updates. Drift is constrained by a KL penalty to the frozen initial generator and periodic supervised flow-matching rehearsal on the original text–motion pairs. The same structurally valid generated references are accumulated and mixed in equal share with public references to update the tracker, so the reward model and the tracker co-evolve; the current trainee never gates its own feedback.","core_discovery":"GenTrack couples a pretrained text-to-motion generator with a pretrained humanoid tracker for the Unitree G1 in an online, alternating loop. Each round, the generator samples robot-space references from training prompts; a tracker frozen from the previous round executes them in closed loop, producing an execution score $S_{\\mathrm{exec}} = (1-c) + [e_j]^2 + [e_t/0.5]^2 + 0.5[e_d/0.5]^2 + 2\\mathbb{I}_{\\mathrm{fall}}$ with $[x]_2=\\min(x,2)$, where $c$ is completion fraction, $e_j$ the maximum wrapped joint error, $e_t$ mean root-trajectory error, $e_d$ root-displacement error, and $\\mathbb{I}_{\\mathrm{fall}}$ a fall indicator. Rewards are normalized within each same-prompt group and applied through FlowGRPO, while a frozen-generator KL penalty and supervised flow-matching rehearsal on the original text–motion pairs limit drift. The structurally valid on-policy generations are accumulated and mixed in equal share with public references to update the tracker. On the SONIC backbone, the loop raises frozen-SONIC generator success from 92.58% to 94.43%, lowers key-body error from 0.410 m to 0.325 m, and improves fall-only zero-shot tracking success on LAFAN1/AMASS-test/Wild-G1 from 85.0/79.0/47.2 to 90.0/79.7/48.0; matched static-replay and one-way filtering controls do not reproduce the gains. The paper concludes that joint online post-training narrows the executability gap between retargeted references and robot-native motion.","pith_inferences":["If the loop's gains hold across morphologies, the same alternating recipe could be applied to other robot platforms and to other generative backbones (latent diffusion or autoregressive motion models), which the paper does not test.","The fall-only success criterion tolerates up to 0.25 m reference-relative pelvis-height deviation; the method's practical value on real hardware would be better probed with a stricter, task- or contact-based success metric that the paper leaves to future work.","Removing the KL and rehearsal anchors raises raw executability but degrades TMR retrieval and FID in the reported ablations, which suggests the approach implicitly treats preserved text–motion alignment as part of \"executability\"; a deployment-focused variant might choose a different trade-off.","The paper's own ablation shows that using the current trainee as sole judge collapses cross-split success, implying a testable extension: varying the lag between tracker updates and reward scoring to find how much judge staleness is needed for stable co-adaptation."],"forward_implications":["Generator executability improves alongside semantics: the SONIC branch raises frozen-SONIC success from 92.58% to 94.43% while preserving or improving TMR-G1 retrieval and FID.","Zero-shot tracker coverage extends beyond the static reference pool: fall-only success on the three frozen splits rises on the SONIC backbone, with the largest relative gain on the most out-of-distribution split.","The recipe transfers across backbone designs: both ProtoMotions (AMP/PPO) and SONIC benefit, so the effect is not tied to one tracker training paradigm.","Equal-budget offline controls—static replay, final-generator replay, reward-weighted SFT, and DPO—fall short of the online loop, indicating that temporal co-adaptation is the active mechanism.","Without new robot data, the loop converts language-conditioned generation and closed-loop execution into a self-sustaining source of tracker supervision."],"supporting_citations":[{"why":"Provides the SONIC tracker backbone, the frozen-SONIC simulator evaluator, and the fall-only success criterion and MPJPE/Eg metric family used to measure gains.","marker":"(Luo et al. 2025)"},{"why":"Supplies the ProtoMotions tracker backbone and AMP/PPO baseline, the second independent instantiation of the loop.","marker":"(NVLabs 2025)"},{"why":"Contributes the FlowGRPO group-relative online-reinforcement objective used to align the generator from execution rewards.","marker":"(Liu et al. 2025a)"},{"why":"Provides the physical-feedback reward structure and the high-level/low-level evaluation decomposition that the generator experiments follow.","marker":"(Yue et al. 2025)"},{"why":"Supplies the AMASS archive that yields one of the three frozen zero-shot tracking splits and part of the public reference pool.","marker":"(Mahmood et al. 2019)"},{"why":"Supplies the LAFAN1 benchmark used as the second frozen zero-shot tracking split and public reference source.","marker":"(Harvey et al. 2020)"},{"why":"Supplies the 6D continuous rotation representation that makes the 38-dimensional robot-native motion parameterization learnable.","marker":"(Zhou et al. 2019)"}],"fun_headline_variants":["GenTrack co-training lifts zero-shot humanoid tracking without new data","Online generator-tracker loop sharpens robot executability and zero-shot reach","GenTrack: robot-native motion from co-trained generation and tracking","Zero-shot humanoid tracking gains from GenTrack's alternating co-training"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claims rest on a single simulated evaluator—closed-loop rollouts of a frozen SONIC policy with a 0.25 m pelvis-height fall threshold—being a faithful stand-in for real robot executability.","fun_headline_variants_meta":{"raw":{"variants":["GenTrack co-training lifts zero-shot humanoid tracking without new data","Online generator-tracker loop sharpens robot executability and zero-shot reach","GenTrack: robot-native motion from co-trained generation and tracking","Zero-shot humanoid tracking gains from GenTrack's alternating co-training"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000602,"raw_usage":{"total_tokens":2910,"prompt_tokens":1141,"completion_tokens":1769,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":757,"completion_tokens_details":{"reasoning_tokens":1694}},"tokens_in":757,"tokens_out":1769,"duration_ms":16194,"temperature":1.0,"reasoning_tokens":1694,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:11:53.771096+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace the frozen SONIC evaluation with a different closed-loop controller (for example the ProtoMotions tracker) or tighten the fall criterion to 0.10 m pelvis-height deviation and re-run the matched GenTrack versus static-replay comparison; if the co-trained generator's success advantage shrinks or inverts, the claim that the loop yields genuinely more executable motion is falsified.","supporting_citations":[],"review_version":2}