{"id":"8059dd01-6284-42f8-a9ff-a388c8223372","arxiv_id":"2507.23053","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A CVAE-based in-between motion generator creates multi-style quadruped gaits from sparse motion data, and the trained controller runs gallop, tripod, trotting, and pacing on a real robot.","lead":"This paper builds a motion generator that fills in robot dog movements, like gallop and tripod, between a starting pose and a target pose, then trains a controller on those generated motions. It shows the controller running on a real Unitree AlienGo quadruped, which matters because less motion capture data could be needed to teach robots many different gaits.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The arbitrary-velocity claim rests on unvalidated phase-manifold extrapolation; no experiment tests generation or tracking outside the training velocity distribution.","rationale":"The reader's weakest_assumption identifies the phase manifold's generalization to arbitrary start/end states and unseen velocities as the key risk. I agree that this is the single most load-bearing concern: the paper's headline contribution is precisely the ability to synthesize motions at arbitrary velocities from sparse data, and that ability is never quantitatively demonstrated across a velocity range. The proposed concrete test would settle whether the concern lands. I considered alternative concerns: the missing no-generated-data baseline for the AMP policy is important, but it is secondary to the arbitrary-velocity claim; even a perfect policy baseline would not establish extrapolation to unseen velocities. The incorrect citation of RSMT [14] is a real but peripheral flaw. The real-robot deployment of gallop and tripod is genuine evidence that the pipeline can work for some velocities, and I credit that, but it does not validate the 'arbitrary' qualifier. Because the central claim is plausible but unverified in its strongest form, the reader's CONDITIONAL verdict remains appropriate; my stress-test does not move it.","tokens_in":7694,"tokens_out":4112,"duration_ms":53358,"concrete_test":"Run a velocity-stratified generation test: train the PAE, sampler, and CMoEs networks on mocap clips restricted to speeds below a threshold (e.g., 1.0 m/s); then generate motions targeting 1.5, 2.0, and 2.5 m/s. Measure foot-skating via Eq. 3, joint-limit violation rate, and terminal-position error as in Table III. Separately, train the AMP policy on the generated references and command these out-of-distribution velocities; if tracking error or foot-skating degrades sharply beyond the training range, the arbitrary-velocity claim fails. If code and data are released, this test can be reproduced directly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Abstract; Sec. V) is that the framework generates motions at arbitrary velocities from sparse motion capture data. The mechanism for this is the PAE phase manifold (Sec. III-A) together with the sampler and CMoEs decoder. However, the PAE manifold is fit only to the original mocap clips, and the sampler/CMoEs networks are trained on those same clips. Nothing in the loss functions (Eqs. 2-5) or the architecture enforces extrapolation to velocities outside the training set. The only quantitative motion-generation evaluation (Table III) reports terminal and full-trajectory L2 errors for the four gaits, but it neither states the velocity range of the test set nor tests generation at velocities beyond the training distribution. Thus 'arbitrary velocities' is asserted, not demonstrated. If the learned phase manifold and expert mixture only interpolate near the training velocities, the generated reference motions used for AMP training will not support the claimed arbitrary-velocity performance, and the real-robot gallop/tripod results (Fig. 6) may only reflect training-distribution velocities. This is load-bearing because the entire data-scarcity motivation depends on synthesizing useful reference motions for velocity regimes where mocap data is missing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a framework for quadruped robot locomotion that combines a CVAE-based in-between motion generator with a periodic autoencoder (PAE) phase manifold and a mixture-of-experts decoder to synthesize multi-gait motions (gallop, tripod, trot, pace) from sparse motion capture data. The generated motions are used as reference data to train an adversarial-motion-prior (AMP) policy in simulation, which is then deployed on a Unitree AlienGo. The paper claims arbitrary-velocity motion generation and improved velocity tracking, supported by simulation results and qualitative real-robot demonstrations of gallop and tripod.","tokens_in":7919,"tokens_out":5665,"duration_ms":65536,"significance":"The proposed approach addresses a real problem—mocap data scarcity for quadruped imitation learning—and the real-robot deployment of gallop and tripod is a notable positive result. If the arbitrary-velocity generalization claim were rigorously validated, the framework could be useful for expanding reference motion datasets. However, the evidence presented is incomplete: the extrapolation claim is not tested, the quantitative comparison is weakened by a citation error and missing statistics, and the contribution of the generated data is not isolated by ablation. The paper does not provide code or data, but the physical-constraint design and phase-manifold idea are interesting.","major_comments":[{"comment":"The central claim that the framework 'is capable of generating motions at arbitrary velocities' is not substantiated. The PAE phase manifold (Section III-A) and the CMoEs/sampler networks are trained exclusively on the original motion capture clips; Eqs. (2)-(5) contain no term that encourages extrapolation beyond the training velocity distribution. Table III reports L2 errors on a test set but does not state the velocity distribution or include any experiment at velocities outside the training range. Please add a quantitative velocity-extrapolation experiment (e.g., generate and track reference motions at velocities above and below the training range) and report the resulting errors.","section":"Section V and Section IV-A"},{"comment":"The comparison baseline and the claimed architectural inspiration are attributed to [14], which is 'RSMT: A remote sensing image-to-map translation model' (Remote Sensing, 2022), a paper unrelated to motion in-betweening. The method described in Section III-A appears to build on DeepPhase [15] and Tang et al. [13] instead. This citation error makes the method's provenance and the baseline comparison unverifiable. Please correct the citation and re-run the benchmark against the actual intended baseline.","section":"Section IV-A and Reference [14]"},{"comment":"The abstract and conclusions credit the generated motion data with 'enhancing controller stability and improving velocity tracking performance,' but no controlled ablation is reported. There is no comparison of an AMP policy trained on the original mocap data alone versus one trained on the original plus generated data. Without this ablation, the improvements cannot be attributed to the motion generator. Please add a baseline policy trained only on the original mocap trajectories and report quantitative velocity-tracking error and stability metrics.","section":"Section IV-B and Section V"},{"comment":"The real-world validation is qualitative (foot contact patterns, joint angles, torques) and covers only gallop and tripod. No numerical velocity-tracking error, stride frequency accuracy, or success rate is reported for hardware, so the conclusion of 'accurate velocity tracking performance through ... deployment on real-world robots' is not quantitatively supported. Please report quantitative metrics from the real-robot trials, ideally with commanded versus measured velocities.","section":"Section IV-B, Figs. 5-6"},{"comment":"The initial phase vector is predicted from a sequence constructed by repeatedly replicating the starting frame. Because the PAE manifold is trained on joint angular velocities (Section III-A1), a static replicated sequence has zero velocity and may fall outside the manifold's valid input distribution. The paper does not describe how this prediction network is trained or validated for such inputs. Since an incorrect initial phase would propagate through the whole generated motion, this is load-bearing for deployment. Please provide training and validation details, and show that the predicted initial phase is consistent with the subsequent generated motion.","section":"Section III-A, 'First Frame Predict'"}],"minor_comments":[{"comment":"The captions contain typos: 'T ripod' and 'T rotting' should be 'Tripod' and 'Trotting'.","section":"Fig. 1 and Fig. 3 captions"},{"comment":"The text says 'matching baseline terminal accuracy during trotting,' but Table III shows the proposed method is better (0.344 vs 0.444 for the last-frame L2 norm); please correct this inconsistency.","section":"Section IV-A, text near Table III"},{"comment":"Table III lacks error bars, sample counts, and units; the reader cannot assess the statistical significance of the reported differences.","section":"Table III"},{"comment":"The regularization reward rows are not clearly formatted; the scale for each term should be listed separately to avoid ambiguity between the term values and their weights.","section":"Table II"},{"comment":"The state definition s = {proot, Rroot, q} is not fully specified; a table of state dimensions and the ranges of q would improve reproducibility.","section":"Section III-A, Data Formatting"},{"comment":"The paper does not include a limitations section; given the open extrapolation concerns, a discussion of the velocity range and generalization limits would be helpful to the reader.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The real-robot gallop and tripod demonstrations are a positive contribution, but the arbitrary-velocity claim is untested, the baseline citation is incorrect, and no ablation isolates the benefit of the generated data. The manuscript is not yet ready for acceptance; these issues should be addressed with additional experiments and corrections."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the short version: this paper does something real — it deploys an AMP policy on a Unitree AlienGo for gallop and tripod, using a CVAE-based motion generator trained on sparse mocap data. That's a worthwhile step. The integration of PAE phase manifolds with CVAE in-betweening and joint-space inputs, plus a joint-limit loss and first-frame phase prediction, is sensible and mostly well presented.\n\nThe good: the pipeline is coherent, the real-robot figures show actual gallop and tripod execution, and the four-gait simulation results are plausible. The authors are honest that trotting and pacing are less demanding. The modifications to the base architecture are technically reasonable — using joint angular velocities instead of estimated skeletal velocities is a practical choice.\n\nThe soft spots are real. The central claim — 'generating motions at arbitrary velocities' — is asserted but not demonstrated. Table III reports L2 errors but never states the velocity range of the test set or tests generation outside the training distribution. The stress-test concern lands: the PAE manifold and CMoEs decoder are trained on the original clips, and nothing enforces extrapolation. The real-robot results may just reflect training-distribution velocities. This is load-bearing because the data-scarcity motivation depends on synthesizing reference motions for velocity regimes where mocap is missing.\n\nThe comparison baseline is a clear citation error: reference [14] is a remote sensing paper, not the motion in-betweening work it's supposed to be. That needs to be fixed before anything else. There are also no error bars or sample counts in Table III, no ablation without generated data, and no code or data release. The style reward in AMP is trained to match the same generated data, which is a minor circularity, but the velocity-tracking metric is external, so that's not fatal.\n\nBottom line: the framework is plausible and the hardware demo is a point in its favor, but the paper overclaims 'arbitrary velocities' without evidence. A serious referee should see it; with a corrected baseline, an ablation, and narrowed claims it could be a solid contribution.","headline":"Plausible integration with a real-robot demo, but arbitrary-velocity claim is unsupported and the baseline citation is wrong.","tokens_in":8450,"tokens_out":3791,"would_cite":false,"duration_ms":36945,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By filling in frames between arbitrary poses, this framework turns sparse motion-capture data into reference trajectories that train a quadruped controller to gallop, tripod, trot, and pace.","keywords":["quadruped locomotion","motion in-between generation","conditional variational autoencoder","periodic autoencoder","phase manifold","adversarial motion priors","imitation learning","velocity tracking"],"falsifier":"Run the generator with a start and target state that require a velocity outside the training set's range, say twice the maximum recorded speed in the sparse capture, and measure the L2 norm of the global position of the final predicted frame, the metric reported in the paper's comparison table. If the error at that out-of-range velocity reaches or exceeds the baseline's error, the phase-manifold generalization premise is not holding, and a policy trained on such data should also fail to track a velocity command at that speed.","tokens_in":7449,"feed_emoji":"🤖","tokens_out":9412,"duration_ms":104563,"temperature":0.7,"pith_summary":"Quadruped locomotion via imitation learning is bottlenecked by scarce, short, velocity-incomplete motion-capture data. This paper claims to break that bottleneck by generating the missing in-between motion: a conditional variational autoencoder synthesizes physically plausible quadruped frames between arbitrary start and end states, using a learned phase manifold for gait-style continuity and joint-limit and foot-skating losses for physical plausibility. The generated sequences serve as reference motions for an adversarial imitation learning controller, and the authors report accurate velocity tracking, stability across four gaits, and real-world execution of gallop and tripod. If the framework works as claimed, sparse motion capture becomes a sufficient source for diverse multi-gait quadruped controllers, removing a central data bottleneck in legged imitation learning.","feed_headline":"Sparse motion data yields four robot gaits via in-between generation","feed_subtitle":"A CVAE fills intermediate frames between poses, and an imitation policy turns them into stable gallop, tripod, trot, and pace.","key_machinery":"The load-bearing mechanism is a three-part motion generator. A periodic autoencoder (PAE) learns a low-dimensional phase manifold from joint-angular-velocity data, encoding where each gait is in its cycle. A conditional mixture-of-experts (CMoEs) decoder, gated by the phase value and conditioned on the current state, target velocity, and a latent style variable, predicts the next state delta $\\Delta s_{t+1}$. A sampler network with an LSTM predictor converts the target state, the current phase, and the difference to the target into the latent, phase, and velocity inputs that let the decoder aim at the destination. The training loss stacks foot-skating, KL, position, rotation, orientation, root-position, and joint-limit terms, so each synthesized frame respects both the task and the robot's hardware limits. A first-frame phase-prediction network supplies the initial phase vector at deployment, where no future frames are available yet.","core_discovery":"The central claim is that an in-between motion generator can turn a sparse motion-capture dataset into dense reference trajectories at essentially arbitrary velocities, and that these synthesized references are good enough to train an adversarial imitation policy that transfers to hardware. The generator's inputs and outputs are the robot's root state and joint poses rather than global skeletal positions, which keeps the synthesized motion compatible with the robot's joint limits and control interface. The paper reports that the generated motions preserve each gait's defining features—aerial phases in gallop, triangular support in tripod, diagonal coordination in trot, same-side coordination in pace—and that the trained policy tracks commanded velocities in simulation and on a real quadruped. In comparison against the baseline in-betweening model, the paper reports lower terminal-frame and whole-clip position error across most gaits.","pith_inferences":["Not claimed in the paper: the same phase-manifold generator could synthesize continuous velocity ramps or gait-to-gait transitions by setting target states outside the recorded clips; the paper demonstrates fixed-gait execution only.","Not claimed in the paper: the data-formatting choice of root state plus joint poses should transfer to other legged morphologies, such as bipeds or hexapods, after re-tuning the phase dimension and joint-limit losses; only a quadruped is tested.","Not claimed in the paper: an ablation that replaces the learned phase manifold with a fixed periodic clock while holding data volume constant would isolate how much of the velocity-tracking gain comes from phase continuity versus simply having more reference frames; the paper does not run this ablation."],"forward_implications":["A single trained policy can execute gallop, tripod, trotting, and pacing, with the reference motions for all four gaits generated from sparse capture data rather than collected separately.","Because the generator outputs root state and joint poses directly, the synthesized motion transfers to the robot's control interface without retargeting or kinematic preprocessing.","Velocity-tracking accuracy against commanded references improves relative to training on sparse data alone, supporting the claim that generated in-between data is load-bearing for controller stability.","The first-frame phase predictor lets the controller start from a single observed frame at runtime, without requiring a history of preceding frames.","Joint-limit and foot-skating penalties in the generator's loss keep the synthesized reference motions inside the robot's hardware capabilities, which the authors report as the reason no kinematic filtering was needed before policy training."],"supporting_citations":[{"why":"Supplies the periodic autoencoder that learns the phase manifold encoding gait style and phase continuity.","marker":"[15]"},{"why":"Provides the conditional variational in-betweening formulation that the generator adapts to quadruped joint-space states.","marker":"[13]"},{"why":"Acts as the base architecture the motion generator is inspired by and as the comparison baseline in the tracking-error table.","marker":"[14]"},{"why":"Introduces the adversarial motion prior used as the style reward that makes the policy imitate generated motion.","marker":"[6]"},{"why":"Shows that adversarial motion priors can replace hand-designed reward functions, motivating the policy's discriminator-based style reward.","marker":"[7]"}],"fun_headline_variants":["CVAE generates four quadruped gaits from sparse pose data","In-between motion synthesis enables multi-gait robot locomotion","Sparse reference data powers real quadruped's four-gait controller","From sparse poses to gallop, trot, pace, tripod on real robot","Multi-style quadruped gait generation via in-between CVAE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the periodic phase manifold learned from the original sparse motion-capture clips stays meaningful for start and end states, and velocities, that never appear in the training data, so the sampler and decoder can fabricate intermediate frames the imitation policy can trust.","fun_headline_variants_meta":{"raw":{"variants":["CVAE generates four quadruped gaits from sparse pose data","In-between motion synthesis enables multi-gait robot locomotion","Sparse reference data powers real quadruped's four-gait controller","From sparse poses to gallop, trot, pace, tripod on real robot","Multi-style quadruped gait generation via in-between CVAE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000545,"raw_usage":{"total_tokens":2569,"prompt_tokens":869,"completion_tokens":1700,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":1609}},"tokens_in":485,"tokens_out":1700,"duration_ms":12799,"temperature":1.0,"reasoning_tokens":1609,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:05:01.111252+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the generator with a start and target state that require a velocity outside the training set's range, say twice the maximum recorded speed in the sparse capture, and measure the L2 norm of the global position of the final predicted frame, the metric reported in the paper's comparison table. If the error at that out-of-range velocity reaches or exceeds the baseline's error, the phase-manifold generalization premise is not holding, and a policy trained on such data should also fail to track a velocity command at that speed.","supporting_citations":[{"cited_title":"Deepphase: Periodic autoen- coders for learning motion phase manifolds,","cited_arxiv_id":null,"evidence_quote":"Supplies the periodic autoencoder that learns the phase manifold encoding gait style and phase continuity."},{"cited_title":"Real- time controllable motion transition for characters,","cited_arxiv_id":null,"evidence_quote":"Provides the conditional variational in-betweening formulation that the generator adapts to quadruped joint-space states."},{"cited_title":"Rsmt: A remote sensing image-to- map translation model via adversarial deep transfer learning,","cited_arxiv_id":null,"evidence_quote":"Acts as the base architecture the motion generator is inspired by and as the comparison baseline in the tracking-error table."},{"cited_title":"Amp: Adversarial motion priors for stylized physics-based character con- trol,","cited_arxiv_id":null,"evidence_quote":"Introduces the adversarial motion prior used as the style reward that makes the policy imitate generated motion."},{"cited_title":"Adversarial motion priors make good substitutes for complex reward functions,","cited_arxiv_id":null,"evidence_quote":"Shows that adversarial motion priors can replace hand-designed reward functions, motivating the policy's discriminator-based style reward."}],"review_version":1}