{"id":"e8395c54-636c-40ca-89b0-d3584d716f82","arxiv_id":"2505.12222","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A centroidal angular velocity reward, combined with actuator operating-region modeling and transmission load penalties, produces the first demonstrated full front flip on a one-leg hopper.","lead":"This paper trains a one-leg hopping robot to perform a full front flip in the real world, using a new reward based on the robot's whole-body angular velocity around its center of mass. The work matters because it offers a recipe for teaching highly dynamic, impact-heavy maneuvers that usually break in the gap between simulation and hardware.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hardware torque-reduction evidence for transmission load regularization rests on an unvalidated effective-inertia approximation; a small bias in Ieff could erase the claimed 21 to 17 Nm reduction.","rationale":"The hardware front flip itself is a direct empirical result and is not threatened by the estimator question, so I would not reject the paper. But the third contribution (transmission load regularization) is supported mainly by the torque estimate and by a small-sample hardware comparison. Since the estimator is validated only in simulation, a systematic bias on the physical closed-loop ankle could explain the entire reported torque reduction. This is the most load-bearing insecurity in the paper: the flip would stand, but the durability mechanism would collapse. I agree with the reader's weakest assumption and keep the verdict CONDITIONAL; no adjustment to the reader's verdict is needed, but the condition should explicitly require either hardware validation of the estimator or a sensitivity analysis of Ieff. The comparisons of BAV/CAM/CAV rewards are internally consistent and the CAV mechanism (rewarding L/I rather than L alone) is a reasonable explanation for the observed difference, though 'necessary' is stronger than what a three-way ablation can prove; that wording overclaim is secondary.","tokens_in":13867,"tokens_out":7644,"duration_ms":84154,"concrete_test":"Using the recorded hardware data behind Fig. 7c-d, recompute the external torque estimates under the extreme plausible effective inertias Ieff = I_rotor and Ieff = I_rotor + I_foot, and compare the peak values of the regularized and baseline policies. If the regularized peak does not remain below the baseline peak under both extremes, or if the separation shrinks below the estimator RMSE reported in Appendix A.5, then the claimed 21 to 17 Nm hardware torque reduction is not robust and the load-regularization benefit should be treated as unverified on hardware.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution most dependent on quantitative evidence is the transmission load regularization claim. Figure 7c-d reports a hardware reduction in estimated peak external ankle torque from about 21 Nm to 17 Nm and attributes the absence of gear fracture to this reduction. The estimator in Appendix A.5 computes tau_ext = tau_input - Ieff * domega_rotor with Ieff = I_rotor + I_foot/2. This effective inertia is an approximation for the closed-loop ankle mechanism and is validated only against simulation, where RMSE is 1.783 Nm during the initial impact phase. The claimed effect (roughly 4 Nm) is only about twice that RMSE, and the two compared policies produce different landing strategies (flat-foot vs rolling), so any bias in Ieff affects the two cases differently. If the true effective inertia on hardware differs from the assumed value, the 21 to 17 Nm reduction may be an estimation artifact, and the load-regularization pillar of the paper is not established. The hardware baseline is also only one successful flip before fracture versus eight for the regularized policy, so the durability conclusion alone is not statistically strong.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a centroidal angular velocity (CAV) reward for learning whole-body rotational maneuvers, combined with Motor Operating Region (MOR) modeling and transmission load regularization for sim-to-real transfer. As a case study, the authors train a one-leg hopper to perform a front flip, evaluate the reward design in simulation against base angular velocity (BAV) and centroidal angular momentum (CAM) rewards, compare MOR-constrained versus unconstrained policies, and present hardware experiments including repeated flips. They report the first hardware front flip on a one-leg hopper and attribute improved hardware durability to transmission load regularization.","tokens_in":14057,"tokens_out":7464,"duration_ms":72059,"significance":"If validated, the CAV reward is an elegant and transferable alternative to link-level rewards: it directly rewards the quantity that must be maximized for rotation while implicitly encouraging inertia reduction. The hardware front flip on a minimal single-foot platform is a notable advance, and the MOR analysis clearly shows the importance of actuator-aware constraints. The paper's strengths include clean ablations (BAV/CAM/CAV), simulation-to-hardware consistency for the MOR-evaluated policy, and reproducible details (reward tables, barrier formulation, domain randomization ranges). The main weaknesses concern the quantitative hardware evidence for transmission load regularization, which relies on an approximate external-torque estimator and a single baseline trial.","major_comments":[{"comment":"The hardware reduction in peak external ankle torque (from about 21 Nm to 17 Nm, Fig. 7c–d) is estimated using τ_ext = τ_input − I_eff·dω_rotor with I_eff = I_rotor + I_foot/2, an approximation for the closed-loop ankle mechanism that is validated only against simulation. The estimator's RMSE in the initial impact phase is 1.783 Nm, which is comparable to the claimed 4 Nm reduction, and the two policies land differently (flat-foot vs rolling), so a small bias in I_eff could alter the two estimates in opposite directions. Please validate the estimator on hardware (e.g., by applying known external torques or using an instrumented ankle) or provide a sensitivity analysis over I_eff to bound the uncertainty of the reported peak-torque reduction.","section":"§5.3, Appendix A.5"},{"comment":"The durability conclusion—that transmission load regularization prevents sun gear fracture—rests on a single successful baseline trial before the fracture (n=1) versus eight trials for the regularized policy. This is anecdotal as reported. If additional baseline trials are not feasible, the paper should either present them or explicitly label the hardware durability evidence as a case observation rather than a demonstrated effect, and give the simulation results (Fig. 7a–b) the primary evidentiary weight.","section":"§5.3, Fig. 8"},{"comment":"The mechanistic claim that CAM rewards fail because they do not incentivize inertia reduction is confounded: the CAM policy also produces substantially lower centroidal angular momentum (4.3 N·s) than the CAV policy (6.0 N·s). Thus the difference in centroidal angular velocity (4.8 vs 10.6 rad/s) could be partly due to lower momentum generation, not solely to the lack of inertia modulation. Please provide a matched-momentum comparison or a quantitative decomposition (e.g., reporting L and I at peak ω) to support the stated mechanism.","section":"§5.1, Fig. 4d–f"}],"minor_comments":[{"comment":"The heading 'Centroidal Momemtum' should be 'Centroidal Momentum'.","section":"Section 2 heading"},{"comment":"In Section 2, reference [11] is cited as 'Zhou et al.' but the reference list assigns [11] to Chignoli et al.; the intended 'Zhou et al.' appears to be [23]. Please correct the citation.","section":"References"},{"comment":"The title and abstract contain 'V elocity' with an extra space; fix the typography.","section":"Title and Abstract"},{"comment":"The caption 'Motor-side torque (without gear reduction) is shown' is ambiguous because the figure plots three quantities; please clarify which curve corresponds to motor-side torque.","section":"Figure 6 caption"}],"recommendation":"major_revision","confidential_remarks":"The 'first hardware front flip on a one-leg hopper' claim should be checked against prior art in one-leg hopper acrobatics that may not be cited. The paper also relies heavily on the authors' own prior MOR [2] and relaxed log barrier [33] formulations; the incremental novelty relative to those should be clarified in revision, although the CAV reward is a distinct contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the headline: this is a solid paper with one genuinely useful idea and one soft spot. The centroidal angular velocity reward is a real contribution. The ablation against base angular velocity and centroidal angular momentum shows exactly why it matters: CAM generates momentum but the policy never reduces inertia, so the rotation stalls; CAV couples momentum generation with inertia shaping, and the resulting mid-air knee fold is visible in the data. That is a clean, reproducible result in simulation, and the hardware flip is a legitimate first for a one-leg hopper.\n\nThe MOR section is also convincing. Training without MOR produces commands that exceed the feasible torque-speed envelope at takeoff and mid-air, and the policy fails when evaluated under MOR. With MOR, the real hardware torques stay inside the envelope and the flip works. That is a coherent sim-to-real story.\n\nThe weak section is transmission load regularization. The hardware evidence depends on an estimated external torque computed as input torque minus I_eff * domega, with I_eff approximated as the rotor inertia plus half the foot inertia. That effective inertia is an approximation for the closed-loop ankle mechanism, and it is validated only in simulation, where the RMSE during the impact phase is 1.783 Nm. The claimed reduction from about 21 to 17 Nm is roughly 4 Nm, only about twice that RMSE. The two compared policies land differently, so any bias in I_eff affects them differently. On its own, the hardware durability result—one successful flip before gear fracture versus eight with regularization—is suggestive but not statistically strong. This does not sink the paper, but it does mean the load-regularization contribution is not as established as the authors claim.\n\nI would also temper the 'general framework' language. The hardware validation is one platform and one maneuver. The appendix shows simulation results for other tasks and a quadruped, which is good, but it is simulation only.\n\nBottom line: the CAV reward and MOR analysis are worth a serious referee. The load regularization claim needs either a hardware-validated estimator or a more cautious write-up. I would send it to review with that requested revision.","headline":"Genuinely useful CAV reward and a real hardware flip, but the transmission-load-regularization evidence rests on an unvalidated torque estimator.","tokens_in":14614,"tokens_out":2659,"would_cite":true,"duration_ms":25683,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A reward built on the whole body's centroidal angular velocity, not base-link spin or total angular momentum, is what makes a one-leg hopper front-flip.","keywords":["reinforcement learning","sim-to-real transfer","centroidal angular velocity","front flip","one-leg hopper","motor operating region","transmission load regularization","legged robots"],"falsifier":"Instrument the physical ankle joint with a strain gauge or torque sensor and compare peak landing loads between regularized and unregularized policies over many trials. The paper's claim predicts a peak external torque drop from roughly 21 N·m to 17 N·m and no sun gear fracture across at least eight flips; if measured peak loads are statistically unchanged or the regularized hardware still fractures, the transmission-load benefit fails.","tokens_in":13642,"feed_emoji":"🤸","tokens_out":8722,"duration_ms":74267,"temperature":0.7,"pith_summary":"The paper claims that to make a legged robot perform a true full-body rotation, the learning objective must reward the system's centroidal angular velocity—the overall rotation rate of the whole body about its center of mass—rather than the base link's angular velocity or the total angular momentum. On a 12.45 kg one-leg hopper, the authors show that maximizing base angular velocity produces only thigh-calf flailing with no takeoff, while maximizing angular momentum generates takeoff but too little rotation; only the centroidal velocity reward yields a complete front flip. The paper also argues that two actuator-aware sim-to-real techniques are needed for hardware transfer: modeling the motor operating region in the torque–speed plane to keep torque commands feasible, and regularizing transmission load so landing impacts do not fracture the ankle's sun gear. With these ingredients, the authors report the first hardware realization of a full front flip on a one-leg hopper, over eight successful trials.","feed_headline":"Centroidal velocity reward makes a one-leg hopper front-flip","feed_subtitle":"First hardware full flip on a minimal hopper, trained in simulation and transferred with actuator-aware tricks.","key_machinery":"The load-bearing object is the centroidal angular velocity reward, defined as $r_{\\mathrm{CAV}} = \\max(\\min(\\alpha^T \\omega_{\\mathrm{com}}, 10), -0.1)$ during the aerial phase, where $\\omega_{\\mathrm{com}}$ is the centroidal angular velocity obtained from the centroidal momentum relation $h_G = I_G v_G$ and $\\omega_{\\mathrm{com}} = I_{\\mathrm{com}}^{-1} L_{\\mathrm{com}}$. This quantity links momentum to posture-dependent inertia, so rewarding it inherently encourages both momentum generation and mid-air inertia reduction. The supporting machinery is made of Motor Operating Region modeling, which clips commanded torques to a trapezoidal torque–speed envelope bounded by the voltage-limit slope and the current limit, and transmission load regularization, which penalizes contact-derived joint loads through a relaxed log barrier and probabilistic episode termination when loads cross a critical threshold.","core_discovery":"The central discovery is that the choice of what quantity is rewarded determines whether whole-body rotation actually emerges. For rotational maneuvers, the centroidal angular velocity $w_{\\mathrm{com}} = I_{\\mathrm{com}}^{-1} L_{\\mathrm{com}}$ couples angular momentum with the posture-dependent composite inertia, so maximizing it rewards both the generation of momentum and the reduction of inertia through configuration change. The paper shows this leads to a policy that extends the leg to build momentum during takeoff, then tucks the knee mid-air to raise spin rate from 7.4 rad/s to 10.6 rad/s and complete a full flip. By contrast, base angular velocity can be gamed by internal joint motion, and momentum-only rewards leave the leg extended and the rotation too slow. Combined with Motor Operating Region clipping and transmission load regularization, this reward is what the authors credit for the successful hardware transfer and repeated flip execution on the one-leg hopper.","pith_inferences":["The BAV-versus-CAV comparison suggests a general diagnostic for acrobatic learning: if a policy spends energy on internal joint motion without global rotation, the reward should be moved from link rates to centroidal rates; the paper only demonstrates this for flips and spins, but the diagnosis is transferable.","Because CAV rewards both momentum and inertia reduction, a momentum-only policy could plausibly be repaired by adding an explicit inertia-reduction bonus; this is a testable variant the paper does not run.","MOR clipping acts as a policy regularization that may benefit other high-torque, high-speed behaviors such as sprinting and jumping, where the torque–speed tradeoff is also decisive; the paper leaves that application unexplored."],"forward_implications":["A CAV-based reward, with MOR and load regularization, produces a complete front flip on real hopper hardware, whereas base-angular-velocity policies never leave the ground and angular-momentum policies undershoot the rotation.","Policies trained with MOR constraints keep torque commands inside the feasible actuator envelope, while policies trained with only box-shaped torque limits issue unattainable commands and fail when evaluated under MOR.","Transmission load regularization cuts the estimated peak ankle external torque from about 21 N·m to 17 N·m and lets the robot complete eight consecutive hardware flips, where the unregularized policy fractured the sun gear on the second trial.","The reward naturally produces a mid-air knee tuck that reduces composite inertia and raises centroidal angular velocity by 43%, from 7.4 rad/s at takeoff to 10.6 rad/s in flight.","The same reward formulation, with the flip axis as a parameter, also learns yaw spins and barrel rolls on the hopper and a backflip on a quadruped in simulation, suggesting the recipe extends beyond the demonstrated maneuver."],"supporting_citations":[{"why":"Defines the centroidal momentum matrix and the relation $h_G = I_G v_G$ on which the centroidal angular velocity reward is built.","marker":"[22]"},{"why":"Supplies the Motor Operating Region modeling approach used to clip torque commands to realistic torque–speed envelopes.","marker":"[2]"},{"why":"Provides the relaxed log barrier function used for transmission load regularization and other hard constraints.","marker":"[33]"},{"why":"Simulator used to model the closed-loop ankle mechanism and to produce the contact and impact simulation results.","marker":"[36]"},{"why":"Concurrent training of the control policy and state estimator gives the actor privileged state predictions and narrows the observation gap.","marker":"[37]"},{"why":"Prior centroidal-angular-momentum reward work that the paper contrasts with its centroidal velocity reward.","marker":"[28]"}],"fun_headline_variants":["One-leg hopper full front flip via centroidal reward","Centroidal angular velocity reward enables hopper flip","Hopper front flip: first hardware success with centroidal reward","Sim-to-real trick unlocks one-leg hopper front flip"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-reduction result depends on the simplified inertia model used to estimate ankle external torque on hardware, $I_{\\mathrm{eff}} = I_{\\mathrm{rotor}} + I_{\\mathrm{foot}}/2$; that estimator is validated only in simulation, so if it is inaccurate on the real closed-loop mechanism, the claimed torque drop and durability gain may not hold.","fun_headline_variants_meta":{"raw":{"variants":["One-leg hopper full front flip via centroidal reward","Centroidal angular velocity reward enables hopper flip","Hopper front flip: first hardware success with centroidal reward","Sim-to-real trick unlocks one-leg hopper front flip"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1291,"prompt_tokens":922,"completion_tokens":369,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":302}},"tokens_in":538,"tokens_out":369,"duration_ms":4111,"temperature":1.0,"reasoning_tokens":302,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:37:56.701873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Instrument the physical ankle joint with a strain gauge or torque sensor and compare peak landing loads between regularized and unregularized policies over many trials. The paper's claim predicts a peak external torque drop from roughly 21 N·m to 17 N·m and no sun gear fracture across at least eight flips; if measured peak loads are statistically unchanged or the regularized hardware still fractures, the transmission-load benefit fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the centroidal momentum matrix and the relation $h_G = I_G v_G$ on which the centroidal angular velocity reward is built."},{"cited_title":"Shin, T.-G","cited_arxiv_id":null,"evidence_quote":"Supplies the Motor Operating Region modeling approach used to clip torque commands to realistic torque–speed envelopes."},{"cited_title":"A Learning Framework for Diverse Legged Robot Locomotion Using Barrier-Based Style Rewards","cited_arxiv_id":"2409.15780","evidence_quote":"Provides the relaxed log barrier function used for transmission load regularization and other hard constraints."},{"cited_title":"Hwangbo, J","cited_arxiv_id":null,"evidence_quote":"Simulator used to model the closed-loop ankle mechanism and to produce the contact and impact simulation results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Concurrent training of the control policy and state estimator gives the actor privileged state predictions and narrows the observation gap."}],"review_version":1}