{"id":"2608d2bf-ce1e-4b36-8a25-93d988a2a235","arxiv_id":"2607.07830","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A proprioceptive humanoid policy trained with slope-adaptive ZMP regularization plus biomechanical reward gating traverses outdoor grass slopes to 32.1° without online exteroception.","lead":"HumoSlope is a two-stage RL framework that lets a blind humanoid walk continuous outdoor grass slopes up to 32.1° by combining a slope-plane ZMP balance prior with biomechanical gait rewards. It matters because steep continuous slopes remain a practical gap for deployable humanoids that cannot rely on vision at runtime.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The load-bearing claim rests on unquantified transfer of Stage-I-fitted soft priors (Eqs. 3–5) to real outdoor grass without online sensing.","rationale":"The reader correctly isolates the transfer of the privileged PCA-gated soft priors to a blind real-world actor as the weakest assumption. The manuscript’s sim evidence (Tables 1–2, Figs. 4–5) is internally consistent and shows that both the slope-adaptive ZMP and BSGA components matter under privileged training. The outdoor claim, however, is supported only by selected continuous traversals and qualitative posture photos; without repeated quantitative outdoor statistics or an ablated real-world control, it remains possible that domain randomization plus Stage-I balance alone, rather than the biomechanical gates, is doing most of the work on grass. That gap justifies keeping the verdict CONDITIONAL rather than upgrading to ACCEPT, and does not warrant REJECT because the engineering contribution and sim results still stand. No stronger internal inconsistency (e.g., in the ZMP derivation or reward formulation) was found.","tokens_in":13871,"tokens_out":621,"duration_ms":35014,"concrete_test":"Deploy the identical proprioceptive actor on the same outdoor grass slopes for N≥20 instrumented trials per grade (including the 32.1° segment), logging success rate, max distance, CoM-height time series, and hip/knee torque profiles; compare against a Stage-I-only checkpoint and a Stage-II policy trained with BSGA rewards ablated. If outdoor SR drops below ~70% or the CoM/torque signatures lose the claimed uphill-hip / downhill-knee asymmetry, the transfer claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the proprioceptive actor, after Stage-II training that gates soft rewards with a privileged five-dimensional PCA descriptor (Eq. 1) and Stage-I-fitted swing/CoM targets (Eqs. 3–5), generalizes to continuous blind outdoor grass slopes up to 32.1°. The paper shows qualitative posture variation (Fig. 6) and successful demos (Fig. 1), but supplies no quantitative real-world success rates, distance statistics, or failure modes across repeated trials, friction/wetness conditions, or abrupt slope transitions. Ablations (Table 2) and held-out sim tracks (Table 1) establish that removing BSGA collapses performance in simulation; they do not establish that the same soft-prior geometry remains sufficiently informative once the height-scan privilege and PCA plane fit are removed and the surface becomes deformable grass. If the fitted hip-pitch trend (Eq. 5) or the asymmetric CoM offsets (Eq. 3) are tuned to rigid sim geometry, the real-world posture adaptation may be incidental rather than caused by the claimed physics-guided mechanism. This is the single softest link between the engineering construction and the headline outdoor result.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper presents HumoSlope, a two-stage physics-guided RL framework for blind humanoid locomotion on continuous steep slopes. Stage I trains a proprioceptive actor–critic with a slope-adaptive ZMP regularizer that evaluates balance deviation on the local inclined support plane rather than a world-horizontal reference, producing a terrain-consistent balance prior. Stage II warm-starts from that actor and introduces the Biomechanical Slope Gait Adapter (BSGA), which uses a training-only five-dimensional PCA terrain descriptor (Eq. 1) extracted from height-scan patches to gate soft reward priors for slope-conditioned CoM height (Eq. 3), uphill/downhill lower-limb asymmetry (Eq. 4), and swing-hip guidance fitted from Stage-I rollouts (Eq. 5). The deployed actor remains purely proprioceptive. On held-out compound slope-track benchmarks the method reaches 77.1% success at 30° and a max grade of 36°, outperforming proprioceptive and one exteroceptive baseline; outdoor Unitree G1 demos show continuous traversal of grass slopes up to 32.1° with qualitative posture adaptation.","tokens_in":14300,"tokens_out":1407,"duration_ms":12372,"significance":"Continuous steep slopes are a distinct and under-studied regime for humanoid RL: they impose a persistent gravitational bias rather than discrete foothold selection, and generic rewards readily induce low-CoM “Groucho” gaits. The combination of a terrain-aligned ZMP prior with biomechanically motivated, descriptor-gated soft rewards is a concrete and transferable design pattern. Strengths include a held-out compound-track protocol with friction-tier normalization, three-checkpoint averaging, a full ablation suite (Table 2), biomechanical diagnostics (Fig. 5), and real outdoor video evidence on deformable grass. If the claimed transfer holds, the work supplies both a practical recipe for extreme-slope humanoid locomotion and a clear demonstration that physics- and biomechanics-informed reward shaping can mitigate posture degeneration without online exteroception.","major_comments":[{"comment":"The headline outdoor claim (continuous blind traversal of grass slopes up to 32.1°) rests on qualitative figures (Figs. 1, 6) and narrative description. Unlike the simulation protocol (Table 1: three checkpoints × three friction tiers, SR/MXD/T_trav), the real-world section supplies no repeated-trial success rates, distance statistics, failure modes, or wetness/friction conditions. Because the load-bearing transfer argument is that Stage-II soft priors (Eqs. 3–5) remain informative once the privileged PCA descriptor is removed and the surface becomes deformable grass, quantitative outdoor metrics (or at least a clear statement of trial counts and observed failure modes) are needed to support the central claim at the strength asserted in the abstract.","section":"§4.2 Real-world experiments / Abstract"},{"comment":"Table 2 shows that removing BSGA collapses SR to 0% at 20°, while Stage I alone retains 100% SR but with markedly lower CoM height and longer traversal time. This establishes necessity of BSGA in simulation, yet leaves open whether the fitted swing-hip trend (Eq. 5) and asymmetric CoM offsets (Eq. 3) are tuned to rigid sim geometry. A sensitivity or re-fit experiment (e.g., perturbing β0, β1 or b_up/b_down and re-evaluating both sim and real posture) would strengthen the causal link between the claimed physics-guided mechanism and the outdoor posture adaptation shown in Fig. 6.","section":"§3.2 Eqs. (3)–(5); Table 2; Fig. 6"},{"comment":"The slope-adaptive ZMP regularizer (Sec. 3.1) is a point-mass apparent-force surrogate with a contact-force-weighted support anchor. The paper correctly notes that a full ZMP/ZML formulation is noisy in massively parallel RL, but does not quantify how often the terrain-aligned intersection falls outside the actual support polygon or how sensitive r_ta_zmp is to the smoothing constant ε and scale σ_zmp. A short diagnostic (distribution of d_ta_zmp on steep segments, or ablation of the force-weighted anchor versus a simple mid-foot anchor) would make the Stage-I contribution more transparent and reproducible.","section":"§3.1 Slope-Adaptive ZMP Regularization"}],"minor_comments":[{"comment":"Reward weights (w_com, w_bio, w_swing), ascent/descent offsets, and the PCA clip threshold θ_clip are free parameters listed only conceptually; a table or appendix of numerical values (and any tuning protocol) would aid reproducibility.","section":"§3.2 / Implementation"},{"comment":"Fig. 5 panels lack error bands or trial counts; given the three-checkpoint averaging used in Table 1, the same protocol should be stated for the biomechanical diagnostics.","section":"Fig. 5"},{"comment":"The friction-tier definition µ = tan(|θ|) + ∆ is clear, but the text should note whether the same tiers were used for the Max Grade sweep column of Table 1.","section":"§4.1 / Table 1"},{"comment":"Minor notation: 1_up / 1_down appear both as indicators and as gating factors; a single sentence clarifying that they are binary regime flags would avoid ambiguity.","section":"Eq. (1) and Eq. (4)"},{"comment":"Related-work discussion of multi-contact ZMP/ZML (Caron et al., Brecelj & Petrič) is appropriate; a brief remark on why those geometric formulations were not used as hard constraints (rather than soft rewards) would help readers unfamiliar with the RL setting.","section":"§2 Dynamic Balance"}],"recommendation":"major_revision","confidential_remarks":"The engineering contribution is solid and the sim evaluation is careful; the main risk is over-claiming the outdoor result relative to the evidence supplied. If the authors can add even modest quantitative outdoor statistics (or a clear multi-trial protocol with failure modes), the paper becomes a strong accept for a robotics venue. Without that, major revision is the proportionate bar. No concerns about citation pattern or scope fit."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The headline result is a fully proprioceptive G1 policy that walks continuous outdoor grass slopes to 32.1° and held-out sim compound tracks to 36°, while avoiding the low-CoM “Groucho” crouch that generic rewards produce. That is a concrete field-robotics advance; continuous steep slopes have been under-treated relative to stairs and parkour.\n\nWhat is new is the two-stage recipe. Stage I evaluates a ZMP-style deviation on the local inclined support plane (force-weighted foot anchor + ray intersection) rather than a world-horizontal plane. Stage II then gates soft, low-weight biomechanical priors (slope-conditioned CoM height, hip-up / knee-down asymmetry, swing-hip target fitted from Stage-I rollouts) with a five-dimensional PCA descriptor of the privileged height scan. The actor stays blind at deployment. The paper shows this cleanly: Table 1 beats URL, FastTD3 and even the depth-based Gallant on the compound track; Table 2 ablations show Stage I alone crouches and removing BSGA collapses success; Fig. 5 and the outdoor photos show the expected posture modulation.\n\nThe soft spot is exactly the one the stress-test flags, and it is real but not fatal. Real-world evidence is qualitative demos and posture snapshots, not repeated-trial success rates, distance statistics or failure modes under wet grass or abrupt transitions. The claim that the Stage-I-fitted soft priors remain the causal mechanism once the height-scan privilege disappears is therefore plausible rather than proven. Free parameters (reward weights, β coefficients, σ_zmp, clip thresholds) are also not fully specified, so exact reproduction will be hard until code appears. None of that overturns the sim evidence or the outdoor capability that was shown.\n\nMath is standard RL reward shaping, not a new derivation; citations cover the biomechanics and multi-terrain baselines fairly. This is for people building or evaluating blind humanoid controllers for outdoor ramps and hillsides. It deserves a serious referee. I would engage with it, cite the outdoor numbers and the slope-plane ZMP idea, and push for public code plus quantitative real-world stats in revision.","headline":"Solid engineering advance on under-served steep-slope humanoid locomotion; the outdoor 32° grass result is real, the soft-prior transfer story is the only soft link.","tokens_in":14914,"tokens_out":539,"would_cite":true,"duration_ms":5576,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A two-stage physics-guided training scheme produces a fully proprioceptive humanoid policy that walks continuous outdoor grass slopes up to 32.1° without collapsing into a crouched low-CoM gait.","keywords":["humanoid locomotion","sloped terrains","reinforcement learning","Zero Moment Point","biomechanical gait adaptation","proprioceptive control","Sim-to-Real","center of mass regulation"],"falsifier":"Deploy the identical proprioceptive actor, trained without the BSGA reward gates or without the slope-adaptive ZMP term, on the same outdoor grass slope of measured grade ≥30°; if that ablated policy still completes continuous traversal while keeping mean CoM height comparable to the full model, the claimed necessity of the two-stage physics-guided adaptation is falsified.","tokens_in":14788,"feed_emoji":"⛰️","tokens_out":1067,"duration_ms":22772,"temperature":0.7,"pith_summary":"Steep continuous slopes put a constant gravitational bias on a humanoid, so ordinary reinforcement-learning rewards often settle on a slow, crouched “Groucho” gait that buys short-term balance at the cost of posture and further slope capability. This paper claims that the problem can be solved by a two-stage procedure called HumoSlope. Stage I first learns a balance prior by measuring Zero-Moment-Point deviation on the local inclined support plane rather than a world-horizontal plane. Stage II then uses a training-only five-dimensional PCA slope descriptor to gate soft biomechanical reward terms that raise CoM height and switch between hip-dominant uphill propulsion and knee-oriented downhill braking. The deployed controller never sees the terrain sensors; it runs on proprioception alone. In simulation the policy finishes compound tracks up to 36°; outdoors it continuously traverses wet grass slopes of 62.7 % grade (32.1°). A sympathetic reader cares because the same gravitational-bias problem appears on everyday ramps and hillsides, and the method shows that physics-aligned balance plus biomechanically motivated reward gates can replace both crouching and online vision.","feed_headline":"Blind humanoid walks 32° grass slopes without crouching","feed_subtitle":"Physics-aligned balance and biomechanical reward gates replace the low-CoM gait ordinary training produces","key_machinery":"HumoSlope: a two-stage framework whose Stage-I slope-adaptive ZMP regularizer (terrain-aligned deviation from a force-weighted support anchor) supplies a balance prior, and whose Stage-II Biomechanical Slope Gait Adapter (BSGA) gates CoM-height, hip/knee, and swing-leg soft rewards from a five-dimensional PCA terrain descriptor available only at training time.","core_discovery":"Generic model-free rewards for humanoid slope locomotion converge to an undesired low-center-of-mass crouched gait. HumoSlope prevents that degeneration by first installing a slope-adaptive ZMP regularizer evaluated on the local support plane, then using a privileged macroscopic terrain descriptor to gate soft biomechanical priors that modulate CoM height and lower-limb coordination. The resulting actor remains purely proprioceptive yet achieves continuous blind traversal of outdoor grass slopes up to 32.1° and simulated compound slopes up to 36°.","pith_inferences":["The same local-plane ZMP construction may transfer to other persistent gravitational biases such as walking under constant external force or on banked curves.","If the PCA descriptor can be replaced by a short history of proprioceptive accelerations, the entire pipeline could become fully unsupervised with respect to height maps.","Abrupt slope transitions will remain a failure mode until some form of look-ahead cue is added, exactly as the paper’s own limitations section anticipates.","Peak knee-torque diagnostics from the ablations suggest that uncontrolled crouching is not merely aesthetic; it is a measurable overload that limits maximum grade."],"forward_implications":["Humanoid policies trained with ordinary tracking-and-survival rewards will systematically prefer low-CoM crouches on continuous inclines unless the balance metric is evaluated on the local support plane.","Training-time macroscopic slope descriptors can encode uphill/downhill joint-work asymmetry without requiring the deployed controller to carry cameras or depth sensors.","Compound uphill–downhill tracks with friction tiers normalized to tan(θ) become a stricter and more informative benchmark than isolated constant ramps.","Once the Stage-I balance prior exists, soft biomechanical gates alone are enough to convert a crouched warm-start into a faster, more upright slope gait."],"fun_headline_variants":["Physics-guided ZMP stops crouched gaits on 32° grass slopes","HumoSlope keeps humanoids upright on blind 32° outdoor climbs","Biomechanical gates enable proprioceptive walks up 32° slopes","Slope-plane ZMP and BSGA prevent low-CoM humanoid crouch","Two-stage priors let blind humanoids traverse 32° grass continuously"],"cache_read_input_tokens":128,"weakest_assumption_plain":"A five-number PCA summary of a privileged height-scan patch, together with soft reward gates fitted from Stage-I rollouts, is rich enough that a purely proprioceptive actor can later walk unseen outdoor grass slopes without any online terrain sensing.","fun_headline_variants_meta":{"raw":{"variants":["Physics-guided ZMP stops crouched gaits on 32° grass slopes","HumoSlope keeps humanoids upright on blind 32° outdoor climbs","Biomechanical gates enable proprioceptive walks up 32° slopes","Slope-plane ZMP and BSGA prevent low-CoM humanoid crouch","Two-stage priors let blind humanoids traverse 32° grass continuously"]},"model":"grok-4.5","effort":"low","cost_usd":0.006976,"raw_usage":{"total_tokens":1768,"prompt_tokens":860,"num_sources_used":0,"completion_tokens":100,"cost_in_usd_ticks":69760000,"prompt_tokens_details":{"text_tokens":860,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":808,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":860,"tokens_out":100,"duration_ms":8904,"temperature":1.0,"reasoning_tokens":808,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T17:09:36.874332+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Deploy the identical proprioceptive actor, trained without the BSGA reward gates or without the slope-adaptive ZMP term, on the same outdoor grass slope of measured grade ≥30°; if that ablated policy still completes continuous traversal while keeping mean CoM height comparable to the full model, the claimed necessity of the two-stage physics-guided adaptation is falsified.","supporting_citations":[],"review_version":1}