{"id":"77119a64-db95-48eb-8f0b-ad07e8227024","arxiv_id":"2505.17627","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A hierarchical framework that maps wrist force/torque into velocity commands and then into stable leg motions lets a humanoid robot carry loads cooperatively with a human using only haptic cues.","lead":"Researchers built a system that lets a humanoid robot follow a human's direction while carrying a load together, using only the forces felt in the robot's wrists. It was tested on a real Unitree G1 robot and matched a human follower on several teamwork metrics, though it was slower overall.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-world payload tests (0–5 kg) exceed the RL training force envelope (Eq. 16, Fmax=15 N/wrist, at most ~3 kg total); without logged wrist forces, the load-adaptive claim at 5 kg is unverified.","rationale":"I read the central claim as requiring a payload-adaptive locomotion policy that transfers zero-shot to real co-manipulation under the tested payloads. The most load-bearing condition is therefore that the RL policy's trained force range covers the forces encountered in deployment. The reader identified exactly this condition, and I agree. The paper does include genuine real-robot experiments and a sim-to-sim test under a 30 N load, which supports the framework at roughly the trained envelope; that is real evidence and prevents me from recommending rejection. However, the missing per-wrist force measurements and the absence of a per-payload breakdown leave the 5 kg claim unsupported. A stratified retest with instrumented wrists would settle the point. If the concern lands, the correct remedy is to retrain with a wider force range or revise the payload claim, not to discard the hierarchical architecture. For these reasons, the reader's CONDITIONAL verdict remains appropriate, and my analysis does not move it.","tokens_in":12125,"tokens_out":11449,"duration_ms":104002,"concrete_test":"Re-run the real-robot evaluation with payloads stratified at 0, 1, 3, and 5 kg (n >= 5 per level), logging both ATI wrist force/torque streams throughout each trial. Compute the per-wrist z-force distribution and compare it to the training support implied by Eq. 16. If 5 kg trials show sustained per-wrist forces above 15 N while Table I metrics are maintained, the generalization concern is resolved. If forces stay below 15 N only because the human leader bears the excess, or if the 5 kg trials degrade, then the 'load-adaptive up to 5 kg' claim is unsupported and the paper must either retrain with a larger force range or soften the payload claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the low-level policy is load-adaptive over the tested payload range depends on sim-to-real generalization beyond its training distribution. Section III-C2 randomizes per-wrist z-forces with Fmax=15 N (Eq. 16), and the text's stated lower bound of -3 N contradicts the symmetric uniform distribution written in the equation, but in any reading the downward support is capped at 15 N per wrist. That corresponds to at most about 3 kg total if both wrists saturate, matching the paper's own '0-3 kg' training description. Section IV-C then evaluates real trials with payloads varying from 0 to 5 kg, which would require up to ~24.5 N per wrist if the robot bore the load evenly, and more under uneven loading. The paper reports no per-payload breakdown and, crucially, no measured wrist force/torque data from the real deployment. Without those logs, it is impossible to tell whether the robot actually experienced forces above the trained envelope or whether the human leader carried the excess load. If the latter occurred, the 'load-adaptive' component is not validated at 5 kg, and the reported 'on par' metrics partly reflect human compensation rather than robot capability. This is an internal inconsistency between the training distribution and the evaluation range, not merely a disagreement with community norms.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes H2-COMPACT, a hierarchical framework for human-humanoid co-manipulation. A diffusion-based behavior-cloning network maps six-axis force/torque signals from dual wrist sensors into planar velocity and yaw-rate commands, and a PPO-trained locomotion policy maps those high-level twists to leg joint targets. The locomotion policy is trained in Isaac Gym with randomized payloads, friction, and external forces, then validated in MuJoCo and deployed zero-shot on a Unitree G1. Training data for the intent model come from dyadic human-human trials collected with only RGB video and F/T streams, processed by SAM2 and WHAM. The paper reports real-world human-humanoid trials with payloads from 0 to 5 kg and compares four metrics (completion time, trajectory deviation, velocity difference, follower force) against a blindfolded human-human baseline, claiming on-par or better performance and the first demonstration of learned haptic guidance fused with full-body legged control.","tokens_in":12388,"tokens_out":6552,"duration_ms":52825,"significance":"If the claims hold, the paper is a meaningful advance: the decoupling of haptic intent inference from whole-body locomotion is a sensible and reusable architecture, the vision-only data-collection pipeline is practical, the real-robot deployment is a genuine effort, and the public release of code and videos supports reproducibility. The contribution would be the first integration of learned wrist F/T-based intent inference with dynamic legged humanoid control for cooperative carrying. However, the central 'on par' and 'load-adaptive over 0-5 kg' claims are currently stronger than the evidence, for reasons detailed in the major comments.","major_comments":[{"comment":"The force-randomization envelope used to train the low-level policy does not cover the payload range on which the load-adaptive claim is evaluated. Equation (16) samples a per-wrist z-force from U([-Fmax, Fmax]) with Fmax = 15 N; the accompanying text says the lower bound is -3 N, which contradicts the symmetric distribution written in the equation, and in either reading the downward force is capped at 15 N per wrist, i.e. at most about 3 kg total when both wrists saturate. Section IV-C reports real trials with payloads varying from 0 to 5 kg, which would require roughly 24.5 N per wrist under an even load split and more under uneven loading. The paper reports no measured wrist F/T data from the real deployment and no per-payload performance breakdown, so it is not possible to verify whether the robot experienced forces above the trained envelope or whether the human partner carried the excess load. This is an internal inconsistency between the training distribution and the evaluation range, not merely a stylistic issue; either the training distribution must be extended to cover 5 kg, or the real trials must be accompanied by wrist-force logs and per-payload results to support the load-adaptive claim.","section":"§III-C2, Eq. (16); §IV-C"},{"comment":"Table I reports only means for each metric, with no standard deviations, confidence intervals, significance tests, or per-trial data, despite the experimental section describing multiple participants and repetitions. The human-human and human-humanoid means for trajectory deviation, velocity difference, and follower force are close (0.1109 vs 0.1294 m, 0.165 vs 0.143 m/s, and 17.355 vs 16.230 N), but without dispersion or a statistical test it is not possible to conclude that these differences are meaningful or that the robot performs 'on par.' The paper should report per-trial distributions and a statistical comparison, or explicitly weaken the on-par claim to a qualitative observation.","section":"§V, Table I"},{"comment":"Completion time is not on par by the paper's own reported means: the human-humanoid dyad took 51.47 s versus 23.78 s for the human-human dyad, a factor of 2.16. The authors attribute this to the G1's 0.8 m/s speed cap, which is plausible, but completion time is still one of the four headline metrics listed in Table I and referenced in the abstract. The paper should either report a speed-cap-normalized completion time, exclude completion time from the on-par statement, or revise the claim to state explicitly that the humanoid is slower while matching the other metrics.","section":"§V-A, Table I"},{"comment":"The simulation experiment intended to demonstrate load adaptation uses a single constant 30 N payload in MuJoCo and does not include a real-hardware comparison between the baseline and the load-adaptive policy. Section V-B shows that the baseline policy falls under a 30 N payload in simulation, but the real-world evaluation does not isolate the load-adaptive component: there is no ablation on the physical robot and no per-payload stability or tracking metrics. As a result, the real-world evidence supports the combined hierarchical system but does not specifically verify the claim that the low-level policy is load-adaptive over the full 0-5 kg range tested in Section IV-C.","section":"§V-B, §IV-C"}],"minor_comments":[{"comment":"The noise term in the forward diffusion is written as epsilon ~ N(0, I4), but y is a velocity sequence in R^{H x 3}; the identity matrix dimension should be consistent with the flattened velocity dimension (e.g., I_{18} for H = 6).","section":"§III-B2, Eq. (4)"},{"comment":"The notation A_iell is used without defining i or the dimension of the attention matrix, and the word 'matrics' should be 'matrices.'","section":"§III-B2, Eq. (6)"},{"comment":"There is a typo in the sentence describing the decoupling: 'orce to high-level whole-body velocities' should read 'force to high-level whole-body velocities.'","section":"§III-A"},{"comment":"References [7] and [28] are the same PPO paper and should be consolidated; this duplication also appears in the related-work citation style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a robotics venue and the real-robot demonstration is valuable. The claimed novelty of learned haptic guidance fused with full-body legged control is plausible, but the evaluation rigor is not yet at the level needed to support the central claims: the training force envelope does not cover the evaluated payload range, and the headline metrics are reported as means only. These issues are fixable with additional experiments, wrist-force logging, statistical reporting, and claim softening, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a genuine system paper: haptic intent inference via a diffusion policy on wrist force/torque feeding a PPO locomotion policy, deployed on a Unitree G1 with real human partners. That combination is new relative to the fixed-base and wheeled co-manipulation work they cite, and the vision-only data collection pipeline is practical. The sim-to-real transfer is real evidence, and the hierarchical decomposition is clean.\n\nThe soft spots are real but not fatal. Table I reports only means, with no variance or significance tests, so the 'on par' claim is statistically thin. Completion time is 2.2x slower for the robot; the speed-cap explanation is plausible, but it undercuts a blanket 'on par'. More importantly, the RL policy is trained with per-wrist forces in U([-3, 15]) N (about 0-3 kg total), yet the real trials use payloads up to 5 kg. Without per-payload results or logged wrist forces, we can't tell whether the robot actually experienced forces beyond its training envelope or the human compensated. That's an internal inconsistency between training and evaluation ranges, not just a missing baseline.\n\nThe intent model also lacks a simple non-learned baseline—something like a linear impedance mapping would show whether the diffusion policy earns its complexity. The text/equation mismatch on the force lower bound is minor but should be fixed.\n\nOverall, the architecture holds up, and the real-robot demonstration is worth taking seriously. The weaknesses are correctable with more data reporting and softened claims. I'd send this to peer review and push for revised statistics, per-payload analysis, and a baseline before accepting.\n\nRecommendation: engage with it; it deserves a serious referee.","headline":"A real humanoid co-manipulation demo with a clean hierarchical design, but the 'on par' claim outruns the evidence and the payload envelope mismatch needs addressing.","tokens_in":12941,"tokens_out":3002,"would_cite":true,"duration_ms":26724,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Wrist forces alone can guide a humanoid robot to carry loads alongside a human partner.","keywords":["human-robot co-manipulation","haptic intent inference","humanoid locomotion","reinforcement learning","diffusion policy","sim-to-real transfer","payload adaptation","force/torque sensing"],"falsifier":"A controlled trial in which a 5 kg box is held so that one wrist bears the entire load, pushing per-wrist vertical force above the 15 N training range, and the robot is expected to follow the leader; if it falls or fails to track, the claim of zero-shot load-adaptive co-manipulation is falsified.","tokens_in":11925,"feed_emoji":"🤖","tokens_out":8333,"duration_ms":61256,"temperature":0.7,"pith_summary":"This paper aims to show that a humanoid robot can act as a cooperative load-carrying partner using only haptic cues, namely six-axis force and torque readings from its wrists, to infer where the human wants to go. The authors build a two-tier learning system: a behavior-cloning network turns wrist forces into whole-body velocity commands, and a reinforcement-learned walking policy turns those commands into stable joint motions under variable payloads. Training data come from human–human dyads captured with ordinary RGB video and force sensors, avoiding the need for motion-capture infrastructure. In real trials, the robot followed a human leader carrying a box with trajectory deviation, velocity synchrony, and follower force comparable to a blindfolded human follower, though slower due to the robot's speed limit. If correct, this is the first demonstration of learned haptic guidance fused with full-body legged control for fluid human–humanoid co-manipulation.","feed_headline":"Wrist forces alone let a humanoid carry loads with a human partner","feed_subtitle":"A learned force-to-motion pipeline matches blindfolded human followers on trajectory and speed, with no vision or speech needed.","key_machinery":"The load-bearing mechanism is the decoupling of intent interpretation from locomotion control. The upper tier is a compact haptic intent model: multi-resolution stationary wavelet coefficients of the six-axis force/torque streams are fed, through learned keys and values, into a multiscale conditional diffusion policy whose Transformer encoder attends across wavelet levels, and deterministic DDIM sampling produces future velocity sequences; the first horizon token is the reference command. The lower tier is a PPO-trained asymmetric actor-critic locomotion policy that takes the command, joint states, and gait phase and outputs target joint angles, with random per-wrist forces applied during training to force load adaptation. The decoupling means the same velocity-command interface could, in principle, be reused with any walking controller.","core_discovery":"The central claim is that a hierarchical policy can make a legged humanoid cooperate with a human at a physical task using force feedback as the sole communication channel. At the upper level, a diffusion-based behavior-cloning model maps the last fraction of a second of dual-wrist force/torque data to planar linear velocity and yaw rate commands, effectively translating the leader's pushes and pulls into motion intentions. At the lower level, a PPO-trained locomotion policy, exposed during training to randomized downward and upward wrist forces up to 15 newtons plus varied friction and mass, converts those commands into joint angles that keep the robot balanced while carrying a load. The authors report that on a real humanoid, this zero-shot deployed pipeline matches a blindfolded human follower on trajectory deviation (0.129 m vs 0.111 m), velocity difference (0.143 m/s vs 0.165 m/s), and follower force (16.23 N vs 17.36 N), while taking longer to finish because of the robot's 0.8 m/s speed cap. The paper's own framing is that this is the first successful fusion of learned haptic intent inference with whole-body legged locomotion for human–humanoid co-manipulation.","pith_inferences":["The same force-to-velocity mapping could be trained for other physical collaboration tasks, such as pushing a cart or guiding an upper-body assist, since it never sees the robot's own dynamics and only the low-level policy would need retraining.","The real trials carried payloads up to 5 kg while training randomized only about 3 kg total across the two wrists, so a direct test at the edge of the payload envelope would clarify whether the zero-shot success reflects true extrapolation or a favorable load distribution.","The inverse mapping, where the robot signals intent to a human through the same wrist sensors, is a natural next direction not addressed by the paper, and the hierarchical design suggests it could be implemented by reversing the direction of the force-to-velocity supervision.","The eight motion primitives used for data collection are planar; testing on slopes or stairs would determine whether the intent inference generalizes beyond flat-ground co-manipulation."],"forward_implications":["Humanoid robots could serve as load-carrying assistants in homes or warehouses guided only by a human's physical lead, with no voice, vision, or joystick commands.","Because the intent model outputs platform-agnostic velocity commands, the haptic inference tier could transfer to other legged or wheeled platforms equipped with the same wrist sensors.","The vision-only training data pipeline removes the need for motion-capture studios when collecting human–human demonstration data for co-manipulation.","Zero-shot deployment from simulation to the real robot works for payloads and speeds within the trained envelope, as shown by stable tracking and forces close to the human–human baseline.","The completion-time gap is attributed to the robot's speed cap rather than haptic miscommunication, so faster humanoids should close that gap without changing the method."],"supporting_citations":[{"why":"Supplies the asymmetric actor-critic humanoid locomotion training recipe and zero-shot sim-to-real principles that the low-level policy builds on.","marker":"[23]"},{"why":"Provides the GPU-accelerated simulation environment where the locomotion policy is trained under randomized payloads and friction.","marker":"[8]"},{"why":"Defines the diffusion policy architecture that the haptic intent inference model adapts for force-to-velocity mapping.","marker":"[34]"},{"why":"Supplies the human–human dyad data and the four evaluation metrics used as the baseline for comparison.","marker":"[5]"},{"why":"Provides the segmentation model that cleans occluders from RGB frames before pose estimation.","marker":"[10]"},{"why":"Provides the 3D human pose and velocity estimator used to produce supervision targets for the intent model.","marker":"[11]"},{"why":"Gives the deterministic DDIM sampling procedure used at inference to generate future velocity commands from the diffusion model.","marker":"[35]"},{"why":"Provides the RSL-RL PPO implementation that trains the locomotion policy.","marker":"[40]"},{"why":"Names the physical robot platform on which zero-shot sim-to-real deployment is evaluated.","marker":"[12]"}],"fun_headline_variants":["Haptic-only control lets humanoid carry loads with a human partner","Force signals alone drive a humanoid's cooperative carrying","Wrist forces teach humanoid to match human in load carrying","Humanoid uses force cues to carry extended loads with humans","No vision needed: humanoid follows human via wrist force"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The locomotion policy was trained with random per-wrist forces only up to 15 newtons (roughly 3 kilograms of payload total), while the real-world trials carried boxes weighing up to 5 kilograms, so the reported zero-shot success presupposes the policy generalizes to forces it never experienced in training.","fun_headline_variants_meta":{"raw":{"variants":["Haptic-only control lets humanoid carry loads with a human partner","Force signals alone drive a humanoid's cooperative carrying","Wrist forces teach humanoid to match human in load carrying","Humanoid uses force cues to carry extended loads with humans","No vision needed: humanoid follows human via wrist force"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1467,"prompt_tokens":1064,"completion_tokens":403,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":320}},"tokens_in":680,"tokens_out":403,"duration_ms":5625,"temperature":1.0,"reasoning_tokens":320,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:43:22.333262+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled trial in which a 5 kg box is held so that one wrist bears the entire load, pushing per-wrist vertical force above the 15 N training range, and the robot is expected to follow the leader; if it falls or fails to track, the claim of zero-shot load-adaptive co-manipulation is falsified.","supporting_citations":[{"cited_title":"Diffusion policy: Visuomotor policy learn- ing via action diffusion,","cited_arxiv_id":null,"evidence_quote":"Defines the diffusion policy architecture that the haptic intent inference model adapts for force-to-velocity mapping."},{"cited_title":"Human-robot planar co-manipulation of extended objects: data-driven models and control from human-human dyads,","cited_arxiv_id":null,"evidence_quote":"Supplies the human–human dyad data and the four evaluation metrics used as the baseline for comparison."},{"cited_title":"Wham: Reconstructing world-grounded humans with accurate 3d motion,","cited_arxiv_id":null,"evidence_quote":"Provides the 3D human pose and velocity estimator used to produce supervision targets for the intent model."},{"cited_title":"Learning to walk in minutes using massively parallel deep reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Provides the RSL-RL PPO implementation that trains the locomotion policy."},{"cited_title":"Unitree g1 humanoid robot","cited_arxiv_id":null,"evidence_quote":"Names the physical robot platform on which zero-shot sim-to-real deployment is evaluated."}],"review_version":1}