{"id":"91ba7c93-01b4-445a-adcc-314252033229","arxiv_id":"2606.27239","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Robot-free VR–UMI demos of sparse whole-body keypoints can be retargeted and executed as deployable Unitree G1 whole-body skills more efficiently than teleoperation.","lead":"HumanoidUMI collects human whole-body demos with portable VR and handheld grippers, then trains a keypoint policy that is retargeted and executed on a humanoid. It offers a cheaper path to humanoid skill data than robot teleoperation, which is a major bottleneck for the field.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Missing teleop-trained policy baseline leaves the transfer claim under-supported relative to the efficiency claim.","rationale":"The reader correctly flags sparse keypoints + fixed λ_leg SKR as a soft morphological assumption and correctly withholds ACCEPT for missing baselines, error bars, and artifacts. I agree the verdict should stay CONDITIONAL. The sharper load-bearing concern for the strongest claim, however, is not only whether SKR under-constrains posture, but whether the paper ever shows that robot-free data yield policies competitive with teleop-trained ones. Throughput (Table I) is cleanly measured; policy quality is only self-compared via ablations. That is the single check that would most tighten or loosen the central claim. Physical success rates still cannot be re-run from the PDF, so confidence remains moderate; no change of verdict category is warranted.","tokens_in":11419,"tokens_out":553,"duration_ms":5751,"concrete_test":"Train the same high-level diffusion policy architecture on an equal number of TWIST2 teleop trajectories for the same five tasks (or at least bimanual vegetable collection and walking coffee delivery), deploy with the same low-level controller, and report success rates over the same 20-trial protocol. If teleop-trained success is within ~10 points of HumanoidUMI, the transfer claim holds; if teleop is substantially higher, the robot-free pipeline’s skill quality remains unproven despite the throughput win.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that robot-free VR–UMI demos can be transformed into deployable G1 whole-body skills that are useful for skill learning, not merely that demos can be collected faster. §IV.A–B report success rates only for HumanoidUMI policies (and SKR / latency / 5-vs-7 keypoint ablations). §IV.C and Table I compare only valid-demo throughput vs TWIST2, not the quality of policies trained on matched teleop data. Without that baseline, high success on five tasks is consistent with either (i) the sparse keypoint + SKR interface truly closing the morphology gap, or (ii) the tasks being easy enough that any reasonable whole-body controller succeeds once a high-level policy is trained. The reader’s weakest assumption (sufficiency of 5/7 keypoints + fixed λ_leg=0.75) is real, but the more load-bearing gap for the paper’s strongest claim is the missing apples-to-apples policy comparison: efficiency is shown; transferable skill quality relative to embodiment-matched teleop data is not.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"HumanoidUMI presents a robot-free pipeline for collecting and deploying humanoid whole-body manipulation demonstrations. Using a PICO VR full-body capture setup and UMI-inspired dual handheld grippers, the system records wrist-view fisheye images, gripper widths, and a sparse five-keypoint (or seven-keypoint) trajectory of pelvis, TCPs, feet (and optionally knees). A diffusion policy predicts future relative keypoint motions and gripper actions conditioned on wrist views and lower-body proprioception; Spatial Keypoint Retargeting (SKR) applies a fixed anisotropic leg scale λ_leg=0.75 and weighted two-stage IK to produce robot-native root/joint references; a learned whole-body controller tracks those references on a Unitree G1. Real-robot experiments cover five tasks (cluttered pick-and-place, bimanual vegetable collection, dynamic ball throwing, under-table waste disposal, walking coffee delivery), with ablations of SKR vs GMR, latency matching, and 5- vs 7-keypoint layouts, plus a throughput comparison against TWIST2 teleoperation.","tokens_in":11782,"tokens_out":1266,"duration_ms":11477,"significance":"If the transfer claim holds, the work offers a practical, low-cost route to scalable humanoid whole-body data without robot-in-the-loop teleoperation, extending the UMI paradigm to coordinated loco-manipulation. Strengths include a clear hierarchical separation of task-space prediction, morphology-aware retargeting, and low-level tracking; real G1 deployment across five qualitatively different tasks; and concrete ablations (SKR, latency matching, knee keypoints) plus a quantified collection-efficiency gain over TWIST2. The sparse keypoint interface and online SKR visualization during collection are useful engineering contributions for the community.","major_comments":[{"comment":"The central claim is that robot-free demos yield transferable, deployable whole-body skills, not only that demos can be collected faster. §IV.A–B report success rates only for HumanoidUMI-trained policies (and internal ablations of SKR, latency matching, and 5- vs 7-keypoint layouts). §IV.C / Table I compare only valid-demo throughput vs TWIST2, not policy quality under matched data budgets or identical evaluation protocols. Without a teleop-trained (or embodiment-matched) policy baseline on the same five tasks, high success is consistent with either successful morphology-gap closure or with tasks that a competent whole-body controller can solve once any reasonable high-level policy is available. An apples-to-apples policy comparison is load-bearing for the transfer claim.","section":null},{"comment":"§III.C, Eqs. (4)–(5): SKR relies on a single free scale λ_leg=0.75 and fixed two-stage weighted IK. The paper does not report sensitivity of success rates to λ_leg, to the IK weights w_i / λ_q, or to alternative morphology-compensation schemes. Given that the weakest technical assumption is that this sparse anisotropic adjustment plus IK sufficiently closes the human–G1 gap for loco-manipulation, a short sensitivity study (or failure-mode analysis when λ_leg is misspecified) is needed to support that the interface is robust rather than tuned to the present demonstrator and robot.","section":null},{"comment":"Success rates in Figs. 6–7 are given as point estimates over 20 trials with no error bars, confidence intervals, or statistical comparison. For the ablations that are used to justify SKR and latency matching, binomial CIs or a simple significance test would make the claimed improvements interpretable and would strengthen the empirical support for the hierarchical design.","section":null}],"minor_comments":[{"comment":"Fig. 6 and Fig. 7 captions refer to “left-side bar plots” of success rates; ensure the published figures include numerical labels or a table of exact rates so readers can recover the values without visual estimation.","section":null},{"comment":"§III.B: the action dimension is stated as 5×9+2=47 (and 65 for seven keypoints). A brief note clarifying that the 6-D rotation is the continuous representation of Zhou et al. and how it is recovered at inference would help reproducibility.","section":null},{"comment":"Latency-matching procedure is only referenced to UMI; a short description of the measured delays (camera, gripper encoder, policy, control) for the G1 setup would make the dynamic-throwing result more self-contained.","section":null},{"comment":"Related work correctly positions HuMI and HoMMI; a one-sentence clarification of how online SKR during collection differs from HuMI’s retargeting coupling would sharpen the novelty claim.","section":null},{"comment":"Minor notation: T_rel and T_pel in Eqs. (1)–(2) would benefit from an explicit statement that all poses are SE(3) and that gripper widths remain absolute scalars after min–max normalization of translations.","section":null}],"recommendation":"major_revision","confidential_remarks":"The efficiency result vs TWIST2 is convincing and well-scoped; the missing teleop-trained policy baseline is the main reason I recommend major rather than minor revision. If the authors can add even a limited matched-budget comparison on two of the five tasks (or a clear argument why teleop data of equal quality cannot be collected under the same protocol), the transfer claim would be substantially stronger. Scope is appropriate for a robotics systems venue."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean systems paper that actually ships a usable robot-free collection stack for humanoid whole-body skills and shows it working on a Unitree G1. The new pieces are the portable PICO + UMI-gripper hardware, the sparse 5/7-keypoint interface, and the explicit Spatial Keypoint Retargeting bridge (anisotropic leg scale + two-stage weighted IK) sitting between a diffusion keypoint policy and a learned whole-body controller. That hierarchy is clearer than HuMI’s more tightly coupled retargeting, and the five real tasks (cluttered pick-place, bimanual, timed throw, under-table bend, walking coffee delivery) plus the SKR / latency / knee-keypoint ablations are enough to show the pipeline can deploy.\n\nWhat it does well: efficiency is real. Table I shows ~2.2× valid-demo throughput vs TWIST2 on non-loco tasks and a much larger gap on walking coffee delivery, including for a novice after 10 minutes of instruction. That is the practical contribution labs will care about. The methods section is concrete (action space, relative pelvis-frame targets, λ_leg=0.75, residual PD tracking), and the citation pattern is honest about UMI, HuMI, TWIST2, and GMR.\n\nSoft spots, in proportion. The stress-test note is right: the strongest claim is transferable skill learning, not just faster demos. Success rates are only for HumanoidUMI policies; there is no matched teleop-trained policy baseline on the same tasks. So we cannot yet tell whether sparse keypoints + fixed SKR close the morphology gap or whether these tasks are simply forgiving once a whole-body controller is in the loop. Minor but real: no error bars/CIs on the 20-trial bars, hand-tuned λ_leg, no public code/data. Those are fixable; they do not sink the work.\n\nWho it is for: anyone collecting humanoid loco-manipulation data or building hierarchical visuomotor stacks. Worth a serious referee. I would engage, cite the collection interface and SKR design, and push for the missing policy-quality comparison in revision.","headline":"Solid portable robot-free humanoid demo pipeline with real G1 results and a clear efficiency win; the transfer claim is under-supported without a teleop-trained policy baseline.","tokens_in":12431,"tokens_out":538,"would_cite":true,"duration_ms":5574,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Portable VR and handheld grippers let people teach humanoids whole-body skills without ever touching the robot.","keywords":["humanoid robot","robot-free data collection","whole-body manipulation","Universal Manipulation Interface","spatial keypoint retargeting","diffusion policy","visuomotor learning","loco-manipulation"],"falsifier":"On the same Unitree G1 and the same five tasks, replace SKR with a naïve global rescaling or pure joint-space retargeting of the identical human demos and measure whether success rates collapse relative to the reported SKR results, especially on under-table waste disposal and walking coffee delivery.","tokens_in":12304,"feed_emoji":"🤖","tokens_out":965,"duration_ms":8716,"temperature":0.7,"pith_summary":"Humanoid robots need coordinated whole-body demos—grasping while bending, walking while handing something over—but collecting them by teleoperating the robot is slow, requires access to hardware, and demands skilled operators. HumanoidUMI instead records natural human motion with a lightweight VR headset, a few body trackers, and UMI-style handheld grippers that also capture wrist-view video and gripper width. The system stores only sparse task-relevant keypoints (pelvis, two tool-center points, two feet, optionally knees), trains a diffusion policy to predict future keypoint motion from the wrist views and lower-body proprioception, then uses Spatial Keypoint Retargeting to turn those predictions into robot-native root and joint references that a learned whole-body controller can track. On a Unitree G1 the pipeline produces working policies for cluttered pick-and-place, bimanual collection, timed ball throwing, under-table waste disposal, and walking coffee delivery, while gathering valid demonstrations roughly twice as fast as a strong teleoperation baseline—and far faster for locomotion-plus-manipulation.","feed_headline":"Teach humanoids whole-body skills without the robot present","feed_subtitle":"VR and handheld grippers yield deployable G1 policies twice as fast as teleoperation","key_machinery":"Spatial Keypoint Retargeting (SKR): a morphology-aware bridge that keeps metric spatial relationships among five (or seven) task-space keypoints, applies only a fixed anisotropic vertical scale (λ_leg = 0.75) to leg-related points to match robot height, then solves a two-stage weighted inverse-kinematics problem to produce executable root pose and joint references.","core_discovery":"Robot-free human demonstrations collected with a portable VR–UMI interface can be turned into deployable humanoid whole-body skills: a high-level diffusion policy predicts sparse keypoint trajectories, Spatial Keypoint Retargeting converts them into feasible robot references, and a learned whole-body controller executes them with balance on a physical Unitree G1 across five real tasks spanning single-arm, bimanual, dynamic, bending, and loco-manipulation behaviors.","pith_inferences":["If the sparse keypoint set proves sufficient, large-scale in-the-wild humanoid datasets could be gathered by ordinary people with consumer VR kits rather than robot labs.","The same hierarchy may transfer to other morphologically mismatched platforms (e.g., different bipedal robots) by changing only the SKR scale and IK weights.","Wrist-view-only sensing may limit long-horizon spatial memory; adding a sparse third-person or head-mounted stream is a natural next stress test."],"forward_implications":["Valid whole-body training data can be collected without the target humanoid present, lowering hardware and safety barriers.","Novice operators can produce usable loco-manipulation demos at rates close to experienced users, reducing skill dependence.","A single sparse keypoint interface plus a general whole-body controller can cover both quasi-static manipulation and timing-sensitive dynamic release.","Hierarchical separation of task-space prediction, retargeting, and low-level tracking makes the same human demos reusable across different humanoid morphologies once SKR is re-calibrated."],"fun_headline_variants":["Robot-free VR demos yield humanoid whole-body skills","Sparse keypoints from handheld grippers train full G1 control","Portable VR-UMI interface collects transferable humanoid demos","Human demos retarget to balanced Unitree G1 whole-body skills","High-level policy turns robot-free trajectories into G1 actions"],"cache_read_input_tokens":128,"weakest_assumption_plain":"A handful of body keypoints plus wrist cameras, after one fixed leg-length scale and weighted inverse kinematics, still carries enough geometric information to close the human-to-robot body gap for coordinated whole-body skills.","fun_headline_variants_meta":{"raw":{"variants":["Robot-free VR demos yield humanoid whole-body skills","Sparse keypoints from handheld grippers train full G1 control","Portable VR-UMI interface collects transferable humanoid demos","Human demos retarget to balanced Unitree G1 whole-body skills","High-level policy turns robot-free trajectories into G1 actions"]},"model":"grok-4.5","effort":"low","cost_usd":0.006164,"raw_usage":{"total_tokens":1578,"prompt_tokens":730,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":61640000,"prompt_tokens_details":{"text_tokens":730,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":769,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":730,"tokens_out":79,"duration_ms":6880,"temperature":1.0,"reasoning_tokens":769,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T11:46:29.954356+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same Unitree G1 and the same five tasks, replace SKR with a naïve global rescaling or pure joint-space retargeting of the identical human demos and measure whether success rates collapse relative to the reported SKR results, especially on under-table waste disposal and walking coffee delivery.","supporting_citations":[],"review_version":2}