{"id":"cdc37169-fd0f-4e40-bf22-6deec4d0194c","arxiv_id":"2606.10614","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Dexterous Point Policy learns dexterous hand policies from human videos using 3D keypoints of hands and objects, achieving 75% success on real-robot tasks compared to 1% for a VLA baseline.","lead":"The paper presents Dexterous Point Policy, a method that trains dexterous robot hand policies directly from human demonstration videos by representing both observations and actions as 3D keypoints. This approach aims to eliminate the need for expensive robot-specific data collection in complex manipulation tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Keypoint alignment between human and robot is asserted without quantitative validation of trajectory distributions or kinematic feasibility.","rationale":"The reader's weakest_assumption directly identifies the same load-bearing premise. Because the review was performed on the abstract, the full manuscript does not appear to contain an explicit quantitative check of keypoint distribution shift or kinematic mismatch; confirming or refuting that single assumption would move the verdict from UNVERDICTED to either ACCEPT or REJECT.","tokens_in":1775,"tokens_out":366,"duration_ms":15931,"concrete_test":"Extract the 3D keypoint time-series from the human training set and from 20 robot rollouts under the deployed policy; compute the average L2 distance between corresponding fingertip velocity histograms and the fraction of predicted actions that violate the robot's joint limits after IK. If either metric exceeds the values observed in successful human demonstrations by >30%, the alignment premise fails to support the reported 75% success rate.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that an autoregressive transformer trained solely on 3D keypoints (wrist + fingertips + objects) extracted from human videos can be deployed zero-shot on a physical robot hand. This rests on the unverified statement that \"at the keypoint level, specifically the wrist and fingertips, human and robot behaviors closely align.\" Different hand morphologies imply that human keypoint velocities, accelerations, and reachable workspaces will differ from those feasible under robot joint limits and link lengths; the policy could therefore output actions that are either unreachable or dynamically mismatched once mapped via IK. No section in the supplied text provides a direct statistical comparison (e.g., KL divergence or success-rate ablation) between human keypoint sequences and the keypoint sequences realized during robot execution.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Dexterous Point Policy, a framework that extracts 3D keypoints (wrist, fingertips, and task-relevant objects) from human demonstration videos, trains an autoregressive transformer policy over these keypoints for both observations and actions, and deploys the resulting policy zero-shot on a physical robot hand without any robot demonstrations. It claims 75% success on real-robot pick-and-place and tool-use tasks versus 1% for a state-of-the-art VLA baseline, plus strong generalization to multi-object scenes and novel object categories, based on the observation that human and robot behaviors align closely at the keypoint level.","tokens_in":1926,"tokens_out":391,"duration_ms":14340,"significance":"If the empirical claims hold after proper validation, the work would be significant for dexterous manipulation: it offers a concrete route to bypass expensive robot data collection by leveraging a unified keypoint representation to close the embodiment gap between human videos and robot hardware.","major_comments":[{"comment":"Abstract: the central zero-shot transfer claim rests on the statement that 'at the keypoint level, specifically the wrist and fingertips, human and robot behaviors closely align,' yet the supplied text provides no quantitative validation (e.g., KL divergence between human and realized robot keypoint trajectories, workspace overlap statistics, or an ablation measuring success-rate drop when the alignment assumption is violated).","section":"Abstract"},{"comment":"Abstract: the headline result (75.0% vs. 1.0% success) is presented without any description of trial count, statistical significance testing, task definitions, baseline implementation details, or potential confounds such as object pose variation or lighting, rendering the magnitude of the improvement impossible to assess from the given material.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the abstract. We agree that additional details on validation and experimental reporting are warranted and will revise the abstract accordingly while preserving its conciseness.","responses":[{"response":"We acknowledge that the abstract presents the alignment as an observation without accompanying quantitative metrics. The full manuscript supports this via successful zero-shot transfer results and qualitative trajectory visualizations in the experiments. To strengthen the presentation, we will incorporate a brief quantitative validation (e.g., mean keypoint trajectory distance or an alignment ablation) into the revised abstract and reference the corresponding analysis in the main text.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central zero-shot transfer claim rests on the statement that 'at the keypoint level, specifically the wrist and fingertips, human and robot behaviors closely align,' yet the supplied text provides no quantitative validation (e.g., KL divergence between human and realized robot keypoint trajectories, workspace overlap statistics, or an ablation measuring success-rate drop when the alignment assumption is violated)."},{"response":"The abstract is intentionally concise, but we agree that the headline numbers require supporting context for proper evaluation. The full manuscript details the evaluation protocol, including trial counts, task definitions, baseline setups, and controls for confounds. We will revise the abstract to include a short clause summarizing the evaluation scale (e.g., number of trials and tasks) and note that full details appear in Section 4.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the headline result (75.0% vs. 1.0% success) is presented without any description of trial count, statistical significance testing, task definitions, baseline implementation details, or potential confounds such as object pose variation or lighting, rendering the magnitude of the improvement impossible to assess from the given material."}],"tokens_in":1416,"tokens_out":406,"duration_ms":14397,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this work trains an autoregressive transformer on 3D keypoints extracted from human hand videos and deploys the resulting policy directly on a robot hand, reporting 75% success on pick-and-place and tool-use tasks against 1% for a VLA baseline, plus some generalization to multi-object and novel-category cases.\n\nWhat stands out is the decision to use the same keypoint representation for both observations and actions, which lets them avoid any robot demonstration data. That framing is distinct from typical fine-tuning pipelines and directly targets the data bottleneck in dexterous manipulation.\n\nThe soft spot is exactly the one flagged in the stress-test note. The abstract asserts that wrist and fingertip keypoints align closely enough between human and robot to enable direct transfer, yet the supplied text contains no quantitative check—no KL divergence on trajectories, no ablation on reachable workspace, no comparison of velocity distributions. Different hand morphologies make it plausible that many predicted keypoints fall outside the robot’s joint limits or produce mismatched dynamics once inverse kinematics is applied. Without that evidence, the performance gap is difficult to attribute to the method rather than implementation details or task selection.\n\nThe experimental protocol is also described at too high a level to assess: no trial counts, variance numbers, or baseline implementation specifics appear. This is the kind of paper that would benefit from a referee who can examine the full methods and results sections.\n\nIt is aimed at researchers working on video imitation for complex manipulation. The idea is practical enough that it deserves peer review to test whether the transfer claim survives scrutiny.","headline":"The paper claims zero-shot transfer of dexterous policies from human videos via 3D keypoints with 75% success, but provides no validation that the keypoint alignment actually holds across embodiments.","tokens_in":2437,"tokens_out":401,"would_cite":false,"duration_ms":16827,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A unified 3D keypoint representation enables training dexterous hand policies from human videos without robot demonstrations.","keywords":["dexterous manipulation","human video demonstrations","3D keypoints","policy learning","embodiment gap","robot hands","autoregressive transformer","real robot evaluation"],"falsifier":"Finding a dexterous task where the policy fails because human and robot fingertip paths diverge significantly even when both are described by the same keypoint extraction process.","tokens_in":2679,"feed_emoji":"🖐️","tokens_out":660,"duration_ms":33253,"temperature":0.7,"pith_summary":"The paper presents Dexterous Point Policy, which trains policies for multi-fingered robot hands by processing 3D keypoints from human demonstration videos. These keypoints serve as both the input observations and the output actions in an autoregressive transformer model. The central idea is that wrist and fingertip positions align sufficiently between humans and robots to allow direct policy transfer. This results in 75 percent success on real-world dexterous tasks such as pick-and-place and tool use, compared to just 1 percent for a leading vision-language-action model. The approach also generalizes to multi-object scenes and novel object categories without additional training data.","feed_headline":"Keypoint policy learns dexterous robot hands from human videos","feed_subtitle":"Dexterous Point Policy reaches 75% success on real pick-and-place and tool tasks by aligning wrist and fingertip points, bypassing robot dat","key_machinery":"Unified 3D keypoint representation used for both observations and actions in an autoregressive transformer trained on human videos.","core_discovery":"By extracting 3D keypoints of objects and hands from raw human videos and training an autoregressive transformer to predict future keypoints, the method creates policies that transfer to robot hands. Human and robot behaviors align at the keypoint level for the wrist and fingertips, so no robot demonstrations are needed. On real-robot evaluations the policy reaches 75.0 percent success across pick-and-place and tool-use tasks while a state-of-the-art VLA baseline achieves only 1.0 percent, and it maintains performance in unseen multi-object and novel-category settings.","pith_inferences":["Similar keypoint bridging could apply to other manipulation tasks or different robot morphologies if alignment holds.","Reducing data collection costs might enable faster iteration on complex dexterous behaviors.","Low-dimensional keypoints may suffice for many manipulation skills, suggesting further compression of visual inputs is viable."],"forward_implications":["Direct policy learning becomes possible for dexterous tasks without collecting robot data.","The policy succeeds on both pick-and-place and tool-use tasks at 75 percent.","Generalization occurs to multi-object environments and novel object categories.","Keypoint alignment allows bypassing the embodiment gap that usually requires fine-tuning."],"fun_headline_variants":["Keypoints from human videos train dexterous robot policies","3D points bridge human hands to robot manipulation","Keypoint transformer learns hand policies from video","Wrist and fingertip points enable direct policy transfer"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Wrist and fingertip keypoint trajectories are similar enough between human demonstrations and robot executions for the learned policy to work without adjustment.","fun_headline_variants_meta":{"raw":{"variants":["Keypoints from human videos train dexterous robot policies","3D points bridge human hands to robot manipulation","Keypoint transformer learns hand policies from video","Wrist and fingertip points enable direct policy transfer"]},"model":"grok-4.3","cost_usd":0.006512,"raw_usage":{"total_tokens":3080,"prompt_tokens":735,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":65124500,"prompt_tokens_details":{"text_tokens":735,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2286,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":735,"tokens_out":59,"duration_ms":17733,"temperature":1.0,"reasoning_tokens":2286,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T13:29:03.218903+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding a dexterous task where the policy fails because human and robot fingertip paths diverge significantly even when both are described by the same keypoint extraction process.","supporting_citations":[],"review_version":1}