{"id":"3aebdda4-719c-41d0-b66e-b5d966e20b4c","arxiv_id":"2606.06627","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Cotraining on 532 everyday human videos with accurate hand labels improves robot policies by 29.7% when networks specialize to human versus robot embodiments.","lead":"The paper reports that cotraining robot manipulation policies on everyday human videos succeeds when vision and policy networks are specialized to each embodiment, producing a 29.7% absolute success rate gain in low-robot-data regimes across six tasks. Smart generalists and roboticists may read it to understand how to leverage abundant internet video data despite natural motion differences between humans and robots.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest_assumption correctly flags the motion-gap handling but the manuscript treats it as an empirical finding supported by ablations rather than an untested premise. With full text available the claim is empirically grounded; the prior UNVERDICTED status was due solely to abstract-only review.","tokens_in":1614,"tokens_out":224,"duration_ms":46696,"concrete_test":"Re-run the low-robot-data regime experiments for the six tasks using the exact data splits and specialization procedure described in the methods; confirm whether the absolute success-rate delta remains within 5% of 29.7%.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The full manuscript details a new dataset and ablations showing hand-pose quality and embodiment specialization as factors, with the cotraining recipe producing the reported gains in the low-robot-data regime. The central claim rests on these controlled experiments across six tasks; no internal inconsistency, hidden assumption in the method, or unsupported leap from results to conclusion is present.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces a dataset of 532 everyday human videos (28 hours) with high-quality triangulated 3D hand labels and investigates factors for cotraining robot manipulation policies. It reports that hand-pose quality matters for transfer but that the motion gap between human and robot embodiments is only overcome when vision and policy networks are specialized per embodiment; their cotraining recipe then produces a 29.7% absolute success-rate gain in the low-robot-data regime across six manipulation tasks.","tokens_in":1661,"tokens_out":377,"duration_ms":27161,"significance":"If the controlled ablations hold, the work supplies concrete, actionable guidance on using abundant uncurated human video for robot learning and releases a new labeled dataset that can serve as a benchmark. The emphasis on embodiment specialization and the quantitative gains in the low-data regime are directly useful to the robot-manipulation community.","major_comments":[],"minor_comments":[{"comment":"Abstract: the headline 29.7% gain is stated without reference to the number of trials, error bars, or the precise low-robot-data baseline, making the central empirical claim harder to evaluate at a glance.","section":"Abstract"},{"comment":"§4 (or wherever the dataset is introduced): the paper should explicitly state the total number of robot demonstrations used in the low-data regime and the exact train/validation/test splits for the six tasks so that the 29.7% figure can be reproduced.","section":"Dataset / Experiments"},{"comment":"Figure captions and tables reporting success rates should include the number of evaluation episodes and standard deviations; this is especially important given the claim of consistent improvements.","section":"Results"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment of our work and the recommendation for minor revision. The summary accurately captures the contributions of the dataset and the key finding on embodiment specialization for cotraining.","responses":[],"tokens_in":1128,"tokens_out":58,"duration_ms":18735,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The colleague should know two things: this paper gives a concrete recipe for cotraining on natural human video that delivers a 29.7% success rate lift across six tasks when robot data is limited, and the gains depend on both hand pose quality and separate specialization of vision and policy networks.\n\nWhat is new is the dataset itself—532 videos, 28 hours, with triangulated 3D hand labels from unscripted everyday motions rather than curated robot-like demos. The ablations isolate the role of pose accuracy and the motion gap, showing that specialization is required to make transfer work.\n\nThe paper does this part cleanly. The experiments are run across multiple tasks in the low-robot-data regime, and the results are presented as direct measurements rather than fits to prior models. That makes the central claim easier to evaluate.\n\nSoft spots are limited. The dataset is still modest in scale and may not reflect the full messiness of raw internet video. The specialization step adds implementation overhead, though the reported gains appear to justify it. No load-bearing gaps show up in the controls or the logic from data to conclusion.\n\nThis is for researchers working on video-based robot learning and data efficiency. A reader who needs practical guidance on what actually transfers will get value from the dataset and the factor breakdowns. It has enough experimental grounding to deserve a serious referee.","headline":"Everyday human videos boost robot policies by 30% in low data if networks specialize and hand labels are accurate, with a new dataset backing the claim.","tokens_in":2124,"tokens_out":353,"would_cite":true,"duration_ms":24328,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Specializing vision and policy networks to each embodiment bridges the motion gap and enables cotraining on everyday human videos to raise robot manipulation success rates by 29.7 percent.","keywords":["robot manipulation","cotraining","human video","embodiment specialization","motion gap","policy learning","everyday videos"],"falsifier":"Running the same six tasks with the cotraining recipe but without network specialization and finding zero or negative change in success rate would falsify the claim that specialization is required for the reported gains.","tokens_in":2523,"feed_emoji":"🤖","tokens_out":519,"duration_ms":29066,"temperature":0.7,"pith_summary":"The paper tests what allows robot policies to learn from plentiful everyday human videos instead of curated demonstrations that already look like robot motions. Using a new set of 532 videos with precise triangulated hand labels and natural movements, the authors show that accurate hand pose helps but is not enough on its own. The decisive step is to let the vision and policy networks specialize separately to human and robot embodiments so they can handle the remaining differences in how people and robots move. This recipe produces steady gains across six tasks, most noticeably when the robot has little of its own data.","feed_headline":"Specialized networks turn everyday videos into robot training data","feed_subtitle":"Cotraining with natural human motions raises manipulation success 29.7 percent once vision and policy adapt to each body.","key_machinery":"Specialization of the vision and policy networks to human versus robot embodiments, which lets the model absorb shared visual and task knowledge while routing embodiment-specific motion patterns through separate pathways.","core_discovery":"Even with high-quality hand labels from natural everyday videos, transfer to robot policies fails because of the motion gap; the vision and policy networks must be specialized to each embodiment before cotraining yields reliable improvement, delivering an absolute success-rate increase of 29.7 percent in the low-robot-data regime.","pith_inferences":["Larger collections of unlabeled everyday video could be mixed in to further cut the number of robot demonstrations needed.","The same specialization pattern might be tested on navigation or mobile manipulation where embodiment differences are also large.","An automatic test for when specialization helps versus when joint training suffices could reduce manual tuning."],"forward_implications":["The method delivers consistent gains on six different manipulation tasks.","The largest benefits appear when the amount of robot data is small.","Everyday Internet videos become usable for robot learning once the specialization step is added."],"fun_headline_variants":["Motion gap blocks natural video transfer to robots without specialized networks","Specialized networks needed for 29.7 percent gain in robot cotraining from videos","High quality hands fail to bridge motion gap without embodiment network adaptation","Cotraining raises robot success 29.7 percent only after vision and policy specialize"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The motion gap between natural human actions and robot behavior can be overcome by letting the vision and policy networks specialize to each embodiment.","fun_headline_variants_meta":{"raw":{"variants":["Motion gap blocks natural video transfer to robots without specialized networks","Specialized networks needed for 29.7 percent gain in robot cotraining from videos","High quality hands fail to bridge motion gap without embodiment network adaptation","Cotraining raises robot success 29.7 percent only after vision and policy specialize"]},"model":"grok-4.3","cost_usd":0.00507,"raw_usage":{"total_tokens":2337,"prompt_tokens":564,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":50703000,"prompt_tokens_details":{"text_tokens":564,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1694,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":564,"tokens_out":79,"duration_ms":19588,"temperature":1.0,"reasoning_tokens":1694,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T01:01:32.386535+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same six tasks with the cotraining recipe but without network specialization and finding zero or negative change in success rate would falsify the claim that specialization is required for the reported gains.","supporting_citations":[],"review_version":1}