{"id":"cdb7e563-3ed8-4546-9053-4804c3b5f9a3","arxiv_id":"2606.04463","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"OSCAR finetunes Cosmos-Predict2.5-2B on a deduplicated multi-embodiment robotics dataset with kinematic skeleton conditioning, claiming better action following and significant correlation between virtual and real robot policy evaluations.","lead":"OSCAR is a video world model finetuned on curated robot and human data using 2D skeleton conditioning to generate action-conditioned videos that generalize across robot bodies. It is positioned to let developers test robot control programs in simulation with reported correlation to real-world results.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"2D skeleton rendering may lose 3D action precision needed for cross-embodiment correlation","rationale":"The reader's weakest assumption directly identifies the load-bearing technical risk for the correlation claim. No other internal inconsistency is visible from the supplied abstract and description; the concern is therefore the same one already flagged.","tokens_in":1734,"tokens_out":271,"duration_ms":18721,"concrete_test":"On a held-out multi-embodiment test set with 3D ground-truth poses, recompute action-following metrics (e.g., joint-angle MSE or end-effector trajectory error) using the paper's 2D skeleton renderer versus an equivalent 3D pose renderer; if 2D error exceeds 3D error by >15% on any embodiment, the conditioning assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (significant virtual-real policy evaluation correlation) requires that 2D kinematic skeleton conditioning preserves precise action information across robot arms and human hands. Projecting 3D joint configurations to 2D images can discard depth and out-of-plane rotation cues; if this occurs, action following degrades for embodiments whose kinematics differ in 3D structure, breaking the claimed generalization and making virtual evaluations unreliable proxies for real-world performance.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents OSCAR, an action-conditioned video world model for robotics that uses 2D kinematic skeleton rendering as a unified conditioning signal across robot arms and human hands. It describes a large-scale data curation pipeline combining robotics and egocentric human datasets, finetuning of the Cosmos-Predict2.5-2B model on a single GH200 GPU, claimed improvements in action following/appearance/motion consistency over baselines, and a significant correlation between virtual policy evaluations (on RoboArena policies) and real-world results.","tokens_in":1819,"tokens_out":449,"duration_ms":16603,"significance":"If the reported virtual-real correlation is supported by quantitative evidence, the work could enable lower-cost policy evaluation by shifting testing into generated worlds. The standardized data pipeline and embodiment-agnostic conditioning approach address practical barriers in current video world models, though the lack of reported metrics prevents gauging the magnitude of these advances.","major_comments":[{"comment":"Abstract: the claims of 'significant improvement on action following' and 'significant correlation between our virtual policy evaluation in OSCAR and real-world evaluation' are presented without any quantitative metrics, baseline comparisons, dataset sizes, statistical details, or evaluation protocol; this absence makes the central data-to-claim link unevaluable from the manuscript.","section":"Abstract"},{"comment":"Method (conditioning representation): the assertion that 2D kinematic skeleton rendering 'generalizes across different robot arms or even human hands' and preserves 'precise action information' is load-bearing for both the cross-embodiment claim and the virtual-real correlation; the manuscript does not address or test whether projection from 3D joint configurations to 2D discards depth/out-of-plane cues that would degrade fidelity for kinematically dissimilar embodiments.","section":"Method"}],"minor_comments":[{"comment":"Abstract: model size (2B) and training hardware (single GH200) are stated, but no corresponding numbers are given for the baselines against which improvements are claimed.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed feedback. We address the two major comments point-by-point below, clarifying where quantitative details appear in the manuscript and acknowledging where additional discussion is warranted.","responses":[{"response":"The abstract is intentionally concise and high-level. Quantitative results—including specific action-following metrics versus baselines, the reported correlation coefficient between virtual and real policy evaluations, dataset sizes after deduplication, and the full evaluation protocol—are provided in the Experiments and Results sections. To make the abstract self-contained, we will revise it to include the key numerical highlights (e.g., correlation value and relative improvements) while preserving brevity.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the claims of 'significant improvement on action following' and 'significant correlation between our virtual policy evaluation in OSCAR and real-world evaluation' are presented without any quantitative metrics, baseline comparisons, dataset sizes, statistical details, or evaluation protocol; this absence makes the central data-to-claim link unevaluable from the manuscript."},{"response":"The 2D skeleton rendering was selected precisely because it yields a compact, embodiment-agnostic signal that can be generated from any 3D joint set. The manuscript demonstrates cross-embodiment generalization through training and evaluation on multiple robot arms plus human-hand data. However, we agree that an explicit analysis of information loss from 3D-to-2D projection (depth/out-of-plane cues) is not present. We will add a dedicated paragraph discussing this potential limitation and its implications for kinematically dissimilar embodiments.","revision_made":"partial","referee_comment":"[Method] Method (conditioning representation): the assertion that 2D kinematic skeleton rendering 'generalizes across different robot arms or even human hands' and preserves 'precise action information' is load-bearing for both the cross-embodiment claim and the virtual-real correlation; the manuscript does not address or test whether projection from 3D joint configurations to 2D discards depth/out-of-plane cues that would degrade fidelity for kinematically dissimilar embodiments."}],"tokens_in":1411,"tokens_out":453,"duration_ms":24500,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main contribution is a data pipeline that pulls together robotics and egocentric human videos, deduplicates them, and trains a 2B video model on the result using 2D kinematic skeletons as the action-conditioning signal. They then run RoboArena policies through the model to generate rollouts and report that those virtual scores line up with real-world scores.\n\nWhat is actually new is the combination of that curation step with direct use for policy ranking on an existing benchmark, plus the choice to start from Cosmos-Predict2.5 rather than training from scratch. The single-GPU fine-tune is also a practical detail that lowers the barrier compared with larger video-model efforts.\n\nThe soft spot is the complete absence of numbers. The abstract says the model improves action following and motion consistency over baselines and that the virtual-real correlation is significant, yet it gives no PSNR, FVD, correlation coefficient, dataset sizes, or statistical test. Without those, it is impossible to tell whether the skeleton conditioning actually preserves enough 3D action information across arms and hands or whether the correlation is strong enough to matter for real deployment.\n\nThe stress-test worry about 2D projection losing depth and rotation cues is therefore still open; the paper would need to show that the correlation survives on embodiments whose kinematics differ in 3D structure.\n\nThis is for people working on scalable robot evaluation who already follow video world models. A reader who wants concrete evidence that virtual testing can replace physical runs will find the current write-up preliminary.\n\nIt deserves peer review because the problem is concrete and the pipeline is reproducible in principle, but any referee will ask for the missing quantitative results and ablations on the conditioning choice.","headline":"OSCAR fine-tunes Cosmos on deduplicated multi-embodiment data with 2D skeleton conditioning and claims virtual-real policy correlation, but the abstract supplies no metrics to judge either the gains or the correlation strength.","tokens_in":2302,"tokens_out":437,"would_cite":false,"duration_ms":19214,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"OSCAR conditions a video world model on 2D kinematic skeletons to evaluate robot policies virtually with strong correlation to real-world results across embodiments.","keywords":["robotics","world model","action-conditioned video","policy evaluation","embodiment generalization","kinematic skeleton","video generation"],"falsifier":"A collection of robot policies whose performance rankings or scores in OSCAR-generated videos differ substantially from the rankings obtained when the same policies are run on physical robots.","tokens_in":2617,"feed_emoji":"🤖","tokens_out":428,"duration_ms":19378,"temperature":0.7,"pith_summary":"The paper introduces OSCAR as an action-conditioned video world model trained on a large standardized dataset combining robotics and egocentric human videos. It uses 2D kinematic skeleton rendering as the conditioning input to achieve better action following and generalization than prior models. The central goal is to enable robot policy evaluation inside generated videos rather than on physical hardware. The authors report that virtual evaluations on RoboArena policies align closely with real-world measurements.","feed_headline":"Video world model correlates virtual and real robot policy tests","feed_subtitle":"By conditioning on 2D skeletons from mixed robot and human data, OSCAR produces virtual evaluations that match physical results.","key_machinery":"2D kinematic skeleton rendering as a unified conditioning representation that carries action information across robot arms and human hands","core_discovery":"OSCAR is a precise action-conditioned video world model finetuned from Cosmos-Predict2.5-2B on a single GH200 GPU using a cleaned joint dataset from diverse robot and human sources. By adopting 2D kinematic skeleton rendering as a unified conditioning representation, the model improves action following, appearance quality, and motion consistency over baselines and produces virtual policy evaluations that show significant correlation with real-world robot performance.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["OSCAR correlates virtual real robot policy tests","2D skeleton OSCAR links virtual and physical robot evaluations","OSCAR from joint data generalizes robot policy evaluation virtually","Cosmos finetuned OSCAR enables virtual robot policy testing correlation"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 2D kinematic skeleton rendering supplies enough precise action detail to work equally well for every robot arm and human hand without embodiment-specific loss of fidelity.","fun_headline_variants_meta":{"raw":{"variants":["OSCAR correlates virtual real robot policy tests","2D skeleton OSCAR links virtual and physical robot evaluations","OSCAR from joint data generalizes robot policy evaluation virtually","Cosmos finetuned OSCAR enables virtual robot policy testing correlation"]},"model":"grok-4.3","cost_usd":0.011381,"raw_usage":{"total_tokens":5004,"prompt_tokens":688,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":113812000,"prompt_tokens_details":{"text_tokens":688,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4252,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":688,"tokens_out":64,"duration_ms":42775,"temperature":1.0,"reasoning_tokens":4252,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T06:32:50.911487+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A collection of robot policies whose performance rankings or scores in OSCAR-generated videos differ substantially from the rankings obtained when the same policies are run on physical robots.","supporting_citations":[],"review_version":1}