{"id":"b39fb5b5-138d-49ba-9440-103fc7e05cc3","arxiv_id":"2608.11204","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Video-only pretraining of a world-action model improves closed-loop dVRK manipulation on SurRoL from 63.5% to 77.8% average success under a fixed action-labeled budget.","lead":"This paper tests whether pretraining a video-world model on action-free surgical footage improves a robot's closed-loop manipulation when action-labeled demonstrations are scarce. On four simulated surgical tasks, the pretrained model raises average success from 63.5% to 77.8%, with the largest gains on contact-rich and bimanual tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal claim rests on an unspecified Stage-1 video corpus: if Cosmos-H-Surgical was trained on SurRoL-like video, the gain is in-domain exposure, not generic action-free video pretraining.","rationale":"The most load-bearing concern is the unverified composition of the Stage-1 video corpus, because the central causal claim ('the gain comes entirely from the video pretraining') requires that the pretraining corpus not be drawn from the same task distribution as the SurRoL evaluation environments. The paper never specifies the corpus, cites He et al. (2026) without describing the data, and performs no leakage or overlap analysis. If the corpus contained SurRoL rollouts, the w/ PT variant would have had additional task-specific visual exposure absent from w/o PT, so the comparison would not isolate the effect of generic action-free video pretraining; it would show that pretraining on the evaluation task's own video helps, a much weaker and less practically significant claim. The JIGSAWS real-video experiment in Section 4.4 is qualitative only and therefore does not quantify cross-domain transfer. The Table 2 non-monotonicity (86% to 34% to 56.5%) is an additional robustness concern, but the corpus issue is the more fundamental threat to the paper's causal narrative. The reader's weakest_assumption identifies exactly this issue; multiple seeds and error bars would address statistical reliability but would not resolve the corpus confound. Hence the verdict remains CONDITIONAL, pending disclosure and a disjoint-corpus rerun. This concern does not change the reader's verdict, so the recommended verdict is UNCHANGED.","tokens_in":12798,"tokens_out":12768,"duration_ms":110854,"concrete_test":"Obtain the training-data specification for Cosmos-H-Surgical (He et al. 2026) and determine whether any SurRoL-generated or SurRoL-like simulated surgical video is present in the Stage-1 corpus. Then rerun the controlled comparison (w/ PT vs w/o PT) using an explicitly disjoint, action-free video corpus (e.g., JIGSAWS real video only, or a separate surgical simulation not used in evaluation), keeping the fixed 10k-demonstration action budget and 80k fine-tuning steps. If the success-rate advantage over w/o PT shrinks or disappears, the central claim must be re-scoped to in-domain video pretraining rather than generic action-free surgical video pretraining.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that action-free video pretraining, and not the fixed action-labeled budget, causes the 63.5% to 77.8% success improvement (Table 1). The Stage 1 initialization is taken from Cosmos-H-Surgical (He et al. 2026), described only as fine-tuned on 'large-scale surgical video' (Section 3.2); the corpus composition is never specified, and no overlap test with the SurRoL evaluation tasks is reported. If that corpus includes SurRoL rollouts or task-specific frames, then the w/ PT variant has already seen the evaluation tasks' visual distribution during pretraining, while w/o PT has not. The comparison would then demonstrate that pretraining on the same tasks' video helps, not that abundant generic surgical video (e.g., real endoscopic recordings) transfers. The paper's practical motivation depends on this distinction: it promises that inexpensive, abundant endoscopic video can substitute for action labels. The Section 4.4 JIGSAWS result is qualitative only and does not quantify cross-domain transfer, so it cannot rescue the main claim. This is a causal-attribution gap, not a mere implementation detail. The Table 2 non-monotonicity (86% at 80k, 34% at 120k, 56.5% at 160k) makes the reported 80k checkpoint fragile, but the corpus question is the more fundamental threat to the causal interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Surgical WAM, a two-stage world-action model built on the Cosmos Policy backbone, in which a model is first pretrained on action-free surgical video and then fine-tuned on a fixed set of action-labeled demonstrations. Closed-loop evaluation on four SurRoL tasks (Needle Pick, Peg Transfer, Needle Regrasp, BiPeg Transfer) reports average success of 63.5% without pretraining and 77.8% with pretraining, with the largest gain (20 absolute points) on PegTransfer. Additional ablations vary the fine-tuning step count (Table 2) and execution horizon (Table 3), and a qualitative experiment on the JIGSAWS real dVRK video is described in Section 4.4. The central claim is that action-free video pretraining, rather than the action-labeled budget, drives the improvement.","tokens_in":13090,"tokens_out":6682,"duration_ms":93525,"significance":"If the central claim is correct, the paper offers a practical route to reducing the demand for teleoperated kinematics in surgical robot learning: abundant unlabeled endoscopic video can be used to learn visual dynamics, while a small action-labeled set grounds the policy. The main strengths are the controlled w/ PT versus w/o PT comparison, which shares architecture, action representation, optimization, and evaluation protocol, and the consistent gains across four tasks that include contact-rich and bimanual behaviors. The two-stage recipe is clearly described, and the dependence of the benefit on fine-tuning budget and execution horizon is explicitly probed. However, the causal interpretation currently rests on an unspecified pretraining corpus and on single-run results, as detailed below.","major_comments":[{"comment":"The causal claim that action-free video pretraining improves closed-loop control is only as strong as the Stage 1 corpus. The paper states that the Stage 1 weights are obtained from Cosmos-H-Surgical (He et al. 2026), fine-tuned on 'large-scale surgical video', but it never describes the composition of that corpus, nor reports any overlap check against the SurRoL evaluation tasks. If the corpus contains SurRoL-style rollouts or task-specific frames, then the w/ PT variant has already seen the evaluation distribution during pretraining, and the 63.5% to 77.8% gain would not demonstrate transfer from generic action-free surgical video. Please specify the corpus, rule out overlap with the evaluation tasks, or add a pretraining variant on a clearly disjoint corpus (e.g., JIGSAWS-only video) with quantitative closed-loop results.","section":"Section 3.2 (Stage 1; Eq. (2))"},{"comment":"The fine-tuning-step ablation is the main evidence for 'faster convergence' and 'higher peak with half the budget', yet the w/ PT curve is severely non-monotonic: 52% at 40k, 86% at 80k, 34% at 120k, and 56.5% at 160k. At 120k and 160k the pretrained model is far worse than the non-pretrained model at the same steps (71.5% and 79%). This pattern is inconsistent with a simple convergence-speed advantage and suggests overfitting, instability, or evaluation noise. The 'Best' row selects checkpoints post hoc. Please report multiple seeds with error bars, explain the collapse, and justify the choice of 80k steps as the default. If the 86% at 80k is not reproducible, the central claim is not supported by this table.","section":"Table 2; Section 4.3"},{"comment":"The execution-horizon ablation reports single runs for each He. The claim that pretraining makes the controller 'robust across horizons' rests on w/ PT values of 86%, 70%, 86%, and 80%, while w/o PT ranges from 50% to 66%. Without variance estimates or additional seeds, it is hard to separate a real robustness effect from checkpoint or seed noise. Add confidence intervals or repeat each configuration.","section":"Table 3; Section 4.3"},{"comment":"The real-data experiment is presented only qualitatively: no quantitative closed-loop success, no description of how JIGSAWS demonstrations are converted to the absolute Cartesian action space used in SurRoL, and no evaluation protocol. This section is used to assert that the benefit 'persists under the visual complexity and manipulation difficulty of real surgical scenes.' Either provide quantitative evidence (e.g., success metrics, or at least prediction or action metrics with standard errors) or weaken the claim to a qualitative observation.","section":"Section 4.4"}],"minor_comments":[{"comment":"The caption states that 'All methods are trained on the same action-labeled dataset with an identical fine-tuning budget,' but DEX (Huang et al. 2023) is a demonstration-guided RL method that likely uses environment interaction, not only fine-tuning on a fixed dataset. Clarify what 'fine-tuning budget' means for each baseline.","section":"Section 4.1, Table 1 caption"},{"comment":"Figure 2(ii) labels 'Large-scale Footage' and 'Internet Videos' as Stage 1 inputs, while Section 3.2 says Stage 1 uses the He et al. (2026) model fine-tuned on unspecified surgical video. Align the figure with the text or describe the actual pretraining corpus.","section":"Figure 2(ii); Section 3.2"},{"comment":"The sentence 'the gain thus comes entirely from the video pretraining, at no extra cost in action supervision' should be rephrased as 'at no extra action-label cost', since Stage 1 incurs compute and data-curation costs; it should also be conditioned on the corpus-overlap caveat raised above.","section":"Section 3.2"},{"comment":"The paper claims 'the first world-action model applied to surgical manipulation'; given the related work on SAW and Cosmos-Surg-dVRK, specify the sense in which Surgical WAM is first (e.g., first to integrate an action head into the same generative model and evaluate closed-loop task execution).","section":"Section 2.3"},{"comment":"The paper does not include a limitations section; the non-monotonicity in Table 2 and the unspecified pretraining corpus are limitations that should be discussed explicitly in the text.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The controlled w/ PT versus w/o PT comparison is well designed and the central question is timely, but the unspecified pretraining corpus and the non-monotonic fine-tuning curve are the main blockers. The authors should be asked to provide a corpus description and leakage check, multi-seed results with error bars, and either a quantitative real-data evaluation or a softened claim about real-world transfer."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read on the Surgical WAM paper. The headline result is that a world-action model instantiated on Cosmos Policy, when initialized from a surgically-pretrained video world model, beats the same model without that initialization on four SurRoL tasks under a fixed action-labeled budget (63.5% to 77.8% average). The comparison is well designed: same architecture, same 10k demos, same 80k fine-tuning budget, and the only difference is the video-pretrained initialization. As a within-subfield empirical result, this is a solid, useful extension—the first closed-loop WAM for surgical manipulation, and it gives a clean answer to a question people in the surgical learning community actually ask.\n\nWhat's good: the controlled w/PT vs w/oPT comparison, the consistent gains across all four tasks, and the extra ablations on fine-tuning steps and execution horizon. The paper is honest about reporting both peak and fixed-step numbers.\n\nThe soft spots are in the causality, not the architecture. First, the pretraining corpus is a black box. Stage 1 is initialized from He et al.'s Cosmos-H-Surgical model, described only as 'fine-tuned on large-scale surgical video.' No composition, no overlap test with SurRoL. If that corpus includes SurRoL rollouts or task-specific frames, then the gain is in-domain exposure, not transfer from abundant generic endoscopic video. That wouldn't make the result worthless, but it would gut the paper's central interpretation and its practical pitch. This is the main thing I'd want fixed before trusting the claim.\n\nSecond, the training curves are non-monotonic in a way that makes the headline number fragile: 86% at 80k on PegTransfer drops to 34% at 120k and recovers to 56.5% at 160k. The paper attributes this to overfitting, and maybe that's right, but with a single seed and no confidence intervals, I don't know whether 86% is a reproducible peak or a lucky checkpoint. Multiple seeds and error bars would settle this.\n\nThird, the real-data JIGSAWS section is qualitative only. It says the trend holds but doesn't quantify closed-loop success, so it doesn't address the cross-domain worry.\n\nOn balance, I'd send this to peer review. The empirical question is important, the controlled comparison is close to the right experiment, and the issues are fixable with more reporting rather than a fundamentally different study. The authors should describe the pretraining corpus, run a leakage check or train on a clearly out-of-domain video corpus, and add seeds/CI. Without those, I'd treat the 14.3-point gain as preliminary.\n\nThat's my take — happy to discuss.","headline":"Well-designed controlled study showing that surgical video pretraining helps, but the central causal claim rests on an unspecified pretraining corpus and a fragile-looking 80k checkpoint.","tokens_in":13624,"tokens_out":6540,"would_cite":true,"duration_ms":52629,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under a fixed action-labeled budget, action-free video pretraining improves closed-loop surgical manipulation, lifting average success from 63.5% to 77.8% on four SurRoL tasks.","keywords":["surgical robotics","world-action model","action-free video pretraining","data-efficient robot learning","closed-loop manipulation","dVRK","SurRoL","imitation learning"],"falsifier":"Collect a Stage 1 corpus that provably excludes all SurRoL task frames or rollouts (e.g., only unrelated surgical recordings or only non-task tissue motion), run the identical two-stage protocol, and compare success rates; if the 14.3-point average gain disappears, the attribution to action-free video pretraining fails.","tokens_in":12577,"feed_emoji":"🦾","tokens_out":4990,"duration_ms":43254,"temperature":0.7,"pith_summary":"Surgical robot policies are starved for action-labeled demonstrations because teleoperating a dVRK with synchronized kinematics is slow and expensive, while endoscopic video is comparatively cheap and abundant. This paper asks whether that abundant video alone can carry the visual-dynamics learning that policies need. It answers yes: a world-action model pretrained on action-free surgical video and then fine-tuned on a fixed budget of demonstrations improves closed-loop task success from an average of 63.5% to 77.8% across four simulated dVRK manipulation tasks, with the biggest gains on contact-rich and bimanual tasks. The result matters because it suggests the data bottleneck for surgical robot learning can shift from costly teleoperation to video that surgery already produces at scale.","feed_headline":"Video pretraining lifts surgical robot success from 63.5% to 77.8%","feed_subtitle":"Closed-loop dVRK manipulation improves without extra action labels, easing the teleoperation bottleneck.","key_machinery":"The central object is the world-action model (WAM), instantiated with Cosmos Policy: a single diffusion transformer whose latent sequence contains slots for past endoscopic frames, proprioceptive state, future-video predictions, and action chunks, all denoised jointly. Stage 1 pretrains only the visual and future-prediction slots on action-free video using the video denoising loss, so weights learned without kinematic labels become the initialization for action prediction. Stage 2 fine-tunes the whole sequence on action-labeled demonstrations with the joint denoising objective, coupling action tokens to the video-derived dynamics. At test time the model is a receding-horizon controller: it predicts a chunk of Hc=16 actions and future frames, executes the first He=4 actions, then replans from the new observation.","core_discovery":"The paper's central claim is that under a fixed budget of action-labeled demonstrations, action-free video pretraining improves closed-loop surgical manipulation. The authors establish this by building Surgical WAM, a world-action model that jointly predicts future endoscopic observations and executable dVRK action chunks, pretraining it on action-free surgical video, and then fine-tuning it on the fixed action-labeled set. With pretraining, average closed-loop success across four SurRoL tasks rises from 63.5% to 77.8%, including a 20-point absolute gain on PegTransfer; ablations show the pretrained model peaks with half the fine-tuning steps and stays robust across execution horizons, and the trend reproduces on real dVRK video from JIGSAWS.","pith_inferences":["If the Stage 1 corpus is truly action-free and disjoint from evaluation tasks, a similar recipe could extend to other surgical platforms and to partially automated data curation, since the bottleneck shifts from teleoperation to video collection and filtering.","The finding predicts a scaling law: success should continue to rise with the volume and diversity of action-free surgical video, which the paper does not directly measure but which is the natural next experiment.","The receding-horizon design suggests that joint sampling of video and action tokens at inference may be replaceable by cheaper open-loop action sampling once the dynamics prior is fixed; the paper discards predicted frames, so ablating joint sampling could cut inference cost without much loss.","A testable extension is to pretrain on out-of-domain video from a different surgical site or simulator to distinguish generic video priors from task-specific leakage, since the current corpus is unspecified."],"forward_implications":["With the same 10k demonstrations, video-pretrained Surgical WAM reaches 77.8% average success versus 63.5% without pretraining, and peaks in 80k fine-tuning steps instead of 160k.","The gain is concentrated where surgical dynamics matter most: contact-rich and bimanual tasks (PegTransfer +20 points, BiPegTransfer 64% vs 42% without pretraining).","The benefit generalizes beyond simulation: the same two-stage protocol on real JIGSAWS dVRK video shows the same qualitative advantage, arguing the prior captures genuine surgical dynamics rather than simulator artifacts.","Generalist vision-language-action pretraining transfers poorly to surgery (one baseline averages 22.3%), whereas the surgical video prior transfers well, suggesting domain-matched dynamics matter more than broad language grounding for precision tasks."],"supporting_citations":[{"why":"Provides the SurRoL benchmark with endoscopic observations, dVRK proprioception, synchronized actions, and the four evaluation tasks.","marker":"Xu et al. 2021"},{"why":"Cosmos Policy is the diffusion-transformer backbone whose shared latent sequence the paper repurposes as its world-action model.","marker":"Kim et al. 2026"},{"why":"Supplies the action-free surgical world model whose weights populate the Stage 1 visual and future-prediction slots.","marker":"He et al. 2026"},{"why":"Cited as evidence that removing video co-training degrades manipulation performance, motivating the two-stage recipe.","marker":"Yuan et al. 2026"},{"why":"JIGSAWS provides real dVRK surgical videos used to show the pretraining benefit is not a simulator artifact.","marker":"Gao et al. 2014"},{"why":"The video diffusion backbone fine-tuned on surgical video in He et al., i.e., the underlying generative model for Stage 1.","marker":"Agarwal et al. 2025"}],"fun_headline_variants":["Video pretraining gives surgical robots a 14-point success boost","Action-free video pretraining improves surgical manipulation","Surgical WAM: video pretraining beats action scarcity","Video-only pretraining raises surgical robot success 14 points","Surgical robot success up 14 points with video pretraining"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The action-free video corpus used in Stage 1 is not described, so the causal claim that gains come from action-free video pretraining assumes this corpus does not contain frames or rollouts from the SurRoL evaluation tasks.","fun_headline_variants_meta":{"raw":{"variants":["Video pretraining gives surgical robots a 14-point success boost","Action-free video pretraining improves surgical manipulation","Surgical WAM: video pretraining beats action scarcity","Video-only pretraining raises surgical robot success 14 points","Surgical robot success up 14 points with video pretraining"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000808,"raw_usage":{"total_tokens":3580,"prompt_tokens":1010,"completion_tokens":2570,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":2490}},"tokens_in":626,"tokens_out":2570,"duration_ms":36381,"temperature":1.0,"reasoning_tokens":2490,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:25.242910+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a Stage 1 corpus that provably excludes all SurRoL task frames or rollouts (e.g., only unrelated surgical recordings or only non-task tissue motion), run the identical two-stage protocol, and compare success rates; if the 14.3-point average gain disappears, the attribution to action-free video pretraining fails.","supporting_citations":[],"review_version":2}