{"id":"a6d9df87-d922-438e-b196-8629fe839ab6","arxiv_id":"2411.14622","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Two reinforcement learning agents, trained in a new fluid-capable surgical simulator and transferred to a real da Vinci Research Kit, autonomously perform irrigation and suction in a simplified bench-top setup.","lead":"The authors built a simulated surgical environment to train two vision-based reinforcement learning agents, one for irrigating a dirty surgical field and one for suctioning liquid out. In bench-top tests with the da Vinci Research Kit, the agents removed most of the contaminant, approaching manual performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 9 force-terminated suction trials are the load-bearing weak point: if they are excluded from the reported means, the quantitative sim-to-real evidence is success-conditioned.","rationale":"The reader's weakest assumption is the same as the concern I would raise, and I find it the most load-bearing issue in the paper. The qualitative claim that two vision-based RL agents can transfer from CRESSim-ML to the real dVRK is plausible and supported by the snapshots and measured weights; the problem is that the central quantitative evidence—the residual grams after suction—may be computed only over trials that did not hit the safety force limit. Because contact is often required to suction liquid, excluding terminated trials would make the reported 2.64, 2.24, and 2.42 g values optimistic. The paper is transparent about the terminations (Section VII.C), which is to its credit, but it never says how they are handled in the statistics. Additional issues—the abstract/text discrepancy (2.21 vs 2.11 g), the absence of significance tests for small samples, and the lack of released code or data—reinforce the need for a conditional verdict rather than full acceptance. I would not move the verdict to REJECT because the core demonstration of a working bench-top irrigation-suction pipeline is not in doubt; I would also not move it to ACCEPT until the full trial distribution is reported.","tokens_in":20210,"tokens_out":5589,"duration_ms":60167,"concrete_test":"Request or reconstruct a per-trial table: for each of the 20 suction-only and 10 combined trials, list initial weight, final weight after autonomous suction, and a flag for force-termination. Recompute the reported means (including the 2.42 g combined residual) with the 9 terminated trials included, using both the measured weight at termination and, as an upper bound, the initial fluid weight. If the means shift by more than about 0.5 g, or if the ANOVA p-value changes materially, the claims in the abstract and Sections VI.C/D must be revised to report the full distribution instead of only completed trials.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section VII.C reports that 9 of roughly 30 suction trials (20 suction-only plus 10 combined) were terminated by the hardcoded force limit, but Sections VI.C and VI.D report residual means (2.64, 2.24, and 2.42 g) without stating whether those trials are included, excluded, or assigned a final weight at termination. This is load-bearing because the main empirical evidence for successful sim-to-real transfer is the residual-liquid mass. Force termination occurs on contact with tissue or liquid, which is exactly the condition needed to suction pooled fluid, so terminated trials are likely to have more residual fluid than completed ones. If they were dropped, the reported means describe only a success-conditioned subset and overstate expected autonomous-suction performance over the full trial distribution; if they were included without a defined terminal measurement, the measurement protocol is underspecified. The paper's own limitation statement (Section VII.C) acknowledges the issue but does not reconcile it with the headline numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes CRESSim-ML, a Unity/PhysX-based learning platform for the da Vinci Research Kit, and uses it to train two vision-based RL agents—one for irrigating a contaminant and one for suctioning liquid—using domain randomization, curriculum learning, and imitation learning. The agents are transferred to a physical dVRK with an EndoWrist Suction/Irrigator and evaluated in irrigation-only (10 trials), suction-only (20 trials), and combined irrigation-suction (10 trials) settings. The headline quantitative claims are that after autonomous irrigation plus manual suction, 2.11 g (abstract: 2.21 g) of contaminant remains versus 1.90 g for manual irrigation, that the suction agent leaves 2.64 and 2.24 g residuals for roughly 20 g and 30 g initial fluid loads, and that combined autonomous irrigation-suction leaves 4.40 g total weight with 2.42 g of suctionable residual. The authors position this as the first automated two-step irrigation-suction policy using raw RGB observations.","tokens_in":20464,"tokens_out":5942,"duration_ms":58628,"significance":"If the empirical results survive scrutiny, this is a useful advance: it extends sim-to-real RL to fluid manipulation in surgery, demonstrates a two-step irrigation-suction pipeline with raw RGB input, and provides a reusable simulation platform. The authors deserve credit for conducting physical experiments with external weight measurements rather than only simulated metrics, for reporting training curves and reward design, and for documenting suboptimal outcomes. The main novelty is the simulation and application contribution; the RL machinery (PPO plus DR/CL/IL) is standard. The quantitative strength of the paper is currently limited by small per-condition sample sizes, the unresolved treatment of force-terminated trials, and an internal inconsistency in the headline irrigation number.","major_comments":[{"comment":"Section VII.C states that 9 of the 30 suction trials (20 suction-only plus 10 combined) were terminated by a hardcoded EE force limit, but Sections VI.C and VI.D report residual means (2.64 +/- 1.87 g, 2.24 +/- 2.24 g, and 2.42 +/- 2.04 g) without stating whether terminated trials were included, excluded, or assigned a terminal weight. Force termination occurs when the EE contacts tissue or liquid, which is exactly the situation in which pooled fluid must be suctioned, so excluding those trials conditions the reported averages on success and likely biases them downward. If the trials were included, the measurement at termination is not defined. This ambiguity is load-bearing for the sim-to-real transfer claim; the authors should report all per-trial outcomes, state the inclusion rule, and provide the mean over all 30 trials under a clearly defined terminal measurement.","section":"Section VII.C; Sections VI.C and VI.D"},{"comment":"The abstract reports that autonomous irrigation leaves 2.21 g after manual suction, while Section VI.B reports 2.11 +/- 0.80 g for the same quantity; one of these values is incorrect, and the inconsistency affects a headline result. In addition, the claim that the agent's performance is 'close to that of a human' (1.90 +/- 0.49 g) is based on overlapping uncertainties with no significance test or confidence interval. With n=10, the difference of 0.21 g is not interpretable as equivalence; report an appropriate statistical comparison or soften the claim to a qualitative observation.","section":"Abstract vs. Section VI.B"},{"comment":"The irrigation-only evaluation measures performance by the contaminant weight remaining after the agent's irrigation followed by a manual suction that is assumed to be optimal. This is an indirect proxy for 'dilution', and the assumption of optimal manual suction is neither validated nor described (number of passes, suction duration, stopping rule). Since the trained irrigation reward is based on simulated particle states, no real-world measurement directly confirms that the agent is diluting the contaminant rather than merely moving fluid around. The proxy is reasonable as a first-order metric, but it should be explicitly framed as a proxy and the manual-suction protocol should be specified so that the irrigation claim is reproducible.","section":"Sections V.A and VI.B"}],"minor_comments":[{"comment":"Figure 11b would be much more informative if force-terminated trials were marked with a distinct symbol, since the current plot is used to support the residual means and the treatment of terminated trials is under discussion.","section":"Figure 11b"},{"comment":"The Pearson and Spearman correlations are computed on n=10 trials; the absence of significant correlation is weak evidence and should be reported as such rather than as support for independence of irrigation and suction performance.","section":"Section VI.D"},{"comment":"There are typos in the manuscript: 'ANOV A' should be 'ANOVA' in Section VI.D, and the acknowledgment heading is printed as 'A CKOWLEDGMENTS'.","section":"Section VI.D and Section IX"},{"comment":"Equation (2) and the adjacent text describe color diffusion with weights that include a velocity factor c||v_j||, but the notation does not make clear whether the weights are normalized over both i and j or only j; clarifying this would make the update well-defined and reproducible.","section":"Equation (2)"},{"comment":"The statement that no quantitative analysis of sim-to-real transfer was performed is an important limitation; it should be reflected more prominently in the conclusion and not only in the discussion.","section":"Section VII.A"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the core result is potentially publishable, but the central quantitative evidence needs to be reported with all trials included and the abstract/main-text discrepancy corrected. If the authors can provide per-trial data showing that the means are robust to the inclusion or exclusion of the 9 force-terminated suction trials, I would support acceptance; otherwise the paper's main empirical claim is not established. The heavy reliance on self-cited prior suction work [19] is acceptable in context, but the authors should state clearly what is new beyond that work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a credible sim-to-real RL result for a genuinely new task—automated two-step irrigation-suction on a dVRK with raw RGB observation. The core demonstration works in a bench-top setup; the question is whether the headline numbers are as clean as they look.\n\nThe genuinely new bits: no prior work automated irrigation or the combined irrigation-suction process; prior suction work used manual feature extraction. The CRESSim-ML platform with particle color diffusion and screen-space rendering is a useful technical contribution. They train with PPO, DR, curriculum, and IL, and they properly ablate the training choices: curriculum helps irrigation, IL helps suction and hurts irrigation. The real-world experiment uses external weight measurements, so there is no circularity issue.\n\nSoft spots: the load-bearing one is the 9 force-terminated suction trials. The results sections (VI.C and VI.D) report residual-liquid means without saying whether terminated trials were included, excluded, or assigned a final weight. Section VII.C acknowledges the terminations but never reconciles them with the reported numbers. Since force termination happens on contact with tissue/liquid, those trials plausibly have more residual fluid. If they are dropped, the averages describe a success-conditioned subset; if they are included, the measurement protocol is underspecified. That's a genuine gap, not a manufactured one. It doesn't kill the paper—the videos and snapshots show real capability—but it means the quantitative transfer claim is not fully supported as stated. A reviewer should ask for trial-level data.\n\nMinor issues: the abstract says 2.21 g for irrigation while the main text says 2.11 g; the human comparison overlaps heavily (2.11±0.80 vs 1.90±0.49) with no significance test; and no code/data is shipped, which limits reproducibility. Also, the 'sim-to-real' evaluation is in a simplified bench-top setting with ketchup and playdough—that's fine as a proof of concept, but the authors themselves acknowledge it's not a surgical scene.\n\nBottom line: this deserves a serious referee. The central claim—that a vision-based RL agent can transfer from this simulator to the real dVRK and do both irrigation and suction—holds up at the qualitative level. The quantitative performance numbers need clarification on the terminated trials, not a redesign. Send it to review with a request for the raw per-trial data and a reconciliation of the terminated trials.","headline":"Credible first demonstration of autonomous irrigation-suction with raw RGB on a dVRK, but the residual-mass numbers are undercut by an unresolved handling of force-terminated trials.","tokens_in":20972,"tokens_out":2453,"would_cite":false,"duration_ms":22574,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two vision-based reinforcement-learning agents trained in a fluid simulator transfer to a physical da Vinci Research Kit and autonomously perform surgical irrigation and suction, leaving contaminant residuals close to those of a human…","keywords":["surgical robotics","irrigation-suction","reinforcement learning","sim-to-real transfer","fluid simulation","domain randomization","vision-based control","da Vinci Research Kit"],"falsifier":"Re-run the suction-only and combined irrigation-suction trials on the physical robot, this time including the nine force-terminated trials in the average as their measured or worst-case residuals; if the mean suctionable residue then rises well above the 2.24-2.64 gram band, the transfer claim is only true for a selected subset of trials. A complementary check is to vary the camera pose, light angle, or contaminant color beyond the trained randomization range and measure the residual after autonomous irrigation and manual suction.","tokens_in":19996,"feed_emoji":"🩺","tokens_out":12747,"duration_ms":112326,"temperature":0.7,"pith_summary":"This paper sets out to establish that the two-step surgical irrigation-suction routine—rinsing contaminants and then vacuuming the fluid away—can be automated end-to-end by reinforcement-learning agents that see only raw RGB camera images, and that these agents will work on a physical da Vinci Research Kit after training in simulation. The authors train separate vision-based policies for irrigation and suction in a custom simulator with visually plausible fluid mixing, then transfer them to the real robot using domain randomization, a two-lesson curriculum, and imitation learning. In real-world trials, the irrigation agent followed by manual suction leaves roughly 2.1-2.2 grams of an initial ~5.5 grams of contaminant, close to the 1.90 grams left by manual human irrigation, while the suction agent leaves 2.64 and 2.24 grams from larger starting volumes. Fully autonomous irrigation followed by autonomous suction leaves an average of 2.42 grams of suctionable liquid behind, which the authors find is not statistically different from suction alone. If correct, this would be the first automated, feature-free demonstration of both fluid-handling subtasks in robot-assisted surgery, with the practical consequence of reducing the surgeon's workload during cleaning.","feed_headline":"Autonomous irrigation-suction works on a real surgical robot","feed_subtitle":"Simulator-trained vision agents clear about half the contaminant, matching human irrigation within 0.3 grams.","key_machinery":"The load-bearing machinery is CRESSim-ML, the paper's simulated surgical-robot learning platform, plus a trained policy pair. The simulation side uses position-based fluid dynamics (PBF) particles with a screen-space rendering pass and a color-diffusion scheme that updates each particle's color as a distance- and speed-weighted average of nearby particles, so irrigated liquid visibly mixes with the contaminant. The control side exploits the fact that the patient-side manipulator with the suction/irrigator tool is a kinematically constrained 5-DoF robot: the policy outputs incremental joint-space targets, avoiding the need for Cartesian inverse kinematics, and observes current joint positions plus a tissue-contact indicator. The transfer side combines domain randomization over colors, lighting, camera pose, tissue shape, and fluid physics with a two-lesson curriculum for irrigation and imitation learning (behavior cloning and a GAIL-style auxiliary reward) that helps the suction agent but not the irrigation one. Together these components let a raw-RGB policy trained in simulation produce the reported physical-world results.","core_discovery":"Stated on the paper's own terms: two agents, one for irrigation and one for suction, map an $84 \\times 84$ RGB frame plus stacked joint positions to incremental 5-DoF joint-space commands at 10 Hz, and after sim-to-real transfer they complete the physical irrigation-suction procedure in most trials. The irrigation agent learns to aim the suction/irrigator at dark contaminant and toggle irrigation only while the tool is close to it, using a curriculum that first teaches approach-and-activate behavior before adding the full dilution reward. The suction agent learns to move toward liquid regions and remove them under a cone-shaped suction force field, with rewards for particles removed and for approaching the nearest liquid. Real-world measurements put irrigation-plus-manual-suction residuals at about 2.1-2.2 grams versus 1.90 grams for human irrigation, suction-only residuals at 2.64 and 2.24 grams for roughly 22-gram and 31-gram starting pools, and combined autonomous trials at 2.42 grams of suctionable liquid left behind out of an initial ~5 grams, with 4.40 grams total remaining because the suction policy leaves some irrigated liquid. An ANOVA finds no statistically significant difference among the three suction-performance groups, and the paper reports no significant correlation between irrigation quality and suction outcome.","pith_inferences":["The paper leaves implicit that a high-level supervisor could decide when to switch from irrigation to suction and when to repeat the cycle; the 2.42 grams of leftover suctionable liquid suggests a second pass would reduce the residual further.","The nine force-limit terminations imply that contact-rich suction needs an explicit recovery behavior, such as lifting the end-effector off the tissue before repositioning, which the current policy lacks.","The color-diffusion and screen-space fluid rendering could be reused for vision-based training in other mixed-fluid surgical contexts, such as bleeding with cauterization smoke, where visual mixing matters more than exact fluid mechanics.","Adding depth or stereo observations would likely shrink the reported sim-to-real gap, but the paper's exclusion of depth is a reasonable design choice because real-world depth is noisy; this trade-off is testable."],"forward_implications":["A template for automating fluid-handling surgical subtasks becomes available: simulate with PBF and color diffusion, randomize appearance and physics, and train a raw-RGB joint-space policy.","The suction agent can be deployed independently, for example in blood-suction tasks, because it was trained on diverse liquid configurations and does not require the irrigation agent.","In the authors' statistical test, the fully autonomous two-step sequence performs no worse than suction alone, meaning an upstream autonomous irrigation step does not measurably harm downstream suction.","The measured irrigation gap to human operation is about 0.2-0.3 grams, which the paper presents as showing agent performance close to that of a human."],"supporting_citations":[{"why":"Supplies the cone-shaped suction force-field model and the blood-suction sim-to-real approach that the suction environment is built on.","marker":"[19]"},{"why":"Is the precedent for vision-based sim-to-real reinforcement learning with a surgical robot that this work extends to fluids.","marker":"[13]"},{"why":"Demonstrates domain-randomized RGB-image policy transfer to the surgical robot, the transfer recipe reused here.","marker":"[12]"},{"why":"Is the earlier simulator on which the proposed platform builds, contributing the patient-side robot model and teleoperation interface.","marker":"[22]"},{"why":"Supplies the reinforcement-learning training loop, the PPO implementation, and the behavior-cloning and GAIL losses used in the paper.","marker":"[45]"},{"why":"Defines the domain randomization technique used to bridge simulation and reality.","marker":"[46]"},{"why":"Is the position-based fluid simulation method underlying the simulated liquid particles and suction behavior.","marker":"[42]"},{"why":"Is the prior image-based blood-suction automation that the paper contrasts with its raw-RGB, learned approach.","marker":"[17]"},{"why":"Introduces the open-source surgical robot research kit used for the real-world evaluations.","marker":"[21]"}],"fun_headline_variants":["RL agents master surgical irrigation and suction on da Vinci","Sim-trained robots clear surgical contaminants on da Vinci","Autonomous irrigation-suction tested on real surgical robot","Irrigation and suction agents transfer from sim to real da Vinci"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported real-world averages come from completed trials only; the performance figures assume that the nine trials halted by the hard-coded force limit would not have added substantially larger residuals had they been allowed to finish.","fun_headline_variants_meta":{"raw":{"variants":["RL agents master surgical irrigation and suction on da Vinci","Sim-trained robots clear surgical contaminants on da Vinci","Autonomous irrigation-suction tested on real surgical robot","Irrigation and suction agents transfer from sim to real da Vinci"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000964,"raw_usage":{"total_tokens":4206,"prompt_tokens":1152,"completion_tokens":3054,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":768,"completion_tokens_details":{"reasoning_tokens":2990}},"tokens_in":768,"tokens_out":3054,"duration_ms":20292,"temperature":1.0,"reasoning_tokens":2990,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:04:51.697447+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the suction-only and combined irrigation-suction trials on the physical robot, this time including the nine force-terminated trials in the average as their measured or worst-case residuals; if the mean suctionable residue then rises well above the 2.24-2.64 gram band, the transfer claim is only true for a selected subset of trials. A complementary check is to vary the camera pose, light angle, or contaminant color beyond the trained randomization range and measure the residual after autonomous irrigation and manual suction.","supporting_citations":[{"cited_title":"Learning nonprehensile dynamic manipulation: Sim2real vision-based policy with a surgical robot,","cited_arxiv_id":null,"evidence_quote":"Demonstrates domain-randomized RGB-image policy transfer to the surgical robot, the transfer recipe reused here."},{"cited_title":"Autonomous blood suction for robot-assisted surgery: A sim-to-real reinforcement learning approach,","cited_arxiv_id":null,"evidence_quote":"Supplies the cone-shaped suction force-field model and the blood-suction sim-to-real approach that the suction environment is built on."},{"cited_title":"Sim2real rope cutting with a surgical robot using vision-based reinforcement learning,","cited_arxiv_id":null,"evidence_quote":"Is the precedent for vision-based sim-to-real reinforcement learning with a surgical robot that this work extends to fluids."},{"cited_title":"A realistic surgical simulator for non-rigid and contact-rich manipulation in surgeries with the da vinci research kit,","cited_arxiv_id":null,"evidence_quote":"Is the earlier simulator on which the proposed platform builds, contributing the patient-side robot model and teleoperation interface."},{"cited_title":"Position based fluids,","cited_arxiv_id":null,"evidence_quote":"Is the position-based fluid simulation method underlying the simulated liquid particles and suction behavior."},{"cited_title":"Autonomous robotic suction to clear the surgical field for hemostasis using image-based blood flow detection,","cited_arxiv_id":null,"evidence_quote":"Is the prior image-based blood-suction automation that the paper contrasts with its raw-RGB, learned approach."},{"cited_title":"An open-source research kit for the da vinci® surgical system,","cited_arxiv_id":null,"evidence_quote":"Introduces the open-source surgical robot research kit used for the real-world evaluations."}],"review_version":1}