{"id":"44380424-8490-49a6-82d1-cf5607bf5eb7","arxiv_id":"2507.07718","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A small pilot study claims haptic-assisted virtual-reality training improves skill transfer in surgical robotics, but n=4 per group and a flawed performance score undermine the finding.","lead":"This paper reports a training study in which novice surgeons practiced robotic surgery tasks in virtual reality, with one group receiving haptic guidance forces from the simulator. The authors claim this guided practice improved later performance on unassisted tasks, but the evidence is weak and the scoring method has internal inconsistencies.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 18's performance index sums lower-is-better error metrics without sign inversion, so higher P means worse performance and the reported transfer advantage may be a sign-convention artifact.","rationale":"I read the paper as an exploratory, small-sample training study whose central claim is that haptic assistance improves performance during training and transfers to unassisted tasks. For that claim to hold, P must be a valid higher-is-better skill score. Eq. 18 is not: it sums normalized distance error, angular error, drops, and repositioning fraction without any sign inversion, so lower errors decrease P. The discussion and results repeatedly treat larger P as better, which is internally inconsistent with the definition. This is the most load-bearing issue because every quantitative conclusion is expressed through P; if the sign is wrong, all reported improvements invert. I also note the absence of inferential statistics with n=8, but the metric defect is prior to statistics. The reader's weakest-assumption analysis identified the same Eq. 18 sign problem, and I agree with that diagnosis. The proposed recomputation test would settle the issue directly by checking both the direction of P against the expert's peak-performance run and by computing Day 7 group differences with the natural sign convention. No change to the reader's REJECT verdict is needed.","tokens_in":9788,"tokens_out":5852,"duration_ms":68660,"concrete_test":"Recompute Day 7 performance from the raw logged metrics (D, A, M, C with F=T=0) under two definitions: (a) exactly as Eq. 18, and (b) an error-only cost with sign flipped so that lower errors yield higher scores. As a validation, check whether the expert surgeon's recorded 'peak performance' run, which should have the smallest D/A/M/C values, receives the maximum or the minimum P under Eq. 18. If the assisted-vs-control ordering changes, or the +21.54% Liver Resection effect reverses, the transfer claim rests entirely on the sign convention.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends entirely on P from Eq. 18. As printed, P = (1/10) Σ w_k X_k(subject)/X_k(surgeon) for k ∈ {D, A, F, T, M, C}. Four components are explicitly error/cost metrics: D is distance error (mm), A is angular error (rad), M is number of drops, and C is fraction of time spent repositioning; for all of these, lower is better. The equation contains no inversion, complement, or negative sign, so better execution yields a smaller P. The Results and Discussion nevertheless interpret higher P as better performance, e.g., a 'reduced median performance' of −2.71% on Nephrectomy is treated as the one negative outcome, and the +21.54% Liver Resection advantage is cited as evidence of transfer. If the natural lower-is-better reading is correct, every reported group difference reverses. A second, related defect is that F and T are assistance feedback magnitudes: during training they are nonzero for the assisted group by construction and zero for controls, so P mechanically rewards receiving assistance; on Day 7 both groups are unassisted and F=T=0, so the transfer comparison reduces to the un-inverted error terms. This is an internal inconsistency in the outcome measure, not a dispute about the field's consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a haptic-enhanced virtual reality simulator for surgical robotics training, including four haptic assistance strategies (trajectory guidance, obstacle avoidance, surface guidance, insertion guidance), eight surgical tasks, and an experimental study with eight novice subjects divided into assisted and control groups. The experiment spans four training days with assistance delivered only to the assisted group, followed by two rest days and a final day of unassisted, never-seen evaluation tasks. The authors claim that haptic assistance improves performance during training and promotes transfer of skills to unassisted scenarios. The quantitative outcome is a performance index P defined in Eq. (18), and results are reported as percentage differences in mean and median P values between the two groups.","tokens_in":10093,"tokens_out":4080,"duration_ms":45802,"significance":"If the central claims were valid, the work would be a meaningful contribution to surgical robotics training: it presents a multi-task simulator integrated with a real dVRK system, implements several haptic assistance formulations, and uses a multi-day protocol with a transfer phase—an advance over ad-hoc single-task evaluations. The authors are also transparent about the exploratory nature and small sample size. However, the quantitative foundation is undermined by an apparent sign error in the performance index, a circularity in the training-phase comparison, and the absence of inferential statistics. These issues affect every quantitative conclusion in the paper, so the study's contribution is currently not established.","major_comments":[{"comment":"Equation (18) defines P as (1/10) Σ w_k X_k(subject)/X_k(surgeon) for k ∈ {D, A, F, T, M, C}. D (distance error), A (angular error), M (number of drops), and C (fraction of time spent repositioning) are all lower-is-better metrics. No inversion, complement, or negative sign is applied, so lower P should correspond to better performance. Section IV and Figure 3 instead interpret higher P as better, treating +21.54% on Liver Resection as an improvement and −2.71% on Nephrectomy as a decrement. Under the literal reading of Eq. (18), every reported group difference reverses, contradicting the abstract's claim of improved performance and transfer. This is an internal inconsistency in the central outcome measure, not a minor notation issue.","section":"Eq. (18), Section III-E"},{"comment":"During the training phase (Days 1–4), the assistance feedback magnitudes F and T are nonzero only for the assisted group and zero for the control group. These terms enter with positive weights in Eq. (18), so the assisted group's P is inflated mechanically. The assistance algorithms are also designed to reduce exactly the error metrics D and A, which carry the largest weights in Table I. The training-phase performance advantage is therefore forced by the construction of the metric rather than by any measured difference in skill acquisition. The only non-circular comparison is Day 7, where both groups are unassisted and F = T = 0 for both; however, that comparison is still compromised by the sign-convention problem in the first comment.","section":"Section III-D, Eq. (18), Table I"},{"comment":"The results section reports only descriptive statistics (means, medians, standard deviations) with no significance tests, confidence intervals, or effect sizes. With n = 4 per group and three repetitions per task, the reported differences (e.g., +21.54% for Liver Resection) are well within sampling variability, and the boxplots in Figure 3 appear to show overlapping distributions. The discussion's conclusion that haptic assistance 'promotes the transfer of the acquired skills' is not supported without at least a permutation test or a mixed-effects model that accounts for repeated measures and the small sample size.","section":"Section IV"},{"comment":"Equation (18) normalizes each metric by X_k(surgeon), the expert surgeon's metric value. For metrics where the expert's value is zero—notably M (number of drops) on Exchange and Suturing, and possibly F and T if the expert's reference performance was recorded without assistance—the ratio is undefined. The manuscript does not state how zero expert values are handled, yet Table I assigns nonzero weights to M for Exchange and Suturing. This affects every computation of P and must be addressed for the index to be valid.","section":"Eq. (18), Metrics defined in Section III-E"}],"minor_comments":[{"comment":"The word 'monitorning' appears to be a typo for 'monitoring'.","section":"Section II"},{"comment":"The text 'Weights are reported in Table III-D' appears to refer to Table I, not a table labeled 'III-D'.","section":"Section III-B"},{"comment":"The word 'assignation' should be 'assignment'.","section":"Section III-D"},{"comment":"The sentence 'The performance distribution shown for Nephrectomy, Liver Resection and Suturing mimic the one reported in Figure 3' should use 'mimics' for grammatical agreement.","section":"Section IV"},{"comment":"The manuscript does not report how many repetitions the expert surgeon performed to obtain the reference metric values, nor whether assistance was active during those reference recordings; this information is needed to interpret the normalization in Eq. (18).","section":"Section III-D"}],"recommendation":"reject","confidential_remarks":"The paper describes a substantial engineering effort, but the central quantitative analysis is not sound as presented. The sign error in the performance index and the circular inclusion of assistance magnitudes in the training-phase comparison are load-bearing; they invert or mechanistically produce the reported effects. Combined with the absence of inferential statistics on a sample of eight subjects, the study does not currently support the claimed conclusions. A revision could potentially fix the analysis, but with only four subjects per group the transfer claim would remain underpowered; new data or a fundamentally different analysis would be needed. I therefore recommend rejection rather than major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the transfer experiment is well-designed in outline, but Eq. 18, the performance index everything rests on, is written as a weighted sum of lower-is-better error metrics with no sign inversion. Literally read, higher P means worse performance, so the reported \"improvements\" would be degradations. That alone sinks the current conclusions.\n\nWhat's genuinely new here is the protocol: four days of training on simple tasks, two days off, then unassisted never-seen surgical tasks on day seven. That is a reasonable way to test transfer, and the dVRK-based simulator with four haptic assistance modes is implemented in enough detail to be reproduced. The authors also deserve credit for flagging the small sample and calling the study exploratory.\n\nThe soft spots are serious. First, Eq. 18. D, A, M, C are all error/cost metrics; F and T are assistance magnitudes, zero for the control group during training and for everyone on day seven. Without inversion, a subject who performs worse gets a larger P. Yet the Results and Discussion treat larger P as better. Either the equation has a typo or the interpretation is backwards; the manuscript as written is internally contradictory. During training the assisted group mechanically accrues nonzero F and T terms, so that comparison is confounded by construction. On the day-seven transfer tasks those terms vanish, and the literal reading of Eq. 18 makes the assisted group's higher P a sign of worse unassisted performance—the opposite of the abstract's claim. Second, with four subjects per group, no significance tests, and only descriptive medians, the evidence is thin even if the metric were fixed. That is appropriate to call a pilot, but not sufficient to support the transfer conclusion. The citation practice is fine; the assistance methods are explicitly traced to prior work.\n\nWho is this for? Researchers building surgical training simulators and designing transfer studies. It deserves a serious referee only if the authors fix the metric and reanalyze; in its current form I would not trust any of the quantitative results. My recommendation: send back for major revision with the metric issue as the gate, or reject if the sign convention cannot be clarified. The protocol is salvageable; the conclusions are not.","headline":"The protocol is well-designed, but Eq. 18 defines performance as a weighted sum of lower-is-better errors with no sign inversion, so the paper's quantitative conclusions invert if read literally.","tokens_in":10579,"tokens_out":4404,"would_cite":false,"duration_ms":46403,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Force-based haptic assistance during virtual-reality surgical training improves performance during training and promotes transfer of acquired skills to unassisted, never-seen surgical evaluation tasks.","keywords":["surgical robotics training","haptic assistance","virtual reality simulator","skill transfer","force feedback","robot-assisted surgery","training curriculum evaluation"],"falsifier":"Recompute the performance index $P$ with the error metrics $D$ and $A$ entered as penalties rather than positive contributions, and with $F$ and $T$ excluded entirely for the unassisted group; if the assisted group then no longer outperforms the control group on the evaluation tasks, the claimed transfer effect does not survive.","tokens_in":9617,"feed_emoji":"🤖","tokens_out":7524,"duration_ms":68908,"temperature":0.7,"pith_summary":"This paper tries to establish that adding force-based haptic assistance to a virtual-reality surgical robotics training curriculum improves performance during training and, more importantly, that the improvement carries over to surgical tasks the trainee has never seen and performs without any assistance. It describes a simulator with eight tasks and four assistance modes that push or pull the trainee's hands toward targets, away from obstacles, along reference trajectories, and inside safe insertion cones. In a study of eight novices split into assisted and unassisted groups, the assisted group scored higher on the training tasks and on three of the four unassisted evaluation tasks after a two-day pause, with the largest median gain on a liver-resection task. The authors take this as evidence that haptic guidance gets integrated into the visuo-haptic motor loop used during teleoperation and continues to shape performance after the assistance is switched off.","feed_headline":"Haptic training lifts later surgery scores, even with help off","feed_subtitle":"In a VR surgical trainer, trainees who got force feedback beat controls on never-practiced, unassisted tasks.","key_machinery":"The machinery is a set of haptic assistance laws that convert distance and angular errors into forces and torques applied to the master manipulators of a surgical teleoperation console. Each law uses a sigmoidal error-mapping function to decide when assistance engages, and combines an elastic component proportional to the error with a viscous component that damps oscillations. Four modes are implemented: trajectory guidance, obstacle avoidance, surface guidance, and insertion guidance. The outcome measure is the performance index $P$, a weighted average of six logged metrics (distance error $D$, angular error $A$, force and torque feedback magnitudes $F$ and $T$, number of drops $M$, and repositioning time fraction $C$), each normalized by the corresponding metric of an expert resident surgeon; higher $P$ is interpreted as better performance.","core_discovery":"The central claim is that force-based haptic assistance during the training phase does not merely improve execution of assisted exercises; it also improves performance on unassisted, never-seen evaluation tasks designed to resemble real surgical procedures. On the final day, after four training days and two days without practice, assisted-group subjects outperformed control subjects on Liver Resection (median +21.54%), Thymectomy (+13.07%), and Suturing (+8.44%), while Nephrectomy showed a small median decrease (−2.71%). The authors attribute this to the integration of haptic guidance into the visuo-haptic motor feedback loop used during teleoperation, so that the benefits persist even when assistance is absent.","pith_inferences":["Beyond the paper: a natural next experiment would test whether fading assistance (reducing force as skill improves) produces larger transfer than constant assistance, because trainees would be forced to internalize the corrections instead of relying on them.","Beyond the paper: the performance index $P$ combines lower-is-better error metrics with feedback magnitudes that are zero for the control group, so re-analysing the data with errors entered as penalties, or with $F$ and $T$ excluded, would show how much of the reported advantage depends on that scoring choice.","Beyond the paper: with eight subjects and a single expert benchmark, the reported percentage gains are likely noisy; a larger pre-registered replication would be needed to estimate the true effect size.","Beyond the paper: because the assisted group trained with help and the control group did not, the higher training scores may partly reflect the assistance doing the work; comparing post-training unassisted performance between groups equated for task exposure would separate genuine skill acquisition from assistance-induced inflation."],"forward_implications":["Trainees who receive force-based assistance during training can be expected to outperform unassisted trainees on surgical tasks they have never practiced, even with assistance switched off.","The transfer benefit is not uniform across tasks: one of the four evaluation tasks showed a slight median decrease, so the effect may depend on the specific skill or task geometry.","Haptic assistance acts as a real-time error-correction strategy, redirecting the instrument toward safer regions, which could reduce errors and invasiveness if applied during clinical teleoperation.","The learning curves of the two groups did not differ significantly, so the advantage shows up in absolute performance level rather than in the rate of improvement."],"supporting_citations":[{"why":"Introduces virtual fixtures, the conceptual origin of the haptic assistance forces used here.","marker":"[4]"},{"why":"Surveys active constraints and virtual fixtures, providing the framework for redirecting the surgeon's motion.","marker":"[5]"},{"why":"Meta-analysis showing VR simulators transfer skills to the operating room, the transfer claim this study extends.","marker":"[10]"},{"why":"Shows that multi-sensory visual and haptic feedback aids robot-assisted surgical training, motivating the assistance design.","marker":"[29]"},{"why":"Prior experimental study of robotic assistance-as-needed during surgical training, the direct predecessor this work builds on.","marker":"[34]"},{"why":"Pilot study of a performance-based adaptive curriculum, which this paper's multi-day protocol extends.","marker":"[35]"},{"why":"Supplies the simulator architecture on which the eight tasks and assistance algorithms are implemented.","marker":"[36]"},{"why":"Provides the vision-assisted virtual-fixture method that inspired the insertion guidance algorithm.","marker":"[37]"},{"why":"Describes the open-source research kit for the surgical robot that provides the motorized manipulators used for haptic feedback.","marker":"[39]"},{"why":"Gives the inverse dynamics model of the manipulators used to generate the assistance forces and torques.","marker":"[40]"}],"fun_headline_variants":["Haptic VR training lifts unassisted surgery scores later","Force feedback in robot surgery training pays off without aid","Haptic aid in VR surgery training helps when unassisted","Surgical robot haptics: training aid helps when turned off","Haptic-assisted surgery training improves unassisted performance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the performance index $P$ in Eq. 18 is a valid higher-is-better measure of surgical skill, because it sums distance and angular errors (where lower is better) together with feedback magnitudes (which are zero for unassisted trainees); if that sign convention is wrong, all reported improvements invert.","fun_headline_variants_meta":{"raw":{"variants":["Haptic VR training lifts unassisted surgery scores later","Force feedback in robot surgery training pays off without aid","Haptic aid in VR surgery training helps when unassisted","Surgical robot haptics: training aid helps when turned off","Haptic-assisted surgery training improves unassisted performance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00265,"raw_usage":{"total_tokens":10070,"prompt_tokens":832,"completion_tokens":9238,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":9156}},"tokens_in":448,"tokens_out":9238,"duration_ms":62131,"temperature":1.0,"reasoning_tokens":9156,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:34:09.701693+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the performance index $P$ with the error metrics $D$ and $A$ entered as penalties rather than positive contributions, and with $F$ and $T$ excluded entirely for the unassisted group; if the assisted group then no longer outperforms the control group on the evaluation tasks, the claimed transfer effect does not survive.","supporting_citations":[{"cited_title":"Virtual fixtures: Perceptual tools for telerobotic ma- nipulation","cited_arxiv_id":null,"evidence_quote":"Introduces virtual fixtures, the conceptual origin of the haptic assistance forces used here."},{"cited_title":"Active constraints/virtual fixtures: A survey","cited_arxiv_id":null,"evidence_quote":"Surveys active constraints and virtual fixtures, providing the framework for redirecting the surgeon's motion."},{"cited_title":"A meta-analysis of the training effectiveness of virtual reality surgical simulators","cited_arxiv_id":null,"evidence_quote":"Meta-analysis showing VR simulators transfer skills to the operating room, the transfer claim this study extends."},{"cited_title":"Multi-sensory guidance and feedback for simulation-based training in robot assisted surgery: A preliminary comparison of visual, haptic, and visuo-haptic","cited_arxiv_id":null,"evidence_quote":"Shows that multi-sensory visual and haptic feedback aids robot-assisted surgical training, motivating the assistance design."},{"cited_title":"Okamura, Andrea Mariani, Edoardo Pelle- grini, Margaret M","cited_arxiv_id":null,"evidence_quote":"Prior experimental study of robotic assistance-as-needed during surgical training, the direct predecessor this work builds on."},{"cited_title":"Design and evaluation of a performance-based adaptive curriculum for robotic surgical training: a pilot study","cited_arxiv_id":null,"evidence_quote":"Pilot study of a performance-based adaptive curriculum, which this paper's multi-day protocol extends."},{"cited_title":"A unity-based da vinci robot simulator for surgical training","cited_arxiv_id":null,"evidence_quote":"Supplies the simulator architecture on which the eight tasks and assistance algorithms are implemented."},{"cited_title":"Oka- mura, and Gregory D","cited_arxiv_id":null,"evidence_quote":"Provides the vision-assisted virtual-fixture method that inspired the insertion guidance algorithm."},{"cited_title":"Fischer, Russell H","cited_arxiv_id":null,"evidence_quote":"Describes the open-source research kit for the surgical robot that provides the motorized manipulators used for haptic feedback."},{"cited_title":"Modelling and identification of the da vinci research kit robotic arms","cited_arxiv_id":null,"evidence_quote":"Gives the inverse dynamics model of the manipulators used to generate the assistance forces and torques."}],"review_version":1}