{"id":"9cf0b53e-932d-4e22-83e9-2442a6b2e713","arxiv_id":"1908.05188","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"An open-source pipeline converts routine clinical MRI and CT scans of brain tumor patients into interactive VR models, with positive but informal feedback from two senior neurosurgeons.","lead":"This paper assembles a virtual reality system for neurosurgical planning from free, open-source tools, then tests it on brain scans from two tumor patients. It reports that the resulting 3D models ran smoothly in a headset and that two senior neurosurgeons judged them accurate and useful.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The clinical value of the VR framework rests on unvalidated manual segmentations; the two-surgeon feedback cannot establish anatomical accuracy.","rationale":"The reader's weakest-assumption analysis points to unvalidated segmentation accuracy, and my reading of Sections 3.2 and 4.4 confirms that this is the most load-bearing gap. The paper delivers a plausible demonstration that open-source components and consumer VR hardware can be assembled into a working visualization pipeline, and the frame-rate results are reasonable evidence for technical feasibility. However, the qualitative claim about usefulness for neurosurgeons depends on the anatomical correctness of the models. Since the critical structures were manually segmented without any independent quantitative validation, the central claim is conditional, exactly as the reader concluded. I also considered the open-source framing: Unity3D is proprietary, so describing the framework as 'based on open-source tools' is imprecise, but that issue affects reproducibility rather than the core anatomical-fidelity assumption, and it does not change the verdict. No change to the reader's conditional verdict is needed.","tokens_in":10209,"tokens_out":3051,"duration_ms":35175,"concrete_test":"Run a blinded inter-rater segmentation study on the same imaging data of P1 and P2: two independent, board-certified neurosurgeons, not involved in the study, each segment the tumor, optic tract, and brainstem nerves from the identical T1/T2/CTA volumes using the same tools. Compute pairwise Dice similarity coefficients and mean/surface Hausdorff distances between the authors' segmentations and each rater, and between the two raters. If the author-vs-rater agreement falls within the inter-rater agreement range (e.g., Dice comparable and boundary error below roughly 2 mm), the accuracy assumption is supported; if it falls outside that range, the claim that the VR models provide reliable additional information is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the framework can provide experienced neurosurgeons with additional information for surgical preparation. That information is only as reliable as the 3D models, and those models inherit the accuracy of the segmentations described in Section 3.2. There, tumors, optic tract, brainstem nerves, and eyes are 'segmented manually using T1 MRA, T2 MRA and T1' in 3D Slicer, while vessels come from semi-automatic region growing and brain structures from FreeSurfer. No ground-truth comparison is reported: no Dice scores, no surface-distance errors, no inter-rater agreement, and no registration to intraoperative findings. The only evaluation in Section 4.4 is a frame-rate measurement and the post-hoc opinion of two senior neurosurgeons who had also suggested interface features. If the segmentations misrepresent the tumor boundary, vessel course, or optic tract, the immersive VR display could mislead rather than help. Because the pipeline is adapted per patient (Fig. 3), the absence of validation also means there is no evidence that the manual steps generalize to other clinical datasets. The paper honestly labels the work preliminary, but the load-bearing condition for its conclusion is anatomical fidelity, and that condition is currently unmeasured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a research framework for virtual reality (VR) neurosurgery built exclusively from open-source tools (3D Slicer, FreeSurfer, FSL, Blender, Unity3D) and consumer hardware (Oculus Rift, GTX 1080 Ti). The framework converts routine clinical MRI and CT data from two tumor patients and one healthy subject into segmented 3D models of brain structures, vessels, tumors, and other relevant anatomy, and presents them in an immersive VR environment with interactive controls. Preliminary evaluation consists of frame-rate measurements (always above 30 FPS) and qualitative feedback from two senior neurosurgeons who reported that the models appeared accurate and useful for surgical preparation. The paper concludes that the framework could provide experienced neurosurgeons with additional information and potentially ease surgical preparation.","tokens_in":10397,"tokens_out":3225,"duration_ms":33033,"significance":"If the framework's anatomical models are accurate, the work would be a valuable open-source, low-cost research platform for VR neurosurgical planning, extending prior work by using routine clinical rather than research-grade imaging data. The demonstration of smooth real-time interaction on consumer hardware is a credible strength, and the paper explicitly positions itself as a preliminary foundation. However, the absence of any quantitative validation of the segmentation accuracy, which the VR models depend on, leaves the clinical usefulness claim unsupported. A structured, blinded evaluation with a clear anatomical ground truth would substantially raise the significance of the work.","major_comments":[{"comment":"The manual segmentation of tumors, optic tract, brainstem nerves, and eyes is treated as ground truth without any validation. No Dice scores, surface-distance errors, inter-rater agreement, or comparison against intraoperative findings are reported. Because the VR models inherit the accuracy of these segmentations, the concluding claim in Section 6 that the framework 'could provide experienced neurosurgeons with additional information' is not supported for the critical anatomical details. The authors should add at least a basic validation, such as a comparison of tumor boundaries against an independent expert segmentation or against surgical observations.","section":"Section 3.2"},{"comment":"The physician evaluation is an unstructured, qualitative assessment by two senior neurosurgeons who had previously suggested interface features (Section 3.4) and who inspected the VR models postoperatively rather than in a blinded planning context. This feedback indicates perceived usefulness but cannot establish the anatomical accuracy of the models, which is the load-bearing condition for the central claim. A structured evaluation (e.g., task-based comparison with 2D planning, independent surgeon ratings, or a quantitative localization task) is needed to support the conclusion.","section":"Section 4.4"},{"comment":"The segmentation pipeline is adapted individually for each patient, so the two patient cases are not replications of a fixed workflow. The paper acknowledges this as work in progress, but as a consequence the results do not yet demonstrate that the framework generalizes to new clinical datasets with different imaging protocols. The conclusion should explicitly state this limitation and avoid implying that the framework is ready as a general pipeline.","section":"Section 3.2, Fig. 3"}],"minor_comments":[{"comment":"In the third paragraph, 'we have started to addressed' should be 'we have started to address'.","section":"Section 1"},{"comment":"'preliminarly' should be 'preliminarily', and 'A NVIDIA GTX 1080 Ti' should be 'An NVIDIA GTX 1080 Ti'.","section":"Section 4.4"},{"comment":"The caption contains a typo: 'asses' should be 'assess'.","section":"Figure 4 caption"},{"comment":"The claim of being 'the first dedicated, research software immersive VR framework' is asserted without a systematic literature review; consider softening to 'to the best of our knowledge'.","section":"Section 2"},{"comment":"The paper provides supplementary videos but does not make the framework source code publicly available; for a paper whose central contribution is an open-source framework, a public repository would considerably strengthen reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an honest preliminary report, and the performance measurements are a useful contribution. However, the missing validation of segmentation accuracy is load-bearing for the clinical value claim, and the physician evaluation is too informal to fill that gap. The revisions needed (adding a validation experiment or substantially softening the conclusion) are achievable within the scope of the paper if the authors have access to additional expert annotation or relevant clinical data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper actually builds the thing it claims: an open-source VR neurosurgery planning pipeline, demonstrated on two tumor patients and one healthy subject, with smooth frame rates on consumer hardware. That alone sets it apart from the many position papers in this space. The framework is a pragmatic integration of FSL, FreeSurfer, 3D Slicer, Blender, Unity, and the Oculus Rift, and the authors are transparent that it is work in progress. The per-patient pipeline differences (Fig. 3) are shown clearly, and the performance measurements are real evidence that the interaction is fluid.\n\nWhat is genuinely new is the specific artifact: a complete workflow that takes routine clinical MRI/CT into an interactive, immersive VR planning environment using only open-source tools and off-the-shelf hardware. Prior work mostly used commercial systems like Dextroscope, ImmersiveTouch, or NeuroTouch, and academic VR visualizations were typically research-grade data. Showing this on routine clinical data, even for just two cases, is a useful proof of concept.\n\nThe soft spots are real and mostly concentrate in one place: segmentation accuracy. Tumors, optic tract, brainstem nerves, and eyes were segmented manually in 3D Slicer with no ground-truth check—no Dice, no surface distance, no inter-rater agreement. Vessels came from semi-automatic region growing, also unvalidated for these patients. The only evaluation of anatomical fidelity is the post-hoc opinion of two senior neurosurgeons who also co-suggested interface features; that is not a substitute for quantitative validation. If the segmentations misrepresent the tumor boundary or the optic tract, the VR display could mislead rather than help. The conclusion that VR could \"reduce time needed for surgery\" goes beyond what the data can support.\n\nTwo smaller issues: the per-patient manual adaptation means there is no evidence the pipeline generalizes to other clinical datasets, and despite the open-source framing, no code or data were released. That limits reproducibility, which is the stated motivation for using open-source tools in the first place.\n\nNone of this kills the paper. The central claim—that a research framework for VR neurosurgery can be built from open-source tools and consumer hardware—holds up. The authors are honest about the preliminary nature of the results. For its intended audience (VR/medical visualization researchers and, secondarily, surgical simulation people), it is a useful reference point.\n\nI would take it seriously as a peer-reviewed systems contribution, but I would expect reviewers to require at least a clear statement that segmentation validation is planned, preferably with some quantitative metric, and a commitment to release the pipeline if the open-source claim is to mean more than \"we used open-source components.\"","headline":"A working open-source VR neurosurgery planning framework with an honest preliminary demo; the missing segmentation validation is the real load-bearing gap.","tokens_in":10936,"tokens_out":1680,"would_cite":false,"duration_ms":19114,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an open-source VR framework can turn routine clinical MRI and CT scans into interactive 3D models for neurosurgical planning, and that this can give experienced neurosurgeons useful additional information before…","keywords":["virtual reality","neurosurgical planning","medical image segmentation","clinical imaging","head-mounted display","open-source software","brain tumor","3D visualization"],"falsifier":"Have independent raters or an automated method re-segment the same two patients' scans and measure how much the tumor, vessel, nerve, and optic-tract boundaries overlap with the paper's models; large disagreement, or a mismatch with intraoperative findings or postoperative imaging, would show that the VR models cannot be trusted for surgical planning.","tokens_in":10029,"feed_emoji":"🧠","tokens_out":7511,"duration_ms":71773,"temperature":0.7,"pith_summary":"The paper sets out to show that a research-grade VR environment for neurosurgical planning can be assembled from open-source software and off-the-shelf consumer hardware, rather than proprietary commercial systems. Using routine clinical MRI and CT scans from one healthy subject and two brain-tumor patients, the authors built interactive three-dimensional models and presented them to two senior neurosurgeons in a head-mounted display. The surgeons judged the models accurate and easy to absorb, and the authors conclude that such a framework could give experienced neurosurgeons additional information and potentially ease surgical preparation. The pilot results support feasibility; the paper positions the framework as a foundation for further research and training tools.","feed_headline":"Open-source VR pipeline turns patient scans into 3D surgery models","feed_subtitle":"Routine MRI and CT of two tumor patients became interactive models that ran smoothly and won over senior neurosurgeons.","key_machinery":"The load-bearing mechanism is a modular imaging-to-VR pipeline. Clinical image files are converted to a common format, co-registered and resliced to a reference scan, then segmented with a mix of partly automatic region growing for skull and vessels, automatic brain tissue parcellation, and manual tracing for pathology such as tumors and cranial nerves. The segmented volumes are meshed, given smoothed surfaces with correct inside/outside orientation, and imported into a VR environment built on a standard game engine, where two controller-based interaction modes let the surgeon move, rotate, scale, toggle layers, and cut cross-sectional planes through the model. Adapting each patient's pipeline to the available sequences is what makes routine clinical data usable; the renderer's stable frame rate is what makes the result usable in practice.","core_discovery":"The central claim is that the existing open-source segmentation pipeline, originally designed for research-grade images of healthy brains, can be adapted to the lower-quality, heterogeneous imaging data acquired in routine clinical care, and the resulting models can be rendered smoothly and inspected interactively in fully immersive VR. The paper demonstrates this on two tumor patients: from routine MRI and CT volumes, the pipeline produced segmented models of skull, vessels, tumor, optic tract, brainstem nerves, eyes, gray and white matter, ventricles, and other structures. Two certified senior neurosurgeons viewed the models postoperatively before seeing the 2D images, reported being surprised by their accuracy, and were able to verify their VR-based spatial understanding against the conventional 2D slices. Frame rates stayed at or above 30 frames per second on a consumer graphics card across all tested model configurations. The authors present this as first evidence that VR presentations of pre-neurosurgical data could provide additional information and potentially ease preparation of surgery.","pith_inferences":["The decisive open question is segmentation fidelity: if the tumor, vessel, and nerve boundaries are not accurate, the immersive display could mislead rather than help, so an independent accuracy benchmark is the logical next test.","A testable extension would be a randomized or crossover study measuring planning time, plan quality, and spatial understanding with VR versus conventional 2D review.","The same pipeline could be pointed at surgical training, letting junior surgeons rehearse approaches on patient-specific anatomy, though the paper only hints at this direction.","Because the two interaction modes were partly suggested by the evaluating surgeons, clinical acceptance may depend on iterative refinement with surgeons rather than generic usability heuristics."],"forward_implications":["Routine clinical scans, not just high-resolution research images, can feed a VR planning environment, so the approach could be applied to normal hospital data.","Neurosurgeons can form a 3D mental model in VR and then check it against 2D slices, suggesting VR could reduce the cognitive load of mental reconstruction.","Because the framework uses open-source tools and consumer hardware, other research groups can reproduce, extend, and customize it without commercial licensing.","Replacing the manual and per-patient-adapted segmentation steps with a fully automated deep-learning pipeline would turn the prototype into a scalable clinical tool, a step the paper explicitly plans.","The positive surgeon feedback motivates a quantitative evaluation of whether VR planning changes surgical decisions or outcomes, which the authors state they are preparing."],"supporting_citations":[{"why":"Supplies the original whole-head segmentation pipeline, built on open-source tools, that is adapted for the clinical patient data.","marker":"[13]"},{"why":"Provides the linear registration method used to align and reslice the clinical volumes to a common reference.","marker":"[19]"},{"why":"Used for automatic cortical surface-based segmentation of the cerebrum in the pipeline.","marker":"[10]"},{"why":"Used for automatic parcellation of the cerebral cortex and related subcortical structures.","marker":"[14]"},{"why":"The image-computing platform used for manual segmentation of tumors, optic tract, brainstem nerves, and eyes.","marker":"[12]"},{"why":"Documents the subject-specific visualization and analysis platform referenced for the manual segmentation workflow.","marker":"[21]"},{"why":"Motivates the approach by showing that 3D visualization helps neurosurgeons detect relevant vessels in related planning tasks.","marker":"[22]"}],"fun_headline_variants":["Open-source VR turns routine scans into 3D surgery models","Senior neurosurgeons validate VR models from routine MRI","Interactive VR of brain anatomy from clinical scans wins over surgeons","Adapted open-source pipeline brings immersive VR to neurosurgery planning","Routine patient scans become VR models neurosurgeons found accurate"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework's usefulness rests on the assumption that the segmented 3D models faithfully match the real anatomy; the manual segmentations of tumors and nerves are treated as ground truth without independent validation, so if they misrepresent the patient, the VR view could mislead rather than assist.","fun_headline_variants_meta":{"raw":{"variants":["Open-source VR turns routine scans into 3D surgery models","Senior neurosurgeons validate VR models from routine MRI","Interactive VR of brain anatomy from clinical scans wins over surgeons","Adapted open-source pipeline brings immersive VR to neurosurgery planning","Routine patient scans become VR models neurosurgeons found accurate"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1455,"prompt_tokens":870,"completion_tokens":585,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":502}},"tokens_in":486,"tokens_out":585,"duration_ms":6588,"temperature":1.0,"reasoning_tokens":502,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:20:17.730790+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent raters or an automated method re-segment the same two patients' scans and measure how much the tumor, vessel, nerve, and optic-tract boundaries overlap with the paper's models; large disagreement, or a mismatch with intraoperative findings or postoperative imaging, would show that the VR models cannot be trusted for surgical planning.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original whole-head segmentation pipeline, built on open-source tools, that is adapted for the clinical patient data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Used for automatic cortical surface-based segmentation of the cerebrum in the pipeline."},{"cited_title":"Fischl, A","cited_arxiv_id":null,"evidence_quote":"Used for automatic parcellation of the cerebral cortex and related subcortical structures."},{"cited_title":"Kikinis, S","cited_arxiv_id":null,"evidence_quote":"Documents the subject-specific visualization and analysis platform referenced for the manual segmentation workflow."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the approach by showing that 3D visualization helps neurosurgeons detect relevant vessels in related planning tasks."}],"review_version":1}