REVIEW 5 major objections 5 minor 19 references
Towards the target and not beyond: 2D vs 3D visual aids in MR-based neurosurgical simulation
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Trainees who practiced external ventricular drain placement with combined 2D and 3D visual aids placed catheters 44% more accurately in unaided testing than untrained control participants.
desk verdict A novel MR simulator and a promising effect, but the 'unaided' testing phase actually provides visual aids, so the headline claim of skill retention does not hold as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the 2D-3D Aid training interface built into NeuroMix, an MR simulator deployed on a Meta Quest 3 headset. It couples a neuronavigation-style Virtual Panel showing axial, coronal, and sagittal CT slices with a real-time 2D projection of the catheter on those slices, and adds a red Guide Line showing the optimal 3D trajectory plus an animation that moves the virtual catheter from its current orientation onto that line. During training the user manipulates a Digital Catheter against a volume-rendered Digital Skull; during testing a tracked Physical Catheter and Physical Skull are registered with the same digital model. What carries the argument is the idea that the 3D guide, by making the ideal depth and angle visible and by demonstrating the corrective motion, encodes the trajectory in a way that can be recalled when no aid is present.
What would settle it
Rerun the testing phase with the Virtual Panel, the red catheter overlay, and the digital skull registration turned off (or physically occluded) so participants see only the phantom and the bare catheter; if the 2D-3D Aid group's mean Error Score no longer differs significantly from Control's, the claim that combined-aid training improves unaided skill retention would be refuted.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the 2D-3D Aid training modality, which adds a red 3D trajectory guide and a context-aware alignment animation to neuronavigation-style CT-slice projections, produced the only statistically robust retention gains in unaided freehand EVD placement. The mean testing-phase Error Score for this group was 1.47 cm (SD 0.60) against 2.64 cm (SD 1.06) for the untrained control group, a 44.3% reduction, and pairwise comparisons showed it differed significantly from No Aid, 2D Aid, and Control. The effect appeared mainly along the axial (Y) direction, meaning insertion depth, and in tilt error, which was 44.1% lower than control. The authors also report that this extra 3D guidance did not increase perceived cognitive workload, while all modalities scored high on usability and acceptance, and trained groups were slower in freehand testing, which they interpret as an accuracy-first strategy.
Load-bearing premise
The paper's retention claim depends on the testing phase being genuinely unaided, yet the testing system still rendered a digital skull overlay, a red stick on the catheter, and real-time insertion depth on the Virtual Panel, so if any of these helped participants, the measured precision is not purely retained skill.
Editorial extensions
If this is right
- If correct, combined 2D-3D training could let residents practice EVD placement in MR and then perform the procedure freehand at the bedside with better accuracy than no training at all.
- The modality's advantage is concentrated in depth and tilt control, so training systems may need to emphasize those components rather than aiming for global guidance.
- Since modality, not number of attempts, predicted testing precision, simulator design choices matter more than simply increasing repetition volume.
- The absence of extra cognitive workload with 3D aids suggests richer guidance can be added without taxing the learner, which is relevant for novice training.
- The longer freehand times of trained groups imply that speed-accuracy trade-offs should be explicitly trained and measured if fast bedside execution is a goal.
Reading between the lines
- Our inference: the testing phase may not be fully unaided; the Virtual Panel's real-time depth readout and the digital red stick overlaid on the catheter could themselves serve as visual aids, so the 44% figure may partly capture continued MR support rather than pure retention. Testing with those overlays disabled would isolate the training effect.
- Our inference: because the 2D-3D group's improvement was strongest along the depth axis and in tilt, the 3D guide may be teaching a specific trajectory template rather than general anatomical understanding; transfer to a skull with different anatomy or a different entry point is untested.
- Our inference: the single CT dataset and single digital skull limit generalizability; a natural next experiment would vary skull anatomy and entry points to see whether the retention effect survives.
- Our inference: the study measures immediate retention; a delayed retest days or weeks later would show whether the combined-aid advantage is a short-lived rehearsal effect or durable skill.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces NeuroMix, a Mixed Reality simulator for External Ventricular Drain (EVD) placement training, and reports a between-subjects experiment with 48 medical residents (or 49, see major comments). Three training modalities are compared: No Aid, 2D Aid, and 2D-3D Aid, followed by a testing phase performed with a physical catheter and a physical skull phantom, plus a Control group that receives no virtual training. The central claim is that participants trained with combined 2D and 3D aids achieve a 44% improvement in precision during supposedly unaided testing relative to the Control group, based on a mean Error Score of 1.47 cm vs. 2.64 cm (pairwise Mann-Whitney p = 0.00014). The paper additionally reports usability, workload, and acceptance questionnaires, and uses ANCOVA to argue that training modality affects precision beyond the number of practice attempts.
Significance. If the central claim were valid, the study would be a useful contribution to the MR-based surgical training literature: it introduces a low-cost, inside-out tracked simulator, compares three interface modalities, and reports a fairly substantial between-group precision benefit with machine-measured outcomes. The authors also provide a pilot study, transparent statistical reporting in most sections, and an explicit discussion of limitations. However, the headline result depends on the testing phase being genuinely unaided, and the manuscript contains a direct internal contradiction on this point. That issue, together with the Control group's pre-testing CT exposure and several statistical reporting inconsistencies, prevents the current version from supporting the paper's main claim as stated.
major comments (5)
- [3.2, 3.4.2] The testing phase is not actually unaided, and the paper contradicts itself. Section 3.2 states that the testing system renders a registered Digital Skull, a red stick registered to the Physical Catheter, and the insertion depth displayed in real-time on the Virtual Panel 'to support the user throughout the EVD placement'; it then asserts that the MR interface was not used to provide any form of assistance. Section 3.4.2 likewise says 'No visual aids or feedback are provided during this phase' after describing the same system. The real-time depth readout is particularly consequential because only the 2D-3D training condition displayed the ideal depth value on the Virtual Panel (Fig. 5d), so the 2D-3D group could memorize that value and reproduce it with the testing-phase readout, thereby improving Error Score (1.47 vs. 2.64 cm) without necessarily acquiring freehand spatial skill. The headline claim of a 44% improvement in unaided testing therefore requires either removing these aids from the testing phase or reframing the result as performance with continued MR support.
- [3.4.2] The Control group received an asymmetric pre-testing intervention: before performing the five EVD placements, Control participants were allowed to visualize the CT scans on a standard computer monitor using Slicer3D, whereas trained participants proceeded directly to the testing phase. This refresher could narrow the gap between Control and trained groups by giving Control participants target localization information that the trained groups had to retrieve from memory. The authors should either justify this procedure as clinically relevant or re-analyze the data with this exposure as a stated limitation; as written, it is a confound in the primary comparison between trained and untrained groups.
- [4.2.2] Section 4.2.2 reports a Kruskal-Wallis H-test p-value of 0.0092 for Rotation Error in the Testing Phase but then states that 'No significant differences were found for the Rotation component.' A p-value of 0.0092 is below the conventional 0.05 threshold, so either the test was misinterpreted, an adjustment (e.g., Bonferroni) was applied but not reported, or the reported p-value is from a different test. This needs correction and consistent reporting of which comparisons were significant.
- [3.4.3] The participant count is inconsistent: the text reports 48 participants, but the demographic data give 21 females and 28 males, which sums to 49. Because all subsequent group sizes (12 per condition) and statistical analyses depend on the correct N, this discrepancy must be resolved.
- [4.1.1, 4.1.3] The TOST equivalence claims are not fully specified. The paper does not state the equivalence margin used for the two one-sided tests, so the reader cannot judge whether the reported 'significant equivalence' for SUS and TAM is meaningful. The authors should report the pre-specified margin (e.g., a half-standard-deviation or a raw score difference) for each TOST analysis.
minor comments (5)
- [Abstract and throughout] The manuscript contains repeated typos and misspellings, including 'catherer' instead of 'catheter', 'procedere', 'Euclinead Distance', 'Avarage', and the running header 'TOW ARDS THE TARGET AND NOT BEYONDPREPRINT'. These should be corrected.
- [3.4.2] In the Familiarization Phase, the paper says EVD depth is 'generally 5 to 7 cm', but this range is not otherwise tied to the specific digital model or target depth used in the experiment; clarifying the actual ideal depth for the test model would help readers interpret the depth-readout concern in the testing phase.
- [4.2.2] The phrase 'revealed significant differences among the No Aid, 2D Aid, 2D-3D Aid and Control groups for the Tilt component (p=0.0042)' is followed by pairwise results; please clarify whether these are post-hoc comparisons with a correction for multiple testing, since the paper otherwise does not mention any multiple-comparison adjustment.
- [Table 3] The row for 'Train / Test (Rotation)' shows +19.21% for No Aid, which is consistent with the reported means, but the corresponding text in Section 4.2.2 should be checked for consistency with the p-value of 0.0092 as noted in the major comments.
- [7] The Conclusions section repeats the claim about 'unaided testing' without acknowledging the testing-phase visual supports described in Section 3.2; the conclusion should be revised to match the actual test conditions or the test should be changed.
Circularity Check
No derivation-level circularity: the 44% improvement is a measured between-group difference (1.47 vs 2.64 cm), not a fitted or derived quantity; a flagged §3.2/§3.4.2 contradiction over real-time depth feedback during 'unaided' testing threatens construct validity but does not reduce the claim to its inputs by construction.
full rationale
The central claim—'participants trained with both 2D and 3D aids achieve a 44% improvement in precision during unaided testing compared to the control group'—is the ratio of measured group means (testing-phase average Error Score 1.47 cm, SD 0.60, for 2D-3D Aid vs 2.64 cm, SD 1.06, for Control, Table 2; pairwise Mann-Whitney p = 0.00014). No parameter is fitted to a subset of the data and then renamed a prediction: ANCOVA, Kruskal-Wallis, Mann-Whitney U, and TOST are standard comparisons, so no step in the statistical chain is equivalent to its inputs by construction. The paper's self-citations (Hajahmadi et al. 2024; Salomoni et al. 2017) include present co-authors but support only background claims about visual-aid roles and interface familiarity, not the retention comparison; no uniqueness theorem or ansatz is imported from the authors' prior work. The one flagged manuscript-level contradiction is located in Section 3.2 (testing system displays 'the insertion depth... in real-time on the Virtual Panel to support the user throughout the EVD placement', plus a digital red stick registered to the Physical Catheter, added in Section 3.3) versus Section 3.4.2 ('No visual aids or feedback are provided during this phase') and Section 3.2's assertion that 'the MR interface was not used to provide any form of assistance.' Since only the 2D-3D training modality showed the ideal depth value (Section 3.1.2) and that group's largest advantage is along the depth (Y) axis (Section 4.2.1), the supposedly unaided test plausibly provided continued MR depth support; this is a construct-validity threat to the headline claim, not a circular derivation. The TOST equivalence statements also omit prespecified bounds, a reporting concern rather than circularity. Overall, the comparison is self-contained data analysis, so circularity is minimal.
Assumptions & free parameters
free parameters (1)
- TOST equivalence margin
assumptions (4)
- domain assumption Controller-based tracking and digital-to-physical registration accurately measure catheter tip position relative to the anatomical target.
- domain assumption The agar-filled phantom skull provides a valid physical proxy for EVD placement.
- domain assumption Performance measured immediately after training reflects skill retention.
- domain assumption A single digital skull and CT dataset supports conclusions about training effectiveness.
Cite this review
Pith. "Pith review of Towards the target and not beyond: 2D vs 3D visual aids in MR-based neurosurgical simulation." pith.science (2026). https://pith.science/paper/S243PILP
@misc{pith2026250605164,
author = {Pith},
title = {Pith review of: Towards the target and not beyond: 2D vs 3D visual aids in MR-based neurosurgical simulation},
year = {2026},
howpublished = {\url{https://pith.science/paper/S243PILP}},
note = {Machine review of arXiv:2506.05164}
}
read the original abstract
Neurosurgery increasingly uses Mixed Reality (MR) technologies for intraoperative assistance. The greatest challenge in this area is mentally reconstructing complex 3D anatomical structures from 2D slices with millimetric precision, which is required in procedures like External Ventricular Drain (EVD) placement. MR technologies have shown great potential in improving surgical performance, however, their limited availability in clinical settings underscores the need for training systems that foster skill retention in unaided conditions. In this paper, we introduce NeuroMix, an MR-based simulator for EVD placement. We conduct a study with 48 participants to assess the impact of 2D and 3D visual aids on usability, cognitive load, technology acceptance, and procedure precision and execution time. Three training modalities are compared: one without visual aids, one with 2D aids only, and one combining both 2D and 3D aids. The training phase takes place entirely on digital objects, followed by a freehand EVD placement testing phase performed with a physical catherer and a physical phantom without MR aids. We then compare the participants performance with that of a control group that does not undergo training. Our findings show that participants trained with both 2D and 3D aids achieve a 44\% improvement in precision during unaided testing compared to the control group, substantially higher than the improvement observed in the other groups. All three training modalities receive high usability and technology acceptance ratings, with significant equivalence across groups. The combination of 2D and 3D visual aids does not significantly increase cognitive workload, though it leads to longer operation times during freehand testing compared to the control group.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[2]
doi:10.3389/fsurg.2023.1241923
ISSN 2296-875X. doi:10.3389/fsurg.2023.1241923. URL https://www.frontiersin.org/articles/10.3389/fsurg.2023. 1241923/full. Ilkay Isikay, Efecan Cekic, Baylar Baylarov, Osman Tunc, and Sahin Hanalioglu. Narrative review of patient-specific 3D visualization and reality technologies in skull base neurosurgery: enhancements in surgical training, planning, and...
-
[4]
ISBN 978-1-68420-504-2. doi:10.1055/b000000751. URL https://medone-neurosurgery. thieme.com/ebooks/cs_20449427#/ebook_cs_20449427_cs28259. Alexander C Flint, Vivek A Rao, Natalie C Renda, Bonnie S Faigeles, Todd E Lasman, and William Sheridan. A simple protocol to prevent external ventricular drain infections.Neurosurgery, 72(6):993–999,
-
[5]
doi:10.1016/j.wneu.2024.05.151
ISSN 18788750. doi:10.1016/j.wneu.2024.05.151. URL https://linkinghub.elsevier.com/retrieve/pii/ S1878875024009136. Hartmut K Gumprecht, Darius C Widenka, and Christianto B Lumenta. Brainlab vectorvision neuronavigation system: technology and clinical experiences in 131 cases.Neurosurgery, 44(1):97–104,
-
[7]
doi:10.1007/s10055-024-01033-9
ISSN 1434-9957. doi:10.1007/s10055-024-01033-9. URL https://link.springer.com/10.1007/s10055-024-01033-9. Frederick Van Gestel, Taylor Frantz, Cédric Vannerom, Anouk Verhellen, Anthony G. Gallagher, Shirley A. Elprama, An Jacobs, Ronald Buyl, Michaël Bruneau, Bart Jansen, Jef Vandemeulebroucke, Thierry Scheerlinck, and Johnny Duerinck. The effect of augme...
-
[10]
doi:10.3171/2023.10.FOCUS23554
ISSN 1092-0684. doi:10.3171/2023.10.FOCUS23554. URL https://thejns.org/view/journals/neurosurg-focus/56/1/ article-pE8.xml. Martin Vychopen, Fabian Kropla, Dirk Winkler, Erdem Güresir, Ronny Grunert, and Johannes Wach. IMAGINER 2—improving accuracy with augmented realIty navigation system during placement of external ventricular drains over Kaufman’s, Kee...
-
[14]
ISSN 1478-5951, 1478-596X. doi:10.1002/rcs.2529. URLhttps://onlinelibrary.wiley.com/doi/10.1002/rcs.2529. Mohamed Benmahdjoub, Abdullah Thabit, Marie-Lise C. Van Veelen, Wiro J. Niessen, Eppo B. Wolvius, and Theo Van Walsum. Evaluation of AR visualization approaches for catheter insertion into the ventricle cavity.IEEE Transactions on Visualization and Co...
-
[15]
ISSN 1077-2626, 1941-0506, 2160-9306. doi:10.1109/TVCG.2023.3247042. URLhttps://ieeexplore.ieee.org/document/10049674/. Simon Skyrman, Marco Lai, Erik Edström, Gustav Burström, Petter Förander, Robert Homan, Flip Kor, Ronald Holthuizen, Benno H. W. Hendriks, Oscar Persson, and Adrian Elmi-Terander. Augmented reality navigation for cranial biopsy and exter...
arXiv 1941
-
[17]
doi:10.1007/s10916-024-02133-4
ISSN 1573-689X. doi:10.1007/s10916-024-02133-4. URLhttps://link.springer.com/10.1007/s10916-024-02133-4. Hung-Jui Guo, Jonathan Z Bakdash, Laura R Marusich, and Balakrishnan Prabhakaran. Augmented reality and mixed reality measurement under different environments: A survey on head-mounted devices.IEEE Transactions on Instrumentation and Measurement, 71:1–15,
Show all 19 references
-
[18]
The factor structure of the system usability scale
18 TOW ARDS THE TARGET AND NOT BEYONDPREPRINT James R Lewis and Jeff Sauro. The factor structure of the system usability scale. InHuman Centered Design: First International Conference, HCD 2009, Held as Part of HCI International 2009, San Diego, CA, USA, July 19-24, 2009 Proce...
2009
-
[684]
URL https://thejns.org/view/journals/neurosurg-focus/51/ 2/article-pE7.xml
doi:10.3171/2021.5.FOCUS20813. URL https://thejns.org/view/journals/neurosurg-focus/51/ 2/article-pE7.xml. Sangjun Eom, Tiffany Ma, Neha Vutakuri, Alexander Du, Zhehan Qu, Joshua Jackson, and Maria Gorlatova. Did you do well? real-time personalized feedback on catheter placeme...
2021 doi
-
[1988]
URL https://www.sciencedirect.com/science/article/pii/S0166411508623869
doi:https://doi.org/10.1016/S0166-4115(08)62386-9. URL https://www.sciencedirect.com/science/article/pii/S0166411508623869. Fred D Davis. Perceived usefulness, perceived ease of use, and user acceptance of information technology.MIS quarterly, pages 319–340,
-
[2013]
doi:10.1097/SIH.0b013e3182662c69
ISSN 1559-2332. doi:10.1097/SIH.0b013e3182662c69. URLhttps://journals.lww.com/01266021-201302000-00006. Kristopher G. Hooten, J. Richard Lister, Gwen Lombard, David E. Lizdas, Samsun Lampotang, Didier A. Rajon, Frank Bova, and Gregory J.A. Murad. Mixed Reality Ventriculostomy ...
-
[2014]
doi:10.1227/NEU.0000000000000503
ISSN 2332-4252. doi:10.1227/NEU.0000000000000503. URLhttps://journals.lww.com/01787389-201412000-00009. Max Schneider, Christian Kunz, Christian Rainer Wirtz, Franziska Mathis-Ullrich, Andrej Pala, and Michal Hlavac. Augmented reality–assisted versus freehand ventriculostomy i...
-
[2019]
doi:10.3171/2018.4.JNS18124
ISSN 0022-3085, 1933-0693. doi:10.3171/2018.4.JNS18124. URL https://thejns.org/view/journals/j-neurosurg/131/ 5/article-p1599.xml. Ronny Grunert, Dirk Winkler, Johannes Wach, Fabian Kropla, Sebastian Scholz, Martin Vychopen, and Erdem Güresir. IMAGINER: improving accuracy with...
1933 doi
-
[2021]
doi:10.3171/2021.5.FOCUS21215
ISSN 1092-0684. doi:10.3171/2021.5.FOCUS21215. URL https: //thejns.org/view/journals/neurosurg-focus/51/2/article-pE8.xml. Sangjun Eom, Tiffany S. Ma, Neha Vutakuri, Tianyi Hu, Aden P. Haskell-Mendoza, David A. W. Sykes, Maria Gorlatova, and Joshua Jackson. Accuracy of routine...
2021 doi
-
[2022]
doi:10.1016/j.wneu.2022.10.002
ISSN 18788750. doi:10.1016/j.wneu.2022.10.002. URL https://linkinghub.elsevier.com/retrieve/pii/ S1878875022014073. Monique J Krabbe-Hartkamp, Jeroen Van der Grond, Frank Erik De Leeuw, Jan Cees De Groot, Ale Algra, Berend Hillen, MM Breteler, and WP Mali. Circle of willis: mo...
2022 doi
-
[2023]
doi:10.1227/ons.0000000000001033
ISSN 2332-4252, 2332-4260. doi:10.1227/ons.0000000000001033. URLhttps://journals.lww.com/10.1227/ons.0000000000001033. 15 TOW ARDS THE TARGET AND NOT BEYONDPREPRINT Kimia Kazemzadeh, Meisam Akhlaghdoust, and Alireza Zali. Advances in artificial intelligence, robotics, aug- men...
-
[2024]
doi:10.3389/fsurg.2024.1427844
ISSN 2296-875X. doi:10.3389/fsurg.2024.1427844. URLhttps://www.frontiersin.org/articles/10.3389/fsurg.2024.1427844/full. Abel J Lungu, Wout Swinkels, Luc Claesen, Puxun Tu, Jan Egger, and Xiaojun Chen. A review on the applications of virtual reality, augmented reality and mixe...
2024
-
[2025]
doi:10.3389/fsurg.2024.1513899
ISSN 2296- 875X. doi:10.3389/fsurg.2024.1513899. URL https://www.frontiersin.org/articles/10.3389/fsurg. 2024.1513899/full. Sangjun Eom, Seijung Kim, Joshua Jackson, David Sykes, Shervin Rahimpour, and Maria Gorlatova. Aug- mented Reality-based Contextual Guidance through Surg...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.