{"id":"1fda1947-5483-4ab8-9dfe-7dee85435dc4","arxiv_id":"2506.22926","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"An XR system couples 2D and 3D medical views with gesture and LLM voice control, and a small user study reports faster task completion and lower workload.","lead":"Researchers built an extended reality system that shows 2D radiology slices next to 3D anatomical models, with hand gestures and voice commands powered by a large language model. In small tests, users finished medical exploration tasks faster and reported lower cognitive load, but the evidence lacks statistical rigor.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"User-study claims lack inferential statistics; with n=5 per group, the observed differences do not establish the central efficiency and workload claims.","rationale":"The reader's stated weakest assumption is the spatial coherence calibration in §3.1, which is a system-correctness precondition: if the bounding-box and percentage-based interpolation misalign VSPs with mesh anatomy, the coordination feature could display mismatched locations. That is a real concern, and the paper indeed provides no quantitative validation of the calibration. However, a calibration failure would make the coordination feature harder to use, so it is not the most load-bearing link in the argument for the headline claim. The more direct vulnerability is the evaluation's inferential gap: even if the calibration is perfect, the study as reported cannot distinguish real effects from chance. With 5 participants per group and no significance tests, the conclusion that the system 'enhances spatial understanding and reduces cognitive load' overstates the evidence. The paper itself signals this in §4.3 ('focusing on descriptive statistics') while §4.4 uses 'confirmed' and 'significantly,' an internal inconsistency. The reader's rationale does mention 'no significance tests' and 'tiny groups,' so we partially agree, but the reader's weakest_assumption field does not identify this as the primary vulnerability. Our recommended verdict remains CONDITIONAL, not REJECT, because the system design is coherent, the paper is honest about the preliminary nature, and the missing statistics are readily repairable in a revision; however, the empirical claims in §4.4 should be softened until the tests are provided.","tokens_in":8851,"tokens_out":5649,"duration_ms":62613,"concrete_test":"Re-analyze Table 1 and Fig. 3 with two-sided Mann-Whitney U or Welch's t-tests for each pairwise comparison (G1-G2, G2-G3, G1-G3), reporting exact p-values, 95% confidence intervals for mean differences, and effect sizes (e.g., Cliff's delta or Hedge's g). If the T3 G3-vs-G2 difference (4.6 s) or the NASA-TLX pairwise differences do not reach p<0.05 (or a corrected threshold accounting for multiple comparisons), the claims that H1-H3 are 'validated' and 'confirmed' must be downgraded to descriptive trends, and the central contribution would need reframing as a preliminary feasibility study rather than demonstrated efficacy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that coordinated 2D-3D visualization plus LLM voice commands improves task efficiency and reduces cognitive load—rests entirely on the between-group comparisons in the 15-participant user study. Section 4.4 states that H1, H2, and H3 are 'validated' and 'confirmed,' and it uses the word 'significantly' for cognitive-load reductions, yet Section 4.3 explicitly says the analysis is 'focusing on descriptive statistics.' No inferential test, confidence interval, or effect-size measure is reported anywhere. With n=5 per group, large descriptive differences can be non-significant: for example, T3 G3-vs-G2 is only 4.6 s (5.6%) with standard deviations near 14 s; a two-sided Mann-Whitney U test would likely not reject the null for that comparison. The NASA-TLX '50-66% reduction' is likewise presented without variance or p-values. The SUS results are internally inconsistent with the component-wise logic: G2 (coordination only) scores lower (74.5) than G1 control (80.5), yet H3 is claimed supported. Without inferential statistics, the observed pattern could easily be sampling noise, and the central claim is not empirically established. This is load-bearing because H1-H3 are the only quantitative evidence for the system's effectiveness; if these tests do not survive significance testing, the conclusions in §4.4 and §5 are unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an XR system for volumetric medical data visualization that combines a coordinated 2D-3D visualization module (Multi-layered Multi-planar Reconstruction with 3D mesh models, synchronized via Volumetric Slice Planes and a bidirectional spatial coherence algorithm) with a multimodal interaction framework (hand gestures plus LLM-enabled voice commands). The evaluation consists of a 15-participant user study comparing three configurations (control, coordination-only, full features) and semi-structured interviews with three experts. The authors report improved task completion times, higher SUS scores for the full-featured group, and reduced NASA-TLX workload, and they conclude that all three hypotheses H1-H3 are validated. The paper also provides a demo application and supplemental materials.","tokens_in":9148,"tokens_out":3622,"duration_ms":40831,"significance":"If the reported results were statistically established, the system would be a useful contribution to XR-based medical visualization and education. The coordinated 2D-3D visualization addresses a real problem in correlating radiological slices with 3D anatomy, the LLM-based voice interaction is a timely and potentially flexible interaction choice, and the public availability of demo and supplemental materials is a concrete strength. The system description is detailed and generally clear, and the authors are appropriately cautious in the conclusion where they call for larger-scale and long-term evaluations. However, the central efficiency and workload claims rest on a descriptive-only comparison of three groups of five participants each, and the manuscript uses language such as \"validated,\" \"confirmed,\" and \"significantly lower\" that is not supported by the reported data. The current evidence is promising but preliminary, so the central claims need revision or additional analysis before the paper can be accepted.","major_comments":[{"comment":"The manuscript claims in §4.4 that H1, H2, and H3 are \"validated\" and \"confirmed,\" and that NASA-TLX shows \"significantly lower cognitive load\" (a 50-66% reduction), but §4.3 explicitly states that the analysis focuses on descriptive statistics, and no inferential test, confidence interval, or effect size is reported anywhere. With n=5 per group, the observed differences may be well within sampling variation; for example, the NASA-TLX overall workload for G3 (M=12.67, SD=14.09) and G2 (M=25.5, SD=9.82) have heavily overlapping standard deviations, and the T3 G3-vs-G2 difference is only 4.6 seconds (5.6%) against standard deviations near 14 seconds. This is a load-bearing issue because H1-H3 are the only quantitative evidence for the paper's central claim that the system improves efficiency and reduces workload.","section":"§4.3 and §4.4"},{"comment":"The spatial coherence algorithm is central to the coordinated 2D-3D visualization, yet the paper does not provide any validation or error analysis of the automatic calibration procedure. The calibration aligns the mesh model's bounding box to the segmentation volume's voxel space using percentage-based interpolation, and it requires pre-segmented anatomical structures; without quantitative evidence that this alignment is accurate, a mismatch between the 2D slices and 3D models would directly undermine any measured improvement in spatial comprehension, making the observed task-time differences an artifact of misalignment rather than evidence for the system's design.","section":"§3.1, Spatial Coherence"},{"comment":"The SUS results are internally inconsistent with the claim that the coordinated 2D-3D feature improves usability. §4.3 reports that the coordination-only group G2 has a lower mean SUS score (74.5) than the control group G1 (80.5), yet §4.4 uses the high SUS of G3 (87.0) as evidence for H3, which concerns the integrated system. Because the full-featured G3 condition adds both coordination and voice commands relative to the control, the G3 advantage cannot be attributed specifically to the 2D-3D coordination, and the G2-vs-G1 comparison actually contradicts that attribution; a statistical analysis separating the factors is needed.","section":"§4.3 and §4.4, SUS results"},{"comment":"The claim in §4.3 and §4.4 that voice commands become more effective when combined with coordinated 2D-3D visualization is not supported by the reported data. For T3, G3 (M=78.0s) is only 5.6% faster than G2 (M=82.6s), a difference of 4.6 seconds, and no variability information or inferential test is provided for this specific comparison. The authors should either report a statistical test or explicitly present the result as a descriptive trend rather than a confirmed hypothesis.","section":"§4.3 and §4.4, Task T3"}],"minor_comments":[{"comment":"The abstract and §5 should be aligned: the abstract states 'comprehensive evaluation' while §5 appropriately calls the user study 'preliminary user studies with 15 participants.' Please revise the abstract to use 'preliminary' or 'exploratory' consistently.","section":"§5 and Abstract"},{"comment":"The y-axis label 'NASA T ask Load Index' contains a typographical error; it should read 'NASA Task Load Index.'","section":"Figure 3"},{"comment":"The participant demographics are heavily unbalanced (13 males, 2 females), and no analysis of gender or prior XR experience is reported; at minimum, the authors should acknowledge this as a limitation in the conclusions.","section":"§4.2"},{"comment":"The task completion time table reports means and standard deviations but no confidence intervals; adding 95% confidence intervals for the group differences would make the variability visible and help readers calibrate the strength of the descriptive trends.","section":"§4.3, Table 1"},{"comment":"The paper references 'the supplemental material' for algorithmic details of the spatial coherence calibration and synchronization procedures; since the supplemental material is not included with the arXiv submission, the authors should ensure it is available at the OSF link and clearly describe the calibration accuracy there.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper on coordinated 2D-3D volumetric visualization in XR. The genuinely new piece is the LLM-driven voice command parsing combined with a coordinated 2D-3D visualization; the rest is a sensible integration of existing techniques. It is not a conceptual breakthrough, but it is a real working system with a coherent architecture: MVVM-based synchronization, bidirectional world-volume mapping, and automatic calibration. The demo and materials are on OSF, which is more than many systems papers do. The self-citations to their own CVHSlicer 2.0 are appropriate context, not the source of the claims. The writing is clear, and the conclusion is appropriately cautious about scale.\n\nThe problem is the evaluation section. Fifteen participants, n=5 per group, and no inferential statistics anywhere. Section 4.3 explicitly says the analysis is descriptive; Section 4.4 then says H2 is 'confirmed' and claims 'significantly lower cognitive load.' The NASA-TLX numbers illustrate the gap: G3 mean 12.67 with SD 14.09, G2 25.5 with SD 9.82, G1 37.5 with SD 21.10. Those distributions overlap substantially. With n=5, a 5.6% T3 difference between G3 and G2 (78.0 vs 82.6) is almost certainly not significant. The SUS outcome is internally awkward: G2 scored lower (74.5) than the control (80.5), yet H3 is claimed supported. Either way, the observed pattern could be sampling noise. That matters because H1-H3 are the only quantitative evidence the system helps.\n\nThe spatial coherence algorithm relies on pre-segmented structures and bounding-box-based calibration; if that alignment is off, the coordination display would mislead. There is no external baseline and no isolated test of the LLM component, so the specific contribution of the voice-command pipeline is not identified.\n\nWhat the paper does well: the system is real, the architecture is sensible, the expert interviews add qualitative context, and the authors are honest that larger long-term evaluations are needed. But the central performance and workload claims are not empirically established by the reported data.\n\nFor a serious venue, I would send it to peer review with a clear expectation of major revision: add inferential statistics (or drop the significance language), report confidence intervals and effect sizes, add an external baseline or a within-subjects control for the voice component, and reword the conclusions to match what descriptive data can support. For your own work, it is worth citing as an example of an applied XR medical visualization system with open materials, but not as evidence of effectiveness.","headline":"A real XR medical visualization system with a sensible design, but with n=5 per group and descriptive statistics only, the efficiency and workload claims are not supported.","tokens_in":9705,"tokens_out":3985,"would_cite":true,"duration_ms":35491,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an XR system synchronizing 2D radiological slices with 3D anatomical models, plus LLM-driven voice commands, makes volumetric medical data exploration faster and less mentally demanding for non-experts.","keywords":["extended reality","volumetric medical imaging","multi-planar reconstruction","2D-3D coordination","multimodal interaction","LLM-enabled voice commands","cognitive load","spatial comprehension"],"falsifier":"Load a mesh with a deliberately introduced scale or offset error, have users point to a 3D landmark and say which slice index contains it, and measure the mismatch: if the error is more than a couple of voxel slices, the coordination benefit claimed here would not survive without manual recalibration.","tokens_in":8672,"feed_emoji":"🩻","tokens_out":6697,"duration_ms":64874,"temperature":0.7,"pith_summary":"The paper claims that an XR system which synchronizes 2D radiological slices with 3D anatomical models makes volumetric medical data easier for non-experts to understand, and that adding large-language-model (LLM) enabled voice commands makes exploration faster and less mentally demanding. The system unites Multi-layered Multi-planar Reconstruction (MLMPR) with 3D mesh models through a bidirectional spatial-coherence mapping, and combines hand gestures with natural-language voice control. In a 15-participant study, coordinated 2D-3D views sped structure identification by 28-41% and the full system was 33.4% faster overall than a basic-UI control, with NASA-TLX workload 50-66% lower. Expert interviews add qualitative support for the design's value in medical education and clinical review. The authors present these as preliminary results that justify larger-scale validation.","feed_headline":"Coordinated 2D-3D and voice control cut XR medical task time 41%","feed_subtitle":"Non-experts finished anatomy tasks faster and reported lower cognitive load.","key_machinery":"The load-bearing mechanism is the bidirectional spatial-coherence algorithm, with world-to-volume mapping $f_{w2v}: \\mathbb{R}^3 \\to \\mathbb{Z}^3$ and volume-to-world mapping $f_{v2w}: \\mathbb{Z}^3 \\to \\mathbb{R}^3$, which keeps the Volumetric Slice Planes (VSPs) and the 3D mesh registered to the same voxel grid. An automatic calibration aligns the bounding boxes of a pre-segmented volume and a mesh model using percentage-based interpolation, so the mapping works without manual coordinate alignment. A reactive MVVM state layer propagates any change from 2D UI, direct 3D manipulation, or voice commands to all views, which is what turns the spatial mapping into synchronized visualization and enables the efficiency gains reported.","core_discovery":"The paper's central claim is that spatial coherence between 2D slices and 3D models, plus natural-language voice control, materially improves non-expert exploration of volumetric medical data in XR. It reports that participants using the coordinated 2D-3D interface were 10% faster at describing spatial relationships among structures and 28-41% faster at matching structures across 2D and 3D views than a control group using a basic UI. Adding LLM-enabled voice commands made structure identification another 18.9% faster and volumetric exploration 5.6% faster than coordination alone, and lowered reported workload by 50-66% on the NASA-TLX. Full-system users also gave a mean System Usability Scale score of 87.0, in the 'excellent' range. The intended consequence is that medical education and clinical review can use this design to reduce the steep spatial-reasoning curve normally associated with reading radiological volumes.","pith_inferences":["An unstated consequence is that the same coordination pipeline could apply to other imaging modalities such as MRI or ultrasound volumes, provided segmentation exists, with the bounding-box calibration becoming the limiting factor.","A testable extension is to compare LLM voice commands against a fixed keyword-command set to isolate whether the benefit comes from natural-language understanding or from voice input itself; the current study does not separate these.","If the workload reduction generalizes beyond the 15 participants tested, the design could transfer to surgical planning and patient consultation, where the interviewed experts saw potential but also requested specialized tools."],"forward_implications":["Users without medical training can locate and name anatomical structures across 2D slices and 3D models faster when the views are synchronized in XR.","Voice commands driven by an LLM reduce the need to memorize interface syntax, and the speed gain appears on top of coordinated visualization rather than replacing it.","Combining coordination with voice input lowers subjective workload by a large margin (50-66% on NASA-TLX) while raising usability into the 'excellent' SUS range.","A stable frame rate of 63-83 FPS at about 1.27 GB memory suggests the approach is feasible on current standalone XR hardware for real-time medical tasks."],"supporting_citations":[{"why":"Supplies the 3D-IRCADb-02 liver dataset used in the user study, grounding the spatial-coherence and interaction results in real volumetric medical data.","marker":"[24]"},{"why":"Provides systematic-review evidence that immersive VR and AR improves anatomy education, motivating the claim that XR coordination can help non-experts.","marker":"[7]"},{"why":"Shows that combining gesture and voice in VR improves manipulation efficiency and lowers cognitive load, the prior result the multimodal interaction design extends.","marker":"[6]"},{"why":"Documents the role of spatial abilities and mental rotation in anatomy learning, motivating the spatial-comprehension task and the non-expert target.","marker":"[8]"},{"why":"Demonstrates a hybrid VR-desktop system for cardiac anatomy that supports spatial understanding, a direct precedent for combining 3D models with radiological views.","marker":"[11]"},{"why":"Supplies the adjective rating scale used to interpret the SUS score of 87.0 as 'excellent'.","marker":"[1]"}],"fun_headline_variants":["XR 2D-3D plus voice: 41% faster anatomy tasks","Voice-guided XR cuts medical task time 41%","Coordinated 2D-3D and LLM voice: 41% faster","XR tool cuts task time 41%, workload 66% for novices","Non-experts 41% faster in XR with 2D-3D + voice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The coordinated views work only if anatomical structures are pre-segmented and if the automatic bounding-box calibration aligns the mesh model and the volume's voxel grid accurately; if that alignment is off, the claimed spatial-understanding gains would be artifacts of mismatched displays.","fun_headline_variants_meta":{"raw":{"variants":["XR 2D-3D plus voice: 41% faster anatomy tasks","Voice-guided XR cuts medical task time 41%","Coordinated 2D-3D and LLM voice: 41% faster","XR tool cuts task time 41%, workload 66% for novices","Non-experts 41% faster in XR with 2D-3D + voice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000742,"raw_usage":{"total_tokens":3292,"prompt_tokens":907,"completion_tokens":2385,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":2281}},"tokens_in":523,"tokens_out":2385,"duration_ms":15725,"temperature":1.0,"reasoning_tokens":2281,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:54:57.916322+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Load a mesh with a deliberately introduced scale or offset error, have users point to a 3D landmark and say which slice index contains it, and measure the mismatch: if the error is more than a couple of voxel slices, the coordination benefit claimed here would not survive without manual recalibration.","supporting_citations":[{"cited_title":"Soler, A","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D-IRCADb-02 liver dataset used in the user study, grounding the spatial-coherence and interaction results in real volumetric medical data."},{"cited_title":"Garc ´ıa-Robles, I","cited_arxiv_id":null,"evidence_quote":"Provides systematic-review evidence that immersive VR and AR improves anatomy education, motivating the claim that XR coordination can help non-experts."},{"cited_title":"Friedrich, S","cited_arxiv_id":null,"evidence_quote":"Shows that combining gesture and voice in VR improves manipulation efficiency and lowers cognitive load, the prior result the multimodal interaction design extends."},{"cited_title":"Guillot, S","cited_arxiv_id":null,"evidence_quote":"Documents the role of spatial abilities and mental rotation in anatomy learning, motivating the spatial-comprehension task and the non-expert target."},{"cited_title":"Huang, J","cited_arxiv_id":null,"evidence_quote":"Demonstrates a hybrid VR-desktop system for cardiac anatomy that supports spatial understanding, a direct precedent for combining 3D models with radiological views."},{"cited_title":"Bangor, P","cited_arxiv_id":null,"evidence_quote":"Supplies the adjective rating scale used to interpret the SUS score of 87.0 as 'excellent'."}],"review_version":1}