{"id":"4e907004-a27d-4cb4-965f-858e2cc29ed4","arxiv_id":"2506.04972","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Fourteen clinicians rated Siemens' Cinematic Reality on the Apple Vision Pro as good to excellent in usability, while requesting segmentation, measurement, and annotation tools for clinical adoption.","lead":"This study put Siemens' Cinematic Reality app for the Apple Vision Pro into the hands of 14 medical experts, who rated its usability and described how they would use it for liver surgery planning. It is an early look at whether photorealistic 3D volume rendering on a consumer headset is ready for clinical workflows.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'confirms potential' conclusion extrapolates from perceived usability to clinical interpretive benefit without measuring task performance or diagnostic accuracy; the key missing link is an objective comparison to standard 2D reading.","rationale":"The reader's conditional verdict already identified the load-bearing weakness: self-reported usability is used as a proxy for clinical workflow benefit. My stress-test confirms that this is the correct point of maximum leverage. The paper is transparent, uses validated questionnaires (SUS, ISONORM), and describes its methods and dataset preparation clearly enough to be reproduced; it also wisely notes that direct clinical deployment was outside scope. However, the concluding sentence overclaims by using the word 'confirms' for clinical potential when the evidence consists only of subjective ratings from 14 experts after a short, unstructured interaction. The absence of a diagnostic task, baseline condition, and outcome measure is not a minor limitation for this claim because 'enhance medical imaging interpretation' is an assertion about interpretive performance, not merely about satisfaction. A usability score cannot settle that assertion. The proposed crossover study would directly test whether the perceived advantages survive an objective reading task and would provide the missing evidence base for moving from a conditional to a more definitive verdict. No ad hominem or methodological fraud is implied; the issue is the logical gap between the data collected and the strength of the conclusion drawn.","tokens_in":8977,"tokens_out":3646,"duration_ms":45535,"concrete_test":"Perform a randomized within-subjects crossover reading study using the same CHAOS and MRCP_DLRecon cases: each of at least 14 clinicians reviews a matched set of cases once on Siemens CR on AVP and once on a standard 2D PACS/radiology workstation, with a washout period and counterbalanced order. For each case, record diagnostic accuracy on predefined findings (e.g., portal vein variant, biliary stricture), time to completion, and confidence; administer SUS after both conditions. If CR-on-AVP is non-inferior on accuracy and time within a prespecified margin and the one-sided lower 95% confidence bound for mean SUS exceeds 68, the 'enhance interpretation' claim is supported; otherwise the conclusion in Section 4 should be weakened to 'perceived usability warrants further study.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sections 3.2-3.4 present SUS, ISONORM, and qualitative feedback as the only evidence for the concluding claim in Section 4 that the system 'confirms the potential' of immersive 3D cinematic rendering to 'enhance medical imaging interpretation.' That inference is load-bearing: no objective outcome related to interpretation is measured. Participants interacted with pre-rendered scenes for roughly 15-20 minutes and then rated usability; they were never asked to perform a diagnostic or planning task, there was no 2D PACS baseline, and no accuracy, time, or decision-quality metric was recorded. The n=14 sample also has wide spread: the SUS IQR is 63.75-91.25, so the lower quartile lies below the 68 'above average' benchmark, and the mean of 77.68 does not by itself establish that most raters found the system usable for clinical work. The paper explicitly disclaims direct clinical deployment in Section 1, yet Section 4 converts 'anticipated integration' feedback into confirmation of clinical potential. Unless usability correlates with interpretive performance, the central claim is only a statement about subjective preference, not about enhanced interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a usability evaluation of Siemens' Cinematic Reality application on the Apple Vision Pro for cinematic volume rendering of liver CT and MRCP data. Fourteen medical experts (11 surgeons, one assistant, two medical students) used pre-rendered scenes from the CHAOS and MRCP_DLRecon datasets and completed the System Usability Scale, the ISONORM 9242-110-S questionnaire, and an open-ended survey. The paper reports a mean SUS of 77.68, positive ISONORM subscale scores, and qualitative feedback identifying strengths, limitations, and requested features for clinical adoption. The authors conclude that the study 'confirms the potential' of immersive 3D cinematic rendering on the AVP to enhance medical imaging interpretation.","tokens_in":9333,"tokens_out":4571,"duration_ms":53770,"significance":"If read strictly as an early usability assessment, the study has clear value: it applies validated instruments with published norms, uses public medical datasets, reports demographic and vision-correction details, and collects concrete workflow-relevant feature feedback from a clinically relevant population. The strongest contributions are the feasibility data and the feature roadmap for CR on the AVP in hepatobiliary contexts. However, the significance for clinical adoption is limited because no objective interpretation task, no 2D baseline, and no decision-quality metric were measured; the headline conclusion goes beyond what the data can establish.","major_comments":[{"comment":"The concluding sentence in Section 4 states that the study 'confirms the potential' of immersive 3D cinematic 3DVR on the AVP 'to enhance medical imaging interpretation.' The evidence in Sections 3.2-3.4 consists of self-reported usability scores, ISONORM ratings, and qualitative statements about anticipated use; no participant performed a diagnostic or planning task, no accuracy, time, or decision-quality metric was recorded, and there was no comparator condition. The qualitative data themselves flag that intraoperative and interventional feasibility was 'hard to assess' (Table 2). The data therefore support a statement about perceived usability and anticipated integration, not about actual interpretive benefit. This inferential gap is load-bearing for the paper's central claim; the conclusion should be rephrased to 'supports the potential' with an explicit limitation, or the authors should add an objective task-based comparison to standard 2D reading.","section":"Section 4 (Conclusion) and Sections 3.2-3.4"},{"comment":"The fixed presentation order (CHAOS CT followed by MRCP, then an optional demo) with no baseline condition, combined with only two participants having prior HMD experience, makes it impossible to separate ratings of the CR application from novelty effects or order effects. Since the central claim relies on the high aggregate SUS and ISONORM scores, this design issue is material. Future work should counterbalance display order and include a 2D reading baseline; at minimum, the manuscript should be revised to state this limitation and soften the causal language in the conclusion.","section":"Section 2.5 (Procedure) and Section 3.2"},{"comment":"The reporting of SUS results does not adequately convey spread. The mean of 77.68 and IQR of 63.75-91.25 imply that at least a quarter of the 14 participants scored below the 68 'above average' threshold, and the sample mixes surgeons, an assistant, and students. The text's 'between good and excellent' characterization therefore overstates the consistency of the ratings. The authors should report the full distribution (for example, the proportion of scores below 68, the median, and the range) and temper the aggregate claims accordingly.","section":"Section 3.2 (SUS results)"}],"minor_comments":[{"comment":"The notation is inconsistent: 'σ = 77.68 (σ = 15.01)' uses the same symbol for the mean and the standard deviation; use μ and SD instead.","section":"Section 3.2"},{"comment":"The text says the SUS results are shown in 'Figure 2,' but the SUS scores appear in Figure 3; Figure 2 is the interaction guide.","section":"Section 3.2"},{"comment":"In the ISONORM list, 'learnability(µ5.5' is missing an equals sign, and the formatting of the statistics is inconsistent across subscales.","section":"Section 3.3"},{"comment":"The phrase 'doctor assistant' should likely be 'physician assistant' or 'doctor's assistant,' and the capitalized 'T ask' is a typo.","section":"Section 2.5"},{"comment":"The text 'Python (version 12)' should be corrected to a valid version identifier such as 'Python 3.12'.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper does not include a conflict-of-interest statement even though two authors are affiliated with Siemens Healthineers and the study evaluates a Siemens product with 'full access' to that product. Please ask the authors to add a COI/funding statement. The manuscript fits the scope of cs.HC as an early usability study, but the conclusion needs the revision described in the main report before it can be considered for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful bottom line: this is the only published usability evaluation I know of for Siemens' Cinematic Reality on the Apple Vision Pro, and it is done carefully enough to be worth a referee's time. It uses validated instruments (SUS, ISONORM 9242-110-S), real retrospective patient data from public datasets (CHAOS and MRCP_DLRecon), and reports the scene preparation pipeline in enough detail that someone could reproduce it. That kind of reproducibility is rare in early XR medical reports, so credit is due. The results are what they say: 14 medical experts, mostly surgeons, interacted with pre-rendered liver CT and MRCP scenes for roughly 15-20 minutes and gave usability ratings. The mean SUS of 77.68 and the ISONORM subscale means are internally consistent, and the qualitative feedback identifies genuinely useful feature requests (measurement tools, segmentation, EMR integration, voice activation). The paper is honest about scope in the intro, stating that direct clinical deployment was beyond the study. The soft spot is the conclusion. The final section states the study 'confirms the potential of immersive 3D cinematic 3DVR on the AVP to enhance medical imaging interpretation.' That is a step beyond the evidence. No diagnostic or planning task was measured, there was no 2D PACS baseline, and no accuracy or decision-quality metric. What the study actually shows is perceived usability and anticipated integration potential. Some participants' SUS scores fell below the 68 'above average' benchmark (IQR lower quartile 63.75), so even the usability claim is not unanimous. The lack of baseline, single-site sample, fixed presentation order, and minimal HMD experience are all real limitations, though typical for a first evaluation. None of these are fatal for an early usability study; the conclusion wording is the main fix needed. The citation pattern looks fine. The single self-citation [15] is a prior Vision Pro survey that does not feed into these results. The public datasets and version numbers are reported, which is good practice. Who gets value: researchers tracking XR for medical imaging, clinicians considering HMD-based 3D rendering, and product teams at Siemens or Apple. I would send it to peer review, not desk reject it. A serious referee should ask for a reworded conclusion and probably a short limitations paragraph, but the empirical core is solid enough to publish. It is an early usability result, not clinical evidence.","headline":"A careful, reproducible first usability evaluation of Siemens' Cinematic Reality on the Apple Vision Pro; the conclusion overreaches by claiming confirmed interpretive benefit when only perceived usability was measured.","tokens_in":9715,"tokens_out":3242,"would_cite":true,"duration_ms":35401,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fourteen medical experts rate Siemens' Cinematic Reality on the Apple Vision Pro as good-to-excellent in usability and see its immersive 3D rendering of CT and MRI volumes as a promising aid for surgical planning and education.","keywords":["Cinematic Rendering","Apple Vision Pro","Volume Rendering","Usability Evaluation","System Usability Scale","ISONORM 9241-110-S","Surgical Planning","Extended Reality"],"falsifier":"A controlled comparison in which surgeons perform a fixed planning task—for example, identifying portal vein variants or measuring tumor-to-vessel distances—in the Cinematic Reality headset versus standard 2D slice viewing, checking accuracy and time; if the immersive condition shows no improvement, or if the SUS scores from a larger multi-centre sample fall below the good-excellent range, the paper's clinical-potential claim would be weakened.","tokens_in":8832,"feed_emoji":"🩻","tokens_out":4703,"duration_ms":47321,"temperature":0.7,"pith_summary":"This paper reports a usability evaluation of Siemens' Cinematic Reality application running on the Apple Vision Pro, a head-mounted display that renders CT and MRI volumes as photorealistic 3D scenes. Fourteen medical experts, mostly liver surgeons, explored venous-phase liver CT and MRCP scans and then rated the system with standard usability instruments. The mean System Usability Scale score was 77.68, which the standard interpretation places between good and excellent, and the ISONORM subscales were positive, with the highest marks for controllability and suitability and the most room for improvement in self-descriptiveness and customizability. The authors conclude from this expert feedback that immersive cinematic 3D volume rendering on the AVP has the potential to enhance medical imaging interpretation, especially for surgical planning and education, and they list the missing features clinicians want before routine adoption.","feed_headline":"Surgeons rate cinematic 3D scan viewer on Vision Pro highly usable","feed_subtitle":"Fourteen experts gave a mean SUS of 77.68 and named the features needed before clinical adoption.","key_machinery":"The carrying instrument is the evaluation protocol: the System Usability Scale (a ten-item questionnaire giving a 0-100 score) paired with the ISONORM 9242-110-S (which measures seven ISO 9241-110 interaction principles such as suitability for the task, controllability, and error tolerance) and an open-ended survey. These are applied after experts freely explore two real medical volumes—a portal venous-phase liver CT from the CHAOS dataset and an MRCP scan from the MRCP_DLRecon dataset—rendered through Siemens' Cinematic Reality on the Apple Vision Pro, with eye gaze acting as pointer and finger pinch as selection.","core_discovery":"The paper's central claim is that, in the judgement of the participating surgeons and medical experts, immersive 3D cinematic volume rendering on the Apple Vision Pro is usable enough and clinically promising enough to warrant further investigation and trials for surgical planning, education, and intraoperative spatial recall. Evidence comes from a mixed-methods study: a mean SUS score of 77.68 (SD 15.01) and ISONORM 9242-110-S subscale means ranging from 5.19 to 5.81 on a 7-point scale, with the strongest ratings for controllability and suitability for the task. Open-ended responses identified concrete strengths (intuitive eye-tracking and pinch interaction, high-resolution photorealistic rendering, fast patient-specific reconstruction) and equally concrete gaps (lack of segmentation, annotation, measurement tools, and integration with patient records). The study does not claim the system is ready for clinical deployment, but that its feasibility and usability are established well enough to motivate clinical trials.","pith_inferences":["A direct extension would be a cross-over trial scoring reading accuracy and time on fixed planning tasks against conventional 2D viewing, which would test whether perceived usability translates into measured performance.","A replication with participants who have regular headset experience, or after repeated sessions, would show how much of the positive impression persists beyond first contact with the device.","The two uncorrected participants' reports of reduced clarity flag vision-correction support as a factor worth controlling in future evaluations."],"forward_implications":["A mean SUS of 77.68 places the system in the 'good to excellent' band used for digital health applications, supporting progression to formal clinical trials of cinematic 3D rendering in surgical planning.","Clinician-requested features—segmentation toggles, measurement and annotation tools, and electronic patient-record integration—define a concrete pre-clinical development checklist.","Setup times under five minutes make pre-operative planning sessions feasible, while participants judged intraoperative use as needing clinical study rather than being ruled out.","The same expert ratings suggest educational use cases, such as anatomy teaching and patient information, are the nearest-term viable applications."],"supporting_citations":[{"why":"Supplies the System Usability Scale questionnaire, the primary quantitative usability instrument.","marker":"[25]"},{"why":"Provides the adjective rating scale used to interpret the SUS score as good to excellent.","marker":"[26]"},{"why":"Supplies the ISONORM 9241/110-S questionnaire measuring seven usability principles.","marker":"[29]"},{"why":"Provides the CHAOS dataset CT volume visualized and evaluated by participants.","marker":"[21]"},{"why":"Provides the MRCP_DLRecon dataset MRCP volume visualized and evaluated by participants.","marker":"[23]"},{"why":"Provides the digital health app SUS benchmarking used to judge the mean score relative to other health applications.","marker":"[32]"}],"fun_headline_variants":["Vision Pro cinematic 3D scans earn high usability from experts","Siemens' cinematic reality on Vision Pro: experts say usable","Vision Pro slice viewer rated 77.7 on usability scale","Experts: cinematic 3D on Vision Pro is clinically promising","Siemens' Vision Pro 3D renders usable for clinical planning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on the premise that usability questionnaires and intended-use statements from 14 experts, most from one hospital and with almost no prior headset experience, will match how the system actually performs in real clinical workflow.","fun_headline_variants_meta":{"raw":{"variants":["Vision Pro cinematic 3D scans earn high usability from experts","Siemens' cinematic reality on Vision Pro: experts say usable","Vision Pro slice viewer rated 77.7 on usability scale","Experts: cinematic 3D on Vision Pro is clinically promising","Siemens' Vision Pro 3D renders usable for clinical planning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000647,"raw_usage":{"total_tokens":2925,"prompt_tokens":850,"completion_tokens":2075,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":1986}},"tokens_in":466,"tokens_out":2075,"duration_ms":20561,"temperature":1.0,"reasoning_tokens":1986,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:28:49.841262+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled comparison in which surgeons perform a fixed planning task—for example, identifying portal vein variants or measuring tumor-to-vessel distances—in the Cinematic Reality headset versus standard 2D slice viewing, checking accuracy and time; if the immersive condition shows no improvement, or if the SUS scores from a larger multi-centre sample fall below the good-excellent range, the paper's clinical-potential claim would be weakened.","supporting_citations":[{"cited_title":"Determining what in- dividual SUS scores mean: Adding an adjective rating scale","cited_arxiv_id":null,"evidence_quote":"Provides the adjective rating scale used to interpret the SUS score as good to excellent."},{"cited_title":"The IEEE Inter- national Symposium on Biomedical Imaging (ISBI), Venice, Italy","cited_arxiv_id":null,"evidence_quote":"Provides the CHAOS dataset CT volume visualized and evaluated by participants."},{"cited_title":"Zenodo.url:https://zenodo.org/records/13912092","cited_arxiv_id":null,"evidence_quote":"Provides the MRCP_DLRecon dataset MRCP volume visualized and evaluated by participants."},{"cited_title":"SystemUsabilityScaleBenchmarkingforDigitalHealth Apps: Meta-analysis","cited_arxiv_id":null,"evidence_quote":"Provides the digital health app SUS benchmarking used to judge the mean score relative to other health applications."}],"review_version":1}