{"id":"c54559a2-72f0-49a7-a857-800a5d17b40b","arxiv_id":"2511.02086","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"In live operating rooms, a markerless depth-only HoloLens pipeline registered CT skin models to feet, ear, and lower leg with median surface errors of 3.2-5.3 mm, though the measurement method is not fully independent.","lead":"A surgical augmented-reality system that aligns HoloLens depth images to CT skin models without fiducials was tested on patients undergoing fibula free-flap and mandibular surgery. It reports median overlay errors around 3-4 mm on feet, ear, and lower leg, but its error-measurement method shares the same sensor calibration as the alignment, so the numbers need independent confirmation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Clinical error metric is circular: the same tracked stylus that fits the Sec 2.3.2 depth-bias correction is used to trace overlay and skin, so common-mode stylus error cancels and the 3-4 mm claim may understate true registration error.","rationale":"The reader's weakest assumption identifies precisely the load-bearing issue: the surface-tracing metric is not independent of the depth-bias correction because both rely on the same tracked stylus. I agree with this concern and find it central. The paper's central claim--that a depth-only markerless HoloLens pipeline achieves ~3-4 mm median error in live surgery--rests entirely on this error metric. If the stylus has any systematic tracking error, the depth-bias correction absorbs it, the registration aligns to the biased surface, and the subsequent skin trace shares the same bias, so the measured overlay-to-skin distance is artificially small. The paper's own limitation statement concedes possible biases but does not quantify them or demonstrate independence. The preclinical validation in Sec 3.1 is not sufficient to address the concern because it validates relative skin-to-bone distances, which are translation-invariant and therefore do not constrain common-mode positioning errors. This is not an accusation of misconduct; it is a structural issue in the evaluation design. The proposed external-tracker landmark test would directly measure absolute registration error without using the same stylus for both correction and evaluation, and would settle whether the reported accuracy is genuine. Because the concern is serious but addressable with an independent validation study, the existing CONDITIONAL verdict remains appropriate; I do not recommend changing it.","tokens_in":8996,"tokens_out":7013,"duration_ms":75782,"concrete_test":"Conduct one additional intraoperative or high-fidelity phantom trial of the same pipeline with 4-6 small CT-visible adhesive skin markers placed before the CT scan and left in place. After markerless AR registration, use a separate optical tracker (e.g., NDI Polaris) with a probe that was never used in the Sec 2.3.2 depth-bias correction to localize these markers in the OR frame. Transform the registered CT model into the OR frame using the existing tracker calibration, and compute point-to-point distances between each physical marker and its corresponding CT landmark. If median independent landmark error is close to the reported 3-4 mm, the concern is resolved. If it is substantially larger (e.g., >8 mm), the stylus-based metric was inflated by common-mode stylus bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central accuracy claim in Table 1 and the abstract's ~3-4 mm median error is measured with a metric that is not independent of the registration. In Sec 2.3.2, a tracked stylus samples the patient's skin and solves a Procrustes transform (Eq. 5) that corrects the AHAT depth cloud before registration. The same stylus is then used in Sec 3.2 to trace both the virtual overlay and the patient's skin, forming nearest-neighbor distances. Any systematic error in the stylus pose is absorbed into the depth correction, shifts the registered overlay by the same amount, and cancels in the overlay-to-skin subtraction. The preclinical validation in Sec 3.1 does not break this coupling: it compares skin-to-bone relative distances, which are invariant to a rigid translation of the traced points, so it cannot detect a common-mode offset. The paper's own Limitations section acknowledges 'small systematic biases from the tool trajectory and from the AHAT depth itself' but does not quantify them. Without an independent absolute reference, the reported accuracy may be substantially better than the true registration error.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a depth-only, markerless AR registration pipeline on HoloLens 2 for surgical guidance, evaluated clinically on feet, ear, and lower leg during live fibula free-flap harvest and mandibular reconstruction. The pipeline comprises per-ROI depth-bias correction (Eq. 5), human-in-the-loop initialization, coarse global alignment (TEASER++), and fine ICP refinement. Preclinical validation shows sub-millimeter agreement between AR-traced skin-to-bone relative distances and CT ground truth. The clinical evaluation reports a pooled median per-point error of 3.9 mm (anatomy-specific medians 3.2–5.3 mm) with 5-mm surface coverage ranging from 72% to 95%, and concludes that these results approach clinical thresholds for moderate-risk tasks without fiducials.","tokens_in":9301,"tokens_out":20887,"duration_ms":234834,"significance":"If the reported accuracy is reliable, the paper provides a meaningful advance: it is among the first clinical evaluations of a markerless HMD-based registration system on multiple anatomies with real patients, and its reported errors are substantially better than the ~10 mm figures of an earlier feasibility study (ARCUS). The depth-bias correction and global-to-local registration design are practical contributions, and the preclinical validation of the tracing metric is a useful methodological step. However, the central accuracy claim rests on a clinical error metric whose absolute reference is not independently verified, and the 'depth-only/markerless' description is qualified by the use of a tracked stylus for depth correction. These issues need to be addressed before the headline accuracy can be fully accepted.","major_comments":[{"comment":"","section":"Sec. 3.2 and Eq. (5)"},{"comment":"","section":"Sec. 2.3.2"},{"comment":"","section":"Sec. 3.2 and Table 1"}],"minor_comments":[{"comment":"","section":"Abstract and Sec. 3.2"},{"comment":"","section":"Sec. 2.3.1, Eq. (3)"},{"comment":"","section":"Sec. 2.3.2"},{"comment":"","section":"Sec. 2.3.5 and Sec. 2.1"},{"comment":"","section":"Sec. 3.2, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about circularity is not fully valid as stated: a constant stylus bias cancels in the overlay-to-skin subtraction, and the reported distance equals the physical overlay displacement. The deeper issue is the absence of an independent absolute reference for the clinical error metric, which the authors themselves acknowledge in the Limitations section. I recommend requiring an independent validation (even on a small subset) to solidify the headline 3–4 mm claim. Also, the 'markerless/depth-only' framing is slightly overstated given the manual stylus-based depth correction step; this should be qualified in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Adam, quick read on arXiv:2511.02086.\n\nThe genuinely new thing here is the data: seven intraoperative trials (feet, ear, lower leg) on a depth-only HoloLens 2 pipeline, with 500+ traced points per trial. No one else has published live-patient accuracy numbers for this kind of fiducial-free, depth-only AR registration. The authors also did a sensible preclinical sanity check: tracing relative skin-to-bone distances with the AR stylus and comparing to CT, and the agreement (sub-mm median difference) is reassuring for their tracing procedures. The paper is honest about its limitations and does not oversell; it frames the result as approaching moderate-risk thresholds, not replacing navigation.\n\nThe soft spot is the one the stress-test flagged, and it does land. The depth-bias correction (Eq. 5) fits a rigid Procrustes transform from AHAT depth points to stylus samples on the patient's skin. The clinical error metric (Sec. 3.2) then traces the overlay and the skin with the same stylus. If the stylus has any systematic pose error, it shifts the corrected depth cloud (and hence the registered overlay) and the skin trace equally, so the measured overlay-to-skin distance is artificially small. The preclinical validation does not break this coupling, because it compares relative distances, which are invariant to a common translation. The paper itself acknowledges 'small systematic biases from the tool trajectory and from the AHAT depth itself' but never quantifies them. Without an independent absolute reference (e.g., an optical tracker fixed in the room, or a known geometry not used in registration), the 3-4 mm median could understate the true error by an unknown amount.\n\nThere is a second, related issue: the pipeline is not fully markerless in the strict sense. The stylus is used to sample the patient's skin to correct depth bias. That is a manual point-based step, which is fine and pragmatic, but it should be reported as such. The 'markerless' claim is a bit strong.\n\nMinor: the statistical comparison (feet vs leg) uses a two-sample permutation test on pooled points, which ignores within-patient/trial correlation. With only 2-3 trials per anatomy, the p-value is overconfident. That's a fixable analysis issue.\n\nBottom line: the clinical dataset is a useful contribution, and the idea is plausible, but the headline accuracy claim is not yet supported by an independent measurement. A serious referee should ask for an external absolute-error validation (even one trial with a tracked reference) and for a revised statistical analysis. I would not desk-reject this; it deserves a chance in revision. It is a good reading-group paper for discussing measurement circularity in clinical AR.","headline":"Live multi-anatomy markerless AR data is a real step forward, but the headline accuracy number rests on a self-referential measurement; treat the 3–4 mm as provisional until an independent absolute reference is used.","tokens_in":9819,"tokens_out":3532,"would_cite":true,"duration_ms":38241,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A depth-only, markerless AR pipeline registered CT models to live surgery with 3–4 mm median error across feet, ear, and lower leg, without fiducials.","keywords":["markerless augmented reality","surgical guidance","HoloLens 2","depth-based registration","point cloud registration","clinical accuracy study","surface tracing","fiducial-free"],"falsifier":"Simultaneously measure the true position of the virtual overlay using an external optical tracker (independent of the HoloLens depth sensor) and compare it to the skin trace used for error evaluation; if the two disagree by substantially more than the reported median error, the claimed ~3–4 mm accuracy is an artifact of common-mode stylus error.","tokens_in":8883,"feed_emoji":"🩺","tokens_out":2367,"duration_ms":26473,"temperature":0.7,"pith_summary":"The paper demonstrates that a markerless, depth-only augmented reality registration system on a head-mounted display can align preoperative CT skin models to small or low-curvature anatomies during real surgery, achieving a pooled median per-point error of 3.9 mm. This approaches the roughly 5 mm error threshold considered acceptable for moderate-risk surgical tasks, while eliminating the workflow burden of fiducial markers. The authors validate their surface-tracing error measurement against CT ground truth in preclinical tests, then report clinical results across seven intraoperative trials on feet, ear, and lower leg. If accurate, this would be a practical step toward routine markerless AR guidance in surgery.","feed_headline":"Markerless AR registration hits 3–4 mm accuracy in live surgery","feed_subtitle":"Depth-only HoloLens pipeline aligns CT skin models to feet, ear, and lower leg without fiducials, approaching clinical error limits.","key_machinery":"The registration pipeline's load-bearing components are: (1) region-specific depth-bias correction using a tracked stylus to solve an orthogonal Procrustes fit between ground-truth surface samples and AHAT depth points; (2) human-in-the-loop ROI initialization, where the user roughly aligns a translucent virtual model to the target anatomy to crop the scene and bound the search region; (3) robust coarse alignment using FPFH features and TEASER++ to handle outliers and large misalignments; and (4) fine point-to-plane ICP with decreasing residual thresholds. The surface-tracing error metric, validated against CT, provides the clinical accuracy numbers.","core_discovery":"The central claim is that a depth-only registration pipeline, combining a brief human-in-the-loop initialization with global (FPFH/TEASER++) and local (point-to-plane ICP) alignment, can achieve clinically relevant accuracy on anatomies that are small or have low curvature. In live surgical settings, the pooled per-point median error was 3.9 mm, with anatomy-specific medians of 3.2 mm (feet), 4.3 mm (ear), and 5.3 mm (lower leg), and 5 mm surface coverage ranging from 72–95%. Preclinical validation showed that stylus-based AR surface tracing reproduces CT-derived skin-to-bone distances with a median discrepancy of about 0.8 mm, supporting the use of this tracing method as an intraoperative a","pith_inferences":["The reported accuracy may understate true overlay displacement if the HoloLens-tracked stylus carries systematic error, since the same stylus is used both for depth-bias correction and for the error measurement; an independent optical tracker would settle this.","The method's dependence on a brief user alignment step means its clinical reliability partly rests on operator skill; automating this step could make the system more robust but may reduce accuracy on symmetric anatomies.","The success on skin surfaces suggests the pipeline might extend to soft-tissue deformation tracking if combined with non-rigid registration, though the current rigid assumption limits that application.","The between-anatomy error differences imply that surfaces with richer curvature (feet) are easier to align, so future work could focus on feature-poor regions like the lower leg with additional constraints (e.g., limb axis priors)."],"forward_implications":["Markerless, fiducial-free AR guidance could meet clinical accuracy thresholds for moderate-risk tasks on small anatomies, reducing OR setup time and invasiveness.","The human-in-the-loop initialization may be generalized to other anatomies where fully automatic global registration is unreliable due to symmetry or low curvature.","The validated surface-tracing metric offers a practical, CT-referenced method for evaluating AR overlay accuracy in situ on patients.","The reported anatomy-specific differences (feet more accurate than leg) suggest that exposure and surface curvature are key factors to optimize in future markerless systems.","If reproduced on more patients, the approach could support visualizing internal structures (fibula, foot bones) directly over the skin without external tracking hardware."],"fun_headline_variants":["Markerless AR on HoloLens hits 3–4 mm in the OR","No-marker AR: 3.9 mm median error in live surgery","Depth-only AR approaches clinical limits in real procedures","HoloLens AR without fiducials: 3–4 mm median accuracy","AR for small anatomies: 3–4 mm error without markers"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The measured accuracy assumes the HoloLens-tracked stylus is an independent ground truth, but the same stylus is used to calibrate the depth-bias correction, so any systematic stylus error would be invisible in the reported overlay-to-skin distances.","fun_headline_variants_meta":{"raw":{"variants":["Markerless AR on HoloLens hits 3–4 mm in the OR","No-marker AR: 3.9 mm median error in live surgery","Depth-only AR approaches clinical limits in real procedures","HoloLens AR without fiducials: 3–4 mm median accuracy","AR for small anatomies: 3–4 mm error without markers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1518,"prompt_tokens":947,"completion_tokens":571,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":691,"completion_tokens_details":{"reasoning_tokens":474}},"tokens_in":691,"tokens_out":571,"duration_ms":6093,"temperature":1.0,"reasoning_tokens":474,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:13:48.002520+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simultaneously measure the true position of the virtual overlay using an external optical tracker (independent of the HoloLens depth sensor) and compare it to the skin trace used for error evaluation; if the two disagree by substantially more than the reported median error, the claimed ~3–4 mm accuracy is an artifact of common-mode stylus error.","supporting_citations":[],"review_version":1}