{"id":"75992cc2-96f0-4612-903b-d459c0d96567","arxiv_id":"2503.16309","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"xvr is a self-supervised patient-specific neural network method for rapid 2D/3D rigid registration that pretrains a foundation model on whole-body scans and finetunes per patient in minutes using physics-based simulation for training data.","lead":"The paper introduces xvr, a self-supervised framework that trains patient-specific neural networks on physics-simulated X-rays generated from a patient's own preoperative CT or MRI scan to align 3D volumes with 2D intraoperative fluoroscopy. This could enable faster, more accurate image-guided procedures across many body regions without large labeled datasets or per-patient hyperparameter tuning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Simulation-to-real domain gap remains the unverified assumption enabling claims of real-fluoroscopy accuracy","rationale":"The reader’s weakest_assumption is identical to the load-bearing risk identified above. Because the reader had access only to the abstract, the concern is unchanged by the availability of the full manuscript placeholder; the simulation fidelity step is still the point at which the real-world performance claim could fail without contradicting any other part of the described method.","tokens_in":1760,"tokens_out":339,"duration_ms":32343,"concrete_test":"Select 10–20 real fluoroscopy frames with known ground-truth poses (from the paper’s evaluation set or a phantom study); recompute registration error after (a) training on the paper’s simulated data only and (b) training on real images with the same patient-specific adaptation protocol; if the median TRE difference exceeds 2 mm or the success rate drops below 80 %, the domain-shift assumption materially affects the headline result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of order-of-magnitude accuracy gains on real intraoperative fluoroscopy across anatomies, modalities, and hospitals rests on patient-specific networks trained exclusively via physics-based simulation from preoperative CT/MRI. For this to hold, the forward model must reproduce the intensity statistics, scatter, noise, and geometric distortions of actual C-arm acquisitions sufficiently closely that no large distribution shift occurs at test time. The abstract provides no quantitative evidence (e.g., histogram matching, perceptual metrics, or phantom-based sim-vs-real registration error) that this fidelity was verified, leaving the generalization step as the least secured link in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents xvr, a self-supervised framework for 2D/3D rigid registration that trains patient-specific neural networks on physics-based simulations generated from a patient's preoperative CT/MRI scan. A foundation model pretrained on thousands of whole-body scans enables 5-minute adaptation per patient; the central claim is that this yields high accuracy in seconds on real intraoperative fluoroscopy across diverse anatomies, modalities, and hospitals, representing an order-of-magnitude improvement over prior intensity-based and learning-based methods, with open-source release.","tokens_in":1890,"tokens_out":440,"duration_ms":82742,"significance":"If the central claims hold, the work would be significant for image-guided interventions by removing the need for manual labels or per-subject hyperparameter tuning while achieving pan-anatomical applicability. The combination of patient-specific simulation-based training with a foundation model and the scale of the real-fluoroscopy evaluation are strengths that could broaden access to accurate registration in clinical and research settings.","major_comments":[{"comment":"The abstract and evaluation sections claim an order-of-magnitude accuracy improvement on real fluoroscopy without reporting quantitative verification that the physics-based forward model reproduces real C-arm intensity statistics, scatter, noise, and geometric properties (e.g., no histogram comparisons, perceptual metrics, or phantom-based sim-vs-real registration error). This assumption is load-bearing for the generalization claim.","section":"Abstract; Evaluation"},{"comment":"The results do not specify data exclusion criteria, exact baseline implementations and hyperparameter settings, or statistical tests supporting the cross-hospital and cross-modality superiority claims, making it impossible to assess whether the reported accuracy gains are robust.","section":"Evaluation"}],"minor_comments":[{"comment":"Notation for the neural network architecture and loss terms could be clarified with an explicit equation reference in the methods.","section":"Methods"},{"comment":"Figure captions should include the number of test cases and exact error metrics shown.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and will revise the manuscript accordingly to improve clarity and rigor.","responses":[{"response":"We agree that direct quantitative validation of simulation fidelity would strengthen the paper. The forward model follows established physics-based principles from prior X-ray simulation literature, and the strong real-fluoroscopy results across sites provide indirect support. However, the manuscript lacks explicit sim-to-real metrics. We will add a supplementary section with intensity histogram comparisons, perceptual metrics, and phantom-based registration error analysis to better substantiate the claims.","revision_made":"yes","referee_comment":"[Abstract; Evaluation] The abstract and evaluation sections claim an order-of-magnitude accuracy improvement on real fluoroscopy without reporting quantitative verification that the physics-based forward model reproduces real C-arm intensity statistics, scatter, noise, and geometric properties (e.g., no histogram comparisons, perceptual metrics, or phantom-based sim-vs-real registration error). This assumption is load-bearing for the generalization claim."},{"response":"We acknowledge that additional methodological transparency is required. The revised manuscript will specify data exclusion criteria, provide exact baseline implementations with hyperparameter settings, and include statistical tests (e.g., paired comparisons with p-values) to support the reported gains across hospitals and modalities.","revision_made":"yes","referee_comment":"[Evaluation] The results do not specify data exclusion criteria, exact baseline implementations and hyperparameter settings, or statistical tests supporting the cross-hospital and cross-modality superiority claims, making it impossible to assess whether the reported accuracy gains are robust."}],"tokens_in":1410,"tokens_out":349,"duration_ms":32690,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that xvr trains a patient-specific network in roughly 5 minutes on simulated X-rays generated from the patient's own preoperative CT or MRI, starting from a model pretrained on thousands of whole-body scans, then uses it for fast rigid 2D/3D registration on real fluoroscopy. It reports strong accuracy across many anatomies and hospitals in what it calls the largest such real-data evaluation so far, with an order-of-magnitude gain over existing methods, and the code is open-sourced.","headline":"xvr's patient-specific adaptation from a whole-body foundation model via physics simulation is a practical synthesis, but the sim-to-real fidelity for real fluoroscopy accuracy is the part that needs direct evidence.","tokens_in":2436,"tokens_out":190,"would_cite":false,"duration_ms":92679,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Orthogonal engineering method for medical image registration; no RS machinery","alignment":"orthogonal","rationale":"The paper's core contribution is a patient-specific NN trained via physics-based differentiable X-ray rendering + gradient refinement using mNCC similarity. This is a standard self-supervised registration pipeline with no reference to J-cost, reciprocal symmetry, φ-ladder, 8-tick periodicity, or the distinction-to-spacetime forcing chain. RS has no theorems on intraoperative fluoroscopy or C-arm pose regression.","tokens_in":57790,"confidence":"high","tokens_out":124,"duration_ms":13043,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Patient-specific neural networks register preoperative 3D volumes to intraoperative X-rays in seconds with order-of-magnitude accuracy gains across anatomies.","keywords":["2D/3D registration","fluoroscopy","patient-specific neural networks","self-supervised learning","image-guided surgery","X-ray to volume registration","intraoperative imaging"],"falsifier":"A comparison on a broad collection of real fluoroscopy cases showing that accuracy does not exceed existing methods by an order of magnitude or that performance collapses on new hospitals or anatomical regions would falsify the central performance claim.","tokens_in":2682,"feed_emoji":"🩻","tokens_out":727,"duration_ms":46248,"temperature":0.7,"pith_summary":"The paper introduces xvr, a self-supervised framework that trains patient-specific neural networks on physics-based simulations derived from each patient's own preoperative CT or MRI scan to align those 3D volumes with 2D fluoroscopy images. Existing intensity-based methods demand per-patient tuning while deep learning approaches require large labeled datasets restricted to narrow anatomies; xvr avoids both by pretraining a foundation model on thousands of whole-body scans and adapting it in five minutes. The largest evaluation on real fluoroscopy data to date reports high accuracy achieved in seconds across structures, modalities, and hospitals, with an order-of-magnitude improvement over prior techniques. A reader would care because precise 2D/3D registration underpins navigation in image-guided interventions and surgical robotics, and the method removes the main barriers to broad clinical use.","feed_headline":"Patient-specific nets register X-ray to 3D volume in seconds","feed_subtitle":"Self-supervised training from preoperative scans eliminates labels and boosts accuracy tenfold on real fluoroscopy data.","key_machinery":"Patient-specific neural network finetuned in five minutes on physics-simulated X-ray projections from the preoperative scan and combined with gradient-based optimization for registration.","core_discovery":"xvr achieves automatic 2D/3D rigid registration by combining patient-specific neural networks with gradient-based optimization, where the networks are trained self-supervised on training data generated through physics-based simulation from the patient's preoperative volume. A foundation model pretrained on thousands of whole-body scans enables adaptation to any anatomical region in five minutes of finetuning. On the largest set of real fluoroscopy cases evaluated to date, the approach reaches high accuracy in seconds across diverse anatomical structures, imaging modalities, and hospitals while improving accuracy over existing methods by an order of magnitude.","pith_inferences":["The approach could support real-time guidance in robotic surgery platforms once inference speed is further optimized for continuous tracking.","Similar simulation-driven patient-specific adaptation might apply to other 2D/3D problems such as ultrasound-to-CT alignment.","If domain shift remains small, the same pretraining strategy could reduce data requirements in related medical image registration tasks."],"forward_implications":["Registration no longer requires careful per-subject hyperparameter tuning of intensity-based optimizers.","Manually labeled datasets specific to each anatomy are no longer needed.","A single foundation model supports pan-anatomical application after brief patient-specific adaptation.","Open-source release makes the method immediately usable by clinical and research communities."],"fun_headline_variants":["xvr enables patient-specific X-ray to volume registration in seconds","Self-supervised nets register fluoroscopy to 3D across anatomy","Five-minute finetuning adapts model to any anatomical region","Neural networks achieve high accuracy X-ray to CT alignment"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The physics-based simulation used to generate training data from preoperative scans produces images sufficiently similar to real intraoperative fluoroscopy that the trained network generalizes without large domain shift.","fun_headline_variants_meta":{"raw":{"variants":["xvr enables patient-specific X-ray to volume registration in seconds","Self-supervised nets register fluoroscopy to 3D across anatomy","Five-minute finetuning adapts model to any anatomical region","Neural networks achieve high accuracy X-ray to CT alignment"]},"model":"grok-4.3","cost_usd":0.01025,"raw_usage":{"total_tokens":4495,"prompt_tokens":736,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":102503000,"prompt_tokens_details":{"text_tokens":736,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3699,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":736,"tokens_out":60,"duration_ms":62339,"temperature":1.0,"reasoning_tokens":3699,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-22T23:08:17.437020+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A comparison on a broad collection of real fluoroscopy cases showing that accuracy does not exceed existing methods by an order of magnitude or that performance collapses on new hospitals or anatomical regions would falsify the central performance claim.","supporting_citations":[],"review_version":1}