Pith. sign in

REVIEW 5 major objections 5 minor 59 references

IXGS-Intraoperative 3D Reconstruction from Sparse, Arbitrarily Posed Real X-rays

T0 review · 5 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Gaussian splatting reconstructs usable 3D spine volumes from sparse arbitrary X-rays.

desk verdict A plausible feasibility study of splatting-based spine reconstruction from sparse arbitrary X-rays, but the CT-defined volume bounds and style-transfer evaluation mean the 'pure X-ray' claim is sturdier on paper than in practice. read the letter →

arxiv 2504.14699 v1 pith:WVSHJHYX submitted 2025-04-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords Gaussiansplattingsparse-viewreconstructionintraoperativeX-raylumbarspinestyletransfersurgicalnavigationC-armfluoroscopyvolumetric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that instance-based Gaussian splatting, normally evaluated on synthetic projections or simple objects captured along circular paths, can reconstruct anatomically useful 3D volumes of the lumbar spine from roughly 20 to 30 arbitrarily posed, real intraoperative X-rays, with no pretraining on patient data. If true, surgical navigation could move from radiation-heavy cone-beam CT sweeps to a short series of standard C-arm fluoroscopy shots, while remaining adaptable to new patients and pathologies. The load-bearing addition is an anatomy-guided standardization step: style transfer maps each real X-ray into a bone-emphasized, appearance-consistent synthetic-like domain before the splatting optimization, and expert ratings confirm that this preprocessing is what makes the volumes navigable. Quantitative 2D metrics remain well below idealized circular-synthetic benchmarks, but the paper argues that those metrics understate clinical utility.

What carries the argument

The central object is a collection of learnable 3D Gaussian kernels, each with position, covariance, and central density, optimized by matching their projected 2D radiographs to the input views. Two adaptations carry the argument: a ray-space projection with an integration-bias rectification factor, extended to arbitrary non-circular poses with random volume initialization, and an anatomy-guided standardization step in which a paired style-transfer network converts real X-rays into synthetic radiograph-like images that emphasize bone. The pipeline converts the optimized Gaussians into a voxel density volume via a differentiable voxelizer, then thresholds and crops.

What would settle it

Take a held-out specimen, estimate the volume bounds purely from the calibrated views by triangulating the principal axes, run the full pipeline, and have the same surgeon rate the slices and volume renderings; if ratings fall below 'Acceptable' or the pedicles and endplates become indiscernible, the clinical claim as stated fails.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that a pretraining-free Gaussian-splatting pipeline can reconstruct a volumetric lumbar-spine density map from as few as 20-30 arbitrarily posed real intraoperative X-rays, at a level an experienced surgeon rates acceptable to very good for pedicle-screw navigation. The discovery is that the bottleneck is not the splatting optimization itself but the inconsistency of real X-ray appearance: mapping each real radiograph into a standardized, bone-emphasized synthetic-like domain with style transfer before reconstruction is what lifts the volumes from unusable to navigable. With 50 views the standardized inputs reach 25.73 dB PSNR versus 23.22 dB for raw X-rays, still below the 39.0 dB of an idealized circular synthetic baseline, and the reconstructed slices show clearly discernible pedicles, spinous processes, and endplates. The same test set is used across all view-count experiments, so the comparison isolates the effect of view number and input domain.

Load-bearing premise

During evaluation the reconstruction volume's center and size come from the ground-truth CT; the clinical promise depends on estimating them from the calibrated X-ray views alone, which the paper does not demonstrate.

Editorial extensions

If this is right

  • If the central claim holds, a C-arm operator could collect 20-30 fluoroscopic shots and obtain a navigational 3D volume without the radiation burden of a cone-beam CT sweep, since the method needs no pretraining and no per-patient training data.
  • The style-transfer standardization becomes a reusable front-end for any real-X-ray Gaussian-splatting reconstruction, since raw real X-rays alone produced volumes the authors judged insufficient for expert assessment.
  • The trend in surgeon ratings and the plateau in PSNR/SSIM past 20 views suggest a practical acquisition protocol: around 20-30 views are enough, and adding views yields diminishing returns.
  • Because the output is a voxel density map with no calibrated Hounsfield units, clinical use would target navigation and visualization, not quantitative CT-like density measurement.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if volume bounds can be estimated from calibrated views, for instance by triangulating principal axes, the whole pipeline becomes CT-free and could run on a standard C-arm in the operating room; this is the natural next validation.
  • Beyond the paper: the same standardization-plus-splatting recipe may transfer to other bony anatomies such as the femur, pelvis, or cervical spine because there is no anatomy-specific pretraining, but the 20-30-view threshold would need re-measuring per site.
  • Beyond the paper: adding a fast initialization, such as a coarse reconstruction or a rough density prior, could cut the reported 13-minute optimization and make intraoperative use practical, since the paper itself notes that convergence can be accelerated.
  • Beyond the paper: replacing PSNR/SSIM with a task-based metric, or overlaying ground-truth segmentations automatically, could turn expert 'Acceptable' ratings into a quantitative acceptance criterion for navigation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The manuscript presents IXGS, a Gaussian-splatting-based framework for volumetric reconstruction of the lumbar spine from sparse, arbitrarily posed real intraoperative X-ray images. It builds on R2-Gaussian, replaces the FDK-based initialization with random initialization inside a reconstruction volume, and introduces an anatomy-guided radiographic standardization step based on Pix2Pix style transfer that maps real X-rays into a synthetic DRR-like domain. The evaluation uses six ex-vivo specimens: an expert surgeon rates 3D volume renderings and slice views for navigation usability, and PSNR/SSIM on held-out views compare reconstructions from synthetic circular, synthetic arbitrary, raw real, and style-transferred inputs. The authors report that clinically acceptable ratings appear at roughly 20-30 views and that standardization improves quantitative scores.

Significance. If the claims are upheld, the paper makes a useful contribution: it is among the first to apply radiative Gaussian splatting to sparse, arbitrary-pose real intraoperative X-rays, and it provides a practical preprocessing recipe that visibly improves bone clarity. The instance-based nature of the splatting optimization is attractive because it avoids anatomy-specific training of the reconstruction network, and the promised release of code and a small dataset is a strength. However, the reported results are conditioned on CT-derived volume bounds, the evaluation lacks quantitative 3D accuracy and error bars, and the test-domain protocol is not fully consistent across conditions. The expert ratings and the 2D trends are encouraging, but the central 'X-ray-only' feasibility claim is currently stronger than the evidence supports.

major comments (5)
  1. [Methods, Initialization] The reconstruction volume center and dimensions are defined from ground-truth CT for all reported experiments, as stated in the text: 'For the evaluations presented in this work, where ground truth CT data was available, the center and dimensions of the reconstruction volume were defined based on the CT.' Because kernel positions are randomly sampled within that volume, every expert rating and every PSNR/SSIM result evaluates a pipeline that has been given a geometric prior unavailable in a pure intraoperative X-ray workflow. This does not invalidate the reconstruction optimization itself, but it means the 'from sparse, arbitrarily posed real X-rays' feasibility claim is not yet demonstrated for the full pipeline. Please either implement and evaluate the volume-estimation approach from calibrated views (e.g., the triangulation cited from prior work) or explicitly reframe the claim and add a sensitivity analysis over shifted or scaled reconstruction volumes.
  2. [Experiments, 3D Evaluation] The paper reports no quantitative 3D accuracy metric, although CT segmentations of L1-L5 are available and are used for overlay visualization. The conclusion that the reconstructions are 'anatomically consistent' rests on expert Likert ratings and 2D PSNR/SSIM, yet the authors themselves argue that such 2D metrics correlate poorly with anatomical plausibility. A voxel-wise or surface-distance comparison between the reconstructed density volume and the CT-derived bone mask is needed to support a volumetric reconstruction claim; the segmentations already in hand make this feasible.
  3. [Table 1 and Figure 6] The quantitative comparison reports only averages over six specimens with no error bars, per-specimen values, or significance tests. Figure 6 shows that one specimen never received a rating above 'Poor', indicating substantial inter-specimen variability; the 2.51 dB PSNR gain of IST over Ireal and the 20-30-view guideline may not be representative. Please report per-specimen results and variance, and state whether the reported differences are statistically significant.
  4. [Experiments, 3D Evaluation and Figure 3] The expert evaluation protocol allowed the surgeon to toggle an overlay of ground-truth CT segmentations while rating anatomical accuracy. If ratings were assigned with the overlay visible, high scores may reflect how well the reconstruction aligns with ground truth available to the rater, rather than standalone clinical usability. Please clarify whether the overlay was visible during scoring or only used after scoring, or perform a blinded evaluation without the overlay.
  5. [Experiments, 2D Quantitative Evaluation and Table 1] The manuscript does not specify whether the held-out test images for the IST condition are the Pix2Pix-transformed versions or the original real X-rays. If the former, the reported IST-vs-Ireal gain partly measures self-consistency of the style-transfer mapping rather than reconstruction fidelity to the original real X-ray content. Please state explicitly which test domain is used for each condition and, ideally, evaluate all reconstructions against the same held-out set, with the remaining domain mismatch discussed.
minor comments (5)
  1. [Throughout] Several cross-references appear as empty 'Section' placeholders (e.g., in Related Work and Methods), and Figure 6's caption refers to 'Section' without a number; these should be resolved before publication.
  2. [Methods, Optimization] The sentence beginning 'The optimize the 3D Gaussian kernels' should read 'To optimize the 3D Gaussian kernels'.
  3. [Results, Table 1] Table 1 reports the Icirc 50-view PSNR as 39.19 dB while the text states 39.0 dB; please make these values consistent.
  4. [Abstract and Methods, Anatomy-guided Radiographic Standardization] The abstract's claim that 'our framework requires no pretraining' should be qualified, since the Pix2Pix standardization step is trained on paired data; the claim is accurate for the Gaussian optimization but not for the full pipeline as stated.
  5. [Related Work] The sentence 'Both techniques and synthetic X-ray-based reconstructions32–37.' is grammatically incomplete and should be revised.

Circularity Check

1 steps flagged · score 4.0 of 10

Quantitative IST improvement is evaluated against the same DRR domain used to train the Pix2Pix style-transfer; central 3D reconstruction is otherwise self-contained.

  1. fitted input called prediction [Methods, 'Anatomy-guided Radiographic Standardization', Eqs.]
    "L_L1(G) = EIreal,IDRR,z[∥ IDRR−G(Ireal,z)∥1] ... After training, the generator G was used to style-transfer Ireal to the synthetic domain, producing IST = G(Ireal). ... The IDRR images — generated directly from the ground truth CT scans using the exact same arbitrary poses as the corresponding Ireal images — serve as a synthetic benchmark and represent the target appearance domain for our proposed IST radiographic standardization step. ..."

    The Pix2Pix generator is fitted by minimizing the L1 distance between G(Ireal) and the synthetic DRR domain IDRR (Eq. 4), so the output domain of the fitted stylization is, by construction, the DRR domain. The quantitative evaluation then identifies IDRR as the target appearance domain for IST and reports the IST-vs-Ireal gain (25.73 vs 23.22 dB) as evidence that the standardization improves reconstruction quality. Because the benchmark rewards closeness to the very domain that Pix2Pix was trained to produce, the measured improvement is partly self-consistency of the fitted stylization rather than an independent measurement of 3D reconstruction fidelity.

full rationale

The central 3D reconstruction pipeline is not circular: the Gaussian splatting is optimized per instance against the input views without anatomy-specific pretraining, rendered from held-out poses, and evaluated both quantitatively and by expert qualitative assessment. The R2-Gaussian backbone is externally published, code-reproduced, and the authors' extension to arbitrary poses with random initialization is a substantive implementation change rather than a renaming of the prior result. The CT-derived volume bounds in the Initialization section are a serious evaluation limitation and are explicitly acknowledged by the authors; however, this is ground-truth conditioning of the test setup, not a derivation that reduces the method to its own inputs, so it does not by itself constitute circularity. The main circularity concern is the quantitative validation of the anatomy-guided standardization: Pix2Pix is trained (Eq. 4) to map Ireal into the IDRR domain, and the benchmark treats IDRR as the target appearance domain for IST, so the reported improvement of IST over Ireal is partly baked into the stylization training objective. This weakens the preprocessing claim but does not undermine the core feasibility conclusion, which also rests on surgeon ratings and held-out novel-view synthesis. Overall score 4 reflects a partial, localized circularity in the standardization evaluation while the central reconstruction claim retains independent content.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical entities or forces. Its load-bearing assumptions are empirical and domain-specific: pose accuracy, style transfer fidelity, convergence of random initialization, and availability of volume bounds. One free parameter, the postprocessing density threshold, is fitted empirically. The central reconstruction method itself is self-contained and requires no pretrained anatomical model.

free parameters (1)
  • Volume density threshold for postprocessing = 80th percentile
    Empirically set in the Output section to suppress cloudy artifacts and isolate the lumbar spine; it influences the visual quality and interpretation of all reported reconstructions.
assumptions (5)
  • domain assumption Fiducial-based calibration, using DLT plus RANSAC, gives camera poses accurate enough for the splatting optimization.
    The entire reconstruction assumes known poses from the Dataset section. Real intraoperative pose estimation may be less accurate than the fiducial-based ex-vivo calibration, and would directly corrupt the volume.
  • domain assumption Pix2Pix style transfer to the DRR domain preserves anatomy and does not hallucinate or erase structures.
    IST is the core preprocessing step introduced in the Methods section. If the generator alters bony or pathological details, the reconstructed volume would be clinically misleading. No quantitative verification of anatomical fidelity of the style-transferred images is provided.
  • domain assumption Random Gaussian initialization converges to an acceptable volume from 20 to 50 arbitrary views.
    The paper replaces FDK initialization with random sampling because arbitrary poses make FDK inapplicable. One specimen remained Poor at all view counts, so this assumption is not universally satisfied.
  • domain assumption Reconstruction volume bounds are known from CT during evaluation and can be estimated from calibrated X-rays in clinical use.
    The Initialization section states that volume center and dimensions were defined based on CT and would need to be estimated in the future. Clinical applicability depends on this estimate being accurate.
  • domain assumption DRR synthesis with a 0 HU threshold is a valid standardized representation of bone for surgical navigation.
    The target domain for Pix2Pix and for quantitative benchmarking is a bone-emphasized synthetic projection. Removing soft tissue may aid clarity but also removes information some surgeons could rely on.

how reviews work

0 comments
Cite this review

Pith. "Pith review of IXGS-Intraoperative 3D Reconstruction from Sparse, Arbitrarily Posed Real X-rays." pith.science (2026). https://pith.science/paper/WVSHJHYX

@misc{pith2026250414699,
  author       = {Pith},
  title        = {Pith review of: IXGS-Intraoperative 3D Reconstruction from Sparse, Arbitrarily Posed Real X-rays},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WVSHJHYX}},
  note         = {Machine review of arXiv:2504.14699}
}
abstract

Spine surgery is a high-risk intervention demanding precise execution, often supported by image-based navigation systems. Recently, supervised learning approaches have gained attention for reconstructing 3D spinal anatomy from sparse fluoroscopic data, significantly reducing reliance on radiation-intensive 3D imaging systems. However, these methods typically require large amounts of annotated training data and may struggle to generalize across varying patient anatomies or imaging conditions. Instance-learning approaches like Gaussian splatting could offer an alternative by avoiding extensive annotation requirements. While Gaussian splatting has shown promise for novel view synthesis, its application to sparse, arbitrarily posed real intraoperative X-rays has remained largely unexplored. This work addresses this limitation by extending the $R^2$-Gaussian splatting framework to reconstruct anatomically consistent 3D volumes under these challenging conditions. We introduce an anatomy-guided radiographic standardization step using style transfer, improving visual consistency across views, and enhancing reconstruction quality. Notably, our framework requires no pretraining, making it inherently adaptable to new patients and anatomies. We evaluated our approach using an ex-vivo dataset. Expert surgical evaluation confirmed the clinical utility of the 3D reconstructions for navigation, especially when using 20 to 30 views, and highlighted the standardization's benefit for anatomical clarity. Benchmarking via quantitative 2D metrics (PSNR/SSIM) confirmed performance trade-offs compared to idealized settings, but also validated the improvement gained from standardization over raw inputs. This work demonstrates the feasibility of instance-based volumetric reconstruction from arbitrary sparse-view X-rays, advancing intraoperative 3D imaging for surgical navigation.

Figures

Figures reproduced from arXiv: 2504.14699 by the authors.

Figure 1
Figure 1. Comparison between conventional circular acquisition paths, typical for CBCT/CT imaging or idealized synthetic DRR generation (left), and the irregular, arbitrary acquisition poses representative of real intraoperative settings (right). anatomical variations, pathological conditions, and implant types encountered in the operating room. Neural scene representation techniques overcome these limitations, as they are ca… view at source ↗
Figure 2
Figure 2. Pipeline Overview. The training of Pix2Pix (blue) uses paired real X-rays (Ireal) and synthetic DRRs (IDRR) to learn style transfer. During inference (green), real X-rays with calibration information are converted to style-transferred images (IST). These images, along with their poses, are passed to the Gaussian splatting network, which outputs 3D reconstructions. The resulting volume can be visualized as slices or … view at source ↗
Figure 3
Figure 3. Comparison of 3D reconstructions of the lumbar spine (axial, coronal, and sagittal slices). Each block shows axial, coronal, and sagittal slices of the reconstructed volume, overlaid with alpha-blended masks of the segmented lumbar spine from the ground truth CT for better accuracy assessment. The columns compare reconstructions using 25 views (left) and 50 views (right). The top row shows reconstructions from Ireal… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Comparison of slices from reconstructed volumes using different inputs and methods. From left to right: VCT: Ground Truth CT volume, Vcirc: Reconstruction from 50 synthetic DRRs generated from circular acquisition, Vreal: Reconstruction from 50 X-rays generated from ar…
Figure 5
Figure 5. Figure 5: Comparison of 3D reconstructions of the lumbar spine. Each block shows AP, lateral, and isometric views of the reconstructed volume. The columns compare reconstructions using 25 views (left) and 50 views (right). The top row displays reconstructions from Ireal baseline…
Figure 6
Figure 6. Figure 6: Evaluation of reconstruction quality using Likert ratings from expert surgeons. a) and b) show ratings over 5 to 50 views, with a) assessing 3D volumes and b) assessing slice representations. See Section for details on the different Likert scales used. Method 50 Views …
Figure 7
Figure 7. Figure 7: Evaluation of novel view synthesis quality from unseen poses using PSNR and SSIM metrics over varying numbers of views. Higher scores indicate better reconstruction quality. Discussion Our IXGS framework demonstrates that 3D volumetric reconstruction from sparse, arbit…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 35 canonical work pages

  1. [1]

    & Seurat, O

    Tonetti, J., Boudissa, M., Kerschbaumer, G. & Seurat, O. Role of 3D intraoperative imaging in orthopedic and trauma surgery. Orthop. & Traumatol. Surg. & Res. 106, S19–S25, DOI: 10.1016/j.otsr.2019.05.021 (2020)

  2. [2]

    Keil, H. et al. Intraoperative revision rates due to three-dimensional imaging in orthopedic trauma surgery: results of a case series of 4721 patients. Eur. J. Trauma Emerg. Surgery: Off. Publ. Eur. Trauma Soc. 49, 373–381, DOI: 10.1007/s00068-022-02083-x (2023)

  3. [3]

    & Hashizume, M

    Hong, J. & Hashizume, M. An effective point-based registration tool for surgical navigation. Surg. Endosc. 24, 944–948, DOI: 10.1007/s00464-009-0568-2 (2010)

  4. [4]

    Ma, Q. et al. Autonomous Surgical Robot With Camera-Based Markerless Navigation for Oral and Maxillofacial Surgery. IEEE/ASME Transactions on Mechatronics 25, 1084–1094, DOI: 10.1109/TMECH.2020.2971618 (2020). Conference Name: IEEE/ASME Transactions on Mechatronics

  5. [5]

    Suenaga, H. et al. Vision-based markerless registration using stereo vision and an augmented reality surgical navigation system: a pilot study. BMC Med. Imaging 15, 51, DOI: 10.1186/s12880-015-0089-5 (2015). 13/17

  6. [6]

    Zheng, J., Miao, S., Wang, Z. J. & Liao, R. Pairwise domain adaptation module for CNN-based 2-D/3-D registration. J. Med. Imaging (Bellingham, Wash.) 5, 021204, DOI: 10.1117/1.JMI.5.2.021204 (2018)

  7. [7]

    & Milone, D

    Ferrante, E., Oktay, O., Glocker, B. & Milone, D. H. On the Adaptability of Unsupervised CNN-Based Deformable Image Registration to Unseen Image Domains. In Shi, Y ., Suk, H.-I. & Liu, M. (eds.)Machine Learning in Medical Imaging , 294–302, DOI: 10.1007/978-3-030-00919-9_34 (Springer International Publishing, Cham, 2018)

  8. [8]

    N., Fan, Y

    Zhang, J. N., Fan, Y . & Hao, D. J. Risk factors for robot-assisted spinal pedicle screw malposition.Sci. Reports 9, 3025, DOI: 10.1038/s41598-019-40057-z (2019). Publisher: Nature Publishing Group

Show all 59 references
  1. [9]

    W., Pirris, S

    Rahmathulla, G., Nottmeier, E. W., Pirris, S. M., Deen, H. G. & Pichelmann, M. A. Intraoperative image-guided spinal navigation: technical pitfalls and their avoidance. Neurosurg. F ocus.36, E3, DOI: 10.3171/2014.1.FOCUS13516 (2014)

  2. [10]

    & Elluru, S

    Venkatesh, E. & Elluru, S. V . Cone beam computed tomography: basics and applications in dentistry.J. Istanbul Univ. Fac. Dent. 51, S102–S121, DOI: 10.17096/jiufd.00289 (2017)

  3. [11]

    Costa, F. et al. Spinal Navigation: Standard Preoperative: Versus: Intraoperative Computed Tomography Data Set Acquisition for Computer-Guidance System: Radiological and Clinical Study in 100 Consecutive Patients. Spine 36, 2094–2098, DOI: 10.1097/BRS.0b013e318201129d (2011)

  4. [12]

    Mendelsohn, D. et al. Patient and surgeon radiation exposure during spinal instrumentation using intraoperative computed tomography-based navigation. The Spine J. 16, 343–354, DOI: 10.1016/j.spinee.2015.11.020 (2016)

  5. [13]

    Villard, J. et al. Radiation Exposure to the Surgeon and the Patient During Posterior Lumbar Spinal Instrumentation: A Prospective Randomized Comparison of Navigated Versus Non-navigated Freehand Techniques. Spine 39, 1004, DOI: 10.1097/BRS.0000000000000351 (2014)

  6. [14]

    Dea, N. et al. Economic evaluation comparing intraoperative cone beam CT-based navigation and conventional fluoroscopy for the placement of spinal pedicle screws: a patient-level data cost-effectiveness analysis. The Spine J. 16, 23–31, DOI: 10.1016/j.spinee.2015.09.062 (2016)

  7. [15]

    & Gradl, G

    Beck, M., Mittlmeier, T., Gierer, P., Harms, C. & Gradl, G. Benefit and accuracy of intraoperative 3D-imaging after pedicle screw placement: a prospective study in stabilizing thoracolumbar fractures. Eur. Spine J. 18, 1469–1477, DOI: 10.1007/s00586-009-1050-5 (2009)

  8. [16]

    & Esfandiari, H

    Jecklin, S., Jancik, C., Farshad, M., Fürnstahl, P. & Esfandiari, H. X23D—intraoperative 3D lumbar spine shape reconstruction based on sparse multi-view X-ray data. J. Imaging 8, 271, DOI: 10.3390/jimaging8100271 (2022). Publisher: MDPI

  9. [17]

    Jecklin, S. et al. Domain adaptation strategies for 3D reconstruction of the lumbar spine using real fluoroscopy data. Med. Image Analysis 98, 103322, DOI: 10.1016/j.media.2024.103322 (2024). Publisher: Elsevier

  10. [18]

    Zha, R. et al. R$^2$-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic Reconstruction. Adv. Neural Inf. Process. Syst. 37, 44907–44934 (2024)

  11. [19]

    Zha, R., Zhang, Y . & Li, H. NAF: Neural Attenuation Fields for Sparse-View CBCT Reconstruction. In Wang, L., Dou, Q., Fletcher, P. T., Speidel, S. & Li, S. (eds.)Medical Image Computing and Computer Assisted Intervention – MICCAI 2022 , 442–452, DOI: 10.1007/978-3-031-16446-0...

  12. [20]

    & Heidrich, W

    Rückert, D., Wang, Y ., Li, R., Idoughi, R. & Heidrich, W. NeAT: neural adaptive tomography.ACM Trans. Graph. 41, 55:1–55:13, DOI: 10.1145/3528223.3530121 (2022)

  13. [21]

    & Miltchev, R

    Rangelov, D., Waanders, S., Waanders, K., van Keulen, M. & Miltchev, R. Impact of Camera Settings on 3D Reconstruction Quality: Insights from NeRF and Gaussian Splatting. Sensors (Basel, Switzerland) 24, 7594, DOI: 10.3390/s24237594 (2024)

  14. [22]

    & Liu, Y

    Ye, S., Dong, Z., Hu, Y ., Wen, Y . & Liu, Y . Gaussian in the Dark: Real-Time View Synthesis From Inconsistent Dark Images Using Gaussian Splatting. Comput. Graph. F orum43, e15213, DOI: 10.1111/cgf.15213 (2024). _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/cgf.15213

  15. [23]

    & Efros, A

    Isola, P., Zhu, J.-Y ., Zhou, T. & Efros, A. A. Image-to-Image Translation with Conditional Adversarial Networks, DOI: 10.48550/arXiv.1611.07004 (2018). ArXiv:1611.07004 [cs]

  16. [24]

    & Kovler, I

    Kasten, Y ., Doktofsky, D. & Kovler, I. End-To-End Convolutional Neural Network for 3D Reconstruction of Knee Bones from Bi-planar X-Ray Images. In Deeba, F., Johnson, P., Würfl, T. & Ye, J. C. (eds.)Machine Learning for Medical Image Reconstruction, 123–133, DOI: 10.1007/978-...

  17. [25]

    Shiode, R. et al. 2D–3D reconstruction of distal forearm bone from actual X-ray images of the wrist using convolutional neural networks. Sci. Reports 11, 15249, DOI: 10.1038/s41598-021-94634-2 (2021). Publisher: Nature Publishing Group. 14/17

  18. [26]

    Ge, R. et al. X-CTRSNet: 3D cervical vertebra CT reconstruction and segmentation directly from 2D X-ray images. Knowledge-Based Syst. 236, 107680, DOI: 10.1016/j.knosys.2021.107680 (2022)

  19. [27]

    & Malik, J

    Kar, A., Häne, C. & Malik, J. Learning a Multi-View Stereo Machine. arXiv:1708.05375 [cs] (2017)

  20. [28]

    Luchmann, D. et al. Spinal navigation with AI-driven 3D-reconstruction of fluoroscopy images: an ex-vivo feasibility study. BMC Musculoskelet. Disord. 25, 925, DOI: 10.1186/s12891-024-08052-2 (2024). Publisher: BioMed Central London

  21. [29]

    Mildenhall, B. et al. NeRF: representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 99–106, DOI: 10.1145/3503250 (2021)

  22. [30]

    & Drettakis, G

    Kerbl, B., Kopanas, G., Leimkuehler, T. & Drettakis, G. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Trans. Graph. 42, 139:1–139:14, DOI: 10.1145/3592433 (2023)

  23. [31]

    Rakotosaona, M.-J. et al. NeRFMeshing: Distilling Neural Radiance Fields into Geometrically-Accurate 3D Meshes. In 2024 International Conference on 3D Vision (3DV) , 1156–1165, DOI: 10.1109/3DV62453.2024.00093 (2024). ISSN: 2475-7888

  24. [32]

    Corona-Figueroa, A. et al. MedNeRF: Medical Neural Radiance Fields for Reconstructing 3D-aware CT-Projections from a Single X-ray. 2022 44th Annu. Int. Conf. IEEE Eng. Medicine & Biol. Soc. (EMBC) 3843–3848, DOI: 10. 1109/EMBC48229.2022.9871757 (2022). Conference Name: 2022 44...

  25. [33]

    Radiative Gaussian Splatting for Efficient X-Ray Novel View Synthesis

    Cai, Y .et al. Radiative Gaussian Splatting for Efficient X-Ray Novel View Synthesis. In Leonardis, A. et al. (eds.) Computer Vision – ECCV 2024, 283–299, DOI: 10.1007/978-3-031-73232-4_16 (Springer Nature Switzerland, Cham, 2025)

  26. [34]

    Gao, Z. et al. DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic V olume Rendering.Adv. Neural Inf. Process. Syst. 37, 39281–39302 (2024)

  27. [35]

    Wysocki, M. et al. Ultra-NeRF: Neural Radiance Fields for Ultrasound Imaging. In Medical Imaging with Deep Learning , 382–401 (PMLR, 2024). ISSN: 2640-3498

  28. [36]

    Awojoyogbe, B. O. & Dada, M. O. Neural Radiance Fields (NeRFs) Technique to Render 3D Reconstruction of Magnetic Resonance Images. In Awojoyogbe, B. O. & Dada, M. O. (eds.) Digital Molecular Magnetic Resonance Imaging , 247–258, DOI: 10.1007/978-981-97-6370-2_10 (Springer Natu...

  29. [37]

    Yang, S. et al. Deform3DGS: Flexible Deformation for Fast Surgical Scene Reconstruction with Gaussian Splatting. In Linguraru, M. G. et al. (eds.) Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 , 132–142, DOI: 10.1007/978-3-031-72089-5_13 (Springer Na...

  30. [38]

    & Wang, A

    Cai, Y ., Wang, J., Yuille, A., Zhou, Z. & Wang, A. Structure-Aware Sparse-View X-ray 3D Reconstruction. 11174–11183 (2024)

  31. [39]

    X-ray Tomographic Datasets Archives

  32. [40]

    VertXNet: an ensemble method for vertebral body segmentation and identification from cervical and lumbar spinal X-rays

    Chen, Y .et al. VertXNet: an ensemble method for vertebral body segmentation and identification from cervical and lumbar spinal X-rays. Sci. Reports 14, 3341, DOI: 10.1038/s41598-023-49923-3 (2024). Publisher: Nature Publishing Group

  33. [41]

    C., Cho, H

    Kim, K. C., Cho, H. C., Jang, T. J., Choi, J. M. & Seo, J. K. Automatic detection and segmentation of lumbar vertebrae from X-ray images for compression fracture evaluation. Comput. Methods Programs Biomed. 200, 105833, DOI: 10.1016/j.cmpb.2020.105833 (2021)

  34. [42]

    & Bojanowski, P

    Darcet, T., Oquab, M., Mairal, J. & Bojanowski, P. Vision Transformers Need Registers, DOI: 10.48550/arXiv.2309.16588 (2024). ArXiv:2309.16588 [cs]

  35. [43]

    Kirillov, A. et al. Segment Anything. 4015–4026 (2023)

  36. [44]

    Wu, J. et al. Medical SAM Adapter: Adapting Segment Anything Model for Medical Image Segmentation, DOI: 10.48550/arXiv.2304.12620 (2023). ArXiv:2304.12620 [cs]

  37. [45]

    Ehab, W., Huang, L. & Li, Y . UNet and Variants for Medical Image Segmentation.Int. J. Netw. Dyn. Intell. 100009–100009, DOI: 10.53941/ijndi.2024.100009 (2024)

  38. [46]

    & Tallamraju, R

    Gupta, S., Lakhotia, S., Rawat, A. & Tallamraju, R. ViTOL: Vision Transformer for Weakly Supervised Object Localization. 4101–4110 (2022). 15/17

  39. [47]

    Coarse X-ray Lumbar Vertebrae Pose Localization and Registration Using Triangulation Correspon- dence

    Yookwan, W.et al. Coarse X-ray Lumbar Vertebrae Pose Localization and Registration Using Triangulation Correspon- dence. Processes 11, 61, DOI: 10.3390/pr11010061 (2023). Number: 1 Publisher: Multidisciplinary Digital Publishing Institute

  40. [48]

    & Zheng, G

    Ye, K., Sun, W., Tao, R. & Zheng, G. A Projective-Geometry-Aware Network for 3D Vertebra Localization in Calibrated Biplanar X-Ray Images. Sensors 25, 1123, DOI: 10.3390/s25041123 (2025). Number: 4 Publisher: Multidisciplinary Digital Publishing Institute

  41. [49]

    & Liao, H.-Y

    Wang, C.-Y ., Bochkovskiy, A. & Liao, H.-Y . M. YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. 7464–7475 (2023)

  42. [50]

    & Efros, A

    Zhu, J.-Y ., Park, T., Isola, P. & Efros, A. A. Unpaired Image-To-Image Translation Using Cycle-Consistent Adversarial Networks. In 2017 IEEE International Conference on Computer Vision (ICCV) , 2223–2232, DOI: 10.1109/ICCV .2017.244 (IEEE, Venice, 2017)

  43. [51]

    Duda, R. O. & Hart, P. E. Use of the Hough transformation to detect lines and curves in pictures. Commun. ACM 15, 11–15, DOI: 10.1145/361237.361242 (1972)

  44. [52]

    I., Karara, H

    Abdel-Aziz, Y . I., Karara, H. M. & Hauck, M. Direct Linear Transformation from Comparator Coordinates into Object Space Coordinates in Close-Range Photogrammetry*. Photogramm. Eng. & Remote. Sens. 81, 103–107, DOI: 10.14358/ PERS.81.2.103 (2015)

  45. [53]

    A., Davis, L

    Feldkamp, L. A., Davis, L. C. & Kress, J. W. Practical cone-beam algorithm. JOSA A 1, 612–619, DOI: 10.1364/JOSAA.1. 000612 (1984). Publisher: Optica Publishing Group

  46. [54]

    & Gross, M

    Zwicker, M., Pfister, H., van Baar, J. & Gross, M. EW A splatting.IEEE Transactions on Vis. Comput. Graph. 8, 223–238, DOI: 10.1109/TVCG.2002.1021576 (2002)

  47. [55]

    & Simoncelli, E

    Wang, Z., Bovik, A., Sheikh, H. & Simoncelli, E. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Process. 13, 600–612, DOI: 10.1109/TIP.2003.819861 (2004). Conference Name: IEEE Transactions on Image Processing

  48. [56]

    I., Osher, S

    Rudin, L. I., Osher, S. & Fatemi, E. Nonlinear total variation based noise removal algorithms. Phys. D: Nonlinear Phenom. 60, 259–268, DOI: 10.1016/0167-2789(92)90242-F (1992)

  49. [57]

    Breger, A. et al. A study on the adequacy of common IQA measures for medical images, DOI: 10.48550/arXiv.2405.19224 (2024). ArXiv:2405.19224 [eess]

  50. [58]

    A., Baltruschat, I

    Dohmen, M., Klemens, M. A., Baltruschat, I. M., Truong, T. & Lenga, M. Similarity and quality metrics for MR image- to-image translation. Sci. Reports 15, 3853, DOI: 10.1038/s41598-025-87358-0 (2025). Publisher: Nature Publishing Group

  51. [59]

    A computer- implemented method, device, system and computer program product for processing anatomic imaging data

    Zhang, R., Isola, P., Efros, A. A., Shechtman, E. & Wang, O. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. 586–595 (2018). Acknowledgements This work has been supported by the OR-X - a Swiss national research infrastructure for translational surgery a...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.