{"id":"5985f59d-4c51-4ce4-9d9c-4edf5bf615b5","arxiv_id":"2505.15235","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A large transformer with fixed-voxel Gaussian splatting reconstructs CT volumes from 6-10 X-ray projections in under a second, substantially beating prior sparse-view methods in simulation.","lead":"X-GRM is a transformer-based model that reconstructs a 3D CT volume from as few as six X-ray images in under one second, using a new representational trick: placing Gaussian blobs at fixed voxel centers. The paper reports large quality gains over existing sparse-view CT methods, but only on synthetic X-ray data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of clinical sparse-view X-ray-to-CT reconstruction is not yet supported because all evaluations use synthetic TIGRE projections under a simplified monochromatic Beer-Lambert model that excludes real-world scatter, beam hardening, and polyenergetic spectra.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: all training and evaluation X-rays are synthetic, rendered from CT volumes with TIGRE under simplified monochromatic Beer-Lambert physics. This is the most foundational issue because it directly affects the central claim's scope. The paper's stated goal is CT reconstruction from X-rays in clinical workflows, and the abstract claims performance on 'various testing inputs' without noting that every input, including the cross-dataset inputs in Sec. 4.3, is generated synthetically. Other concerns the reader noted — the 16-layer versus 12-layer inconsistency, the undefined 'ReconX-16K' ablation dataset, and missing error bars — are real but secondary: they affect transparency and reproducibility but do not undermine the core empirical result within the synthetic benchmark. The physics gap is different: if real X-ray acquisition deviates from the assumed model, the reported performance may not transfer at all, which would invalidate the central claim's practical relevance. A targeted real-data (or realistic simulation) evaluation would decisively test this. Since the reader already reached CONDITIONAL and the concern is the same, I recommend no change to the verdict.","tokens_in":15081,"tokens_out":6671,"duration_ms":64625,"concrete_test":"Evaluate the released X-GRM on real cone-beam CT projection data with corresponding CT ground truth (e.g., a public CBCT dataset or a clinically acquired phantom study). Use 6, 8, and 10 views and compute volumetric PSNR/SSIM, comparing to the synthetic results in Table 2. If the real-data PSNR drops by more than 3 dB or introduces structured artifacts such as cupping or streaking, the unqualified central claim fails and the model would need domain adaptation or retraining on more realistic physics before the clinical claim is credible. Alternatively, if real data are unavailable, run a validated polyenergetic Monte Carlo simulation (e.g., Gate or a spectral TIGRE extension) with scatter included; the same threshold applies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim — that X-GRM directly predicts CT volumes from sparse X-ray projections and drastically outperforms SOTA in quality and speed — rests on an evaluation pipeline in which every input X-ray is rendered from a CT volume using the TIGRE toolbox (Sec. 4.1). The paper's own Sec. A.1 limits the 3DGS/X-ray equivalence to 'a simplified imaging model that accounts solely for isotropic absorption (per Beer-Lambert law).' The added Gaussian and Poisson noise (Sec. 4.1) does not model polyenergetic spectra, scatter, beam hardening, detector blur, or system calibration offsets that are present in real CT acquisitions. The cross-dataset generalization tests (Sec. 4.3) use unseen CT datasets (PENGWIN, FUMPE) but still generate projections with the same TIGRE pipeline, so they demonstrate anatomical domain shift, not physical domain shift. Consequently, the abstract's unqualified statements about 'various testing inputs' and clinical deployment are not warranted by the reported experiments. If real X-ray projections differ substantially from the synthetic renders, the reported 28-29 dB PSNR could drop significantly, and the central claim would not transfer to the intended clinical setting. The paper does not provide any real-X-ray validation, a known and explicitly acknowledged simplification, yet the central claim is stated without this caveat.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes X-GRM, a feed-forward transformer-based model that reconstructs a 3D CT volume from 6, 8, or 10 sparse X-ray projections in about one second. The model uses a DINO-initialized ViT encoder to tokenize each projection, a fusion ViT to exchange information across views, and a novel Voxel-based Gaussian Splatting (VoxGS) representation in which isotropic Gaussians are placed at fixed voxel centers. Training combines an MSE loss on the extracted volume with L1 and D-SSIM losses on rendered X-rays, using sub-volume sampling to reduce memory. Experiments compare X-GRM with traditional, feedforward, and self-supervised baselines on a collected 14,972-volume dataset assembled from eight public CT datasets, and report consistent PSNR/SSIM gains, faster inference, and additional cross-dataset and novel-view-synthesis results.","tokens_in":15371,"tokens_out":4239,"duration_ms":39512,"significance":"If the reported results hold, X-GRM is a meaningful advance in sparse-view CT reconstruction: it combines a large-capacity transformer with a differentiable Gaussian volume representation, and the reported gains over strong baselines are consistent across 6/8/10-view settings. The paper contributes a sizable public-data-derived training set, a clean feed-forward formulation, and a promised code release, all of which are valuable for reproducibility and follow-up work. The main limitation is that every input X-ray in every experiment is synthesized with the TIGRE toolbox under a simplified monochromatic Beer-Lambert model; this makes the core technical contribution convincing as a proof of concept on synthetic data, but it does not by itself support the abstract's unqualified claims about 'various testing inputs' and clinical deployment.","major_comments":[{"comment":"All X-ray inputs across training, test, cross-dataset, and novel-view experiments are rendered from CT volumes with the TIGRE toolbox, and the equivalence between 3DGS rasterization and X-ray imaging is explicitly limited in A.1 to 'a simplified imaging model that accounts solely for isotropic absorption (per Beer-Lambert law).' The added Gaussian and Poisson noise in §4.1 does not model polyenergetic spectra, scatter, beam hardening, detector blur, or calibration offsets. As a result, the central claims in the Abstract and §1 that the model handles 'various testing inputs' and is suited to clinical workflows are stronger than the evidence supports. I would like to see the claims restricted to synthetic monochromatic projections, or ideally a validation on real paired X-ray/CT data (or at least a realistic polyenergetic scatter-inclusive simulation) to test physical domain shift.","section":"§4.1, §A.1, Abstract"},{"comment":"The ablation study is described as performed on the 'ReconX-16K dataset,' but this dataset is never defined anywhere in the paper or appendix. Its source, number of volumes, split, resolution, and projection parameters are unknown, so the reader cannot determine whether the ablation is run on the same scale as the main experiments or whether the reported component rankings (e.g., 0.28 dB for pose, 0.55 dB for VoxGS, 0.52 dB for attention) are stable. This should be specified exactly, or the ablation should be moved to the main test split.","section":"§4.5, Table 6"},{"comment":"The claim that X-GRM 'drastically outperforms' prior methods is based on single-run PSNR/SSIM numbers with no error bars, multiple seeds, or significance tests. Since feed-forward models are trained with stochastic optimization, run-to-run variance of several tenths of a dB is plausible at these resolutions, and some of the reported margins (e.g., 0.28 dB in Table 6a) are within that range. In addition, Table 2 reports timings on an A100 GPU while Table 3 uses an RTX 4090Ti, so the speed comparisons across tables are not directly comparable. Please report mean±std over at least three seeds and state the GPU configuration for each timing measurement.","section":"§4.2, Tables 2 and 3"},{"comment":"The cross-dataset experiments on FUMPE and PENGWIN demonstrate generalization to unseen anatomies, but because the projections are still generated with the same TIGRE rendering pipeline, they do not demonstrate generalization to new acquisition physics. The text in §4.3 and the Abstract's phrase 'out-domain X-ray projections' suggest a broader domain shift than the experiment actually tests. Please rephrase these claims as anatomical-domain generalization, or add an experiment with a different forward model to support physical-domain generalization.","section":"§4.3, Table 4"}],"minor_comments":[{"comment":"Table 3 cites R2-Gaussian as [76], but the reference list places R2-Gaussian at [75]; moreover, the same work appears to be duplicated as references [74] and [75]. Please reconcile the numbering and deduplicate.","section":"References, Table 3"},{"comment":"The main text says novel-view synthesis is evaluated on 30 distinct CT samples, while §A.3 and Table 8 describe the 'sampled test set (40 samples)'. Please clarify which number is correct.","section":"§4.4 vs. §A.3"},{"comment":"The rendering loss weights λ_L1 and λ_SSIM are not reported. Please give their values, as well as the sub-volume sampling factor, so that the training objective is fully reproducible.","section":"§3.5, Eq. (10)"},{"comment":"Reference [78] is the object-detection DINO paper, but the text says the encoder is initialized from DINO pre-trained weights, which normally refers to the self-supervised ViT-DINO of Caron et al. Please correct the citation.","section":"References, §3.3"},{"comment":"The notation for sub-volume sampling uses K for both the number of views and the depth dimension (M/4×N/4×K/4), while the volume is earlier defined as M×N×L. Please use consistent dimensional notation.","section":"§3.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid systems-style contribution to feed-forward sparse-view CT reconstruction, and the synthetic-only evaluation is the main gating issue. If the authors add a real-X-ray validation or clearly restrict all claims to the simplified monochromatic model, the work could be suitable for publication. The undefined 'ReconX-16K' ablation dataset and the missing error bars should be fixed regardless, since they currently weaken confidence in the component ablations and the headline performance claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: X-GRM is a serious large-model sparse-view CT reconstruction paper with a representation change that matters. Pinning 3D Gaussians to voxel centers and reading opacity directly (VoxGS) beats plain voxels by about 0.5 dB and shiftable Gaussians by about 2.4 dB in their ablations, and the whole system reconstructs a 256^3 volume in under a second. The gains over the feedforward baselines are consistent (3.5–4 dB over FreeSeed and DIF-Gaussian at 6/8/10 views) and hold on two unseen anatomical datasets. The numbers are believable.\n\nWhat is actually new is VoxGS. Original 3DGS and R2-Gaussian-style representations leave Gaussian centers free, which makes volume extraction unstable in a feedforward setting; fixing centers to the voxel grid removes the instability and lets the rendering loss constrain training. That is a clean, simple idea, and the ablation supports it. The transformer scaling is borrowed from X-LRM and DeepSparse, but combining it with this representation is a reasonable incremental step. The dataset is large and public (14,972 volumes from 8 sources), and the paper is written clearly enough to reproduce.\n\nThe soft spots, in order. First, all evaluations use synthetic X-ray projections rendered with TIGRE under a Beer-Lambert model with no scatter, beam hardening, or polyenergetic spectra. The cross-dataset tests (PENGWIN, FUMPE) still render through the same pipeline, so they demonstrate anatomical shift, not physical shift. The abstract's phrase 'various testing inputs' and the forward-looking clinical language overstate what the evidence supports. The authors do acknowledge the simplified model in Appendix A.1, but that caveat never reaches the abstract. This is the standard gap in this literature, and it does not invalidate the method comparison, but it does mean the clinical claim is not yet tested. Second, the ablation table (Tab. 6) reports results on a 'ReconX-16K' dataset that is never defined anywhere in the text, so those numbers are hard to audit. Third, minor transparency issues: the fusion ViT is called 16-layer in Sec. 3.3 but 12-layer in the implementation details, the GPU in Setting 2 is listed as 'RTX 4090Ti', and there are no error bars anywhere. None of these break the central result, but they clutter the paper.\n\nMy verdict: this deserves serious peer review. The representation is worth testing, the comparison is fair, and the main claims are empirically grounded on synthetic data. Reviewers should push for a defined ablation set, error bars or at least seed variance, and an abstract that does not promise clinical readiness without real X-ray validation. I would accept a revised version along those lines.","headline":"Competent large-model sparse-view CT paper with a genuinely useful fixed-center Gaussian representation, but the clinical claim runs ahead of the synthetic-only evidence.","tokens_in":15888,"tokens_out":2778,"would_cite":true,"duration_ms":26538,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"X-GRM is a large feedforward Transformer that reconstructs a full 3D CT volume from six to ten X-ray projections in about one second, using fixed-position voxel-based Gaussians to make the volume both extractable and differentiable.","keywords":["computed tomography","sparse-view CT reconstruction","X-ray imaging","3D Gaussian Splatting","feedforward transformer","volume rendering","novel view synthesis","medical imaging"],"falsifier":"Run the trained X-GRM on a real cone-beam CT study with measured polyenergetic projections, or on simulated projections that add scatter and beam hardening, and compare its PSNR and SSIM against per-sample optimization methods; if the quality gap shrinks sharply or anatomical detail develops new artifacts, the paper's equivalence between 3DGS rasterization and X-ray imaging would be broken for practical CT.","tokens_in":14920,"feed_emoji":"🩻","tokens_out":10601,"duration_ms":85243,"temperature":0.7,"pith_summary":"X-GRM is a large feedforward Transformer that takes six to ten 2D X-ray projections of a body region and outputs a full 3D CT volume in about one second. The paper's central claim is that a scalable cross-view transformer paired with a new volume representation, Voxel-based Gaussian Splatting, gives better reconstruction quality and far faster inference than existing sparse-view CT methods, including per-sample optimization approaches that take minutes to hours. If that claim holds, sparse-view CT could shift from iterative, time-consuming reconstruction to instant learned reconstruction from many fewer projections, lowering radiation exposure and enabling time-sensitive uses. The experiments support the claim on a large curated set of 14,972 CT volumes with synthetic X-ray projections, plus cross-dataset and novel-view tests.","feed_headline":"Full CT volumes from six X-ray views in under one second","feed_subtitle":"A Transformer with voxel-based Gaussian splatting beats per-sample optimization on speed and quality in synthetic sparse-view CT tests.","key_machinery":"The load-bearing object is Voxel-based Gaussian Splatting (VoxGS): a set of 3D Gaussians whose centers are locked to voxel centroids, each carrying only opacity $\\alpha_i$, scale $s_i$, and rotation $r_i$. Locking positions lets the CT volume be extracted by direct indexing, $V(x,y,z)=\\alpha_i$, with no trilinear interpolation, and dropping color is consistent with X-ray attenuation being a scalar line integral. The differentiable rasterizer for VoxGS supplies a rendering constraint during training, while the encoder-and-fusion ViT, with per-view patch tokens and all-to-all self-attention across views, supplies the capacity and cross-view reasoning that the paper argues prior CNN and voxel-grid models lack.","core_discovery":"The paper proposes X-GRM, a one-pass model that maps sparse X-ray projections with their camera matrices to a voxelized density field $V \\in \\mathbb{R}^{M \\times N \\times L}$. Each projection is tokenized by a DINO-initialized ViT, given ray geometry through camera-ray-modulated adaptive layer norm, and all views are fused by a 16-layer all-to-all self-attention transformer. The fused tokens are decoded into Voxel-based Gaussian Splatting (VoxGS) attributes: every voxel center hosts a 3D Gaussian with opacity, scale, and rotation but no color, making CT extraction a direct opacity lookup and making X-ray rendering differentiable. The model is trained with a volume MSE loss plus a rendering loss combining L1 and D-SSIM, and it reports PSNR of 28.39, 28.86, and 29.21 dB for 6, 8, and 10 input views on the 680-volume test set, with SSIM of 0.873, 0.879, and 0.886. These numbers exceed the best feedforward baseline by roughly 3.6 to 3.8 dB while running about twice as fast, and exceed per-sample optimized methods by 4 to 5 dB while running hundreds to thousands of times faster. The same model also synthesizes unseen X-ray views with higher reported fidelity than NeRF- and 3DGS-based per-sample methods.","pith_inferences":["Beyond the paper's explicit claims, VoxGS's opacity-as-density reading suggests the same fixed-lattice Gaussian head could be applied to other tomographic inverse problems whose forward operator is a line integral, such as PET or ultrasound computed tomography.","The paper's synthetic-only evaluation leaves an immediate stress test implicit: re-running the model on projections with beam hardening and scatter would quantify how much of the reported margin over per-sample optimization survives real scanner physics.","The model was trained and tested only with uniformly spaced views; an untested extension is non-uniform or limited-angle trajectories, where all-to-all cross-view attention may behave differently and missing angular coverage may expose the fixed voxel lattice's limits."],"forward_implications":["With 6 to 10 input projections, a $256^3$ CT volume is reconstructed in about 0.9 seconds, a regime per-sample optimization methods cannot reach.","Because VoxGS supports differentiable X-ray rendering, the trained model can also synthesize novel projection views; the paper reports 49.44 dB PSNR on held-out views at 0.02 seconds per projection.","A single model trained with variable view counts (6, 8, or 10) serves different sparsity levels without retraining for each setting.","On unseen chest and pelvis datasets, the model retains a quality advantage over feedforward baselines and matches or beats per-sample optimization while being about 500 times faster, indicating out-of-distribution generalization."],"supporting_citations":[{"why":"Supplies the synthetic projection rendering used to generate X-ray images from every CT volume in the dataset.","marker":"[3]"},{"why":"Defines 3D Gaussian Splatting and its differentiable rasterization, which VoxGS adapts for CT.","marker":"[28]"},{"why":"Provides the CT-oriented Gaussian rasterizer and the opacity/scale/rotation kernel formulation that VoxGS builds on.","marker":"[75]"},{"why":"Introduces the X-LRM task setup and the dataset collection and evaluation protocol that this paper follows.","marker":"[77]"},{"why":"Defines the ViT architecture used for both the per-view encoder and the cross-view fusion transformer.","marker":"[15]"},{"why":"Provides the pretrained DINO ViT weights that initialize the encoder.","marker":"[78]"},{"why":"The strongest 3D feedforward baseline, representing the voxel-grid direct regression approach the paper argues is inflexible.","marker":"[35]"},{"why":"A self-supervised NeRF baseline whose per-volume optimization time and reconstruction quality are central comparison points.","marker":"[6]"}],"fun_headline_variants":["Full CT from six X-rays in under one second","One-pass CT: sparse X-rays to volume in a second","X-GRM: six views, full CT, no iteration","Sparse X-ray CT via Voxel Gaussian Splatting","Transformer turns 6 X-rays into CT in 1 second"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire evaluation rests on synthetic X-ray projections rendered from CT volumes under simplified X-ray physics with no scatter or beam hardening, so the reported one-second reconstruction gains may not transfer to real clinical scanners if actual projection physics differ.","fun_headline_variants_meta":{"raw":{"variants":["Full CT from six X-rays in under one second","One-pass CT: sparse X-rays to volume in a second","X-GRM: six views, full CT, no iteration","Sparse X-ray CT via Voxel Gaussian Splatting","Transformer turns 6 X-rays into CT in 1 second"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000431,"raw_usage":{"total_tokens":2242,"prompt_tokens":1027,"completion_tokens":1215,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":1130}},"tokens_in":643,"tokens_out":1215,"duration_ms":10293,"temperature":1.0,"reasoning_tokens":1130,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:21:14.701647+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained X-GRM on a real cone-beam CT study with measured polyenergetic projections, or on simulated projections that add scatter and beam hardening, and compare its PSNR and SSIM against per-sample optimization methods; if the quality gap shrinks sharply or anatomical detail develops new artifacts, the paper's equivalence between 3DGS rasterization and X-ray imaging would be broken for practical CT.","supporting_citations":[{"cited_title":"Tigre: a matlab-gpu toolbox for cbct image reconstruction","cited_arxiv_id":null,"evidence_quote":"Supplies the synthetic projection rendering used to generate X-ray images from every CT volume in the dataset."},{"cited_title":"R2-gaussian: Rectifying radiative gaussian splatting for tomographic reconstruction","cited_arxiv_id":null,"evidence_quote":"Provides the CT-oriented Gaussian rasterizer and the opacity/scale/rotation kernel formulation that VoxGS builds on."},{"cited_title":"Learning 3d gaussians for extremely sparse-view cone-beam ct reconstruction","cited_arxiv_id":null,"evidence_quote":"The strongest 3D feedforward baseline, representing the voxel-grid direct regression approach the paper argues is inflexible."},{"cited_title":"Structure-aware sparse-view x-ray 3d reconstruction","cited_arxiv_id":null,"evidence_quote":"A self-supervised NeRF baseline whose per-volume optimization time and reconstruction quality are central comparison points."}],"review_version":1}