{"id":"699b2dea-e83b-4c1b-b2ab-1dfacfb32165","arxiv_id":"2608.04752","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Splat-based CT artifacts under sparse views are traced to pose inaccuracy, and a joint pose-volume refinement substantially improves reconstruction quality.","lead":"This paper shows that streaky artifacts in sparse-view, splat-based CT come mainly from camera pose errors, not from having too few X-ray views. It adds a self-calibrating pose refinement step that jointly corrects geometry while reconstructing the volume, and this reduces artifacts in real and simulated scans.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Pose-attribution claim rests on an FDK pseudo-GT and a reproduction experiment that do not exclude systematic geometry bias or non-pose forward-model mismatches; the causal evidence remains incomplete.","rationale":"The paper's synthetic experiments with known pose perturbations convincingly show that pose errors can produce the observed needle-like artifacts and that joint pose refinement improves reconstruction, and the proposed method is a plausible lightweight adaptation of existing self-calibrating splatting. The reader's weakest assumption, the FDK pseudo-GT, is real and important, but the more load-bearing issue is that the entire real-data argument attributes the observed artifacts to pose specifically, whereas the experiments as described cannot distinguish pose errors from other forward-model mismatches such as scatter, beam hardening, or detector response. The supplement's Eq. 8-10 justification for the pseudo-GT additionally assumes independent zero-mean per-view pose errors, an assumption that is often violated by systematic calibration errors in real CT systems, so the pseudo-GT itself may carry the very bias under investigation. The Sec. 7 reproduction experiment is the natural decisive test, but its one-sentence description omits how pose errors were estimated; if those estimates come from the same reconstruction objective, the test is circular and cannot exclude alternative explanations. Because the core causal claim about real data is not yet secured, the paper should remain conditional on a direct test with independent geometry information. This does not change the reader's verdict, but it sharpens the condition that must be met before the attribution claim can be accepted as established.","tokens_in":17247,"tokens_out":9336,"duration_ms":121158,"concrete_test":"Acquire or use an existing real cone-beam dataset with a calibration phantom containing fiducial markers to obtain independent per-view ground-truth poses (e.g., by marker triangulation or encoded stage positions), without using the reconstruction objective. From the 721-view FDK pseudo-GT volume, generate two synthetic 75-view projection sets, one with ideal poses and one with the independently measured pose errors. Reconstruct both with the baseline R2-Gaussian pipeline and compare artifact patterns with the real 75-view reconstruction. The claim is supported only if the measured-pose synthetic set reproduces the real needle/stripe artifacts and the ideal-pose set does not; repeating with scatter-reduction or monochromatic acquisition would further separate pose errors from beam-hardening and scatter effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that real sparse-view splat artifacts are caused primarily by pose inaccuracies. The decisive evidence is the Fig. 2-3 comparison and the supplement's Sec. 7 reproduction. The Fig. 2 comparison assumes the 721-view FDK volume is a faithful pseudo-GT; the supplement's justification (RMSE proportional to 1/sqrt(N), Eqs. 8-10) requires independent, zero-mean per-view pose perturbations. Real CT calibration errors are frequently systematic (detector tilt, source offset, stage tilt), in which case the FDK average does not remove the bias and the pseudo-GT inherits the geometry error. Even with a perfect pseudo-GT, the real-versus-synthetic asymmetry only shows that real projections are inconsistent with an ideal forward model of a single volume; it does not identify pose as the cause. Scatter, beam hardening, detector response, and cone-beam approximation also produce edge-correlated directional residuals. The Sec. 7 reproduction experiment could resolve this, but it does not state how pose errors were estimated from real data; if they are estimated by the same joint reconstruction objective, the reproduction is circular. The load-bearing condition, that pose error rather than another model mismatch is the primary cause, is therefore unsupported by the current evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies splat-based (Gaussian) cone-beam CT reconstruction under sparse-view conditions. The authors observe that real CT scans produce strong streak/strip artifacts that are absent in synthetic sparse-view reconstructions, and they argue that these artifacts are primarily caused by camera pose inaccuracies rather than by view sparsity. They support this with an analysis experiment (Section 3) that compares reconstructions from real 75-view projections against reconstructions from 75 synthetic projections derived from an FDK-based pseudo ground truth, and with a supplementary reproducibility experiment. The method contribution is a self-calibrating splat-based pipeline that jointly optimizes Gaussian volume parameters and per-view camera poses via a stable gradient formulation, without TV regularization. Experiments on 15 synthetic scenes with injected pose noise show consistent PSNR/SSIM improvements over prior joint calibration and reconstruction methods, and qualitative results on a real dataset show artifact reduction. The central claim is that pose error, not sparsity, is the dominant artifact source, and that the proposed joint refinement substantially improves reconstruction fidelity.","tokens_in":17440,"tokens_out":3527,"duration_ms":44859,"significance":"If the causal claim is correct, the paper identifies a practical bottleneck for splat-based CT and offers a lightweight, differentiable self-calibration method that can be integrated into existing pipelines. The synthetic evaluation is broad: 15 scenes, multiple pose-noise levels, view-count ablations at 75/50/25 views, pose-estimation RMSE tables, and a controlled perturbation-generation scheme with an explicit Lie-algebra derivation. The method also demonstrates a clear quantitative advantage on synthetic data over strong baselines (NeAT, Thies et al.). The paper is explicit about its limitation in extreme sparse-view settings and the trade-off with TV regularization, which adds credibility. However, the evidence for the primary causal attribution is weaker than the language of the abstract suggests: the real-data artifact analysis relies on an FDK pseudo ground truth and a reproduction experiment whose estimation procedure is not fully specified. The central contribution is defensible, but the load-bearing artifact-attribution evidence needs strengthening.","major_comments":[{"comment":"The FDK pseudo-ground-truth assumption is load-bearing for the claim that artifacts are due to pose errors rather than view sparsity. The support offered in Supplement Section 2, Eqs. (5)-(10), assumes independent zero-mean Gaussian per-view pose perturbations and an approximately constant Jacobian across views. Real CT calibration errors are often systematic (e.g., detector tilt, source offset, stage wobble), in which case the FDK average over 721 views does not remove the bias, and the pseudo-GT inherits the geometry error. With such systematic bias, the synthetic 75-view projections are generated from a biased volume, so the reduced artifacts in Figure 3(b) do not cleanly demonstrate the absence of sparsity-induced artifacts. The paper should either justify that the real dataset's geometry errors are independent across views, or validate the pseudo-GT against an independent high-quality reference (e.g., a phantom scan with known geometry).","section":"Section 3; Supplement Section 7"},{"comment":"The reproduction experiment that establishes a 'direct causal link' between pose errors and needle-like artifacts is incompletely specified. The text states that pose errors are 'estimated from real sparse-view data' and applied to generate simulated projections from the ground-truth volume, but it does not state how these pose errors are obtained. If the pose estimates come from the same joint reconstruction objective that produces the artifact-laden reconstruction, then the reproduction is circular: the pose estimates would already be biased by the method's assumptions. To avoid circularity, the pose errors should be measured independently (e.g., from known phantom geometry, from a separate calibration scan, or from a method that does not use the splat reconstruction objective), and the reproduction should demonstrate that the simulated artifacts match the real ones both qualitatively and quantitatively (e.g., by comparing artifact orientation, location, and magnitude).","section":"Supplement Section 7"},{"comment":"The directional edge-correlated bias in the real-data error maps is evidence of model mismatch, but it does not by itself identify pose inaccuracy as the cause. Scatter, beam hardening, detector response nonuniformity, and cone-beam approximation also produce edge-correlated directional residuals under sparse-view sampling. The paper's conclusion that the bias 'reveals the inaccuracy of camera poses' (Section 3, final sentence) is stronger than what the evidence supports. A control experiment is needed that injects independently measured pose errors into a forward model with a known ground-truth volume and checks that the resulting error maps match the real-data pattern; alternatively, a phantom with known geometry and no other physical effects (scatter, beam hardening) could be scanned to isolate pose as the sole variable.","section":"Section 3, Figure 4"},{"comment":"The real-world evaluation is qualitative only. The central claim concerns behavior under 'real-world sparse-view conditions,' yet no quantitative metric is reported on the real dataset (e.g., against a high-quality reference volume, a calibration phantom, or a consistent image-quality measure). Given that the synthetic results already show strong quantitative gains, the paper would be strengthened by a quantitative real-data evaluation, even if the reference is an FDK volume from dense projections (with the caveat discussed above) or a known phantom. This would also make the 'substantially improves reconstruction fidelity' claim in the abstract measurable rather than visual.","section":"Section 5.3, Figure 7"}],"minor_comments":[{"comment":"The mapping function φ(·) is defined as returning a 3D vector, but the projection ψ(·) is described as R^3→R^2; please clarify the relationship between φ and ψ and the role of the third component in the projection of the Gaussian center.","section":"Section 4, Eq. (6)"},{"comment":"The statement that translation noise σ_trans is '1.0 in unit length (corresponding to the voxel size)' should be stated more precisely: is the unit length exactly one voxel of the reconstruction volume, and is the same scaling used across all scenes? The phrase 'may vary with dataset scaling' is ambiguous.","section":"Section 5.2"},{"comment":"Table 3 reports only two scenes (Beetle and Head), while the text says 'Reconstruction performance with camera noise varies.' Please either report per-scene results for all scenes or state that the table is representative and include the full results in the supplement.","section":"Table 3"},{"comment":"The description of the Thies et al. comparison states that the score network was not used and instead an L2 loss against an FBP reference was employed. This is a significant deviation from the original method and should be disclosed in the main text, not only in the supplement, to avoid misleading readers about the strength of the baseline.","section":"Supplement Section 4"},{"comment":"The orientation error formula θ = (1/N) sqrt(arccos(tr(R)-1)/2) is missing parentheses; as written, it is ambiguous whether the square root applies to the entire fraction. Please revise to θ = (1/N) sqrt((arccos(tr(R)-1))/2) or an equivalent clear expression.","section":"Section 5.3, Eq. (12)"}],"recommendation":"major_revision","confidential_remarks":"The paper's method contribution is solid and the synthetic evaluation is strong, but the central causal claim ('artifacts primarily originate from pose inaccuracies') rests on an FDK pseudo-GT and a reproduction experiment that are not yet rigorous enough. The authors should be asked to either strengthen the artifact-attribution evidence (e.g., with an independent, non-circular pose-error estimate and a control experiment) or temper the causal language in the abstract and conclusions. This is a correctable issue within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the engineering result here is solid and useful. Jointly refining pose parameters during splat-based CT reconstruction visibly cleans up real sparse-view data, and the artifact-attribution comparison between real and synthetic projections is genuinely new. The abstract's causal claim—that artifacts 'primarily originate from pose inaccuracies... rather than view sparsity'—is a bit stronger than the evidence supports, but not by a huge margin.\n\nWhat's new: adapting the self-calibrating 3DGS recipe (incremental quaternion/translation updates, L1+SSIM loss) to CT is an incremental step over [7,33,49], and the paper doesn't oversell that. The new value is the diagnostic pipeline in Fig. 2: comparing splat reconstructions from real 75-view projections vs. synthetic 75-view projections re-rendered from an FDK pseudo-GT, plus the projection-error maps showing directional bias only for real data. The synthetic evaluation is broad—15 scenes, pose RMSE tables, ablations at 75/50/25 views, and direct comparisons to NeAT and Thies et al. under the same pose perturbations. The supplemental pose-calibration on/off table is a nice control. I believe the central practical claim is probably true: pose errors are a major contributor to the needle artifacts, because optimizing poses alone removes most of them on real data.\n\nSoft spots, in order. First, the FDK pseudo-GT assumption. The supplement's 1/sqrt(N) derivation assumes zero-mean, independent, per-view pose perturbations. Real CT calibration errors are often systematic (detector tilt, source offset, stage tilt), so the FDK average can inherit bias and the pseudo-GT may not be faithful. Even with a perfect pseudo-GT, the real-vs-synthetic asymmetry only shows inconsistency with an ideal forward model; scatter, beam hardening, detector response, and cone-beam approximation also produce edge-correlated directional residuals. Second, the supplement's artifact reproduction experiment is suggestive but doesn't state how pose errors were estimated from real data; if they come from the same joint objective, the demonstration is partially circular. Third, minor: the SSIM weight lambda is unreported, there are no error bars anywhere, no code is released, and real-data evaluation is qualitative only.\n\nNone of these are load-bearing enough to sink the paper. The method works, the synthetic evidence is strong, and the practical takeaway—calibration matters more than sparsity for splat CT—is well supported by the fact that pose optimization alone removes most real artifacts. The authors should either soften the causal language or tighten the pseudo-GT analysis with a systematic-error experiment (e.g., inject detector tilt into the pseudo-GT pipeline). Who this is for: anyone working on splat-based or neural-field CT reconstruction, and people building online calibration for cone-beam systems. It deserves a serious referee; I'd send it to review rather than desk-reject, expecting a revise-and-resubmit.","headline":"Solid self-calibrating splat CT with a genuinely useful diagnostic experiment; the pose-attribution claim is slightly stronger than the evidence, but the method works and deserves a serious referee.","tokens_in":18007,"tokens_out":3087,"would_cite":true,"duration_ms":36614,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Pose inaccuracies, not view sparsity, cause the streak artifacts in splat-based CT under sparse views, and joint geometric refinement during reconstruction removes them.","keywords":["cone-beam CT","sparse-view reconstruction","Gaussian splatting","geometric calibration","pose refinement","self-calibrating tomography","CT artifact analysis"],"falsifier":"Reconstruct a known calibration phantom with a real cone-beam system, obtain a high-precision pose estimate from independent phantom-based calibration, then deliberately add known pose perturbations. If splat-based reconstruction from 75 views with deliberately wrong poses reproduces the needle and streak artifacts while the same views with correct poses do not, the causal role of pose errors is directly confirmed; if artifacts persist when a high-precision independent calibration is used, the paper's attribution would be falsified.","tokens_in":17011,"feed_emoji":"🩻","tokens_out":6598,"duration_ms":71215,"temperature":0.7,"pith_summary":"Splat-based CT, which reconstructs a volume as a cloud of anisotropic 3D Gaussians, produces pronounced streak and strip artifacts when applied to real sparse-view X-ray data. This paper argues that the artifacts are caused not by having too few projection views, but by small inaccuracies in the camera poses of the acquisition geometry. The argument is supported by a controlled comparison: when 75 projections are re-synthesized from a dense 721-view FDK reconstruction, the streaks vanish, while the same number of real projections still show them. The paper then derives a stable, differentiable refinement of pose parameters that is optimized jointly with the Gaussian volume, showing that this self-calibrating pipeline suppresses the artifacts and improves reconstruction fidelity under synthetic and real sparse-view conditions.","feed_headline":"Pose errors, not sparse views, create CT streaks","feed_subtitle":"Refining camera geometry during reconstruction removes streaks without blurring fine detail.","key_machinery":"The central object is the splat-based forward-projection operator: the attenuation volume is a continuous sum of 3D Gaussians, and each ray's intensity is computed analytically by integrating the projected Gaussian density along the ray. The paper's mechanism for stable self-calibration is the explicit gradient flow through the pose parameters: each camera pose is parameterized as a small incremental quaternion and translation relative to its initial value, and the backward pass is tracked through both the perspective projection matrix $\\mathbf{P}_k$ and the rotation matrix $\\mathbf{W}_k$ via the intermediate matrix $\\mathbf{M}_{i,k} = \\mathbf{J}_{i,k} \\mathbf{W}_k$. The key correction is the term $\\partial \\mathcal{L}/\\partial \\mathbf{W}_{k,b}$, which previous splatting-based calibration omitted. This formulation adds only seven pose parameters per view and allows the volume and geometry to be optimized together with a simple L1 plus SSIM loss, making total-variation regularization unnecessary.","core_discovery":"The central discovery is that splat-based CT is far more sensitive to acquisition-geometry errors than to view sparsity alone. Under pose perturbations typical of real rotating systems, the anisotropic Gaussian representation propagates small rotations and translations into large errors in the rendered attenuation, producing needle-like streaks that are far more pronounced than in conventional reconstruction methods. The paper identifies the missing piece in prior splat-based CT: the gradient through the rotation matrix of the camera pose was neglected, and once the full Jacobian flow through both the projection matrix and the rotation matrix is tracked, joint pose-and-volume optimization becomes stable and accurate. The resulting method recovers camera poses to about 0.6 degrees and 0.7 voxel of translation on average, removes the artifacts, and does so without the total-variation regularization that prior splat-based methods needed, which had over-smoothed fine structure.","pith_inferences":["If pose error dominates sparse-view artifacts in splat-based CT, the same joint-calibration principle should transfer to other explicit differentiable volume representations such as voxel grids or triplane features that currently rely on regularization to hide misalignment.","The paper's artifact-attribution diagnostic—comparing reconstructions from real projections against resimulated projections of a pseudo ground truth and looking for a directional bias in the projection-error map—could be adopted as a general calibration-quality test in other CT systems.","A natural extension is an analytical model of the pose-error tolerance of a splat volume as a function of Gaussian covariance and ray length, letting practitioners know the maximum mechanical wobble that a scan can tolerate before streaks appear.","If confirmed, the results suggest that in clinical and industrial CT, investing in stage mechanical stability may yield larger image-quality gains than adding more views."],"forward_implications":["Real-world sparse-view CT with splat-based reconstruction can become artifact-free by correcting camera poses, instead of relying on TV smoothing that blurs fine detail.","The estimated pose corrections are reusable: they can be exported and applied to other tomographic reconstruction methods, not just the Gaussian splatting pipeline.","Self-calibration removes the need for offline phantom-based geometric calibration, allowing reconstruction on systems with mechanical imperfections or drift during scanning.","The method remains effective as view count drops from 75 to 25 views, extending the usable sparsity range of splat-based CT before regularization is needed.","Because pose error, not sparsity, is the dominant failure source, system-design efforts can focus on the geometric stability of the rotation stage rather than on merely acquiring more views."],"supporting_citations":[{"why":"The splat-based CT baseline whose streak artifacts motivate the analysis; supplies the Gaussian volume representation and the forward-projection model that the paper's derivative analysis extends.","marker":"[48]"},{"why":"FDK reconstruction from 721 real projections provides the pseudo ground-truth volume used in the artifact-attribution experiment.","marker":"[8]"},{"why":"Real cone-beam CT dataset (walnut, seashell, pine) used to observe artifacts and to evaluate the method under real sparse-view conditions.","marker":"[36]"},{"why":"Introduces 3D Gaussian splatting, the differentiable rendering framework on which the CT volume representation is built.","marker":"[17]"},{"why":"Defines volumetric EW A splatting whose ray-space mapping, including the third component, is used in the forward-projecting equation for X-ray attenuation.","marker":"[50]"},{"why":"Provides the quaternion-and-translation incremental pose parameterization and self-calibrating formulation that the paper adapts to CT.","marker":"[7]"},{"why":"TIGRE toolbox used to generate synthetic cone-beam projection data with controlled pose perturbations for quantitative evaluation.","marker":"[3]"},{"why":"Supplies the Lie-algebra exponential and logarithmic maps used to sample unbiased SO(3) rotational perturbations in the synthetic dataset.","marker":"[35]"},{"why":"NeAT is a state-of-the-art joint volume-and-pose optimization baseline that the paper compares against in reconstruction quality and calibration accuracy.","marker":"[30]"},{"why":"Thies et al. is a gradient-based head-motion-compensation baseline for cone-beam CT that the paper's calibration quality and PSNR results are directly compared with.","marker":"[39]"}],"fun_headline_variants":["Splat CT streaks traced to pose error, not view sparsity","Fixing camera pose removes CT streaks without blur","Pose correction key to clean sparse-view CT","Splat-based CT sensitive to geometry errors, new fix","Revealing true culprit: geometry, not sparsity, in CT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The artifact-attribution experiment assumes that FDK reconstruction from 721 real projections is accurate enough to serve as pseudo ground truth, because all synthetic comparisons are generated from that volume.","fun_headline_variants_meta":{"raw":{"variants":["Splat CT streaks traced to pose error, not view sparsity","Fixing camera pose removes CT streaks without blur","Pose correction key to clean sparse-view CT","Splat-based CT sensitive to geometry errors, new fix","Revealing true culprit: geometry, not sparsity, in CT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000277,"raw_usage":{"total_tokens":1624,"prompt_tokens":890,"completion_tokens":734,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":652}},"tokens_in":506,"tokens_out":734,"duration_ms":8234,"temperature":1.0,"reasoning_tokens":652,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:25:20.056152+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reconstruct a known calibration phantom with a real cone-beam system, obtain a high-precision pose estimate from independent phantom-based calibration, then deliberately add known pose perturbations. If splat-based reconstruction from 75 views with deliberately wrong poses reproduces the needle and streak artifacts while the same views with correct poses do not, the causal role of pose errors is directly confirmed; if artifacts persist when a high-precision independent calibration is used, the paper's attribution would be falsified.","supporting_citations":[{"cited_title":"13th conference on industrial computed tomography (iCT) , year=","cited_arxiv_id":null,"evidence_quote":"The splat-based CT baseline whose streak artifacts motivate the analysis; supplies the Gaussian volume representation and the forward-projection model that the paper's derivative analysis extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"TIGRE toolbox used to generate synthetic cone-beam projection data with controlled pose perturbations for quantitative evaluation."}],"review_version":1}