REVIEW 2 major objections 5 minor 34 references
FORCE-Interior shows that a Poisson-flow generative prior, anchored to the measured sinogram at every sampling step and started from a full-FOV reconstruction, improves structural and perceptual quality in interior tomography under severe R
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 02:26 UTC pith:7YVMDPYI
load-bearing objection A competent and mostly honest extension of FORCE to interior CT, but the headline p<0.001 claims are overstated because hyperparameters were tuned on test-patient slices and the sample is small and correlated. the 2 major comments →
FORCE-Interior: A Poisson Flow Generative Prior for Interior Tomography Reconstruction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
FORCE-Interior solves the ROI-truncated inverse problem y=MAx+n by sampling from a Poisson-flow generative prior (PFGM++) while enforcing data consistency. Two design choices carry the argument: (1) initializing the sampler with a full-FOV OS-SART reconstruction rather than an ROI-restricted one, so that out-of-ROI attenuation is accounted for and cupping/bias artifacts are reduced; (2) at each denoising step, applying OS-SART updates conditioned on the truncated sinogram before the ODE step, keeping the trajectory anchored to the measurements. Experiments report that at ROI radii 96 and 128 pixels, FORCE-Interior achieves the best ROI PSNR, SSIM, and LPIPS among OS-SART, SART-TV, ROI-CT-CNN
What carries the argument
The load-bearing mechanism is the combination of a full-FOV OS-SART warm-start and per-step data-consistency conditioning on the ROI-truncated sinogram, inside a Poisson-flow generative sampler (an EDM-based PFGM++ that learns a denoiser/score for full-dose CT images). The full-FOV initialization accounts for out-of-ROI attenuation, while per-step OS-SART updates enforce agreement with the measured truncated projections at every denoising iteration, preventing the sampling trajectory from drifting to measurement-inconsistent anatomy. A lightweight total-variation proximal step moderates noise amplification from repeated OS-SART corrections.
Load-bearing premise
The reported gains rest on the assumption that the hyperparameters (TV weight lambda=0.003 and starting noise level t_start=0.6) chosen by ablations on a small random subset are not effectively tuned to the structure-rich test slices, so the improvements are not an artifact of selection on the evaluation set.
What would settle it
A concrete check: retune the TV weight and starting noise level on a truly separate validation set disjoint from the evaluation slices, then re-run the comparison; if the PSNR/SSIM/LPIPS margins at radii 96 and 128 shrink to non-significance, the claim of out-of-the-box improvement is weakened. Additionally, applying the method to a different public CT dataset with varied anatomy and no structure-rich filtering would test whether the gains generalize beyond the specific test pool.
If this is right
- If the central claim holds, generative-model-based CT reconstruction can be extended to interior tomography by explicitly modeling the truncation mask in the data-consistency update.
- The full-FOV warm start implies that out-of-ROI attenuation must be accounted for even when only the ROI is of interest, challenging ROI-restricted initialization practices.
- Per-step data consistency keeps projection-domain residuals near the noise floor; removing it raises residuals by roughly 24x at radius 128, indicating that the anchoring is essential.
- The method retains small inserted lesions (false-negative rate 0) while a generative prior without data consistency misses all lesions, suggesting the anchoring preserves clinically relevant structures.
- The advantage grows as truncation becomes more severe, so the approach is most relevant for tightly targeted ROIs where the interior problem is hardest.
Where Pith is reading between the lines
- The same warm-start-plus-per-step-DC recipe could likely be applied to other generative priors (e.g., score-based diffusion) to test whether the gains stem from the PFGM++ prior specifically or from the anchoring scheme generally.
- The noise-level sweep shows no per-dose retuning is needed, suggesting a path toward protocol-agnostic reconstruction, but the fixed TV weight chosen at the main dose may slightly oversmooth high-dose data, as the paper itself notes for LPIPS.
- A testable extension would evaluate the method on volumetric or cone-beam geometries and on anatomy beyond the structure-rich subset, since the reported statistics are slice-level and from two patients, which may not capture full clinical variability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes FORCE-Interior, a generative reconstruction framework for interior tomography that combines a pre-trained PFGM++ (Poisson-flow generative) prior with a full-FOV OS-SART warm start and per-step data-consistency conditioning against truncated ROI measurements. The method is evaluated on the 2016 NIH-AAPM-Mayo Low-Dose CT dataset at three ROI radii (96, 128, 160 px) under Poisson noise, against OS-SART, SART-TV, ROI-CT-CNN, and DDS. The central claim is that FORCE-Interior achieves improved structural and perceptual quality (ROI PSNR, SSIM, LPIPS) at the two smaller ROI sizes with Holm-adjusted p<0.001, and competitive results at the largest, while preserving projection-domain consistency. Ablations and additional analyses (warm-start variants, TV-weight sweep, noise-level sweep, projection-residual consistency, small-lesion retention) are provided.
Significance. If the central claim holds, the paper makes a useful methodological contribution: it demonstrates how a foundation-style generative prior can be adapted to the interior-tomography setting through a measurement-anchored warm start and per-step data consistency, rather than treating generative sampling as a plug-in post-processor. The experimental design is careful in several respects: paired Wilcoxon tests with Holm correction, multiple ROI radii, a dose sweep without per-dose retuning, projection-domain consistency metrics, and a lesion-retention analysis. These features make the evaluation substantially more informative than a simple PSNR table. However, the headline significance claims are weakened by the unresolved provenance of the hyperparameter-selection subset and by slice-level statistics over two patients, as detailed below.
major comments (2)
- [Section IV-D3, Fig. 5, Table V, Appendix A] The two key hyperparameters (TV weight λ=0.003 and starting noise level t_start=0.6) are selected via ablations on a 'random 20-slice subset.' The manuscript never states that this subset is disjoint from the structure-rich test set (N=34 at r=128) on which the headline comparisons in Table I are made. Appendix A mentions a separate pool of 20 held-out slices used for the radius sweep, but does not state whether the hyperparameter ablations use that pool or a random subset of the test partition. If the latter, the test data have influenced the model configuration, making the subsequent paired Wilcoxon p<0.001 results in Table VIII partially circular for those hyperparameters. Please clarify the provenance of the ablation subset, and ideally re-select hyperparameters on a truly held-out validation set and re-run the benchmark.
- [Appendix A, Statistical analysis; Section IV-A] The significance claims rest on slice-level paired Wilcoxon tests with N=18, 34, and 16 slices drawn from only two patients. The appendix concedes that 'adjacent slices from the same patient may not be fully independent,' but no correction is made for within-patient correlation. Treating correlated slices as independent inflates the effective sample size; with two patients, the reported Holm-adjusted p<0.001 are not credible as evidence about patient-level generalization. Please provide patient-level summaries (e.g., per-patient mean differences), a mixed-effects analysis, or a conservative cluster-based correction, and temper the abstract's significance claim accordingly.
minor comments (5)
- [Abstract/Introduction] Typo: 'proposeFORCE-Interior' should be 'propose FORCE-Interior'.
- [Table I caption] The heading 'POISSON-NOISE RECONSTRUCTION ATr∈{96,128,160}PX' appears to have a formatting error ('ATr' should be 'at r'). Also, the ROI-CT-CNN caveat (trained for r=128 only, evaluated under a shifted distribution) should appear in the table caption since the table includes it at r=128.
- [Fig. 5 caption / Section IV-D3] Please clarify which 'random 20-slice subset' is used for the TV-weight and t_start ablations and whether it is the same as the 'separate pool of 20 held-out slices' described in Appendix A. This is important for the reader to assess potential leakage.
- [Algorithm 3 / Eq. (14)] The notation M_g A_{S_g} and the update equation could be defined more explicitly; currently it is not immediately clear how the subset mask M_g interacts with the projection rows A_{S_g}. A brief sentence connecting Eq. (14) to the truncation mask M in Eq. (10) would help.
- [Section IV-C2 / Appendix A] The lesion-retention analysis uses '20 structure-rich test slices.' It would be useful to confirm whether these are the same 20 slices used for hyperparameter tuning or the radius-sweep pool, and to report the overlap explicitly.
Circularity Check
No significant circularity: central claims rest on held-out empirical comparisons, not on a fitted parameter or self-citation chain; noted caveats are statistical, not constructional.
full rationale
The paper does not derive its headline result from its own inputs by construction. FORCE-Interior is defined by explicit equations (Algorithm 3; Eqs. (13)-(14)) combining a pretrained PFGM++ prior, full-FOV OS-SART warm start, per-step OS-SART data consistency, and TV proximal steps, and it is evaluated against OS-SART, SART-TV, ROI-CT-CNN, and DDS on two held-out Mayo patients. The central metric improvements are empirical outcomes, not identities. The reliance on the authors' own FORCE/PFGM++ prior is a normal self-citation: the prior's training objective (Eq. (7)) does not include the interior-tomography benchmark, and the measured gains come from the experimental protocol rather than from the prior's objective. Likewise, the projection-consistency results are not circular: they are presented as a check of the enforced constraint, with an ablation (w/o per-step DC) showing a 24x residual increase, so the analysis has independent content. The appendix explicitly notes that adjacent slices from the same two patients may not be fully independent, and the TV weight/t_start were selected on a random 20-slice subset (Section IV-D3, Table V); these are statistical validity caveats (potential dependence/selection on evaluation data) rather than a reduction of prediction to input. No equation in the paper equates a fitted parameter with the reported outcome by construction, and no load-bearing claim rests solely on an unverified self-citation. Hence no significant circularity; score 1 reflecting only minor self-citation and mild statistical caveats.
Axiom & Free-Parameter Ledger
free parameters (4)
- TV weight lambda =
0.003
- Starting noise level t_start =
0.6
- Per-step OS-SART settings =
G=16 subsets, 3 iterations, relaxation omega=1
- Warm-start OS-SART iterations =
30 iterations, 16 subsets
axioms (4)
- standard math OS-SART converges as a data-consistency operator for truncated projections
- domain assumption The PFGM++/EDM denoiser approximates the score of the normal-dose CT image distribution
- ad hoc to paper A full-FOV OS-SART warm start from truncated projections captures out-of-ROI attenuation
- domain assumption Slice-level paired statistics on adjacent slices approximate independent samples
Cite this review
Pith. "Pith review of FORCE-Interior: A Poisson Flow Generative Prior for Interior Tomography Reconstruction." pith.science (2026). https://pith.science/paper/7YVMDPYI
@misc{pith2026260714320,
author = {Pith},
title = {Pith review of: FORCE-Interior: A Poisson Flow Generative Prior for Interior Tomography Reconstruction},
year = {2026},
howpublished = {\url{https://pith.science/paper/7YVMDPYI}},
note = {Machine review of arXiv:2607.14320}
}
read the original abstract
Interior tomography reconstructs a region of interest (ROI) from truncated projection measurements. However, projection truncation makes the inverse problem severely ill-posed, leading to non-unique solutions, oversmoothing, and truncation-induced artifacts when conventional reconstruction methods are directly applied to interior tomography. Existing learning-based interior CT methods have shown promising performance, but their generalization across different truncation patterns, ROI sizes, and noise levels remains an important challenge. Meanwhile, current generative model-based reconstruction methods are primarily designed for non-interior tomography settings and do not directly address ROI-based projection truncation. Moreover, without sufficient data-consistency constraints, generative sampling may yield anatomically plausible but measurement-inconsistent structures. To address these challenges, we propose FORCE-Interior, a Poisson-flow generative reconstruction framework for interior tomography. FORCE-Interior combines a full-FOV measurement-constrained initialization with per-step data consistency for ROI-truncated measurements, anchoring generative sampling to the acquired measurements throughout reconstruction. Experiments show that FORCE-Interior achieves improved structural and perceptual reconstruction quality at the two more severely truncated ROI sizes, with competitive results at the largest, while maintaining projection-domain consistency.
Figures
Reference graph
Works this paper leans on
-
[1]
The meaning of interior tomography,
G. Wang and H. Yu, “The meaning of interior tomography,”Phys. Med. Biol., vol. 58, no. 16, pp. R161–R186, 2013
2013
-
[2]
Compressed sensing based interior tomography,
H. Yu and G. Wang, “Compressed sensing based interior tomography,” Phys. Med. Biol., vol. 54, no. 9, pp. 2791–2805, 2009
2009
-
[3]
Image reconstruction for sparse-view CT and interior CT: Introduction to compressed sensing and differentiated backprojection,
H. Kudo, T. Suzuki, and E. A. Rashed, “Image reconstruction for sparse-view CT and interior CT: Introduction to compressed sensing and differentiated backprojection,”Quant. Imaging Med. Surg., vol. 3, no. 3, pp. 147–161, 2013
2013
-
[4]
A practical local tomography reconstruction algorithm based on a known sub-region,
P. Paleo, M. Desvignes, and A. Mirone, “A practical local tomography reconstruction algorithm based on a known sub-region,”Journal of Synchrotron Radiation, vol. 24, no. 1, pp. 257–268, 2017
2017
-
[5]
A. C. Kak and M. Slaney,Principles of Computerized Tomographic Imaging. New York, NY , USA: IEEE Press, 1988
1988
-
[6]
FBP and the interior problem in 2D tomography,
A. Bilgot, L. Desbat, and V . Perrier, “FBP and the interior problem in 2D tomography,” in2011 IEEE Nuclear Science Symposium Conference Record, 2011, pp. 4080–4085
2011
-
[7]
Simultaneous algebraic reconstruction technique (SART): A superior implementation of the ART algorithm,
A. H. Andersen and A. C. Kak, “Simultaneous algebraic reconstruction technique (SART): A superior implementation of the ART algorithm,” Ultrason. Imaging, vol. 6, no. 1, pp. 81–94, 1984
1984
-
[8]
Ordered-subset simultaneous algebraic recon- struction techniques (OS-SART),
G. Wang and M. Jiang, “Ordered-subset simultaneous algebraic recon- struction techniques (OS-SART),”J. X-Ray Sci. Technol., vol. 12, no. 3, pp. 169–177, 2004
2004
-
[9]
Bone-induced streak artifact suppression in sparse-view CT image reconstruction,
S. O. Jin, J. G. Kim, S. Y . Lee, and O.-K. Kwon, “Bone-induced streak artifact suppression in sparse-view CT image reconstruction,”BioMed. Eng. OnLine, vol. 11, p. 44, 2012
2012
-
[10]
Artifact reduction methods for truncated projections in iterative breast tomosynthesis reconstruction,
Y . Zhang, H.-P. Chan, B. Sahiner, J. Wei, C. Zhou, and L. M. Hadjiiski, “Artifact reduction methods for truncated projections in iterative breast tomosynthesis reconstruction,”J. Comput. Assist. Tomogr ., vol. 33, no. 3, pp. 426–435, 2009
2009
-
[11]
A diffusion-based truncated projection artifact reduction method for iterative digital breast tomosynthesis reconstruction,
Y . Lu, H.-P. Chan, J. Wei, and L. M. Hadjiiski, “A diffusion-based truncated projection artifact reduction method for iterative digital breast tomosynthesis reconstruction,”Phys. Med. Biol., vol. 58, no. 3, pp. 569– 587, 2013
2013
-
[12]
Image reconstruction in circular cone-beam computed tomography by constrained, total-variation minimization,
E. Y . Sidky and X. Pan, “Image reconstruction in circular cone-beam computed tomography by constrained, total-variation minimization,” Phys. Med. Biol., vol. 53, no. 17, pp. 4777–4807, 2008
2008
-
[13]
Prior image constrained compressed sensing (PICCS): A method to accurately reconstruct dynamic CT images from highly undersampled projection data sets,
G.-H. Chen, J. Tang, and S. Leng, “Prior image constrained compressed sensing (PICCS): A method to accurately reconstruct dynamic CT images from highly undersampled projection data sets,”Med. Phys., vol. 35, no. 2, pp. 660–663, 2008
2008
-
[14]
Solving inverse problems using data-driven models,
S. Arridge, P. Maass, O. ¨Oktem, and C.-B. Sch ¨onlieb, “Solving inverse problems using data-driven models,”Acta Numer ., vol. 28, pp. 1–174, 2019
2019
-
[15]
Low-dose ct reconstruction via edge-preserving total variation regularization,
Z. Tian, X. Jia, K. Yuan, T. Pan, and S. B. Jiang, “Low-dose ct reconstruction via edge-preserving total variation regularization,” Physics in Medicine and Biology, vol. 56, no. 18, p. 5949–5967, Aug. 2011. [Online]. Available: http://dx.doi.org/10.1088/0031- 9155/56/18/011
doi:10.1088/0031- 2011
-
[16]
Deep learning interior tomography for region-of-interest reconstruction,
Y . Han, J. Gu, and J. C. Ye, “Deep learning interior tomography for region-of-interest reconstruction,” 2018. [Online]. Available: https://arxiv.org/abs/1712.10248
Pith/arXiv arXiv 2018
-
[17]
One network to solve all rois: Deep learning ct for any roi using differentiated backprojection,
Y . Han and J. C. Ye, “One network to solve all rois: Deep learning ct for any roi using differentiated backprojection,” 2019. [Online]. Available: https://arxiv.org/abs/1810.00500
Pith/arXiv arXiv 2019
-
[18]
End-to-end deep learning for interior tomography with low-dose x-ray ct,
Y . Han, D. Wu, K. Kim, and Q. Li, “End-to-end deep learning for interior tomography with low-dose x-ray ct,” 2025. [Online]. Available: https://arxiv.org/abs/2501.05085
Pith/arXiv arXiv 2025
-
[19]
Solving inverse problems in medical imaging with score-based generative models,
Y . Song, L. Shen, L. Xing, and S. Ermon, “Solving inverse problems in medical imaging with score-based generative models,” inProc. Int. Conf. Learn. Represent. (ICLR), 2022
2022
-
[20]
Dolce: A model-based probabilistic diffusion framework for limited-angle ct reconstruction,
J. Liu, R. Anirudh, J. J. Thiagarajan, S. He, K. A. Mohan, U. S. Kamilov, and H. Kim, “Dolce: A model-based probabilistic diffusion framework for limited-angle ct reconstruction,” 2022. [Online]. Available: https://arxiv.org/abs/2211.12340
Pith/arXiv arXiv 2022
-
[21]
Generative modeling in sinogram domain for sparse-view ct reconstruction,
B. Guan, C. Yang, L. Zhang, S. Niu, M. Zhang, Y . Wang, W. Wu, and Q. Liu, “Generative modeling in sinogram domain for sparse-view ct reconstruction,” 2022. [Online]. Available: https://arxiv.org/abs/2211.13926
Pith/arXiv arXiv 2022
-
[22]
Tomographic foundation model – force: Flow-oriented reconstruction conditioning engine,
W. Xia, C. Niu, and G. Wang, “Tomographic foundation model – force: Flow-oriented reconstruction conditioning engine,” 2025. [Online]. Available: https://arxiv.org/abs/2506.02149
Pith/arXiv arXiv 2025
-
[23]
Decomposed diffusion sampler for accelerating large-scale inverse problems,
H. Chung, S. Lee, and J. C. Ye, “Decomposed diffusion sampler for accelerating large-scale inverse problems,” 2024. [Online]. Available: https://arxiv.org/abs/2303.05754
Pith/arXiv arXiv 2024
-
[24]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” 2020. [Online]. Available: https://arxiv.org/abs/2006.11239
Pith/arXiv arXiv 2020
-
[25]
Score-based generative modeling through stochastic differ- ential equations,
Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differ- ential equations,” inProc. Int. Conf. Learn. Represent. (ICLR), 2021
2021
-
[26]
Elucidating the design space of diffusion-based generative models,
T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” inProc. Adv. Neural Inf. Process. Syst., 2022
2022
-
[27]
Poisson flow generative models,
Y . Xu, Z. Liu, M. Tegmark, and T. Jaakkola, “Poisson flow generative models,” inProc. Adv. Neural Inf. Process. Syst., 2022
2022
-
[28]
PFGM++: Unlocking the potential of physics-inspired generative mod- els,
Y . Xu, Z. Liu, Y . Tian, S. Tong, M. Tegmark, and T. Jaakkola, “PFGM++: Unlocking the potential of physics-inspired generative mod- els,” inProc. Int. Conf. Mach. Learn. (ICML), 2023, pp. 38 566–38 591
2023
-
[29]
Image quality assessment: From error visibility to structural similarity,
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,”IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004
2004
-
[30]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2018, pp. 586–595
2018
-
[31]
A first-order primal-dual algorithm for convex problems with applications to imaging,
A. Chambolle and T. Pock, “A first-order primal-dual algorithm for convex problems with applications to imaging,”J. Math. Imaging Vis., vol. 40, no. 1, pp. 120–145, 2011
2011
-
[32]
Tigre v3: Efficient and easy to use iterative computed tomographic reconstruction toolbox for real datasets,
A. Biguri, T. Sadakane, R. Lindroos, Y . Liu, M. Sabat´e Landman, Y . Du, M. Lohvithee, S. Kaser, S. Hatamikia, R. Bryll, E. Valat, S. Wonglee, T. Blumensath, and C.-B. Sch ¨onlieb, “Tigre v3: Efficient and easy to use iterative computed tomographic reconstruction toolbox for real datasets,”Engineering Research Express, vol. 7, no. 1, p. 015011, mar
-
[33]
DM4CT: Benchmarking dif- fusion models for computed tomography reconstruction,
J. Shi, D. M. Pelt, and K. J. Batenburg, “DM4CT: Benchmarking dif- fusion models for computed tomography reconstruction,”arXiv preprint arXiv:2602.18589, 2026
arXiv 2026
-
[2025]
Available: https://doi.org/10.1088/2631-8695/adbb3a
[Online]. Available: https://doi.org/10.1088/2631-8695/adbb3a
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.