{"id":"9bda269f-12cc-4d7e-a9ed-962b0f00632c","arxiv_id":"2412.11641","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"cSART outperforms FBP and UTR on objective image quality metrics in low-dose phase-contrast breast CT, but radiologists prefer FBP images, likely because they weight contrast more heavily than noise.","lead":"This study compared three reconstruction algorithms on low-dose phase-contrast breast CT scans of ten mastectomy samples. It found that cSART scored best on objective metrics, but radiologists preferred FBP images, a discrepancy attributed to radiologists weighting contrast over noise.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Shannon-information claim is not established because Q_s = SNR/Res^{1.5} (Eq. 6) is only proven for linear systems, while cSART includes nonlinear bilateral filtering; the metric can be inflated by denoising in the flat ROIs where it is measured.","rationale":"The reader's weakest assumption was parameter tuning. I agree that tuning cSART on the evaluation data and metrics is a serious fairness problem, and Section 2c.ii indeed selects cSART parameters using the same SNR and resolution metrics on slices from the same projection-count groups used later. However, the more load-bearing problem is that the headline interpretation of Q_s as Shannon information per photon is only valid for linear shift-invariant systems, and cSART is nonlinear. The paper itself states that nonlinear processing can change SNR/Res^{1.5}. Since the bilateral filter is an edge-preserving denoiser, it can improve Q_s in flat ROIs without adding information about the object. The resolution estimate from the noise spectrum is also affected by the filter's noise shaping. Thus even if cSART's parameters were tuned on independent data, a higher Q_s would not prove more Shannon information. The subjective preference for FBP is interesting and likely robust, but the objective/information-theoretic conclusion requires either a nonlinear-information-theoretic validation or a ground-truth simulation. Until then, the central claim should be treated as conditional: the paper should be revised to either validate Q_s for cSART or soften the claim to 'nonlinear denoising yields higher measured SNR and resolution proxies.' The verdict remains CONDITIONAL because the data and observer study are valuable, but the central objective claim is not yet supported.","tokens_in":19105,"tokens_out":6761,"duration_ms":67542,"concrete_test":"Simulate a numerical breast phantom with known ground truth under the same 32 keV, 2 mGy, 1200/2400-projection geometry. Reconstruct with cSART (parameters optimized on a separate training phantom), UTR, and FBP; compute Q_s as in Eq. (6) and a direct ground-truth information metric, such as mean squared error or mutual information between reconstruction and truth. If cSART's Q_s advantage is not mirrored by the direct metric, or if Q_s rises when cSART is applied to object-free noise with zero true information, the central Shannon-information claim is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 2e interprets Q_s = SNR/Res^{1.5} (Eq. 6) as proportional to Shannon information per incident photon, citing linear-systems theory (Gureyev et al., 2014; 2016). The same section states that this ratio 'can be changed only by means of non-linear processing.' cSART is precisely such a process: Section 2c.ii describes a periodic bilateral 3D filter during iterations, an edge-preserving nonlinear denoiser, plus iterative regularization. SNR is measured in flat adipose ROIs and spatial resolution is estimated from the noise power spectrum in the same flat regions (Eqs. 7-8). A bilateral filter is designed to suppress noise in flat areas while preserving edges, so it can raise the measured SNR and reshape the noise spectrum without adding true object information. Therefore the observed Q_s advantage of cSART over UTR and FBP does not by itself demonstrate that cSART extracts more Shannon information per photon; it may only reflect nonlinear denoising in the metric's own ROIs. This is compounded by the parameter selection in Section 2c.ii: cSART parameters were chosen on a 1 mm slice from the same acquisitions using the same SNR and spatial-resolution criteria used in the evaluation, whereas UTR and FBP used fixed, unoptimized parameters. As a result, the objective superiority claim is not yet supported by a fair, metric-valid comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares three CT reconstruction algorithms—FBP, UTR, and cSART—combined with Paganin phase retrieval for synchrotron-based phase-contrast breast CT at 2 mGy on ten fresh mastectomy specimens. Objective metrics (contrast, SNR, CNR, spatial resolution, and Q_s = SNR/Res^{1.5}) indicate that cSART outperforms UTR and FBP, while a seven-observer radiological study finds that FBP images are preferred overall, reportedly because radiologists weight contrast more than noise. The authors interpret the objective advantage as evidence that cSART extracts more Shannon information per incident photon, and they propose that this makes cSART attractive for automated image analysis despite the subjective preference for FBP.","tokens_in":19372,"tokens_out":5345,"duration_ms":49311,"significance":"The study is valuable as a realistic, clinically oriented comparison at a fixed low dose, with real mastectomy samples, a consistent acquisition protocol, a reproducible public implementation of UTR, and a structured subjective evaluation with ICC and VGC analyses. If the central objective claim were supported by a fair comparison and a valid information-theoretic metric, the finding that cSART provides more useful information per photon would be important for the design of synchrotron breast CT reconstruction pipelines and for AI-based image analysis. However, the present evidence for that claim is weakened by two load-bearing issues: circular parameter optimization for cSART and the use of a metric whose information-theoretic interpretation is proven only for linear systems.","major_comments":[{"comment":"cSART's parameters were selected on a 1-mm slice from the same datasets using the same SNR, spatial-resolution, and noise-power-spectrum criteria used in the evaluation, whereas UTR (noise-to-signal ratio 0.05) and FBP (Hamming filter) were used with fixed, unoptimized parameters. This means that the reported superiority of cSART on the objective metrics is partly a consequence of fitting those metrics on the same data. The comparison should be repeated with a proper train/test split, or the authors should demonstrate that the selected parameters are optimal for independent data and that the conclusion is unchanged.","section":"2c.ii, 3a"},{"comment":"The quantity Q_s defined in Eq. (6) is derived from linear-systems theory, and Section 2e states that this ratio 'can be changed only by means of non-linear processing.' cSART is a nonlinear reconstruction method (periodic bilateral 3D filter plus iterative regularization, Section 2c.ii), and SNR and resolution are measured in flat adipose ROIs where a bilateral filter suppresses noise. Therefore the larger Q_s for cSART may reflect nonlinear denoising within the ROIs rather than additional Shannon information per incident photon. The claim that 'cSART-reconstructed slices objectively contained more measurable (Shannon) information' (Section 3a) is not supported without an information-theoretic analysis that is valid for nonlinear reconstruction, or at least a demonstration that the Q_s advantage is not due to the bilateral filter.","section":"2e, Eq. (6), 3a (Figure 2e)"},{"comment":"The subjective assessment was performed on 3-mm axial slices created by the 30-pixel binning procedure that, as acknowledged in Section 3b, produced calcification-like artefacts from amplified bright noisy pixels, with an example shown for cSART in Fig. 3(c). Since these artefacts appear to affect the very images that were rated, the observed lower subjective scores of cSART on calcification visibility and overall image quality may be confounded by the binning method rather than by the reconstruction algorithms themselves. The statement that this effect is not expected to be major is not quantitatively supported; the authors should report the number and extent of affected images and, if possible, re-analyze the subjective data with a binning method that does not amplify noise.","section":"2d, 3b"}],"minor_comments":[{"comment":"The pair 'cSART vs. FFBP' contains a typo; it should read 'cSART vs. FBP.'","section":"3a"},{"comment":"The typesetting of the threshold value and of Eq. (1) is garbled; the numerical value of the threshold and the mapping formula should be rendered cleanly so that the binning procedure is reproducible.","section":"2d"},{"comment":"Two panels are labeled (b); the labels should be corrected to (a)–(f) so that the panels are unambiguous.","section":"Figure 5"},{"comment":"The notation Res is used for both the spatial-resolution measure and its reciprocal relationship in Eq. (8); the definitions should be consolidated to avoid confusion.","section":"2e"},{"comment":"The statement that radiologists assign low weight to SNR is an informal interpretation; the VGC analysis only compares pairs of algorithms and does not directly estimate attribute weights, so this conclusion should be framed as a hypothesis rather than a demonstrated cause.","section":"3c"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about artefacts and limitations, and the UTR code is publicly available, which is commendable. The main risk is that the objective comparison is both circular and based on a metric that is invalid for nonlinear algorithms; this needs thorough revision. The authors include two groups who have proprietary or co-authored algorithms (cSART and UTR), which is legitimate but requires a particularly rigorous fairness check. The manuscript also has obvious typesetting problems in equations that need editorial attention before it can be published."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for two reasons: it is the first direct three-way comparison of FBP, UTR, and cSART on fresh mastectomy samples at 2 mGy, and it documents a clean split where radiologists prefer FBP while the objective metrics favor cSART. That split is probably real and worth taking seriously. But the paper's stronger claim—that cSART extracts more Shannon information per photon—is not supported by the evidence as presented.\n\nWhat is genuinely new: ten fresh mastectomy samples, synchrotron phase-contrast CT at a clinically relevant dose, both objective metrics (contrast, SNR, CNR, resolution, Q_s) and a seven-observer VGC study. The observer study is reasonably careful: randomized panel order, ICC reported, bootstrapped VGC. The UTR code is public, and the objective measurements are implemented in X-TRACT, so the numbers are reproducible. The finding that radiologists weight contrast more heavily than SNR is a useful empirical observation, consistent with the broader literature on observer preferences.\n\nTwo problems weaken the central objective claim. First, cSART parameters were optimized on a 1 mm slice from the same datasets using the same SNR and resolution criteria later used for evaluation; FBP and UTR got fixed, unoptimized settings. That makes the objective ranking a fitting report, not a prediction. Second, the \"Shannon information per photon\" interpretation of Q_s = SNR/Res^1.5 relies on linear-systems theory. The paper itself notes that this ratio can only be changed by nonlinear processing. cSART applies a bilateral 3D filter, which is exactly a nonlinear denoiser. SNR is measured in flat adipose ROIs, where a bilateral filter suppresses noise without adding object information. So the Q_s advantage may reflect denoising in the metric's own ROIs rather than true information extraction. The coronal-slice contrast finding (FBP highest, cSART lowest) is probably robust, and the subjective preference for FBP is credible. But the headline \"cSART clearly superior objectively\" needs a fairer test.\n\nMinor issues: the binning procedure created calcification-like artifacts in a few images; the authors acknowledge this and it may have affected the calcification visibility scores, though probably not the overall ranking. Also, objective and subjective metrics were measured on different slice types (thin coronal vs 3mm binned axial), which adds a small confound.\n\nWho this is for: people working on phase-contrast breast CT or reconstruction evaluation methodology. It deserves a serious referee; the dataset is valuable and the observer-preference observation is worth publishing even if the Shannon-information claim is toned down. Recommendation: send it to peer review, but the objective-superiority claim needs reframing or re-analysis—either by tuning cSART on independent training data, comparing default parameters, or restricting the information-capacity interpretation to linear algorithms. With that revision, I would be comfortable with it.","headline":"Useful three-way reconstruction comparison with a real objective/subjective split, but the objective superiority claim for cSART is undercut by parameter tuning on the same data and by using a linear-systems metric on a nonlinear algorithm.","tokens_in":20048,"tokens_out":2119,"would_cite":true,"duration_ms":19579,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"At a fixed 2 mGy dose, the cSART reconstruction algorithm yields more Shannon information per photon in phase-contrast breast CT than UTR or FBP, yet radiologists still prefer the higher-contrast FBP images.","keywords":["phase-contrast breast CT","cSART","filtered back projection","unified tomographic reconstruction","image quality assessment","observer study","Shannon information capacity","synchrotron radiation"],"falsifier":"A clean test would freeze cSART's parameters after training on one set of mastectomy samples and apply them, without any re-tuning, to a newly scanned separate set, comparing against FBP and UTR with their standard settings on all five objective metrics; if cSART no longer dominates on SNR, resolution, and $SNR/Res^{1.5}$, the central claim fails. A second check is to rerun the observer study with per-algorithm window/level optimization: if the FBP preference persists even when all three image sets are displayed at their individually optimal contrast settings, the contrast-weighting explanation is supported rather than being an artifact of a global rescaling.","tokens_in":18863,"feed_emoji":"🩻","tokens_out":7398,"duration_ms":60079,"temperature":0.7,"pith_summary":"The paper compares three CT reconstruction algorithms — filtered back projection (FBP), unified tomographic reconstruction (UTR), and a customized simultaneous algebraic reconstruction technique (cSART) — on phase-contrast breast CT scans of ten mastectomy specimens at 32 keV and a mean glandular dose of 2 mGy. Measured objectively, cSART clearly beats UTR and FBP on signal-to-noise ratio, contrast-to-noise ratio, spatial resolution, and on a ratio $SNR/Res^{1.5}$ that the authors take as a proxy for Shannon information extracted per detected photon. But seven human readers, without knowing which algorithm produced which image, rated FBP best for perceptible contrast, sharpness, and overall quality, and cSART only best for noise and calcification visibility. The authors argue the discordance is explained by a heavy subjective weighting of image contrast and a lighter weighting of noise. If the paper is right, the choice of reconstruction algorithm should depend on whether the images are destined for human reading or for automated analysis.","feed_headline":"cSART packs more breast-CT information per X-ray photon","feed_subtitle":"It wins on objective quality metrics, but radiologists prefer FBP's contrast; the best algorithm depends on who reads.","key_machinery":"The load-bearing identity is the objective image-quality characteristic $Q_s = SNR/Res^{1.5}$ at fixed dose, which is proportional to the Shannon information capacity of a linear imaging system and is invariant under linear filtering; the paper measures it at fixed 2 mGy dose and interprets it as the information extracted per detected photon. The other half of the machinery is the paired three-panel observer protocol using visual grading characteristics (VGC) with bootstrapping, run on 3-mm thick axial slices, with intraclass correlation to confirm reader agreement. Contrast is defined as $(\\beta_{\\max}-\\beta_{\\min})/(\\beta_{\\max}+\\beta_{\\min})$ from a five-bin histogram across adipose–glandular interfaces, and the contrast measured this way is the one objective metric that tracks the subjective preference.","core_discovery":"The central claim is that, at a clinically relevant dose of 2 mGy, the customized SART algorithm reconstructs phase-contrast breast CT images that contain more measurable Shannon information per incident photon than either UTR or FBP, as quantified by the ratio $SNR/Res^{1.5}$ among the three algorithms, while the same images receive lower perceptual scores from trained readers. This apparent contradiction is resolved by a hypothesis: in subjective radiological evaluation, image contrast carries far more perceptual weight than noise suppression, so the lower-contrast but higher-SNR cSART images lose out to FBP despite FBP having the poorest objective information content. The paper presents this as an argument rather than a settled conclusion, and it frames the methodology as a template for combining physics-based objective metrics with visual grading characteristics analysis.","pith_inferences":["An untested extension of the paper's contrast-weighting explanation is that applying a contrast-enhancement or unsharp-masking post-process to cSART volumes would close the subjective gap; this could be tested with a similar observer study.","If cSART provides more Shannon information at the same dose, it may also allow dose reduction for automated screening tasks while preserving diagnostic information for machine readers—a testable prediction.","The fixed linear rescaling from 32-bit floating point to 12-bit integers with a global window may itself disadvantage cSART, whose noise statistics and histogram shape differ from FBP's; an observer test with per-algorithm window/level optimization would be a fairer perceptual comparison.","The paper's own observation that calcification artifacts appear in some cSART images suggests the thresholding/binning step used to make thick axial slices may interact with the reconstruction algorithm; refining that step could change the subjective ranking."],"forward_implications":["If cSART truly puts more Shannon information per photon into its images, then for fixed dose it is the better input for automated and AI-based diagnostic tools, since those tools rely on objective information content rather than human perceptual preferences.","The observed preference for FBP suggests that radiologist-facing image display, windowing, or post-processing of cSART images would need to be adjusted to make the objective advantage perceptually visible.","For clinical deployment of phase-contrast breast CT at synchrotron facilities, algorithm choice will depend on the reading workflow: human reading may favor FBP or UTR, while automated analysis favors cSART.","The methodology—pairing SNR, CNR, contrast, resolution, and $SNR/Res^{1.5}$ with a VGC observer study—can be applied to any CT reconstruction or denoising algorithm, not just these three."],"supporting_citations":[{"why":"Supplies the cSART algorithm and the parameter-optimization procedure whose parameters are transferred to this study.","marker":"(Donato et al., 2022)"},{"why":"Defines the UTR algorithm and its noise-to-signal parameter used in the comparison.","marker":"(Gureyev et al., 2022)"},{"why":"Provides the homogeneous transport-of-intensity phase retrieval applied to all projections before reconstruction by any of the three algorithms.","marker":"(Paganin et al., 2002)"},{"why":"Establishes that $SNR/Res^{3/2}$ is proportional to Shannon information capacity, the theoretical basis for the objective quality metric.","marker":"(Gureyev et al., 2016)"},{"why":"Supplies the statistical-optics framework (ergodicity, Fourier-based resolution from noise) used to measure SNR and spatial resolution.","marker":"(Goodman, 2000)"},{"why":"Provides the half phase retrieval gamma value and the radiological assessment methodology the subjective study builds on.","marker":"(Taba et al., 2019)"},{"why":"The textbook reference for the FBP algorithm used as the baseline.","marker":"(Natterer, 2001)"},{"why":"The X-TRACT software implementing the objective quality measurements.","marker":"(Gureyev et al., 2011)"}],"fun_headline_variants":["cSART wins on metrics, FBP wins with radiologists","Radiologists prefer FBP despite cSART's objective edge","Objective cSART beats FBP; subjective verdict flips","Contrast wins: why radiologists pick FBP over cSART","At 2 mGy, the best algorithm depends on who's reading"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The head-to-head superiority of cSART rests on parameters that were tuned on a thin slice drawn from the same datasets and evaluated with the same SNR and resolution metrics that later declared cSART the winner, while FBP and UTR were used with fixed, unoptimized settings.","fun_headline_variants_meta":{"raw":{"variants":["cSART wins on metrics, FBP wins with radiologists","Radiologists prefer FBP despite cSART's objective edge","Objective cSART beats FBP; subjective verdict flips","Contrast wins: why radiologists pick FBP over cSART","At 2 mGy, the best algorithm depends on who's reading"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001091,"raw_usage":{"total_tokens":4602,"prompt_tokens":1032,"completion_tokens":3570,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":3488}},"tokens_in":648,"tokens_out":3570,"duration_ms":26104,"temperature":1.0,"reasoning_tokens":3488,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:43:47.493350+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A clean test would freeze cSART's parameters after training on one set of mastectomy samples and apply them, without any re-tuning, to a newly scanned separate set, comparing against FBP and UTR with their standard settings on all five objective metrics; if cSART no longer dominates on SNR, resolution, and $SNR/Res^{1.5}$, the central claim fails. A second check is to rerun the observer study with per-algorithm window/level optimization: if the FBP preference persists even when all three image sets are displayed at their individually optimal contrast settings, the contrast-weighting explanation is supported rather than being an artifact of a global rescaling.","supporting_citations":[],"review_version":1}