{"id":"28699c6b-c6ae-4422-8ab5-59e24b544edf","arxiv_id":"2501.01341","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A U-Net trained on Monte Carlo simulations estimates scatter sinograms for a long-axial field-of-view PET scanner, matching or beating single scatter simulation on phantom and clinical tests.","lead":"This study tests a deep learning system that estimates scattered radiation in long-axial field-of-view PET scanners directly from raw detector data. The method matches or beats the standard scatter correction on simulated and patient scans, which could improve image quality and dose efficiency in total-body PET.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The simulated-data evidence for DLSE superiority may be inflated by a slice-level train/test split: the 18 GATE simulations are split into sinogram slices (11559 per simulation), so test slices share anatomy and noise with training slices; a simulation-level split is needed before the accuracy…","rationale":"The reader's weakest_assumption correctly flags the Monte-Carlo simulator as the sole source of ground truth and notes the risk of slice-level leakage, but it places the emphasis on simulator bias and single-instance targets. My stress-test isolates the more immediately decisive issue: the paper's own data-allocation description (Section 2.2.3) is most naturally read as a slice-level split of the same 18 simulations. Because DLSE is trained on two-thirds of the slices and tested on the remaining one-sixth of the same simulations, the test examples are not independent acquisitions; the reported NRMSE advantage over SSS—whose error does not benefit from such leakage—is therefore not a valid measure of generalization. This does not require assuming the GATE model is wrong; even a perfect simulator would produce this artifact. The proposed check, a simulation-level split with held-out complete GATE runs, directly settles whether the concern lands. I keep the reader's CONDITIONAL verdict because the issue is resolvable by re-analysis and does not, on the current evidence, warrant outright rejection; however, the condition should be made explicit: the authors must demonstrate that the phantom results survive a simulation-level split, and should report the single-instance noise floor.","tokens_in":14901,"tokens_out":6548,"duration_ms":64413,"concrete_test":"Ask the authors to redo the simulated-data evaluation with a strict simulation-level split: hold out all 11,559 slices from at least two complete GATE runs (e.g., the large phantom at the +30% dose and the small phantom at the standard dose), train and validate on the remaining 16 simulations, and recompute the sinogram and image-domain NRMSE in Figs. 3 and 5 using only the held-out simulations. If the DLSE-vs-SSS gap shrinks or reverses, the reported superiority is inflated by slice-level leakage; if the gap persists, the central claim is supported. As a secondary check on the single-instance target, compare DLSE predictions on a held-out simulation to the mean scatter of several independently seeded GATE runs to quantify the noise floor of the training target.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 2.2.3 states that each of the 18 GATE simulations (3 morphologies × 6 doses) yields 11,559 sinogram slices, giving P = 208,062 realizations, \"allocated as follows: 2/3 for training, 1/6 for validation and 1/6 for testing.\" This describes a slice-level split, not a simulation-level split. Adjacent sinogram slices from the same phantom and dose realization are highly correlated in anatomy, activity, and Monte-Carlo noise; training on two-thirds of these slices means the test slices are near-duplicates of training data. The central quantitative claim—DLSE has lower NRMSE than SSS on phantom data (Figs. 3 and 5)—compares a trained network (whose test data are intra-simulation) against SSS (whose error is independent of the split), so the reported DLSE advantage is not an unbiased out-of-sample comparison. The sentence \"The phantoms used for evaluation were not included in the training\" does not resolve this, because the size/dose curves in Figs. 3 and 5 are labelled with the same three morphologies and six dose levels used for training. A second compounding issue is Eq. (4)/(Section 2.2.2): the target s is a single MC instance, not the expectation s̄, so the network is trained to a noisy target and evaluated against the same type of single-instance noise; this further inflates apparent accuracy. Clinical comparisons to SSS alone cannot serve as ground truth. Thus the load-bearing condition for the paper's central claim—that DLSE outperforms SSS on LAFOV—is an unbiased simulation-level evaluation, and the current text does not establish it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes and evaluates DLSE, a U-Net-based scatter estimation method for long-axial field-of-view PET, trained on GATE Monte-Carlo simulations of XCAT phantoms to estimate scatter sinograms from emission and ACF sinograms. The method is compared against the clinical standard SSS on simulated phantom data (sinogram and image domain) and on 14 clinical datasets (7 [18F]-FDG, 7 [18F]-PSMA). The authors report that DLSE outperforms SSS on simulated data in terms of NRMSE, robustness to patient size and dose, and lesion contrast recovery, and that clinical results are promising, with improved lesion contrasts on FDG and consistent behavior on PSMA data despite no PSMA training data.","tokens_in":15223,"tokens_out":5590,"duration_ms":58896,"significance":"If the findings are valid, DLSE would offer a fast, accurate scatter estimate for LAFOV PET, directly addressing a known limitation of SSS in systems with wide acceptance angles and multiple scatter contributions. The study is timely and uses a realistic scanner model (Siemens Vision Quadra) with multiple anthropomorphic phantoms and dose levels, and it tests cross-radiopharmaceutical generalization. The clinical evaluation is a useful feasibility check. However, the central simulated-data claim rests on a data-split procedure that may introduce slice-level leakage, which would materially inflate the reported DLSE advantage over SSS. The clinical data cannot independently validate the method because they are compared only against SSS, not against a ground truth.","major_comments":[{"comment":"The train/validation/test split is performed at the sinogram-slice level, not at the simulation level. Section 2.2.3 states that each of the 18 GATE simulations yields 11,559 sinogram slices, giving P = 208,062 realizations, allocated 2/3, 1/6, 1/6 to training, validation, and testing. Because the three morphologies and six dose levels are exactly the ones reported in the evaluation (Figs. 3 and 5), the test slices are drawn from the same 18 simulations as the training slices. Test slices from a given simulation share the same phantom anatomy, activity distribution, and correlated Monte-Carlo noise with training slices from that simulation. The sentence in Section 2.3.1 that 'The phantoms used for evaluation were not included in the training' is therefore misleading: only the particular 1/6 of slices were held out, not the simulations themselves. Under this split, the NRMSE comparisons in Figs. 3 and 5 do not measure generalization to unseen phantoms or dose levels, and the comparison to SSS, which has no training component, is not a fair out-of-sample evaluation. Please redo the evaluation with a simulation-level split (e.g., leave out entire morphology/dose combinations for testing) or with newly generated independent GATE simulations for testing, and report the resulting metrics.","section":"Section 2.2.3 and Section 2.3.1"},{"comment":"The training target in Eq. (4) is a single Monte-Carlo instance s of the scatter sinogram, with the explicit assumption (Section 2.2.2) that s ≈ s̄. The evaluation in Section 3.1 also uses this same single-instance s as the ground truth for NRMSE. Because s contains Poisson/Monte-Carlo noise, the reported NRMSE values include this noise component, and the network may partially learn to predict the specific noise realization of the training slices. Even if the slice-level split were corrected, the absolute accuracy numbers would still depend on the noise level of the single MC instance. Please quantify the MC noise (for example, by generating several MC instances for at least one phantom/dose combination and computing the variance of s) and discuss how the use of a single noisy target affects the DLSE-vs-SSS comparison and the interpretation of the reported NRMSE.","section":"Section 2.2.2, Eq. (4)"}],"minor_comments":[{"comment":"Please clarify whether the lesion phantom simulations (six spherical lesions, three contrasts) were included in the training data or held out entirely, and if held out, describe how the split was performed.","section":"Section 2.3.1"},{"comment":"The clinical results are comparisons only against SSS-corrected and uncorrected images; there is no independent ground truth. The abstract's phrase 'improving lesion contrasts' should be tempered to 'improving lesion contrasts relative to SSS-corrected images' to avoid overstating the evidence.","section":"Section 3.2 and Abstract"},{"comment":"The NRMSE definition uses the range of the estimate (ˆsmax − ˆsmin) in the denominator; standard NRMSE typically uses the range of the ground truth. Please justify this choice or switch to the GT range for easier comparability with other studies.","section":"Eq. (6)"},{"comment":"There are minor typographical issues: 'respecfully' should be 'respectively' in the text after Eq. (6), 'leaded' should be 'led' in Section 2.3.2, and the abstract inconsistently uses '7' and 'seven' for the number of clinical datasets.","section":"General"},{"comment":"The manuscript states that scattering from outside the axial FOV was included by considering activity along the entire body, but it does not specify how the activity outside the FOV was modeled in the MC simulation or whether the same activity distribution was used for training and evaluation; please clarify.","section":"Section 2.2.3"}],"recommendation":"major_revision","confidential_remarks":"The slice-level data split is the main technical concern and is load-bearing for the paper's central simulated-data claim. I would require a simulation-level split or independent MC test simulations before the phantom results can be accepted. The single-instance MC target is a secondary but related issue. The clinical evaluation is appropriately cautious but does not compensate for the simulated-data leakage. The paper is well written and the topic is relevant; with the split corrected, the work would be a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new parts here are the LAFOV (Vision Quadra) evaluation and the cross-tracer PSMA generalization test. The U-Net itself is unchanged from your earlier paper, so this is an incremental extension of an established program. That is fine for the field, and the work is well positioned: LAFOV scatter is a real problem, and a fast sinogram-based DL estimator would be useful. The paper has real strengths. The GATE simulation setup is described carefully, with three morphologies, six dose levels, and attention to scatter fractions. The evaluation is thorough in both sinogram and image domains, and the computational-speed discussion is practical. The clinical results on 14 patients, including the PSMA generalization and one arms-down acquisition, are genuinely promising even if mostly qualitative. But the stress-test concern is correct and load-bearing. Section 2.2.3 says the 18 GATE simulations were split into slices (11,559 each) and allocated 2/3 training, 1/6 validation, 1/6 testing. That means test slices come from the same phantoms and dose levels as training slices. The claim in Section 2.3.1 that the phantoms used for evaluation were not included in the training does not square with that. Figures 3 and 5 report NRMSE on this test split, so the DLSE-versus-SSS comparison is not an unbiased out-of-sample test. SSS is not trained at all, so the comparison is unfair: DLSE has seen that anatomy, SSS has not. The central claim that DLSE outperforms SSS on phantom data is therefore not established. The single-MC-instance issue (Section 2.2.2) is acknowledged, but it compounds the problem: the network is trained to predict a noisy target and evaluated against the same type of single-instance noise. That is a lesser issue than the split leakage, but it still means the reported error bars are not what they seem. The clinical comparisons only use SSS as the reference, so they cannot provide ground truth. The improved lesion contrast statements are suggestive, not proof. Who is this for? Researchers working on scatter correction for LAFOV PET. It deserves a serious referee, but not acceptance as is. I would send it back for major revision: perform a simulation-level split (train on a subset of the 18, test on the rest, or ideally generate new phantoms), report the phantom results under that split, and soften the abstract claims accordingly. The method is plausible and the clinical results are encouraging, so the revision is worth pursuing. Concretely: I would not desk-reject, but I would condition acceptance on fixing the evaluation protocol.","headline":"Useful LAFOV extension of DLSE, but the phantom superiority claim is undermined by a slice-level train/test split that leaks anatomy; needs a simulation-level split before the numbers can be trusted.","tokens_in":769,"tokens_out":888,"would_cite":false,"duration_ms":46563,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep learning network trained purely on Monte Carlo simulations estimates scatter in long-axial-field-of-view PET more accurately than the standard single scatter simulation, with better robustness to body size and dose, and transfers…","keywords":["positron emission tomography","scatter correction","scatter estimation","deep learning","U-Net","long-axial field-of-view PET","sinogram domain","Monte Carlo simulation"],"falsifier":"Acquire a physical phantom with known activity concentrations on the same model of LAFOV scanner, estimate scatter with DLSE and with SSS, and compare reconstructed concentration maps against the known values, for example using a line source in a water cylinder with a beam blocker to measure scatter directly; the central claim collapses if SSS matches or beats DLSE in mean error or contrast recovery on real, non-simulated data.","tokens_in":14668,"feed_emoji":"🩻","tokens_out":7813,"duration_ms":76639,"temperature":0.7,"pith_summary":"Long-axial-field-of-view PET scanners catch more oblique and multiple-scattered photons, and the standard single scatter simulation (SSS) correction struggles with them. This paper argues that a convolutional U-Net trained only on simulated FDG acquisitions can estimate the full scatter sinogram directly from measured emission and attenuation sinograms, capturing multiple scatters and oblique planes that SSS handles poorly. On simulated phantom data the network beat SSS in sinogram and image accuracy, stayed more accurate across body sizes and injected doses, and recovered lesion contrast better. On 14 clinical datasets the network matched or improved on SSS correction, including PSMA scans the network had never seen. If the simulation-based training is faithful to real scanner physics, DLSE offers a fast, accurate scatter correction for LAFOV PET without needing clinical ground truth for training.","feed_headline":"Deep-learning scatter map beats standard SSS on long-axis PET","feed_subtitle":"Trained only on simulated FDG scans, the U-Net improves lesion contrast and transfers to PSMA patients.","key_machinery":"The load-bearing object is the U-Net function $f_\\theta(y,b)$ that takes as input the random-free emission sinogram $y$ and the attenuation correction factor sinogram $b$ and outputs an estimate $\\hat{s}$ of the expected scatter sinogram $\\bar{s}$. It is trained by minimizing the mean squared error $\\frac{1}{I}\\|f_\\theta(y,b)-s\\|_2^2$ against Monte Carlo simulated scatter sinograms $s$, on 208,062 slice pairs from 18 simulated FDG acquisitions spanning three body morphologies and six dose levels. The architectural choice of processing each 520 by 50 sinogram slice independently with 2D convolutions, via a five-level U-Net, is what lets oblique-plane scatter contributions from the full-angle acceptance mode enter directly. The key simplifying assumption is that a single Monte Carlo instance $s$ approximates the expected scatter $\\bar{s}=\\mathbb{E}[s]$.","core_discovery":"The paper's central claim is that a deep learning scatter estimator (DLSE), a U-Net mapping the measured emission sinogram and attenuation correction factor sinogram to a scatter sinogram, can serve as an accurate scatter correction for a 106-cm axial-field-of-view PET scanner. Because the network processes raw sinogram slices independently, it incorporates multiple-scatter and oblique-plane contributions directly rather than scaling a single-scatter model. Trained on Monte Carlo simulations of three anatomies and six dose levels of FDG distributions, it produced scatter sinograms closer to the Monte Carlo ground truth than SSS across phantom sizes and doses, with sinogram NRMSE ranging from 0.153 to 0.163 for DLSE versus 0.215 to 0.243 for SSS, and reconstructed images with lower error in lungs and brain and closer lesion contrasts. On FDG patient data DLSE gave lesion contrasts equal to or better than SSS, and on PSMA data it agreed closely with SSS despite never being trained on PSMA distributions. The method's prediction time of roughly 381 seconds per whole 3D sinogram, about 33 ms per slice, is presented as practical, with input downsampling projected to cut this below 50 seconds.","pith_inferences":["If the simulator's scatter physics are accurate, DLSE's advantage should be largest where SSS's tail-scaling is weakest: large patients, high multiple-scatter fractions, and oblique planes; a direct comparison on physical phantoms with known activity would test this.","Because the network is slice-wise, it could be applied to time-of-flight PET by feeding each time-bin sinogram separately, a natural extension the paper mentions implicitly; the same architecture may then transfer to other scanners with matching sinogram geometry.","The single-instance MC training target means the accuracy ceiling is set by Monte Carlo variance and bias; averaging several MC instances or adding measured scatter data as targets might improve the already reported margins.","The PSMA generalization hints that the network learns the physical relation between attenuation, emission, and scatter rather than tracer-specific uptake; if so, cross-scanner transfer may be achievable with only sinogram-dimension adaptation."],"forward_implications":["In LAFOV PET, DLSE can replace or supplement SSS as the scatter estimate inside iterative reconstruction, reducing image error in low-count regions such as lungs while improving lesion contrast recovery.","Scatter correction no longer requires per-patient tail-fitting or clinical training data, because training uses only Monte Carlo simulations; deployment would need only emission and attenuation sinograms.","DLSE's insensitivity to body size and injected dose means one trained model can serve a wide range of patient morphologies and acquisition protocols on the same scanner geometry.","A model trained only on FDG appears to transfer to a different tracer, PSMA, so retraining may not be needed for new radiopharmaceuticals on the same scanner.","With input downsampling, scatter estimation time can drop below 50 seconds, making the method clinically practical alongside reconstruction."],"supporting_citations":[{"why":"Introduces the DLSE U-Net method on a conventional PET scanner; this paper extends it to LAFOV geometry.","marker":"[22]"},{"why":"Supplies the digital anatomical phantom morphologies and activity distributions from which all training and testing simulations were generated.","marker":"[30]"},{"why":"Defines the performance characteristics and geometry of the long-axial-field-of-view scanner model used in the Monte Carlo simulations.","marker":"[5]"},{"why":"Defines the single scatter simulation baseline against which DLSE is compared in sinogram and image domains.","marker":"[7–9]"},{"why":"Documents that scatter fraction and single-to-multiple scatter ratio vary with oblique plane angle in long-axial PET, motivating a method that handles oblique planes.","marker":"[16]"},{"why":"Validates that Monte Carlo models of the scanner agree with experimental data and provides the simulation-to-list-mode conversion chain, supporting use of simulated scatter as training ground truth.","marker":"[33, 34]"},{"why":"The Monte Carlo simulation toolkit used to generate the training and testing scatter sinograms.","marker":"[35]"}],"fun_headline_variants":["Deep learning scatter beats SSS on long-axis PET","U-Net scatter correction outperforms SSS on total-body PET","Simulation-trained CNN improves PET scatter, matches PSMA","Fast DL scatter estimation for long-axial FOV PET"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Monte Carlo model of the scanner is faithful to true scatter and that a single simulated scatter instance is close enough to the expected scatter to serve as training truth; if the simulation is biased, the reported accuracy advantage over SSS may not transfer to real scanners.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning scatter beats SSS on long-axis PET","U-Net scatter correction outperforms SSS on total-body PET","Simulation-trained CNN improves PET scatter, matches PSMA","Fast DL scatter estimation for long-axial FOV PET"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1468,"prompt_tokens":1140,"completion_tokens":328,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":756,"completion_tokens_details":{"reasoning_tokens":262}},"tokens_in":756,"tokens_out":328,"duration_ms":3787,"temperature":1.0,"reasoning_tokens":262,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:29:48.195049+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire a physical phantom with known activity concentrations on the same model of LAFOV scanner, estimate scatter with DLSE and with SSS, and compare reconstructed concentration maps against the known values, for example using a line source in a water cylinder with a beam blocker to measure scatter directly; the central claim collapses if SSS matches or beats DLSE in mean error or contrast recovery on real, non-simulated data.","supporting_citations":[{"cited_title":"Physics in Medicine & Biology 68(6), 065004 (2023) https://doi.org/ 10.1088/1361-6560/ac9a97","cited_arxiv_id":null,"evidence_quote":"Introduces the DLSE U-Net method on a conventional PET scanner; this paper extends it to LAFOV geometry."},{"cited_title":"Physics in Medicine and Biol- ogy 49(19), 4543–4561 (2004) https://doi.org/10","cited_arxiv_id":null,"evidence_quote":"The Monte Carlo simulation toolkit used to generate the training and testing scatter sinograms."}],"review_version":1}