{"id":"8da363f9-3d96-48db-a08d-d9a7004ddb5a","arxiv_id":"2504.20698","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A simulated high-granularity proton CT with track discrimination achieves sub-1 percent RSP accuracy and 1.1 mm spatial resolution at 0.16 mGy by filtering out perturbed proton tracks.","lead":"This simulation study proposes a new proton CT scanner that uses a silicon pixel tracker and a scintillator range telescope, with algorithms to discard proton tracks disturbed by nuclear interactions or scattering. It reports sub-1 percent stopping-power accuracy and about 0.5 mm spatial resolution at standard dose, and about 1.1 mm resolution at a very low dose of 0.16 mGy, all in Geant4 simulation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is conditional on the nPE model: the CNN and Bortfeld discriminators are trained and tuned on noise-free Geant4 light-yield maps, while the paper defers crosstalk/noise simulation and prototype validation to future work, leaving the low-dose RSP accuracy dependent on an untested…","rationale":"The central claim has two layers: the detector and algorithms can achieve sub-1% RSP accuracy at ultra-low dose, and this will transfer to a physical prototype. The first layer is supported by a coherent, though idealized, Geant4 study; the second is untested. Among candidate weaknesses, the nPE model transfer is the most load-bearing because both the low-dose results and the abstract's '<3 mm per-track WEPL' claim are mediated by discriminators whose inputs and labels come from the same simulation. A mismatch in light yield, crosstalk, or dark count would degrade classification before any reconstruction or CT algorithm could compensate. The paper acknowledges this by labeling nonideal simulation and crosstalk/noise mitigation as future work (Sec. IV B) and by noting that the process-tagged CNN requires MC-data nPE consistency (Sec. III A). Other issues, such as air being excluded from the RSP-accuracy claims, PP spatial-resolution instability, in-sample threshold tuning, and the absence of code and data, are real but secondary; they affect interpretation without invalidating the simulation's internal logic. The reader's weakest-assumption identification matches my own, and the appropriate verdict remains CONDITIONAL until the proposed nonideal-simulation or prototype check is performed.","tokens_in":14454,"tokens_out":5684,"duration_ms":65039,"concrete_test":"Re-run the full pipeline with a nonideal detector model added to Geant4: introduce inter-bar optical crosstalk of a few percent, SiPM dark-count and electronic noise, and digitization thresholds consistent with the MPT2321 readout; retrain the range-tagged CNN on this augmented data and recompute the track-level WEPL STD (Fig. 5b) and low-dose RSP accuracy (Table 2) at 2e7 protons. If the WEPL STD rises above roughly 3 mm or any non-air RSP accuracy exceeds 1%, the headline claim fails in the direction most likely to occur in hardware.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The experimentally relevant version of the central claim (sub-1% RSP at 2e7 protons, 0.16 mGy) is produced by the range-tagged CNN pathway, not the process-tagged CNN, which the paper itself marks as experimentally infeasible without MC-data nPE agreement (Sec. III A, Sec. IV B). The CNN is trained on nPE maps generated by sampling from a standalone single-bar light-yield model with 10% coupling smearing, with noise and crosstalk intentionally excluded from the full-system Geant4 simulation (Sec. II B). Because the range tags are derived from reconstructed ranges in this same idealized simulation, both labels and input features can encode the same simulation-specific artifacts; a real detector's inter-bar crosstalk, SiPM dark counts, threshold nonlinearities, and position-dependent light collection would shift the nPE images away from the training distribution. The theta_s and Edep thresholds are also selected in-sample from Fig. 2 and Fig. 5 to balance STD versus selection rate, so the reported STD suppression is not a held-out estimate of hardware performance. This is not an internal inconsistency: internally the simulation is coherent, and the paper explicitly schedules nonideal simulation, crosstalk/noise mitigation, and prototype testing as future work. The load-bearing assumption is therefore that the idealized nPE model is faithful enough for the discriminators to transfer; until a prototype or a nonideal simulation is compared, the low-dose performance claim is conditional.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a simulation study of a proton CT system combining a CMOS pixel tracking system with a 100-layer segmented scintillator range telescope. The central idea is to use track discrimination—via Bortfeld-function fitting or a CNN classifier on light-yield images—to reject protons that undergo nuclear interactions or large-angle scattering, thereby reducing WEPL uncertainty and, in turn, the number of protons needed for clinically acceptable RSP imaging. The authors report that in Geant4 simulation both discrimination methods reduce per-track WEPL uncertainty to below 3 mm, that the standard protocol (4×10^8 protons, 3.2 mGy) achieves roughly 0.5 mm spatial resolution with sub-1% RSP accuracy, and that an ultra-low-dose protocol (2×10^7 protons, 0.16 mGy) achieves sub-1% RSP accuracy (excluding air) with roughly 1.1 mm resolution. The paper also includes systematic studies of detector digitization, scintillator thickness and width, and proton energy, and it identifies a prototype and nonideal simulation as future work.","tokens_in":14786,"tokens_out":6187,"duration_ms":67455,"significance":"If the reported simulation results transfer to a physical system, the proposed multi-filter architecture would be a meaningful step toward low-dose proton CT for treatment planning and repeated positioning. The paper is careful in several respects: it explicitly distinguishes the experimentally infeasible process-tagged CNN benchmark from the range-tagged CNN that could be used in practice; it acknowledges that noise and crosstalk are excluded from the present simulation; it provides a clear three-stage filtering workflow; and it includes a range of detector configuration studies that give the results internal coherence. The main significance is therefore conditional: the central claim of ultra-low-dose feasibility rests on the untested fidelity of the light-yield model and on in-sample optimization of thresholds and parameters. The paper is not internally inconsistent, but the gap between simulation and hardware is the key load-bearing issue.","major_comments":[{"comment":"The central low-dose claim (Table 2) depends on the range-tagged CNN, which is trained and evaluated on nPE maps generated by sampling from a standalone single-bar light-yield model with 10% coupling smearing, while noise and crosstalk are intentionally excluded from the full-system Geant4 simulation. The range labels are derived from reconstructed ranges in this same idealized simulation, so both the input features and the labels can encode the same simulation-specific artifacts. Because the paper defers nonideal simulation, crosstalk/noise mitigation, and prototype validation to future work, the transfer of the reported sub-3 mm WEPL STD and sub-1% RSP accuracy to a physical detector is not yet supported. I recommend either adding a nonideal simulation with realistic noise, crosstalk, and position-dependent light collection and retraining/testing the CNN, or substantially qualifying the clinical-feasibility statements so that they are explicitly limited to the idealized simulation.","section":"Sec. II B and Sec. IV B"},{"comment":"The θs and Edep thresholds are selected by balancing reconstructed WEPL STD and selection rate on the same simulated data that is then used to report the filtered STD values, and the Bortfeld parameters are fitted to MC unperturbed tracks and fixed during the reported evaluation. This makes the 3 mm track-level WEPL STD and the subsequent low-dose protocol in-sample estimates rather than held-out or hardware-transferable estimates. A validation protocol should be reported: for example, thresholds and Bortfeld parameters could be chosen on one simulated dataset and evaluated on an independent phantom or on data with different noise conditions, and the sensitivity of the low-dose conclusion to threshold variation should be quantified.","section":"Sec. III A and Sec. II C"},{"comment":"The ultra-low-dose protocol at 2×10^7 protons is demonstrated only with the range-tagged CNN, and the paper itself notes that this method exhibits sample-dependent variations and truncation anomalies at 0 mm and 140 mm thickness. The abstract's broad statement of 'sub-1% RSP accuracy' at low dose should be qualified by specifying that it refers to the range-tagged CNN, to the specific phantom materials, and to the exclusion of the air insert, whose relative error is large (Table 2: +64±141%) because of its near-zero RSP. Without this qualification, the headline claim overstates the generality of the low-dose result.","section":"Sec. III C and Table 2"},{"comment":"The abstract and conclusion suggest that the 10 MHz proton detection rate enables real-time image guidance and a 2-second scan at low dose, but the reconstruction framework is currently limited to single-proton events and the paper states that multi-proton reconstruction is under development. The detector electronics may support 10 MHz rates, but the imaging throughput claim is not yet demonstrated at the reconstruction level. Please separate the demonstrated detector-rate capability from the realized imaging time and soften the real-time claim accordingly.","section":"Abstract and Sec. IV C"}],"minor_comments":[{"comment":"Eq. (2) with σ_R=3 mm, R=260 mm, and p=1.77 gives σ_E/E ≈ 0.65%, not 0.6% as stated in the text; please update the number.","section":"Sec. III A and Eq. (2)"},{"comment":"The text says 'Applying this filter to only 8% of proton events reduces the energy loss STD from 16.7 MeV to 3.3 MeV' but later reports a 97% selection rate for the θs filter; please clarify whether 8% is the rejection fraction for the θs<10° example and how the optimized threshold differs.","section":"Sec. II C 1 and Sec. III A"},{"comment":"The text states that all materials show spatial resolution of about 0.5 mm except PP, but Table 1 lists PP with range-tagged CNN as 0.27 mm; please clarify whether the quoted resolution is per-method or an average and explain the PP instability more precisely.","section":"Sec. III B and Table 1"},{"comment":"There are some language errors, e.g., 'Such goals are empirical achieved' and 'varying by less that 2%'; a careful proofread would improve readability.","section":"Introduction and Sec. IV A"},{"comment":"The CNN training description lists hyperparameters (10 epochs, batch size 128, Adam at 1e-3) but does not state the final training set size, the class balance between good and bad tracks, or whether early stopping was used; please add these details for reproducibility.","section":"Sec. II C 2"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and internally coherent simulation study, and the authors are appropriately explicit about many of its limitations. The main risk is that the abstract and conclusion present ultra-low-dose imaging feasibility and real-time imaging in stronger terms than the current evidence supports, given that the CNN and threshold choices are optimized on an idealized simulation with no noise or crosstalk and no experimental nPE validation. I do not see grounds for rejection because the paper's claims are explicitly simulation-based and the authors schedule prototype and nonideal-simulation work; however, the load-bearing transfer assumption needs to be addressed by additional simulation or by a clearly scoped revision of the claims. The paper fits the journal's scope well, but the clinical-feasibility language should be moderated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read on arXiv:2504.20698. The paper proposes a pCT design that combines CMOS pixel tracking with a 100-layer scintillator range telescope, and uses either Bortfeld fitting or a CNN to reject proton tracks corrupted by nuclear interactions or large-angle scattering. The genuinely new bit is the integration and the explicit low-dose protocol: 2e7 protons, 0.16 mGy, sub-1% RSP accuracy and ~1.1 mm resolution in simulation. That is a meaningful target if it transfers to hardware.\n\nInternally the simulation is coherent. The WEPL calibration is standard, the MLP tracking is well grounded, and the comparison of the three discrimination methods (process-tagged CNN, range-tagged CNN, Bortfeld) is informative. The design studies on scintillator thickness, width, ADC bit depth, and proton energy are a useful engineering contribution. The paper is honest about what is not done: noise and crosstalk are excluded, the nPE model is a standalone single-bar sampling, and prototype validation is future work. The authors explicitly note that the process-tagged CNN requires MC-data nPE agreement to be experimentally usable.\n\nThe soft spots are exactly where the reader flagged. The CNN and the Bortfeld parameters are trained and tuned on the same idealized simulation used for evaluation; thresholds (theta_s, Edep, energy leakage weighting) are chosen on the same data. So the reported WEPL STDs and low-dose RSP accuracy are not held-out performance estimates for a physical detector. The range-tagged CNN, which is the experimentally viable one, shows anomalies at the edges of the training range (0 and 140 mm), and the low-dose protocol uses only that range-tagged CNN, not the stronger process-tagged version. The air insert is excluded from the headline accuracy, which is defensible given the relative error blow-up but should be stated more prominently.\n\nNone of this is a fatal flaw. The claims are explicitly simulation-based, and the paper schedules the missing steps. The load-bearing assumption is that the nPE sampling and noise-free geometry sufficiently capture the real detector response. Until a prototype or nonideal simulation is compared, treat the numbers as predictions, not demonstrated performance.\n\nWho is this for? Researchers working on pCT instrumentation and reconstruction, especially those thinking about range telescopes and track-level filtering. It deserves a serious referee: the architecture is plausible, the simulation is careful, and the clinical motivation is real. I would send it to review with the expectation that the authors clarify the in-sample fitting and add a nonideal simulation or at least a sensitivity analysis.\n\nRecommendation: send to peer review. My own verdict would be conditionally accept after revision, with the conditional on making the simulation-to-experiment gap explicit and adding robustness checks.","headline":"Coherent simulation study of a high-granularity pCT with track discrimination; the low-dose numbers are plausible in silico but rest on an unvalidated light-yield model and in-sample tuning.","tokens_in":15373,"tokens_out":2066,"would_cite":true,"duration_ms":20905,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multi-stage track filter lets proton CT reach sub-1% relative stopping power accuracy at 0.16 mGy dose in simulation.","keywords":["proton computed tomography","range telescope","track discrimination","relative stopping power","Bortfeld function","convolutional neural network","ultra-low-dose imaging","Monte Carlo simulation"],"falsifier":"Run the proof-of-concept prototype, send 100 or 200 MeV protons through a 100 mm water phantom, apply the same scattering-angle and energy-deposition filters, and compare the measured track-level WEPL spread; if it exceeds about 3 mm, or if the reconstructed RSP error at 2×$10^{7}$ protons exceeds 1%, the central claim fails. A quicker check is to compare the light output along the scintillator bar against the single-bar model, since the CNN's practical 'range tag' labels depend on that model being accurate.","tokens_in":14239,"feed_emoji":"⚛️","tokens_out":8170,"duration_ms":73058,"temperature":0.7,"pith_summary":"The paper claims that a proton CT scanner can reach clinically useful accuracy at roughly twenty times lower radiation dose than standard protocols by discarding corrupted proton tracks before image reconstruction. The proposed system replaces the usual precision energy measurement with a high-granularity range telescope whose energy-deposition profiles double as track-quality classifiers, using either a fit to the Bortfeld Bragg curve or a convolutional neural network. In Monte Carlo simulation this multi-filter chain lowers per-track water-equivalent path length (WEPL) uncertainty below 3 mm, which in turn lets a 4×$10^{8}$-proton scan (3.2 mGy) reach about 0.5 mm spatial resolution with sub-1% relative stopping power (RSP) accuracy, and a 2×$10^{7}$-proton scan (0.16 mGy) keep sub-1% RSP accuracy at about 1.1 mm resolution. If these numbers survive prototype tests, the design would enable low-dose pediatric imaging and real-time image guidance for proton therapy.","feed_headline":"Proton CT hits sub-1% accuracy at 0.16 mGy dose","feed_subtitle":"Track discrimination filters noisy protons before reconstruction, cutting the required proton count 20-fold.","key_machinery":"The load-bearing mechanism is a three-stage filter chain that runs from raw detector hits to the final RSP map. The first stage cuts on the reconstructed proton scattering angle (θs, around 10°), removing tracks deflected by large-angle elastic scattering; the second stage is the track discriminator in the range telescope, where each track's energy-deposition profile is fitted with the Bortfeld function (an analytical Bragg-curve approximation with parameters fixed by simulation, leaving only the proton range R0 free) or scored by a CNN fed with two orthogonal nPE projection images; the third stage is a pixel-wise 2σ filter on WEPL values grouped in 0.5 mm pixels, applied under the assumption of uniform WEPL within a pixel. The Bortfeld function and the CNN act as classifiers of 'good' versus 'perturbed' tracks, and it is this classification that lets the pixel-wise filter operate with as few as 11 protons per pixel in the low-dose protocol.","core_discovery":"The authors' central claim is that a pCT calorimeter should be treated as a range telescope whose measured energy-deposition profile is itself a discriminator of track quality, rather than as a spectrometer that must measure each proton's residual energy precisely. They demonstrate that two discrimination algorithms—a physics-based fit using the Bortfeld function and a CNN classifying raw scintillation-light projections in two orthogonal views—both separate 'unperturbed' tracks from tracks that underwent nuclear interactions or large-angle scattering. With this discrimination, WEPL standard deviation for individual tracks drops below 3 mm, and the subsequent pixel-wise 2σ filter in WEPL-map construction needs far fewer protons per pixel. The consequence is a standard imaging protocol at 3.2 mGy achieving about 0.5 mm resolution with RSP error below 0.5% for non-air inserts, and an ultra-low-dose protocol at 0.16 mGy keeping RSP error below 1% at about 1.1 mm resolution after a count-driven sinogram correction.","pith_inferences":["If the prototype confirms the simulation, the same filter chain could be retrofitted to existing range-telescope pCT designs that currently rely on precision energy measurement, since the discrimination uses the energy-deposition profile itself.","The count-driven sinogram interpolation is a basic fix; a statistical reconstruction that jointly models WEPL values and proton counts per pixel could recover additional resolution at low dose.","The CNN's range-tag labels depend on the training WEPL window (0–150 mm), so clinical use would likely require retraining per anatomical site or a wider calibration phantom range.","The Bortfeld discriminator, with fixed physics parameters and no training data, could serve as a sanity check for the CNN on real detector data, since the two methods are expected to agree on distal-tail events."],"forward_implications":["At the standard 4×10^8-proton protocol, simulated RSP accuracy stays below 0.5% for polypropylene, Teflon, and bone-equivalent inserts with about 0.5 mm spatial resolution.","At 2×10^7 protons (0.16 mGy), simulated RSP accuracy remains below 1% for tissue-like inserts and spatial resolution degrades only to about 1.1 mm.","Because only about 40 of the original 56 protons per pixel survive filtering, the discrimination chain directly cuts the proton statistics needed for pixel-wise filtering, which is what enables the low-dose protocol.","A 10 MHz detection rate makes a 2-second scan possible, opening the door to real-time image guidance during radiotherapy.","WEPL standard deviation grows only about 2% when the ADC digitization is reduced from 12 bits to 4 bits, so the architecture tolerates cheaper, faster readout electronics."],"supporting_citations":[{"why":"Defines the range telescope concept for proton CT that this design extends to include track discrimination.","marker":"[18]"},{"why":"Supplies the analytical Bortfeld approximation of the Bragg peak used for energy-deposition fitting.","marker":"[26]"},{"why":"Provides the convolutional-neural-network approach for classifying particle tracks in range telescopes.","marker":"[28]"},{"why":"Establishes the fast-scanner design and pixel-wise filtering used in the WEPL map reconstruction.","marker":"[15]"},{"why":"Classifies proton interaction processes (ionization, elastic scattering, nuclear reactions, bremsstrahlung) that define perturbed versus unperturbed tracks.","marker":"[24]"},{"why":"Gives the maximum-likelihood path formalism used to estimate proton positions in the phantom.","marker":"[25]"},{"why":"Supplies the RSP accuracy definition and the experimental comparison methodology used to report results.","marker":"[31]"},{"why":"Exemplifies the iMPACT calorimeter range-telescope paradigm that motivates the detector architecture.","marker":"[17]"}],"fun_headline_variants":["Track discrimination cuts proton CT dose 20-fold","CNN and Bortfeld fit discriminate pCT tracks, dose drops 20x","0.16 mGy proton CT: sub-1% RSP via track filtering","Track discrimination enables 20x lower dose proton CT"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole performance story rests on the assumption that the simulation, which leaves out electronic noise and light leaking between neighboring detector channels and generates scintillation light from a standalone single-bar calibration, faithfully predicts how the physical range telescope and its discrimination algorithms will behave.","fun_headline_variants_meta":{"raw":{"variants":["Track discrimination cuts proton CT dose 20-fold","CNN and Bortfeld fit discriminate pCT tracks, dose drops 20x","0.16 mGy proton CT: sub-1% RSP via track filtering","Track discrimination enables 20x lower dose proton CT"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001006,"raw_usage":{"total_tokens":4293,"prompt_tokens":1023,"completion_tokens":3270,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":3195}},"tokens_in":639,"tokens_out":3270,"duration_ms":24928,"temperature":1.0,"reasoning_tokens":3195,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:22:40.073232+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the proof-of-concept prototype, send 100 or 200 MeV protons through a 100 mm water phantom, apply the same scattering-angle and energy-deposition filters, and compare the measured track-level WEPL spread; if it exceeds about 3 mm, or if the reconstructed RSP error at 2×$10^{7}$ protons exceeds 1%, the central claim fails. A quicker check is to compare the light output along the scintillator bar against the single-bar model, since the CNN's practical 'range tag' labels depend on that model being accurate.","supporting_citations":[{"cited_title":"Esposito, C","cited_arxiv_id":null,"evidence_quote":"Defines the range telescope concept for proton CT that this design extends to include track discrimination."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the analytical Bortfeld approximation of the Bragg peak used for energy-deposition fitting."},{"cited_title":"Poludniowski, N.M","cited_arxiv_id":null,"evidence_quote":"Establishes the fast-scanner design and pixel-wise filtering used in the WEPL map reconstruction."},{"cited_title":"Pozzobon, F","cited_arxiv_id":null,"evidence_quote":"Classifies proton interaction processes (ionization, elastic scattering, nuclear reactions, bremsstrahlung) that define perturbed versus unperturbed tracks."},{"cited_title":"Granado-González et al., A novel range telescope concept for proton CT","cited_arxiv_id":null,"evidence_quote":"Gives the maximum-likelihood path formalism used to estimate proton positions in the phantom."},{"cited_title":"Newhauser and R","cited_arxiv_id":null,"evidence_quote":"Supplies the RSP accuracy definition and the experimental comparison methodology used to report results."}],"review_version":1}