{"id":"f2c42c5f-0c1f-4a90-8dc9-05aba7a4fc73","arxiv_id":"2502.07630","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Explicit correction models trained with a redefined timing residual improve TOF-PET timing resolution from 371 to 281 ps and are robust to single-position calibration data.","lead":"The authors introduce a new way to define timing residuals for machine-learning-based PET timing calibration, training models to output explicit correction values instead of corrected time differences. The approach removes the need for many source positions during calibration and improves coincidence time resolution from 371 ps to 281 ps on clinical-style detector stacks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The single-position transfer claim depends on the untested assumption that post-calibration per-event skew is independent of source position z; the appendix's linearity argument is circular, and the headline 281 ps result is not from the 1x1 single-source model.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing point: the single-source calibration works only if the residual skew after first-order analytical correction is independent of the source position z. That independence is not derived; it is an empirical property of this particular detector pair, and the appendix's linearity argument is circular because Eq. (25) assumes the model already predicts the z-independent labels. I agree with the reader that this is the main thing that could break the central claim. The paper does provide meaningful empirical support: in the in-plane distribution study, models trained only at z = 0 still improve CTR over no ML and pass the physics-based checks on transaxial test data, and the explicit models retain linearity across the full z-range even with sparse z-sampling. These are real, falsifiable results, not artifacts. However, the support is for one detector stack pair, and the quoted 281 ps headline is obtained with an extended 33 × 33 in-plane grid, not with a single point source. None of this demonstrates a known error, but it does mean the 'single source position' claim is conditional on a z-independence that could plausibly fail with DOI-related or angle-dependent skews in thicker crystals or different geometries. The proposed concrete test uses data the authors already have and would directly estimate the z-dependence of the residual; it would settle whether the single-position model transfers. Since the reader's verdict is already CONDITIONAL and my concern matches theirs, no adjustment is needed; the condition should explicitly require either this residual-slope analysis or a similar cross-z validation on the existing data, plus ideally a second independent detector pair.","tokens_in":23899,"tokens_out":9891,"duration_ms":109665,"concrete_test":"Using the existing transaxial test dataset (35 z-positions, 430–590 keV window), take the explicit model trained only at z = 0 with the 1 × 1 in-plane distribution (EM01x01) and compute, per z-bin, the mean residual r_i − p_i between the explicit label and the model prediction. Fit the binned means against z; if the slope is statistically inconsistent with zero, or if the per-z-bin CTR grows significantly with |z| relative to the reported fit uncertainties, then δ depends on z and the single-position transfer claim is falsified for this detector geometry.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the explicit residual definition r_i = (Δtm,i − ΔtE(zi))/2 makes the learned corrections independent of the source position z, so a model trained at z = 0 transfers to all z and data acquisition can be reduced to a single z-plane. For this to hold, after the first-order analytical time-skew calibration the remaining per-event skew δ_i must be independent of z: the label equals (δ_a − δ_b + noise)/2 only if the ΔtE(z) term cancels exactly, and a model trained at z = 0 has no z information in its features, so it can only apply a z-independent correction. If δ_a or δ_b depends on z — for example through DOI or incidence-angle effects in 20 mm LYSO:Ce,Ca crystals — the correction will be systematically wrong away from the training plane, and CTR should degrade with |z|. The appendix does not resolve this: Eq. (25) simply assumes the predictions equal the labels, and the subsequent algebra only shows E[Δt_corr] = E[Δt], which asserts rather than proves the needed independence. The empirical transfer evidence is limited to one detector pair, and the paper itself lists cross-stack stability as future work. Additionally, the abstract's headline improvement to 281 ± 5 ps is achieved by EM33x33, which uses 33 × 33 = 1089 in-plane source positions at z = 0; the genuine single-point-source model EM01x01 reaches 306 ± 4 ps (Table 3), so the 'single source position' simplification is less dramatic than the headline number suggests, though still better than the 371 ± 6 ps no-ML baseline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an explicit timing-residual formulation for machine-learning-based TOF calibration in PET detectors. The label is defined as half of the difference between the measured time difference and the expected time difference from the known source position, r_i = (Δt_m,i − Δt_E(z_i))/2, which yields a continuous label distribution that is claimed to be independent of the number of source positions along the transaxial axis. The authors compare this explicit-correction approach with their earlier implicit-correction models using a three-stage evaluation (MAE, linearity/ε, and CTR) on real data from two 4×4 LYSO:Ce,Ca detector blocks coupled to Broadcom NUV-MT SiPMs and read out with TOFPET2. They report an improvement in CTR from (371±6) ps to (281±5) ps for 430–590 keV coincidences, robustness to spatial undersampling, preserved linearity over the full test range, and a large reduction in model size suitable for FPGA deployment.","tokens_in":24278,"tokens_out":5815,"duration_ms":55548,"significance":"If the results hold, the explicit residual formulation is a practically valuable contribution: it removes the need for dense transaxial source sampling, preserves linearity over the full test range, and yields compact tree ensembles suitable for high-throughput PET applications. The three-stage evaluation is thorough, uncertainties are reported, and the comparison with implicit models is informative. The main weaknesses are that the central transfer claim—that a model trained at a single z-plane corrects events at all z—rests on an assumption about the z-independence of residual skews that is not proven, and the headline CTR is obtained with a 33×33 in-plane grid rather than the single-point-source configuration emphasized in the abstract.","major_comments":[{"comment":"The linearity argument assumes predictions equal labels (p_i ≈ l_i) and then derives E[Δt_corr] = E[Δt]. This is circular because the property to be established is precisely that the model's predictions, which cannot depend on z since z is not a feature, equal the label l_i = (Δt_m,i − Δt_E(z_i))/2 for all z_i. For a model trained only at z=0, l_i reduces to Δt_m,i/2, and no argument in the appendix shows that this equality survives at z≠0. Please provide a non-circular derivation or a direct empirical demonstration that the residual skew δ_i is independent of z after the first-order calibration.","section":"Appendix, Eq. (25)"},{"comment":"The headline improvement to (281±5) ps is obtained by model EM33x33, which uses 33×33 = 1089 in-plane source positions; the single-point-source model EM01x01 reaches (306±4) ps (Table 3). Since the paper's key simplifying claim is that acquisition reduces to a single source position, presenting the 281 ps value as the headline conflates the best-possible calibration with the single-source calibration. The abstract and Section 4 should either report the single-source value as the headline for the simplification claim or explicitly separate the two achievements.","section":"Abstract and Table 3"},{"comment":"The single-position transfer result rests on the assumption that, after the analytical first-order time-skew calibration, the per-event residual skew δ_i is independent of z (DOI and incidence-angle effects in 20-mm crystals could violate this). The current evidence is limited to one detector pair, and the paper itself defers cross-stack stability to future work. To support the central claim, please add a quantitative analysis of CTR as a function of |z| for a model trained only at z=0, including the range over which the single-source model stays within a pre-specified tolerance of the full multi-position model.","section":"Section 2.6.2 and Section 5"},{"comment":"Hyperparameters (tree depth) are selected from a grid evaluated on the same test data used to report CTR. The CTR differences among depths are small (e.g., 281±5 vs 287±5 ps), so selection-on-test can bias the reported improvement. Use a separate validation set for hyperparameter selection, or report the full distribution of CTR across depths and a correction for multiple comparisons.","section":"Section 3.1.3 and Tables 1–3"}],"minor_comments":[{"comment":"The sentence \"we propose to redefine the residuals\" should refer to redefining the labels; also \"high coverade\" is misspelled (should be \"high coverage\").","section":"Section 2.3"},{"comment":"The caption would benefit from a verb: \"the implicit model IM10,12 and the explicit model EM10,4 are used\" reads more clearly.","section":"Figure 3 caption"},{"comment":"The notation would benefit from an explicit statement that µ is the fitted mean of the prediction distribution and that ε is the parameter being estimated.","section":"Section 2.4.2, Eq. (11)"},{"comment":"The phrase \"capable of proving good results\" should be \"capable of providing good results\" or \"showing good results\".","section":"Section 4"},{"comment":"The expectation expression contains unmatched brackets and non-standard delimiters from the LaTeX source; please correct the typesetting.","section":"Appendix, Eq. (26)"},{"comment":"The placement of the CTR tables could be improved: Table 3 (the main in-plane result) appears before the extended energy-window tables, but the text in Section 3.2.3 refers to them in a way that is hard to follow; consider renumbering or adding in-text pointers.","section":"Tables 1, 3, 5, 6"}],"recommendation":"major_revision","confidential_remarks":"The experimental work is careful and the three-stage evaluation is a strength. However, the central transfer claim is not convincingly proven: the appendix's linearity argument is circular, and the single-source model does not achieve the headline CTR. These are fixable with additional analysis or revised claims, hence major revision rather than rejection. I would also encourage the editor to ask for a statement on how model-selection bias was handled, since hyperparameters were chosen on the test set."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the explicit residual definition is a genuinely useful reformulation that delivers on its main promise—single-position training works and CTR improves from 371 to 306 ps with a true point source, and to 281 ps with a dense in-plane grid. The paper is solid applied ML physics.\n\nWhat's new: instead of predicting the expected time difference (implicit correction), they define labels r_i = (Δt_m,i − Δt_E(z_i))/2, so the model outputs explicit timestamp corrections. This makes the label distribution continuous and, crucially, independent of where along z the source sits, so you don't need multi-position scans or a motorized stage. They show the explicit models are robust to large training step widths (50/100 mm) where the implicit models oscillate and lose linearity. The three-stage evaluation (MAE, linearity/ε, CTR) is thorough, uncertainties are reported, and the CTR gains are consistent across tree depths and energy windows. The model-size reduction (depth 4 vs 18) is a real plus for FPGA deployment.\n\nSoft spots: the appendix's linearity argument is circular—Eq. (25) just assumes predictions equal the labels, and the rest reduces to E[Δt] = E[Δt]. It doesn't prove z-independence. The z-independence itself is an empirical property, validated on one detector pair; if the residual skew δ depends on DOI or incidence angle, a model trained at z=0 may not transfer to other z positions. The paper lists cross-stack stability as future work, which is honest. Also, the abstract's headline 281 ± 5 ps is from EM33x33, i.e., 1089 in-plane positions; the genuine single-point-source model reaches 306 ± 4 ps. Still a clear improvement over 371 ± 6, but the 'single source position' statement overstates the headline number. Finally, no code, data, or trained models are released, and the best model is chosen from a hyperparameter grid evaluated on test data, with no cross-validation or pre-registration. That weakens reproducibility but is fixable.\n\nOverall, the central claim holds up: explicit correction removes the step-width dependence, enables simplified acquisition, and improves CTR. The flaws are addressable rather than fatal. This deserves a serious referee—send it out, but ask the authors to release artifacts and clarify the source-grid distinction.","headline":"Explicit residual TOF correction genuinely improves on implicit models and enables single-position training; the 281 ps headline, however, comes from a 33×33 in-plane grid, not a true point source.","tokens_in":24821,"tokens_out":2585,"would_cite":true,"duration_ms":24565,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["87.57.uk","29.40.Mc"],"model":"deepseek-v4-flash","headline":"The paper redefines the timing residual as half the measured-minus-expected time difference, enabling explicit per-event TOF corrections that improve CTR from (371±6) ps to (281±5) ps with a single source position.","keywords":["TOF-PET","timing calibration","coincidence time resolution","gradient-boosted decision trees","explicit TOF corrections","single-source calibration","TOFPET2 ASIC","machine learning"],"falsifier":"Move a point source to several $z$-positions, train an explicit model using only the $z=0$ data, and compare its predicted corrections with those of a model trained on all positions: if the per-event residual $\\delta_i = (\\Delta t_{m,i} + 2 z_i / c)/2$ shifts with $z$ by more than the detector's timing resolution, the single-position assumption fails.","tokens_in":23696,"feed_emoji":"⏱️","tokens_out":7385,"duration_ms":60595,"temperature":0.7,"pith_summary":"This paper proposes that a PET detector's remaining timing errors, after a standard analytical skew calibration, can be corrected by machine-learning models that directly predict timestamp corrections, if the training labels are defined as $r_i = (\\Delta t_{m,i} - \\Delta t_E(z_i))/2$ rather than as the expected time difference itself. With this redefinition, the label distribution becomes continuous and translationally symmetric, so a model trained with a radiation source at one position along the axis between detectors also works at other positions. The authors show on a pair of LYSO:Ce,Ca/SiPM detector blocks read out by the TOFPET2 ASIC that explicit-correction gradient-boosted trees improve coincidence time resolution from $(371 \\pm 6)$ ps to $(281 \\pm 5)$ ps for 430–590 keV coincidences, preserve linearity over the full axial range, and shrink model size enough for FPGA implementation. If correct, this removes the need for a motorized multi-position calibration setup and makes machine-learning TOF calibration practical for full PET scanners.","feed_headline":"New PET timing formula cuts resolution from 371 to 281 ps","feed_subtitle":"One source position now suffices to train TOF corrections, preserving linearity and shrinking the model for on-chip use.","key_machinery":"The central object is the redefined timing residual label $r_i = (\\Delta t_{m,i} - \\Delta t_E(z_i))/2 = (\\Delta t_{m,i} + 2 z_i / c)/2$, used as the regression target for gradient-boosted decision trees. It carries the argument by converting the calibration problem from predicting an absolute expected time difference (implicit) to predicting a per-event skew correction (explicit), which makes the label distribution independent of source-position sampling and lets a model trained at $z = 0$ provide corrections for all $z$. The rest of the machinery is the two-stage residual physics scheme: a first-order analytical timestamp skew calibration removes position-independent offsets, and the ML model is trained only on the remaining higher-order, event-dependent deviations, with the appendix's linearity argument assuming that predictions equal labels.","core_discovery":"The paper's central claim is that the residual physics-based calibration concept reaches its full practical value when the timing residual is defined as $r_i = (\\Delta t_{m,i} - \\Delta t_E(z_i))/2$, because then a model's output is itself the timestamp correction ($t_{a,i} \\leftarrow t_{a,i} - r_i$, $t_{b,i} \\leftarrow t_{b,i} + r_i$) instead of a predicted corrected time difference. This restores a translational symmetry to the labeling: shifting the source shifts the expected time difference and the measured time difference by the same amount, so the label distribution no longer encodes source position. The authors demonstrate that explicit-correction gradient-boosted trees are immune to the step-width collapse observed for implicit models, remain linear across the full $\\pm170$ mm test range, and improve timing resolution from $(371 \\pm 6)$ ps to $(281 \\pm 5)$ ps for 430–590 keV coincidences; a model trained on a single centered point source still reaches $(306 \\pm 4)$ ps, and a $33 \\times 33$ in-plane distribution reaches $(281 \\pm 5)$ ps. They interpret the small degradation relative to the best implicit model as acceptable given the simplification in acquisition and the exponential reduction in model size (tree depth 4 instead of 18).","pith_inferences":["Editorial inference: the independence of the explicit label from source position is a symmetry argument, not a proof that all per-event skews are position-independent; the transferability of a $z=0$-trained model to off-center positions should be re-tested for detectors with stronger depth-of-interaction-dependent skews.","Editorial inference: the comparison between 1×1 and 33×33 in-plane distributions is confounded by a $33^2$ difference in training statistics, so the observation that 3×3 performs worse than 1×1 is the cleanest evidence that source arrangement itself matters; a matched-statistics experiment would separate the two effects.","Editorial inference: because the explicit residual is defined from the measured time difference and the geometric expected time difference, the same label construction could be applied to any coincidence pair with a known geometric time-difference model, such as dual-sided readout or monolithic detectors, provided the single-position assumption holds.","Editorial inference: a testable extension is to train explicitly on data from one detector stack and apply it to an unseen stack of the same design; the paper lists this as future work and expects feature-based robustness."],"forward_implications":["Explicit-correction models trained with coarse axial sampling (50 mm or 100 mm step width) generalize to unseen 10 mm-sampled positions, whereas implicit models oscillate and fail linearity checks.","A single centered source position is enough to train an explicit model, removing the requirement for a motorized multi-position translation stage in calibration.","The explicit formulation preserves linearity over the full $\\pm170$ mm test range, so TOF information remains interpretable for image reconstruction.","Because shallow trees (depth 4) suffice, model memory shrinks by a factor of about $2^{18-4}$, making the correction suitable for FPGA-based, high-throughput PET readout.","Both correction approaches correct time-walk effects well: enlarging the energy window from 430–590 keV to 300–700 keV degrades CTR only slightly (about 2\\%) for explicit models, versus about 20\\% without ML."],"supporting_citations":[{"why":"Onishi et al. single-source-position waveform approach that the explicit residual definition is oriented on.","marker":"[43]"},{"why":"Authors' proof-of-concept implicit correction models and the depth-18 best model that the explicit approach is compared against.","marker":"[44]"},{"why":"Follow-up study documenting the step-width dependence and model collapse that motivated the new residual definition.","marker":"[46]"},{"why":"Convex time-skew calibration that provides the first-order analytical correction preceding ML training.","marker":"[59]"},{"why":"Provides the tree-boosting implementation used for all gradient-boosted tree models.","marker":"[56]"},{"why":"FPGA implementation of gradient tree boosting that motivates the model-size reduction claim.","marker":"[53]"},{"why":"Characterization of the LYSO:Ce,Ca/SiPM detector timing performance used in the experiments.","marker":"[5]"}],"fun_headline_variants":["Explicit TOF corrections push PET resolution to 281 ps","PET timing under 300 ps via explicit residual corrections","One source position trains PET TOF model for 281 ps","Machine learning PET calibration: 371→281 ps, 4× smaller model","Simpler PET timing calibration hits 281 ps with explicit method"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that, after the first analytical time-skew calibration, whatever timing error remains in an event does not depend on where the radiation source sits between the detectors, so a model trained with the source at one position works for all positions.","fun_headline_variants_meta":{"raw":{"variants":["Explicit TOF corrections push PET resolution to 281 ps","PET timing under 300 ps via explicit residual corrections","One source position trains PET TOF model for 281 ps","Machine learning PET calibration: 371→281 ps, 4× smaller model","Simpler PET timing calibration hits 281 ps with explicit method"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1566,"prompt_tokens":1160,"completion_tokens":406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":776,"completion_tokens_details":{"reasoning_tokens":318}},"tokens_in":776,"tokens_out":406,"duration_ms":4073,"temperature":1.0,"reasoning_tokens":318,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T12:06:19.789879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Move a point source to several $z$-positions, train an explicit model using only the $z=0$ data, and compare its predicted corrections with those of a model trained on all positions: if the per-event residual $\\delta_i = (\\Delta t_{m,i} + 2 z_i / c)/2$ shifts with $z$ by more than the detector's timing resolution, the single-position assumption fails.","supporting_citations":[{"cited_title":"Holistic evaluation of a machine learning-based timing calibration for PET detectors under varying data sparsity","cited_arxiv_id":null,"evidence_quote":"Follow-up study documenting the step-width dependence and model collapse that motivated the new residual definition."},{"cited_title":"Analysis of a convex time skew calibration for light sharing-based PET detectors","cited_arxiv_id":null,"evidence_quote":"Convex time-skew calibration that provides the first-order analytical correction preceding ML training."},{"cited_title":"XGBoost: A Scalable Tree Boosting System","cited_arxiv_id":null,"evidence_quote":"Provides the tree-boosting implementation used for all gradient-boosted tree models."},{"cited_title":"Timing advances of commercial divalent-ion co-doped LYSO:Ce and SiPMs in sub-100 ps time-of-flight positron emission tomography","cited_arxiv_id":null,"evidence_quote":"Characterization of the LYSO:Ce,Ca/SiPM detector timing performance used in the experiments."}],"review_version":1}