{"id":"9d9e02a3-2c51-4d49-ac4c-4a7680febaae","arxiv_id":"2608.06518","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Gamma-neutron discrimination in a water Cherenkov detector is demonstrated using a statistical energy threshold plus a soft-voting ML ensemble, with 0.816 accuracy and 0.921 AUC.","lead":"A team at Bariloche shows that a water Cherenkov detector can separate gamma and neutron sources by combining a 3-sigma energy cutoff with a machine-learning ensemble on pulse shapes. The two-stage method reaches 0.816 accuracy and 0.921 ROC AUC, which could support water-based radiation monitors for nuclear security without helium-3 detectors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Neutron-labeled training window (4500–8000 ADU, i.e. 2.06–3.46 MeV) lies above the ~1.99 MeV Compton edge of the 2.22 MeV capture gamma, so the 'neutron' class is residual 4.44 MeV gamma leakage, not neutrons.","rationale":"Good-faith reading: the paper is an experimental feasibility study of source identification in WCDs, with a two-stage threshold+ML method. The threshold stage is straightforward and the ML pipeline is standard, but the central claim depends on the training labels representing the physical classes named. That condition fails. The reader's weakest_assumption identified the same class-label fragility: in a WCD, the observable neutron signal is the 2.22 MeV capture gamma, so a classifier trained on Pb-shielded events versus Cd/borated-paraffin events may be separating two gamma energies rather than neutrons from gammas. My analysis sharpens this with an internal quantitative inconsistency: the selected charge window, 4500–8000 ADU, corresponds to 2.06–3.46 MeV by the paper's Eq. 3, while the 2.22 MeV capture gamma has a maximum Compton-electron energy of ~1.99 MeV (~4330 ADU). Therefore, the Pb-labeled 'neutron' events in the training window cannot be capture-gamma events; they are dominated by residual 4.44 MeV gamma leakage through the lead shield. The Cd/B-Paraffin 'gamma' class is direct 4.44 MeV source gamma. The ML results thus reflect differences in gamma energy or shielding history, not particle type. Because the experimental dataset as constructed cannot support the headline claim of gamma-neutron discrimination, the verdict should move from CONDITIONAL to REJECT unless the authors redefine the task (e.g., detecting neutron-induced capture gammas) and validate with simulations or a pure 2.22 MeV gamma source. I do not question the DAQ engineering or the internal consistency of the ML code; the physics interpretation of the labels is the load-bearing issue.","tokens_in":13583,"tokens_out":9730,"duration_ms":95947,"concrete_test":"Run a Geant4 simulation of the same WCD with ground-truth labels: class N = 2.22 MeV gammas from neutron capture in the water (only those), class G = 4.44 MeV source gammas, and apply the same 4500–8000 ADU selection used in Section 5.2. Train the same Bagging/CatBoost/MLP ensemble on the simulated pulse shapes and report test AUC. If the AUC drops to ~0.5 when the deposited-energy distributions of the two classes are matched, the experimental AUC of 0.921 is an energy artifact, not gamma-neutron discrimination.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's central claim is that the ML stage (Section 5) discriminates neutron from gamma pulses. For this to be true, the training labels must correspond to those physical classes. They do not. Section 2.2 states that neutrons in a WCD are observed only through the 2.22 MeV gamma from 1H radiative capture. Using the paper's own calibration (Eq. 3), the charge window chosen in Section 5.2, 4500–8000 ADU, corresponds to 2.06–3.46 MeV. A 2.22 MeV gamma cannot deposit more than ~1.99 MeV in a Compton interaction (its maximum Compton-electron energy), which corresponds to ~4330 ADU. Thus, the Pb-shielded 'neutron' pulses that survive the window are not neutron-capture events; they are dominated by the residual 4.44 MeV gamma leakage that Section 2.4 explicitly says remains after 10 cm of lead. The Cd/B-Paraffin 'gamma' class is the direct 4.44 MeV source gamma. The ensemble is therefore separating two gamma populations that differ by shielding/attenuation history, not by particle type. This also explains Table 1: the Pb and Pb/P-Paraffin cutoffs (9000 and 8700 ADU) are essentially equal because both are set by 4.44 MeV gamma leakage. Consequently, the reported accuracy and AUC 0.921 do not demonstrate gamma-neutron discrimination in a WCD.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims to demonstrate gamma-neutron discrimination in a water Cherenkov detector using a two-stage method: a 3-sigma cutoff to establish an energy threshold and a soft-voting machine-learning ensemble to classify pulses. The experimental data come from a 241AmBe source under lead, borated-paraffin/cadmium, lead/paraffin, and unshielded configurations, with 60Co and 137Cs used for calibration. The authors report a linear ADU-to-MeV calibration, a physically grounded neutron threshold, and an ML ensemble with accuracy 0.816 and ROC-AUC 0.921, and they apply the trained model to decompose the unshielded mixed field into neutron and gamma components.","tokens_in":13847,"tokens_out":8546,"duration_ms":81709,"significance":"If the claims were valid, the work would provide a practical, 3He-free route to neutron-gamma discrimination in water Cherenkov detectors, relevant to nuclear security and radiation monitoring. The paper contains a detailed experimental description, transparent uncertainty propagation for the cutoff determination, and a reproducible ML pipeline with explicit hyperparameters and cross-validation. These are concrete strengths. However, the central claim is not established because the training labels do not correspond to the physical particle classes, and the reported ML performance is affected by test-set leakage. The significance of the work as a demonstration of gamma-neutron discrimination is therefore much lower than claimed.","major_comments":[{"comment":"The training labels for the ML classifier are not physically valid. The paper states in Section 2.2 that neutrons in a WCD are detected only through the 2.22 MeV gamma from hydrogen capture. Using the calibration in Eq. (3), the selected charge window 4500-8000 ADU corresponds to 2.06-3.46 MeV. A 2.22 MeV gamma can deposit at most 2.22 MeV (approximately 4900 ADU), so any pulse in the window above that energy cannot be a neutron-capture event; those pulses must originate from residual 4.44 MeV source gammas that survive the lead shield, as acknowledged in Section 2.4. Thus the Pb-shielded 'neutron' training class contains a substantial fraction of non-neutron events, and the classifier may be separating gamma populations with different shielding/attenuation histories rather than neutron pulses from gamma pulses.","section":"Sections 2.2, 5.2, Eq. (3)"},{"comment":"The reported accuracy is optimistically biased by threshold selection on the test set. The text says the ensemble was evaluated on the test set across a range of decision thresholds and the optimal threshold (0.52) was selected at the intersection of TP and TN rates; the same test set is then used to produce the confusion matrix and the reported accuracy of 0.816. This is a form of test-set leakage. The threshold should be tuned on a validation fold or with nested cross-validation, and the final metrics should be reported on a held-out test set that is never used for threshold or model selection. The ROC-AUC may also be inflated if the threshold is not the only quantity tuned on the test set.","section":"Section 5.4, Fig. 9"},{"comment":"The energy calibration is internally inconsistent and does not support the claimed physical interpretation. Table 1 lists the Pb configuration cutoff as 9000±1300 ADU with an associated energy of 2.22 MeV, but Eq. (3) converts 9000 ADU to 3.86 MeV. Similarly, the Cd/B-Paraffin cutoff of 14900±3000 ADU is associated with 4.44 MeV, but Eq. (3) gives 6.22 MeV. Since a single gamma cannot deposit more than its full energy, these cutoffs cannot represent the maximum energy deposition of the stated gamma lines. The linear fit in Eq. (3) is therefore not a valid mapping between charge and deposited energy in the region used for the threshold and the ML window. The acknowledgement in Section 3.2 that the reduced chi-square exceeds unity does not resolve this systematic inconsistency.","section":"Table 1, Section 3.2"},{"comment":"The unshielded 241AmBe dataset is described as an independent test, but it has no ground-truth labels for individual pulses. The 'predicted neutron and gamma components' in Fig. 5 are simply the model's assignments under the learned shielding-dependent classes; they cannot validate the classifier's ability to identify neutrons in a mixed field. To demonstrate feasibility for mixed radiation fields, the authors would need tagged events (for example, time correlation with the AmBe alpha particle or a well-characterized neutron source such as 252Cf) or a validated Monte Carlo simulation with known particle identities.","section":"Section 5.3, Fig. 5"}],"minor_comments":[{"comment":"The abstract states 'achieving an accuracy of 0.816 and an area under the Receiver Operating Characteristic (ROC) curve' but omits the numerical AUC value (0.921, reported later). The AUC should be stated in the abstract for completeness.","section":"Abstract"},{"comment":"The text says the classifier is most effective in the mid energy region (1.5-2.5 MeV), but the selected charge window 4500-8000 ADU corresponds to 2.06-3.46 MeV via Eq. (3). The stated energy range is inconsistent with the window; this should be reconciled or corrected.","section":"Section 5.2"},{"comment":"The accuracy threshold of 0.75 used to select 12 classifiers should state explicitly whether this selection is based on validation-fold performance or test-set performance. If it is based on the test set, it introduces additional selection bias beyond the threshold tuning discussed above.","section":"Section 5.3"}],"recommendation":"reject","confidential_remarks":"The central difficulty is not stylistic but experimental: the training labels are defined by shielding configuration, not by particle type, and the chosen charge window excludes most of the neutron-capture signal while admitting residual source gammas. This cannot be corrected by local edits; it would require new data with tagged neutron events or a fundamentally different validation scheme. The overlap with the companion NIMA paper [32] on the same two-stage framework may also deserve editorial attention regarding incremental novelty."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a well-documented experimental study with a clear description of the detector, the shielding configurations, and the uncertainty budget. The 3-sigma cutoff method and the traffic-light screening idea are sensible operational tools, and the ML pipeline is thorough: class balancing, hyperparameter search, ensemble diversity, and an honest note about gain drift. I want to give credit for that. The problem is the central claim and it is not a minor one. Section 2.2 says neutrons are detected only through the 2.22 MeV capture gamma. The ML charge window is 4500–8000 ADU, which with the paper's own Eq. 3 corresponds to 2.06–3.46 MeV. A 2.22 MeV gamma cannot deposit more than about 2.22 MeV even if fully absorbed, which maps to roughly 4900 ADU; the single-scatter Compton edge is about 1.99 MeV (4330 ADU). So the upper two-thirds of the chosen window cannot contain 2.22 MeV capture-gamma events at all. The Pb-shielded 'neutron' pulses surviving that window must be dominated by the residual 4.44 MeV gamma leakage that the paper admits remains after 10 cm of lead. The Cd/B-paraffin 'gamma' class is also 4.44 MeV gammas. The classifier is therefore learning to separate two gamma populations that differ by shielding history, not by particle type. This also explains why the Pb and Pb/P-paraffin cutoffs in Table 1 are nearly equal: both are set by the same 4.44 MeV leakage. The internal inconsistency shows up in the calibration itself: Table 1 labels the Pb cutoff as 2.22 MeV at 9000±1300 ADU, but Eq. 3 maps that to about 3.9 MeV. The reduced chi-square being above unity is not enough to paper over a 74% energy offset. A second serious problem is that the decision threshold (0.52) is tuned on the test set and then performance is reported on the same set, so the 0.816 accuracy is optimistic. Finally, the relationship to reference [32] by the same authors must be clarified; the preprinted title and method look like the same result, which would make the preprint a duplicate rather than an extension. The experimental effort and the uncertainty treatment suggest a serious group, but the load-bearing labeling assumption does not hold. This is a paper a referee could help fix, but only if the classes are redefined, a genuine validation set is used for threshold selection, and the 2.22 MeV contribution to the window is quantified. I would not cite it in its current form, but I would send it to peer review because the questions are important and the experimental data are not worthless.","headline":"Careful experimental work, but the gamma/neutron labeling is likely circular and the ML classifier appears to separate two gamma populations, not neutrons from gammas.","tokens_in":14463,"tokens_out":5931,"would_cite":false,"duration_ms":55738,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a single-photomultiplier water Cherenkov detector can separate neutron and gamma signals by combining a 3-sigma energy threshold with a soft-voting machine-learning ensemble.","keywords":["Water Cherenkov detector","gamma-neutron discrimination","pulse shape analysis","machine learning ensemble","nuclear security","energy calibration","3-sigma cutoff","americium-beryllium source"],"falsifier":"Run the trained ensemble on pulses from a neutron-only source that produces capture gammas inside the water at energies different from 2.22 MeV, with no 4.44 MeV gamma present; if accuracy falls to chance, the discrimination is energy-based rather than neutron-based.","tokens_in":13297,"feed_emoji":"☢️","tokens_out":8559,"duration_ms":72409,"temperature":0.7,"pith_summary":"The paper claims that a standard water Cherenkov detector can tell neutron-producing sources from pure gamma sources without helium-3 counters. It proposes a two-stage method: first, a statistical 3-sigma cutoff in the charge spectrum yields an energy calibration that screens low-energy gamma emitters; second, a soft-voting machine-learning ensemble trained on 32-bin pulse waveforms classifies the remaining sources. The ensemble reaches an accuracy of 0.816 and an ROC AUC of 0.921 on held-out test events, and it separates an unshielded mixed americium-beryllium field into neutron and gamma components. If the claim holds, water tanks already built for cosmic-ray physics could double as cheap, safe radiation monitors for nuclear security.","feed_headline":"Water-tank detector separates neutrons from gamma rays, AUC 0.921","feed_subtitle":"Two-stage method identifies radioactive sources with only water and one photomultiplier, no 3He needed.","key_machinery":"The central mechanism is the cutoff point (CP), the first charge-spectrum bin where the source becomes statistically indistinguishable from background, $|N_{\\rm src}-N_{\\rm bkg}|<3\\sigma$ with $\\sigma=\\sqrt{N_{\\rm src}+N_{\\rm bkg}}$. The CP both anchors the linear ADU-to-MeV calibration and defines a 'traffic-light' threshold: red below the neutron cutoff (pure gamma), green at the cutoff (neutron presence confirmed), yellow above it (high-energy gamma, neutrons unproven). The second mechanism is the soft-voting ensemble, which combines a Bagging classifier, CatBoost, and a multilayer perceptron over the 32 time bins of each PMT pulse; per-model weights are set proportional to validation F1 scores, and the decision threshold is tuned to 0.52 to balance false positives and false negatives. The ensemble operates inside a 4500–8000 ADU charge window chosen so the neutron signal dominates background.","core_discovery":"On the paper's own terms, the discovery is that a single-photomultiplier water Cherenkov detector carries enough pulse-shape information to separate neutron-induced from gamma-induced signals once the energy axis is calibrated. The 3-$\\sigma$ cutoff procedure anchors a linear charge-to-energy relation, $E = (4.00\\pm0.19)\\times10^{-4} Q + 0.26$ (MeV, ADU), and places a neutron threshold near 9000 ADU, so sources ending below it are identifiable as pure gamma emitters. Above that threshold, a soft-voting ensemble of three classifiers—a Bagging classifier, CatBoost, and a multilayer perceptron—trained on 32 time bins of each pulse achieves accuracy $0.816$ and ROC AUC $0.921$ at decision threshold $0.52$, with balanced true-positive and true-negative rates. Applied to unshielded $^{241}$AmBe pulses never used in training, the model separates the mixed field into neutron and gamma spectral components consistent with the shielded calibrations. The authors conclude that this two-stage architecture gives water-based detectors an operational radiation-identification capability.","pith_inferences":["Editorial inference: because the WCD detects neutrons only through the 2.22 MeV capture gamma, the machine-learning model may actually be separating 2.22 MeV capture-gamma pulses from 4.44 MeV source-gamma pulses; the paper does not test transfer to other neutron energies or geometries.","Editorial inference: the cutoff-based calibration could work as a self-calibration tool for other water Cherenkov installations, using any mixed source with a known high-energy gamma endpoint as the anchor.","Editorial inference: a direct test of particle-based discrimination would be to train on one neutron source and test on a different neutron source (different energy or moderation geometry); maintaining high AUC would show the features are neutron-specific.","Editorial inference: with ensemble weights all near one-third and similar F1 scores, a simple majority vote of the three classifiers might match the soft-voting performance, which would simplify field deployment; the paper does not report this comparison."],"forward_implications":["A source whose 3-sigma cutoff falls below the neutron threshold (~9000 ADU, about 2 MeV) can be screened as pure gamma without pulse-level analysis.","For sources above the threshold, the trained ensemble can decompose a mixed field into neutron and gamma components, as demonstrated on unshielded 241AmBe pulses not used in training.","Water Cherenkov detectors could serve as scalable, helium-free radiation monitors for nuclear security, using the same tanks as cosmic-ray observatories.","The method is detector-specific: any change in PMT gain, geometry, or electronics shifts the calibration, so the paper requires recalibration and possible classifier retraining before field deployment.","Longer acquisition times improve the reliability of the 3-sigma cutoff, so weak or distant sources need longer measurements than the 5-minute lab runs."],"supporting_citations":[{"why":"Establishes that a pure-water single-PMT WCD detects neutrons through the 2.22 MeV capture gamma, the physical basis for labeling neutron events.","marker":"[33]"},{"why":"Supplies the detector design, the 262 keV Cherenkov threshold, and prior neutron-detection calibration that this work extends.","marker":"[35]"},{"why":"Prior demonstration of neutron tagging in a water Cherenkov detector that motivates pulse-level neutron/gamma separation.","marker":"[39]"},{"why":"Earlier machine-learning approach to neutron identification in water Cherenkov detectors that this ensemble builds upon.","marker":"[22]"},{"why":"Companion two-stage gamma-neutron classification work that this paper continues and validates with the ensemble.","marker":"[32]"},{"why":"Frank-Tamm theory of Cherenkov emission, the physics underlying the detector response.","marker":"[17]"},{"why":"Reference for the gamma energies of 60Co and 137Cs used as calibration anchors and for radiation detection practice.","marker":"[23]"}],"fun_headline_variants":["Single-PMT water detector tells neutrons from gammas, AUC 0.92","Water Cherenkov tank identifies radioactive sources via pulse shapes","Neutron vs gamma: water detector with ML hits 0.92 AUC","One photomultiplier, water, ML: source ID reaches 0.92 AUC"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that pulses recorded behind lead shielding can be labeled neutron signals and pulses behind cadmium-lined borated paraffin can be labeled gamma signals, so that the pulse shapes learned by the classifier generalize to an unknown mixed radiation field.","fun_headline_variants_meta":{"raw":{"variants":["Single-PMT water detector tells neutrons from gammas, AUC 0.92","Water Cherenkov tank identifies radioactive sources via pulse shapes","Neutron vs gamma: water detector with ML hits 0.92 AUC","One photomultiplier, water, ML: source ID reaches 0.92 AUC"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2440,"prompt_tokens":1069,"completion_tokens":1371,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":685,"completion_tokens_details":{"reasoning_tokens":1285}},"tokens_in":685,"tokens_out":1371,"duration_ms":25087,"temperature":1.0,"reasoning_tokens":1285,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:17:02.285835+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained ensemble on pulses from a neutron-only source that produces capture gammas inside the water at energies different from 2.22 MeV, with no 4.44 MeV gamma present; if accuracy falls to chance, the discrimination is energy-based rather than neutron-based.","supporting_citations":[{"cited_title":"Sidelnik, H","cited_arxiv_id":null,"evidence_quote":"Establishes that a pure-water single-PMT WCD detects neutrons through the 2.22 MeV capture gamma, the physical basis for labeling neutron events."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the detector design, the 262 keV Cherenkov threshold, and prior neutron-detection calibration that this work extends."},{"cited_title":"Using machine learn- ing to improve neutron identification in water cherenkov detectors.Frontiers in big Data, 5:978857, 2022","cited_arxiv_id":null,"evidence_quote":"Earlier machine-learning approach to neutron identification in water Cherenkov detectors that this ensemble builds upon."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Companion two-stage gamma-neutron classification work that this paper continues and validates with the ensemble."},{"cited_title":"Coherent visible radiation of fast electrons passing through matter","cited_arxiv_id":null,"evidence_quote":"Frank-Tamm theory of Cherenkov emission, the physics underlying the detector response."},{"cited_title":"Radiation detection and measurement.John & Wiley Sons Inc, 2010","cited_arxiv_id":null,"evidence_quote":"Reference for the gamma energies of 60Co and 137Cs used as calibration anchors and for radiation detection practice."}],"review_version":1}