{"id":"3b5cbb97-6353-4a44-8e03-c9735161db17","arxiv_id":"2507.02810","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A technical status report on ML-based event selection, fast detector simulations, TPC simulations, and readout modeling for the HIBEAM-NNBAR experiment, with several methodological caveats.","lead":"The HIBEAM-NNBAR collaboration reports on computing and simulation advances for a proposed neutron-antineutron oscillation search, covering machine learning event selection, fast calorimeter simulations, TPC simulations, and readout electronics modeling. The paper presents preliminary results but contains a mis-specified ML metric and circular validation of the fast simulation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 100% background-rejection claim is statistically uninterpretable because Table 1 gives no test-set class counts or confidence intervals; with few background events, a perfect rejection is vacuous.","rationale":"The reader's verdict is REJECT, and the stress-test pass supports that conclusion, but through a more immediate internal blocker than the reader's stated weakest assumption. The central claim—that ML models achieve 100% background rejection with high efficiency—rests on Table 1. That table is uninterpretable without knowing the background class size in the test set and the associated binomial uncertainty. If the test set contains few background events, a perfect rejection rate is statistically vacuous, regardless of simulation fidelity. This concern is concrete and can be settled by a single reporting and calculation step. In addition, Eq. (1) does not implement the promised false-positive penalty: with beta=1000, the coefficient of FP is 0.999, nearly equal to the coefficient of FN. This makes the stated training objective internally inconsistent with the stated experimental requirement. Neither issue is an ad hominem or an external consensus disagreement; both are internal to the paper's stated results and method. The concrete test proposed—class counts plus a Clopper-Pearson interval—would determine whether the perfect rejection is a robust finding or an artifact of a small or unrepresentative test set. Until that test is performed, the reader's REJECT verdict remains appropriate; the same conclusion holds if the simulation-fidelity concern is later investigated, but the present lack of statistical grounding is the most load-bearing weakness. Therefore the verdict should be unchanged.","tokens_in":5115,"tokens_out":8426,"duration_ms":99015,"concrete_test":"Report the number of background events N_bkg in the test set and compute the Clopper-Pearson 95% confidence interval for the background rejection rate (or equivalently the false-positive rate). If the lower bound is not above 99%, replace the headline '100% rejection' with an upper-limit statement; if N_bkg is not reported and cannot be provided, the claim in Table 1 should be treated as unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 1 reports an unconditional 'Rejection 100' for every model, but the paper never states how many background events are in the 50% test split or the uncertainty on the zero false-positive count. If that test set contains only a handful of background events, a 0/N_bkg rejection rate is compatible with a large true false-positive rate; the headline claim then says nothing about the actual separation problem. This is an internal reporting gap, not an external physics assumption: the claim '100% background rejection is achievable' is only meaningful if the test set is large enough in the background class. The issue is compounded by Eq. (1): with beta=1000, the false-positive penalty is (beta-1)/beta = 0.999, numerically almost identical to the false-negative weight of 1, so the stated mechanism 'penalizing false positives' is not implemented as described. The perfect rejection may be an artifact of threshold selection on a small or highly separable sample rather than a robust property of the trained models.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes recent progress in computing and simulation techniques for the HIBEAM-NNBAR experiment at the European Spallation Source. Section 2 presents machine-learning models (Random Forest, XGBoost, LightGBM) trained on a simulated dataset of neutron-antineutron annihilation signals and CRY-generated cosmic-ray backgrounds, with a custom metric (Eq. 1) and threshold optimization on a validation set; Table 1 reports 100% background rejection with efficiencies up to 98.71% on a 50% test split. Section 3 introduces fast parametric simulations for the electromagnetic calorimeter, with resolution parametrizations for photons, charged pions, and protons, validated against WASA SEC data. Section 4 describes a TPC simulation framework using Garfield++, COMSOL, and Matlab, with a preliminary study of three zigzag pad geometries using 21 simulated muons. Section 5 outlines a readout simulation framework that models Poisson-distributed event timing, Geant4 energy deposition, electronics shaping, and Gaussian noise.","tokens_in":5358,"tokens_out":6925,"duration_ms":71885,"significance":"The paper reports a set of computational tools under development for the HIBEAM-NNBAR experiment. The strongest contribution is the ML event-selection study, which, if the perfect-separation claim could be substantiated with proper statistics, would be valuable for the experiment's background rejection strategy. The fast calorimeter simulation is a practical speed-up for detector design, but its predictive value depends on independent validation. The TPC and readout simulations are early-stage frameworks that will likely evolve. The paper is honest about several limitations (e.g., no shower development in fast simulation), which is a credit, but the current reporting does not yet support the headline claims.","major_comments":[{"comment":"The headline claim that XGBoost, LGBM, and Random Forest models achieve 100% background rejection on the test set is not statistically supported as reported. Table 1 gives no class counts for the 50% test split, so a perfect rejection may be vacuous if the background test sample is small; the authors should report the number of signal and background events in the test set and provide a confidence interval (e.g., Wilson) for the zero false-positive count. In addition, the custom metric in Eq. (1), with β=1000, assigns a false-positive weight of (β−1)/β = 0.999, essentially identical to the false-negative weight of 1, so the stated objective of 'penalizing false positives' is not realized by the equation as written; either the equation or the text needs correction.","section":"§2, Table 1 and Eq. (1)"},{"comment":"The validation of the fast simulation is circular for the photon parametrization: Eq. (2) is taken from reference [10], the WASA SEC detector paper, and the validation compares the fast simulation with real data from that same detector, so agreement is largely by construction. The charged-pion and proton parametrizations, Eqs. (3) and (4), are fitted to data from other detectors ([12] and [13]) and are not validated at all against the WASA SEC. The authors should either validate the fast simulation against independent data (e.g., a different calorimeter) or explicitly frame the model as an unvalidated interpolation tool rather than claiming 'good agreement' as evidence of fidelity.","section":"§3"},{"comment":"The TPC geometry study is based on only 21 simulated muons, with no statement about how these are distributed across the three pad designs. The residuals in Table 2 (e.g., real-path z residuals of 0.002±0.027, 0.003±0.014, and 0.005±0.017 cm for the prototype, double-pad, and double-angle designs) are all mutually consistent within their quoted uncertainties, so the conclusion that 'increasing pad count improves z-resolution' is not justified by the data. The authors should increase the sample size, report the per-configuration track counts, and apply a statistical test before drawing conclusions about pad geometry.","section":"§4, Table 2"}],"minor_comments":[{"comment":"The text states that PCA reduced the input variables to 14, but Table 1 lists 17 variables for the LGBM (Optuna) + PCA model; this inconsistency should be corrected.","section":"§2, Table 1"},{"comment":"The phrase 'to find he best hyperparameters' contains a typo and should read 'to find the best hyperparameters'.","section":"§2"},{"comment":"The fast simulation does not simulate shower development or position smearing; this limitation is stated in the text but should be more prominent in the abstract or conclusion so that the model's scope is not misrepresented.","section":"§3"},{"comment":"The captions of Figures 3 and 5 do not specify the number of simulated muons or the pad configuration; adding these details would aid reproducibility.","section":"§4, Figures 3 and 5"},{"comment":"The readout simulation framework is described without any comparison to measured electronics signals; the authors should state the current validation status of this framework.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The paper is in scope for physics.ins-det and is a useful progress report, but the ML background-rejection claim, which is the paper's headline, requires substantially stronger statistical reporting to be credible. The fast-simulation validation issue is also fundamental to the claimed performance. I believe these are fixable within a revision, so I recommend major revision rather than rejection, provided the authors address the specific points in the report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a collaboration status report, and its most striking claim—100% background rejection with up to 98.7% efficiency—does not hold up as written. The rest of the paper is a reasonable, if preliminary, set of detector studies.\n\nWhat is genuinely useful: the TPC simulation work is concrete, with three pad geometries compared and residuals reported; the framework combining Matlab, COMSOL, and Garfield++ is sensible, though 21 muons is a very small sample. The fast calorimeter simulation uses published resolution parametrizations from WASA, KEK, and proton data, and the check against WASA pi0 data is a real cross-check, not a circular fit to the same dataset. The readout simulation framework is straightforward and clearly described. These parts are solid enough for a proceedings paper.\n\nThe soft spot is the ML section. Table 1 reports 'Rejection 100' for every model, but the paper never states how many background events are in the 50% test split or gives any confidence interval. With a handful of background events, 0/N rejection is compatible with a large true false-positive rate. The custom metric in Eq. 1 is also mis-specified: with beta = 1000, (beta-1)/beta is 0.999, so false positives and false negatives have nearly equal weight. The text says the metric penalizes false positives, but numerically it barely does. On top of that, thresholds are tuned on the validation set, so the test numbers are a fit, not an independent prediction. The ML claim needs test-set counts, error bars, and a metric that actually implements the stated penalty.\n\nThere are also smaller issues: the fast simulation treats photons as well-separated by construction, so shower overlaps are ignored; the TPC sample of 21 muons is too small for the residual comparisons to mean much; and no code or data are provided, which limits reproducibility. None of these are fatal for a status report, but they reinforce that the paper is a progress update, not a definitive result.\n\nI would send this to peer review, but with a clear expectation of heavy revision. The collaboration is doing serious work for a proposed experiment, and a referee can push them to quantify the ML claim properly. As is, the headline result is unsupported, but the underlying detector studies are worth engaging with.","headline":"Useful status report with an overclaimed ML headline; the detector studies are fine, but the 100% rejection claim needs proper counting and a metric that does what it says.","tokens_in":5876,"tokens_out":1972,"would_cite":false,"duration_ms":25102,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine-learning models trained on simulation can reject 100% of cosmic-ray background while keeping up to 98.71% of neutron-annihilation signal in the HIBEAM-NNBAR test set, and fast parametric simulations of the calorimeter and TPC are…","keywords":["neutron-antineutron oscillation","baryon number violation","machine learning event selection","background rejection","fast detector simulation","time projection chamber","readout electronics simulation","cosmic-ray background"],"falsifier":"Take the trained classifiers and expose them to background events generated by a second, independent simulation or by real cosmic-ray calibration data; any event that passes the chosen threshold as a false positive would show that the 100% rejection is tied to the specific training simulation rather than to the background itself.","tokens_in":4959,"feed_emoji":"⚛️","tokens_out":9447,"duration_ms":93930,"temperature":0.7,"pith_summary":"The paper reports on practical computing tools for the proposed HIBEAM-NNBAR neutron-antineutron search: machine learning for event selection, fast parametrized detector simulation for design studies, and a readout-electronics simulator. Its central quantitative claim is that, on a simulated sample of 369,569 events, gradient-boosted tree models and a random forest trained with a false-positive-penalizing metric achieve 100% background rejection while keeping up to 98.71% of signal. The same approach reduces the feature set from 49 high-level variables to 26 with essentially no loss in efficiency. If these results transfer to real data, the experiment could rely on essentially background-free event selection with most of the signal preserved, which matters because the expected annihilation rate is extremely small. The remaining sections give faster approximations of calorimeter response and a detailed simulation chain for the time projection chamber and readout electronics to support detector development.","feed_headline":"ML rejects all background, keeps 98.7% of signal","feed_subtitle":"Boosted-tree models trained on simulated neutron-annihilation events hit 100% background rejection in the HIBEAM-NNBAR test set.","key_machinery":"The central object is the custom asymmetric metric $C_{\\mathrm{metric}} = \\mathrm{FN} + (\\beta-1)\\mathrm{FP}/\\beta$ with $\\beta=1000$, which makes a single false positive 999 times more costly than a single false negative; models are tuned and thresholded on this metric, and that is why they reach 100% background rejection. For the simulation sections, the carrying objects are parametric resolution functions such as $\\sigma_E/E = 0.05/\\sqrt{E[\\mathrm{GeV}]}$ for photon energy deposits, a Gaussian-smearing wrapper that turns particle kinetic energy into calorimeter signals, and a modular readout chain combining Poisson event timing, energy deposition, convolution with an electronic impulse response, and Gaussian noise.","core_discovery":"The authors claim that a custom cost metric of the form $C_{\\mathrm{metric}} = \\mathrm{FN} + (\\beta-1)\\mathrm{FP}/\\beta$ with $\\beta=1000$, applied to tree-based classifiers, produces a decision threshold with zero false positives on the test set for several model families, with signal efficiencies of 95.81%, 98.71%, and 98.69%. Correlation-based feature reduction to 26 variables keeps 100% rejection at 98% efficiency, while PCA reduction to 14 variables drops efficiency to 74%. The paper further claims that a fast simulation that smears energy deposits with the measured calorimeter resolution reproduces real data from a lead-glass calorimeter, and that a detailed gas-detector simulation reconstructs simulated muon tracks with residuals at the level of a few tenths of a millimeter or better.","pith_inferences":["Editorial inference: because the background model is a single cosmic-ray generator, the claimed 100% rejection is a proof-of-principle for the custom metric rather than a guarantee about real running conditions; a different background source, such as beam-related neutrons, could break it.","Editorial inference: the custom metric with $\\beta=1000$ encodes the physics priority of never letting background through, so the same training recipe should transfer to other experiments with extreme background-suppression requirements.","Editorial inference: the fast simulation ignores position smearing and shower overlap, so its reconstructed pion-mass peak is likely narrower than what a full simulation or real data would show; testing overlap effects is a natural next step.","Editorial inference: if the readout simulation is faithful, it could be used to train classifiers on digitized waveforms rather than high-level variables, potentially improving rejection and making selection robust to electronics variations."],"forward_implications":["Event selection for HIBEAM-NNBAR could use machine-learned scores instead of hand-crafted cuts, with zero simulated background accepted at thresholds that keep roughly 98% of annihilation signal.","The 26-variable subset performs nearly as well as the full 49-variable set, suggesting the analysis can be simplified or made more interpretable without much loss.","The fast calorimeter simulation is accurate enough for design studies, predicting about 73% acceptance for neutral pions from annihilation at around 150 MeV, which shortens detector-optimization cycles.","TPC pad geometry can be tuned in simulation: more pads and steeper zigzag angles improve z-resolution, while y-resolution is set by the number of readout rows.","The readout simulation framework can show how electronic noise and pulse shaping affect physics variables, allowing electronics specifications to be evaluated before construction."],"supporting_citations":[{"why":"It is the earlier HIBEAM-NNBAR simulation work that produced the event dataset used here for training and testing the classifiers.","marker":"[6]"},{"why":"It supplies the cosmic-ray background events used with the annihilation signal to form the training and test sets.","marker":"[7]"},{"why":"It provides the hyperparameter optimization used to tune each machine-learning model.","marker":"[8]"},{"why":"It gives the electromagnetic calorimeter energy-resolution formula that the fast simulation uses for photon energy deposits.","marker":"[10]"},{"why":"It provides the charged-pion energy-loss measurements used to parametrize pion response in the fast simulation.","marker":"[12]"},{"why":"It provides the proton energy-loss data used to parametrize proton response in the fast simulation.","marker":"[13]"},{"why":"It supports the particle-transport simulation that underlies both the fast-simulation validation and the energy-deposition step of the readout framework.","marker":"[11]"},{"why":"It is the gas-detector simulation code used to compute electron drift and reconstruct tracks in the TPC studies.","marker":"[14]"},{"why":"It is the comparable readout-simulation effort whose methodology this work extends to HIBEAM-NNBAR detectors.","marker":"[18]"}],"fun_headline_variants":["ML rejects 100% background, keeps 98.7% signal","Zero background in ML: 98.7% signal efficiency","Fast sim and ML sharpen HIBEAM-NNBAR searches","HIBEAM ML: all background gone, signal stays at 98.7%","Neutron experiment: ML perfect rejection, high signal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 100% background rejection is a statement about the simulated test set: if the simulated cosmic-ray events do not match the real background the detector will see, the rejection rate will be lower on actual data.","fun_headline_variants_meta":{"raw":{"variants":["ML rejects 100% background, keeps 98.7% signal","Zero background in ML: 98.7% signal efficiency","Fast sim and ML sharpen HIBEAM-NNBAR searches","HIBEAM ML: all background gone, signal stays at 98.7%","Neutron experiment: ML perfect rejection, high signal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3146,"prompt_tokens":780,"completion_tokens":2366,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":396,"completion_tokens_details":{"reasoning_tokens":2272}},"tokens_in":396,"tokens_out":2366,"duration_ms":22409,"temperature":1.0,"reasoning_tokens":2272,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:20:09.524922+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the trained classifiers and expose them to background events generated by a second, independent simulation or by real cosmic-ray calibration data; any event that passes the chosen threshold as a false positive would show that the 100% rejection is tied to the specific training simulation rather than to the background itself.","supporting_citations":[{"cited_title":"A Computing and Detector Simulation Framework for the HIBEAM/NNBAR Experimental Program at the ESS","cited_arxiv_id":"2106.15898","evidence_quote":"It is the earlier HIBEAM-NNBAR simulation work that produced the event dataset used here for training and testing the classifiers."},{"cited_title":"Hagmann et al., CRY: Cosmic-Ray Shower Generator , in IEEE Nucl","cited_arxiv_id":null,"evidence_quote":"It supplies the cosmic-ray background events used with the annihilation signal to form the training and test sets."},{"cited_title":"Akiba et al., Optuna: Hyperparameter Optimization Framework, in Proc","cited_arxiv_id":null,"evidence_quote":"It provides the hyperparameter optimization used to tune each machine-learning model."},{"cited_title":"Bargholtz et al., Nucl","cited_arxiv_id":null,"evidence_quote":"It gives the electromagnetic calorimeter energy-resolution formula that the fast simulation uses for photon energy deposits."},{"cited_title":"Yamazaki et al., Nucl","cited_arxiv_id":null,"evidence_quote":"It provides the charged-pion energy-loss measurements used to parametrize pion response in the fast simulation."},{"cited_title":"Merchez et al., Nucl","cited_arxiv_id":null,"evidence_quote":"It provides the proton energy-loss data used to parametrize proton response in the fast simulation."},{"cited_title":"Schindler, Garfield++ (2010), https://garfieldpp.web.cern.ch/ garfieldpp/","cited_arxiv_id":null,"evidence_quote":"It is the gas-detector simulation code used to compute electron drift and reconstruct tracks in the TPC studies."},{"cited_title":"Luna et al., FPGA-Based Simulator for ATLAS TileCal Readout , in Proc","cited_arxiv_id":null,"evidence_quote":"It is the comparable readout-simulation effort whose methodology this work extends to HIBEAM-NNBAR detectors."}],"review_version":1}