REVIEW 4 major objections 5 minor 34 references
Automatizing the search for mass resonances using BumpNet
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A single neural network, BumpNet, turns arbitrary invariant-mass histograms into per-bin resonance-significance maps, matching ideal likelihood-ratio tests within about 0.5–0.8σ and spotting injected new-physics signals across tens of…
desk verdict A solid, honest generalization of the DDP bump hunter that deserves a referee, with the caveat that its unbiasedness claim is conditional on training coverage—a point the paper itself makes in Sec 3.2.2. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is BumpNet, a convolutional neural network with four parallel stacks of kernel sizes 3, 9, 15, and 25, a skip connection that preserves raw bin counts, and a per-bin multilayer perceptron head. It is trained on roughly three million samples built by adding one-bin-wide Gaussian signals to background histograms drawn from eleven analytic functional forms and from smoothed fits to simulated LHC-like events, with the target being the true likelihood-ratio significance computed from the known injected signal and background. The network's job is to compress the background shape and signal location into a per-bin significance value, so that at inference time no fitting or functional-form choice is needed.
What would settle it
A direct test: take a background-only invariant-mass histogram from a real LHC selection whose kinematic cuts produce a sharply depleted high-mass tail, inject a Gaussian bump of known strength, and compare BumpNet's predicted significance against the exact likelihood-ratio significance. The paper predicts a systematic negative ΔZmax in just those sparse high-mass bins; if the same bias appears on real backgrounds outside the training set, the unbiasedness claim would be falsified in the regime that matters.
Extended reading notes
Core claim
The paper's central claim is that one network, BumpNet, predicts the bin-by-bin statistical significance of a resonant bump in an invariant-mass histogram without knowing the signal or background model, and that these predictions approximate the ideal likelihood-ratio test. On nominal test histograms the difference between the predicted maximum significance and the true one has a central value near zero, with standard deviations of 0.53σ for analytic backgrounds and 0.75σ for simulated LHC-like backgrounds; the false-positive rate for a 5σ threshold is 0.048% and 0.129%, respectively. On real data, BumpNet reproduces the published local significances for the Higgs-to-two-photon distribution, predicting 4.5σ versus a likelihood-ratio significance of 4.2σ, and follows the reported dilepton-resonance significances within its known variance. Finally, when applied to about 40,000 histograms with injected beyond-Standard-Model signals, BumpNet finds each signal at the expected mass and object combination, and the paper's Global Analysis Algorithm removes most of the look-elsewhere false positives.
Load-bearing premise
The load-bearing premise is that the background shapes used to train BumpNet, namely eleven analytic curves plus fits to simulated LHC-like events, are representative of the backgrounds real LHC selections will produce; if a real selection depletes the high-mass tail more severely than anything in training, the paper's own systematic-distortion tests show BumpNet will under-predict significance there, and the claimed unbiasedness will not transfer.
Editorial extensions
If this is right
- A single trained BumpNet can replace per-analysis background fits for narrow-resonance searches across many final states, since it handles variable bin counts, dynamic ranges down to sparse bins, and background shapes from analytic, simulated, and real sources.
- Resonance searches can be carried out on tens of thousands of histograms at once: the paper demonstrates scanning about 40,000 LHC-like histograms and finding injected BSM signals at the expected masses.
- The look-elsewhere effect can be tamed without a full trials-factor computation by requiring a family of correlated histograms—same mass, same object combination—to claim an excess, which reduces false-positive histograms from 28 to 9 in the channel-3 background-only scan.
- BumpNet's significance predictions are slightly biased for signals wider than one bin, with 2- and 3-bin Gaussian widths showing a growing negative bias, so re-binning histograms to detector resolution is essential for its sensitivity to narrow resonances.
Reading between the lines
- Inference: Because the failure mode identified in the paper is concentrated in sparse high-mass bins under standard and reverse bias distortions, a practical deployment would likely include a per-bin minimum background count cut or a separate network trained on depleted high-mass shapes.
- Inference: The same network architecture could be retrained on signal templates other than one-bin Gaussians, such as Breit–Wigner shapes or multi-bin decay chains, extending the approach from bump hunting to general spectral anomaly detection without changing the core likelihood-ratio target.
- Inference: The Global Analysis Algorithm's correlation criterion—same mass, same object type, statistically uncorrelated histograms—looks transferable to other broad scans, but its thresholds, such as 5σ seeds and two-bin tolerance, would need to be re-tuned on each dataset before use as a formal discovery procedure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents BumpNet, a convolutional neural network that maps invariant-mass histograms to per-bin statistical significances for resonant bumps. BumpNet is trained in a supervised manner on a large set of histograms built from eleven analytic background functions and from Dark Machines simulated backgrounds, with Poisson fluctuations and injected Gaussian signals; the targets are likelihood-ratio significances computed from the known underlying signal and background shapes. The network is validated on held-out nominal histograms, on systematically distorted background shapes, on ATLAS H→γγ and di-lepton data, and on Dark Machines BSM signals injected into SM backgrounds. The claimed results are near-zero mean bias (ΔZmax ≈ 0.1σ) with small spread (0.53–0.75σ) on nominal tests, a false-positive rate near 0.1% at 5σ, reproduction of published ATLAS significances, and successful identification of several BSM signals when combined with a Global Analysis Algorithm (GAA) to address the look-elsewhere effect.
Significance. If the central claims hold, BumpNet would be a valuable tool for the data-directed paradigm, enabling rapid scans of the large space of invariant-mass histograms that are not currently examined by dedicated LHC analyses. The paper has real strengths: the training target is an external, well-defined likelihood-ratio significance, so the network is a supervised surrogate rather than a circular fit; the validation includes truly held-out distributions and real ATLAS data; and the application to BSM signals in a realistic simulated sample with O(10^4) histograms is a useful proof of concept. At the same time, the strongest claimed properties — unbiasedness and transfer to arbitrary LHC-like backgrounds — are conditional on training-distribution coverage, and Section 3.2.2 demonstrates a concrete violation of that condition. The paper is honest about this limitation, but the limitation directly affects the paper's central message and needs to be addressed quantitatively before the claims can be taken as established.
major comments (4)
- [Sec. 3.2.2, especially Figs. 11 and 12] The standard and reverse bias transformations produce a systematic negative ΔZmax, concentrated in high-mass, low-statistics bins; the paper attributes this to the training examples predominantly having more background events at high mass and to the impossibility of negative Poisson fluctuations. This is not a peripheral robustness check: it is exactly the regime that can arise in real LHC selections whose cuts deplete the high-mass tail, and the paper's own unbiasedness claim in Sec. 3.1 is stated for the nominal training distribution. The manuscript should either provide a quantitative coverage test (e.g., a measure of how far an application background can deviate from the training manifold before the bias exceeds a stated tolerance) or explicitly restrict the unbiasedness claim to backgrounds that satisfy the training distribution's tail behaviour. A qualitative statement that the bias appears only in sparse bins is insufficient, because the intended application is precisely a scan of many histograms, some of which will fall in this regime.
- [Sec. 3.1, Fig. 5] The decision to exclude the first 10% of bins from the definition of Zmax is introduced after observing the disagreement in that region and is then applied to all subsequent analysis. Because this exclusion is a post-hoc choice made on the same test set used to report the unbiasedness and false-positive rates, it should be validated as a generalizable procedure rather than a fitted modification. For example, the authors could justify it from the training distribution (e.g., by showing that the first 10% region is underrepresented in training) or demonstrate that the exclusion also reduces bias on an independent test set that was not used to motivate it. As written, the reported false-positive rates and the mean-bias numbers are conditioned on this choice, which weakens the claim that BumpNet is unbiased on untouched histograms.
- [Sec. 3.2.3, Fig. 13] For injected signal widths of 2 and 3 bins, the mean ΔZmax becomes approximately -1.1σ and -1.9σ, respectively, i.e., BumpNet systematically underestimates the likelihood-ratio significance. The paper argues that the primary purpose is identification rather than precise significance, but the network's output is explicitly a significance, and the realistic BSM signals used in Sec. 4.2 are broader than one bin. The manuscript should quantify how this bias propagates into the real-data and BSM-signal results (e.g., whether the 4.5σ Higgs prediction or the >5σ claims in Sec. 4.2 would be reduced if the signal width is not exactly one bin). Without this, the agreement with ATLAS results and the >5σ BSM findings are not fully established for signals that are not perfectly narrow.
- [Sec. 4.2, Figs. 19 and 20] The LEE handling relies on the GAA with several thresholds and tolerances (5σ seed threshold, 5σ family threshold, two-bin mass tolerance, and the requirement of at least two histograms in a family) that are chosen without a systematic optimization or a study of their dependence. The claim that the GAA 'manages' the look-elsewhere effect is based on a single set of choices: it reduces false positives from 25 to 18 in channel 2b and from 28 to 9 in channel 3, but does not eliminate them, and the surviving false positives are disposed of by additional ad hoc observations. Moreover, the W′→qqνν signal is not recovered by the GAA because it appears in only one histogram. The paper should present a more systematic evaluation of the GAA—e.g., false-positive and true-positive rates as functions of its thresholds—or soften the conclusion that the method provides a practical solution to the LEE.
minor comments (5)
- [Sec. 3.2.3, last paragraph] There is a typo: 'As will be shown will be shown' should read 'As will be shown'.
- [Sec. 4.1, Fig. 16 caption] The caption states that the BumpNet prediction is the dashed line and the published significance is the solid line, but the main text describes the opposite assignment. These should be made consistent.
- [Sec. 2.2.2, 'Smoothing procedure'] The phrase 'aformentioned procedure' contains a typo; it should be 'aforementioned procedure'.
- [Sec. 4.2, first paragraph] The paper states that 25 and 28 histograms show a significance above 5σ, corresponding to a false-positive rate 'on the order of 0.1%'. For channel 2b, 25 out of 8,104 is 0.31%, which is not 0.1%; please quote the exact rates or state the range.
- [Sec. 2.2.2, 'Signal regions'] The text describes categories with 'exactly two charged leptons' and later subdivides them, but the definition of the parent categories and the role of the Z-candidate veto could be stated more explicitly; a diagram or a precise enumeration of the category tree would aid reproducibility.
Circularity Check
No material circularity: BumpNet is a supervised surrogate for the analytic likelihood-ratio significance, validated on held-out, transformed, simulated, and real data.
full rationale
BumpNet's central claim is that a single network can reproduce the likelihood-ratio significance z_LR computed from known injected signals and known background shapes. The training targets are generated by an external, closed-form definition, not by the network itself; the evaluation compares predictions on held-out histograms, unseen DM distributions, systematically distorted backgrounds, ATLAS-style background functions, real H→γγ and dilepton data, and injected BSM signals. None of these comparisons reduces by construction to the training objective, and the paper documents concrete generalization failures (e.g., Sec. 3.2.2 standard/reverse bias underestimates significance in sparse high-mass bins, and Sec. 3.2.3 wider signals degrade performance). The only self-reference is the prior bump-hunt DDP [21], which motivates the architecture but does not supply the validated results; the Sec. 3.2.2 coverage limitation is a robustness concern, not a circular reduction. The paper is therefore essentially self-contained, with a score of 1 reflecting only the minor, non-load-bearing self-citation to the authors' earlier DDP work.
Assumptions & free parameters
free parameters (5)
- First-10% bin exclusion =
10% of histogram bins
- GAA thresholds =
5σ seed, 5σ family threshold, 2-bin tolerance
- Histogram quality cuts =
≥100 events, >30 bins
- Bin-width factor =
width(m) = 0.5*σ(m)
- Network hyperparameters =
4 stacks, kernels 3/9/15/25, 64 channels, MLP 128/64/32, batch 5000, 300 epochs
assumptions (5)
- domain assumption Smoothly falling background after the histogram peak
- domain assumption 4th-order log-polynomial fits define the true background
- domain assumption 1-bin Gaussian signals represent expected resonances after resolution binning
- domain assumption Dark Machines samples emulate real LHC backgrounds
- standard math Likelihood-ratio test statistic is the correct significance measure
Cite this review
Pith. "Pith review of Automatizing the search for mass resonances using BumpNet." pith.science (2026). https://pith.science/paper/BKPFU7OK
@misc{pith2026250105603,
author = {Pith},
title = {Pith review of: Automatizing the search for mass resonances using BumpNet},
year = {2026},
howpublished = {\url{https://pith.science/paper/BKPFU7OK}},
note = {Machine review of arXiv:2501.05603}
}
read the original abstract
The search for resonant mass bumps in invariant-mass distributions remains a cornerstone strategy for uncovering Beyond the Standard Model (BSM) physics at the Large Hadron Collider (LHC). Traditional methods often rely on predefined functional forms and exhaustive computational and human resources, limiting the scope of tested final states and selections. This work presents BumpNet, a machine learning-based approach leveraging advanced neural network architectures to generalize and enhance the Data-Directed Paradigm (DDP) for resonance searches. Trained on a diverse dataset of smoothly-falling analytical functions and realistic simulated data, BumpNet efficiently predicts statistical significance distributions across varying histogram configurations, including those derived from LHC-like conditions. The network's performance is validated against idealized likelihood ratio-based tests, showing minimal bias and strong sensitivity in detecting mass bumps across a range of scenarios. Additionally, BumpNet's application to realistic BSM scenarios highlights its capability to identify subtle signals while managing the look-elsewhere effect. These results underscore BumpNet's potential to expand the reach of resonance searches, paving the way for more comprehensive explorations of LHC data in future analyses.
Reference graph
Works this paper leans on
- [1]
-
[2]
G. Aad et al. (ATLAS), Phys. Rev. D109, 112008 (2024), arXiv:2402.16576 [hep-ex]
work page Pith review arXiv 2024
- [3]
-
[4]
Pursuit of paired dijet resonances in the Run 2 dataset with ATLAS
G. Aad et al. (ATLAS), Phys. Rev. D108, 112005 (2023), arXiv:2307.14944 [hep-ex] . – 29 –
work page Pith review arXiv 2023
-
[5]
A. Tumasyan et al. (CMS), Phys. Rev. D110, 012013 (2024), arXiv:2402.11098 [hep-ex]
work page Pith review arXiv 2024
- [6]
-
[7]
A. Tumasyan et al. (CMS), Phys. Rev. D108, 012009 (2023), arXiv:2205.01835 [hep-ex]
arXiv 2023
-
[8]
J. H. Kim, K. Kong, B. Nachman, and D. Whiteson, JHEP04, 030, arXiv:1907.06659 [hep-ph]
work page Pith review arXiv 1907
Show all 34 references
-
[9]
A. M. Sirunyanet al. (CMS), JHEP 07, 208, arXiv:2103.02708 [hep-ex]
- [10]
- [11]
-
[12]
A. M. Sirunyanet al. (CMS), JHEP 05, 033, arXiv:1911.03947 [hep-ex]
1911 arXiv
-
[13]
Aaboud et al
M. Aaboud et al. (ATLAS), Phys. Lett. B775, 105 (2017), arXiv:1707.04147 [hep-ex]
2017 arXiv
-
[14]
Aad et al
G. Aad et al. (ATLAS), Phys. Lett. B822, 136651 (2021), arXiv:2102.13405 [hep-ex]
2021 arXiv
-
[15]
A. M. Sirunyanet al. (CMS), Phys. Rev. D98, 092001 (2018), arXiv:1809.00327 [hep-ex]
2018 arXiv
-
[16]
ATLAS Collaboration, ATLAS Public Results on Searches for New Phenomena (2024), [Online; accessed 19-December-2024]
2024
-
[17]
Aad et al
G. Aad et al. (ATLAS), Phys. Rev. Lett.125, 131801 (2020), arXiv:2005.02983 [hep-ex]
2020 arXiv
-
[18]
Aad et al
G. Aad et al. (ATLAS), Phys. Rev. Lett.132, 081801 (2024), arXiv:2307.01612 [hep-ex]
2024 arXiv
-
[19]
S. V. Chekanov, Estimation of the chances to find new phenomena at the LHC in a model-agnostic combinatorial analysis (2023), arXiv:2311.09012 [hep-ph]
2023 arXiv
-
[20]
Belis, P
V. Belis, P. Odagiu, and T. K. Aarrestad, Rev. Phys.12, 100091 (2024), arXiv:2312.14190 [physics.data-an]
2024 arXiv
-
[21]
Volkovich, F
S. Volkovich, F. De Vito Halevy, and S. Bressler, Eur. Phys. J. C82, 265 (2022), arXiv:2107.11573 [hep-ex]
2022 arXiv
-
[22]
Birman, B
M. Birman, B. Nachman, R. Sebbah, G. Sela, O. Turetz, and S. Bressler, Eur. Phys. J. C82, 508 (2022), arXiv:2203.07529 [hep-ph]
2022 arXiv
-
[23]
Bressler, I
S. Bressler, I. Savoray, and Y. Zurgil, Phys. Rev. D110, 095004 (2024), arXiv:2401.09530 [hep-ex]
2024 arXiv
-
[24]
Cowan, K
G. Cowan, K. Cranmer, E. Gross, and O. Vitells, Eur. Phys. J. C71, 1554 (2011), [Erratum: Eur.Phys.J.C 73, 2501 (2013)], arXiv:1007.1727 [physics.data-an]
2011 arXiv
-
[25]
Nair and G
V. Nair and G. E. Hinton, inProceedings of the 27th International Conference on Machine Learning (ICML-10) (2010) pp. 807–814
2010
-
[26]
D. P. Kingma and J. Ba, Adam: A method for stochastic optimization (2017), arXiv:1412.6980 [cs.LG]
2017 arXiv
-
[27]
Aarrestad et al., SciPost Phys.12, 043 (2022), arXiv:2105.14027 [hep-ph]
T. Aarrestad et al., SciPost Phys.12, 043 (2022), arXiv:2105.14027 [hep-ph]
2022 arXiv
-
[28]
Alwall, R
J. Alwall, R. Frederix, S. Frixione, V. Hirschi, F. Maltoni, O. Mattelaer, H. S. Shao, T. Stelzer, P. Torrielli, and M. Zaro, JHEP07, 079, arXiv:1405.0301 [hep-ph]
-
[29]
Sjöstrand, S
T. Sjöstrand, S. Ask, J. R. Christiansen, R. Corke, N. Desai, P. Ilten, S. Mrenna, S. Prestel, C. O. Rasmussen, and P. Z. Skands, Comput. Phys. Commun.191, 159 (2015), arXiv:1410.3012 [hep-ph] . – 30 –
2015 arXiv
-
[30]
de Favereau, C
J. de Favereau, C. Delaere, P. Demin, A. Giammanco, V. Lemaître, A. Mertens, and M. Selvaggi (DELPHES 3), JHEP02, 057, arXiv:1307.6346 [hep-ex]
-
[31]
Aaboud et al
M. Aaboud et al. (ATLAS), Eur. Phys. J. C80, 1104 (2020), arXiv:1910.04482 [hep-ex]
2020 arXiv
-
[32]
Virtanen, R
P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, S. J. van der Walt, M. Brett, J. Wilson, K. J. Millman, N. Mayorov, A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C. J. Carey, İ. Polat, Y. Feng,...
2020
- [33]
- [34]
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.