REVIEW 2 major objections 8 minor 11 references
Biased rate estimates in bump-hunt searches
T0 review · 2 major / 8 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read In a bump-hunt search with unknown mass, reporting the most significant excess overestimates the signal rate by roughly 10% and underestimates the mass uncertainty by 10–20%, so upgrading a 3σ hint to 5σ requires about 960 inverse…
desk verdict Real effect, honest toy study, but the headline 13% rate bias is a mean-calibration inversion and needs a conditional treatment before I'd trust the number. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the heuristic mapping from true to observed significance, $y=\left(x^n+c^n\right)^{1/n}$ with $c=2.04$ and $n=3.06$, calibrated from toy fits over a scan range of $\pm8\sigma$ that mimics diphoton Higgs searches. The toy model is a binned negative-log-likelihood fit of a Gaussian signal on a uniform known background (1000 events per bin, signal width $\sigma=0.5$), with the mass and rate left free and the most significant point across the scanned mass grid reported. The mapping converts an observed 3σ into the underlying 2.66σ used for luminosity extrapolation, and it is the reason the required additional luminosity becomes 960 inverse femtobarns rather than 720. For the mass bias, the machinery is the pull distribution of the fitted mass from the greedy scan, compared against the minimizer's error estimate and the $-2\Delta\ln L=1$ error convention.
What would settle it
Repeat the same greedy mass-scan simulation with a non-uniform background (for example, an exponential fall) and with the background normalization and slope fitted rather than fixed; if the observed-to-true significance ratio for a genuinely 3σ signal stays within a few percent of unity, the claimed 13% rate bias and 20% mass-error inflation would not generalize.
Extended reading notes
Core claim
The core discovery is that the greedy choice of the most significant excess in an unknown-mass scan is a biased estimator: background fluctuations near the true signal are merged with the fitted peak, inflating the measured rate by about 13% when the reported significance is 3σ and degrading the mass error by 10–20%, with non-Gaussian tails beyond 5σ. In the toy model, the relation between the significance expected at the true mass, x, and the mean observed maximum significance, y, is well described by y = (x^n + c^n)^(1/n) with c = 2.04 and n = 3.06, so an observed 3σ corresponds to an underlying 2.66σ and an observed 5σ to an underlying 4.9σ. Inverting this relation, an observed 3σ hint requires roughly 2.4 times the original dataset as additional data (960 inverse femtobarns rather than 720 after an initial 400) to reach a 5σ discovery, and the mass uncertainty remains underestimated until the significance is at least 5σ.
Load-bearing premise
The quantitative 10–13% rate bias, 20% mass-error inflation, and 2.4-times luminosity factor come from a toy model with a flat known background, a Gaussian signal of known width, and no systematic errors, so they carry over to real experiments only if those idealizations are representative.
Editorial extensions
If this is right
- An observed 3σ bump-hunt excess corresponds to an underlying signal of only about 2.66σ at the true mass, so fitted signal rates from evidence-level peaks are biased high by roughly 13%.
- The mass of a particle claimed at 3–4σ carries an underestimated uncertainty: the core error is too small by 10–20%, and there are tails with pulls beyond 5, so early mass measurements should be quoted with enlarged errors.
- If the signal width or detector resolution is not known perfectly, the rate bias grows and can reach 10% even at 5σ, so fitting the width freely makes the overestimate worse.
- Turning a 3σ hint into a 5σ discovery requires about 960 inverse femtobarns of additional data after an initial 400 inverse femtobarns, not the 720 a naive scaling suggests.
- The bias is unavoidable whenever the mass is unknown and the reported significance is the maximum over a scan; it is absent only when the signal mass is fixed in advance.
Reading between the lines
- The same greedy-selection mechanism should apply to any search that optimizes over nuisance parameters—unknown width, unknown line shape, or unknown location in a multidimensional space—so analogous biases should appear in axion-line, dark-matter, and gravitational-wave searches, with magnitudes growing with the number of scanned hypotheses.
- A practical correction is possible: use the inverse of the heuristic mapping to deflate an observed significance before setting a luminosity target, effectively a trials correction for the signal strength rather than the p-value.
- For real evidence-level excesses, such as the ~95 GeV diphoton excess, the measured signal rate should drift downward as more data accumulate and the significance should grow more slowly than the square root of luminosity; tracking those two quantities over time would provide a direct empirical test of the bias.
- Combining multiple channels or datasets at evidence level should inflate each channel's rate and mass error by the channel-specific bias before averaging, otherwise the combined result will inherit a low mass error and a high rate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the selection ('greedy') bias introduced when a bump-hunt search scans over an unknown mass and reports the most significant excess. Using an idealized toy model with a uniform background of known rate and a Gaussian signal of known width, the authors simulate many experiments with injected signals of various strengths, scan over a mass range, and record the maximum fitted signal significance. They find that the mean observed significance as a function of the true (single-mass) significance is well described by y=(x^n+c^n)^(1/n) with c=2.04 and n=3.06 for a scan range mimicking H→γγ. Inverting this curve, they claim an observed 3σ excess corresponds to a true significance of about 2.66σ, i.e., a rate overestimate of ~13%, and that reaching 5σ requires 2.4 times the original dataset (e.g., 960 fb^{-1} vs 720 fb^{-1}). They also quantify the mass-measurement pull in ensembles conditioned on the observed significance, finding that the Gaussian core of the mass pull has width ~1.19 at 3σ and non-Gaussian tails, supporting the claim that mass uncertainties are underestimated by ~20%. The main quantitative rate-bias claims are, however, based on inverting the mean calibration curve rather than computing the conditional distribution of the true rate given an observed significance; this is the major issue.
Significance. The paper addresses a real and often overlooked selection effect: quoting the most significant mass hypothesis in a search with unknown particle mass positively biases the measured signal rate and underestimates mass errors. The toy model is transparent and the fit validation is meticulous (pull distribution with RMS 0.9998, non-closure of only 0.4%). The mass-pull analysis in Section 4 is methodologically sound because it conditions on the observed significance using an explicit prior on the true signal rate. If the rate-bias calibration is corrected to the same standard, the paper will provide a useful quantitative caution for the interpretation of 3σ hints and for luminosity planning. The qualitative conclusion is robust; the quantitative LHC extrapolation is currently not directly supported and should be revised.
major comments (2)
- [Section 3, Eq. (2), Fig. 3(left)] The paper fits the heuristic y=(x^n+c^n)^(1/n) to the mean observed significance y as a function of the true significance x, then inverts this relation to state that an observed 3σ excess corresponds to a true significance of 2.66σ and a rate overestimate of 13%. The abstract and Section 5 make claims about an observed 3σ excess, but the quantity implied by those claims is the conditional expectation E[x|y=3] (or a posterior interval), not g^{-1}(3) with g(x)=E[y|x]. Because Fig. 2 shows substantial scatter of y for fixed x and g is nonlinear, these two quantities can differ materially; for the uniform prior on x used in Section 4, the concavity of g^{-1} suggests E[x|y=3] is smaller than 2.66σ, so the rate bias may be larger than 13%. The paper does not compute this conditional quantity anywhere in Sections 3 or 5, despite doing exactly this type of conditioning in Section 4 for the mass-pull distribution. Please re-derive the rate-bias claim by generating an ensemble with a prior on true rates, conditioning on observed significance, and reporting E[x|y] (or the median and a credible interval).
- [Section 5, luminosity projection] The statement that a 3σ evidence at 400 fb^{-1} requires 960 fb^{-1} rather than 720 fb^{-1} relies on the same inverted curve. There are two distinct quantities being conflated: the true significance x such that the *median* observed significance is 3 (i.e., g(x)=3), and the conditional expected true significance given an observed 3σ. The forward mapping g(x)=5 is the right target for planning a median 5σ reach, but the starting point "an observed 3σ evidence" is not the same as the x that yields median 3σ. After correcting the calibration as suggested in the previous comment, the luminosity projection should be re-derived; it may increase substantially (e.g., if E[x|y=3]≈2.4σ, the required total luminosity would be about (5/2.4)^2≈4.3 times the original rather than 3.4 times). The current 960/fb number is therefore not directly supported by the analysis.
minor comments (8)
- [Section 2, paragraph 2] The phrase "this ensures that so that there is negligible signal density" contains a duplicated "that so that" and should be corrected.
- [Section 2, paragraph 3] The word "whuch" is a typo and should read "which".
- [Eq. (2)] The notation "n√" is malformed; it should be displayed as an n-th root, e.g., \sqrt[n]{x^n+c^n}.
- [Section 3, near Fig. 2] The sentence "The other data show the impact of scanning a range for the highest excess" is ambiguous; it should read "The other curves show".
- [Section 5, paragraph 1] The phrase "has the advantage that that they are independent" contains a repeated "that" and should be corrected.
- [Abstract] The phrase "the data-doubling period is now measured in years" is colloquial; consider "the time to double the dataset is now measured in years" for clarity.
- [Section 4] The prior on true signal rates (uniform between 0 and 8σ) is a strong assumption; the authors test robustness to the prior upper limit and to conditioning on expected significance, but the manuscript should state explicitly that the quoted mass-pull widths are conditional on that prior choice.
- [References] Reference [5] is an internal CDF note from 2000; if a more accessible version or DOI exists, it should be cited; otherwise, the URL should be checked for stability.
Circularity Check
No significant circularity: the bias claims are Monte Carlo measurements with injected truth, and Eq. (2) is an empirical calibration curve rather than a target input.
full rationale
The paper's derivation chain is self-contained and driven by toy simulations with fully known injected signal rates, masses, and background levels. The fitted function of Eq. (2), y = (x^n + c^n)^(1/n), is a heuristic parameterization of the simulated mean observed significance as a function of the true expected significance; c and n are fitted to those simulated data, not to the claims being made. The inversion of y = 3 to obtain x = 2.66, and the resulting 13% rate bias, is an interpolation through this calibration curve, and the 960/fb versus 720/fb luminosity projection follows arithmetically from that same curve. Whether one should instead condition on the observed significance when estimating the true rate is a statistical-validity concern about the estimator, not a circularity: the paper does not assume the conclusion it is trying to measure. The width-uncertainty results come from separate toy fits with floating width, and the Section 4 mass-pull analysis explicitly constructs an ensemble conditioned on observed significance. The only prior work on the effect, the CDF note by Dorigo and Schmitt, is cited as independent context obtained late in the project and is not load-bearing for Eq. (2) or the toy results. The paper also explicitly states that the numerical importance depends on the shapes of signal and background distributions, which is a scope caveat rather than a circular move. In sum, the quantitative conclusions are Monte Carlo measurements with known injected truth; no load-bearing step reduces to its own input.
Assumptions & free parameters
free parameters (2)
- c =
2.04
- n =
3.06
assumptions (5)
- domain assumption The background is uniform with a known rate.
- domain assumption The signal shape is Gaussian with known width (for the rate bias part).
- domain assumption MIGRAD errors approximate p-values.
- domain assumption The toy model is representative of real LHC bump-hunt searches.
- ad hoc to paper Eq. 2 provides an adequate description of the bias.
Cite this review
Pith. "Pith review of Biased rate estimates in bump-hunt searches." pith.science (2026). https://pith.science/paper/7HT2VGMX
@misc{pith2026250601505,
author = {Pith},
title = {Pith review of: Biased rate estimates in bump-hunt searches},
year = {2026},
howpublished = {\url{https://pith.science/paper/7HT2VGMX}},
note = {Machine review of arXiv:2506.01505}
}
read the original abstract
The cleanest way to discover a new particle is generally the "bump-hunt" methodology: looking for a localised excess in a mass (or related) distribution. However, if the mass of the particle being discovered is not known the procedure of quoting the most significant excess seen is "greedy", and random fluctuations from the background can be merged with the signal. This means that an observed 3{\sigma} evidence probably has the rate over-estimated by of order 10% and the mass uncertainty underestimated by more like 20%. If the width of the particle being discovered is also unknown, or the experimental resolution uncertain, the effect grows larger. In the context of LHC, the data-doubling period is now measured in years. If evidence for a genuine signal is obtained at the 3{\sigma} level, the true signal rate probably corresponds to a lower significance. The additional data required to produce a 5{\sigma} discovery may be hundreds of inverse femtobarn more than might be naively expected.
Figures
Reference graph
Works this paper leans on
-
[1]
ATLAS collaboration. Observation of a new particle in the search for the Standard Model Higgs boson with the ATLAS detector at the LHC. Physics Letters B, 716(1):1–29, 2012
work page 2012
-
[2]
Observation of a new boson at a mass of 125 GeV with the CMS experiment at the LHC
CMS Collaboration. Observation of a new boson at a mass of 125 GeV with the CMS experiment at the LHC. Physics Letters B, 716(1):30–61, 2012
work page 2012
-
[3]
Procedure for the LHC Higgs boson search combination in Summer 2011
ATLAS and CMS Collaborations. Procedure for the LHC Higgs boson search combination in Summer 2011. CMS- NOTE-2011-005, ATL-PHYS-PUB-2011-011 https: // cds. cern. ch/ record/ 1379837? ln= en, 8 2011
work page 2011
-
[4]
Trial factors for the look elsewhere effect in high energy physics
Eilam Gross and Ofer Vitells. Trial factors for the look elsewhere effect in high energy physics. The European Physical Journal C, 70(1–2):525–530, October 2010
work page 2010
-
[5]
T Dorigo and M. Schmitt. On the significance of the dimuon mass bump and the greedy bump bias. CDF/DOC/TOP/CDFR/5239 http: // www. pd. infn. it/~ dorigo/ cdf5239_ Sign_ mm_ bump_ and_ GBB. ps, 2000
work page 2000
-
[6]
CMS Collaboration. Search for a standard model-like Higgs boson in the mass range between 70 and 110 GeV in the diphoton final state in proton-proton collisions at √s=8 and 13 TeV. Physics Letters B, 793:320–347, 2019
work page 2019
-
[7]
The infamous 95 GeV excess at LEP: two b or not two b? JHEP, (223), October 2024
Patrick Janot. The infamous 95 GeV excess at LEP: two b or not two b? JHEP, (223), October 2024
work page 2024
-
[8]
root-project/root: v6.18/02 https://doi.org/10.5281/zenodo.3895860, August 2019
Rene Brun et al. root-project/root: v6.18/02 https://doi.org/10.5281/zenodo.3895860, August 2019
Show all 11 references
-
[9]
F. James. MINUIT: Function Minimization and Error Analysis Reference Manual, 1998. CERN Program Library Long Writeups https://cds.cern.ch/record/2296388
1998
-
[10]
Measurement of the Higgs boson mass from the H → γγ and H → ZZ ∗ → 4ℓ channels in pp collisions at center-of-mass energies of 7 and 8 TeV with the ATLAS detector
ATLAS Collaboration. Measurement of the Higgs boson mass from the H → γγ and H → ZZ ∗ → 4ℓ channels in pp collisions at center-of-mass energies of 7 and 8 TeV with the ATLAS detector. Phys. Rev. D, 90:052004, Sep 2014
2014
-
[11]
Study of the Mass and Spin-Parity of the Higgs Boson Candidate via Its Decays to Z Boson Pairs
CMS Collaboration. Study of the Mass and Spin-Parity of the Higgs Boson Candidate via Its Decays to Z Boson Pairs. Phys. Rev. Lett., 110:081803, Feb 2013. 6
2013
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.