Pith. sign in

REVIEW 2 major objections 8 minor 11 references

Biased rate estimates in bump-hunt searches

T0 review · 2 major / 8 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read In a bump-hunt search with unknown mass, reporting the most significant excess overestimates the signal rate by roughly 10% and underestimates the mass uncertainty by 10–20%, so upgrading a 3σ hint to 5σ requires about 960 inverse…

desk verdict Real effect, honest toy study, but the headline 13% rate bias is a mean-calibration inversion and needs a conditional treatment before I'd trust the number. read the letter →

arxiv 2506.01505 v1 pith:7HT2VGMX submitted 2025-06-02 hep-ph

classification hep-ph
keywords bump-huntgreedybiaslook-elsewhereeffectmassscansignalrateuncertaintyluminosityscalingresonancesearch
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that when a resonance search scans over unknown mass and reports the most significant excess, random background fluctuations near the true signal are absorbed into the fitted peak. In a toy model with a uniform, known background and a Gaussian signal of known width, an observed 3σ excess corresponds to a true signal of only about 2.66σ, meaning the rate is overestimated by roughly 13% and the mass uncertainty is underestimated by 10–20%. The paper derives a heuristic mapping between true and observed significance and uses it to show that turning a 3σ hint into a 5σ discovery requires about 960 inverse femtobarns of additional data after an initial 400 inverse femtobarns, not the 720 that a naive scaling suggests. If the signal width or detector resolution is also unknown, the rate bias grows and can reach 10% even at 5σ. The numerical magnitudes depend on the shapes of the signal and background, but the mechanism is universal.

What carries the argument

The load-bearing object is the heuristic mapping from true to observed significance, $y=\left(x^n+c^n\right)^{1/n}$ with $c=2.04$ and $n=3.06$, calibrated from toy fits over a scan range of $\pm8\sigma$ that mimics diphoton Higgs searches. The toy model is a binned negative-log-likelihood fit of a Gaussian signal on a uniform known background (1000 events per bin, signal width $\sigma=0.5$), with the mass and rate left free and the most significant point across the scanned mass grid reported. The mapping converts an observed 3σ into the underlying 2.66σ used for luminosity extrapolation, and it is the reason the required additional luminosity becomes 960 inverse femtobarns rather than 720. For the mass bias, the machinery is the pull distribution of the fitted mass from the greedy scan, compared against the minimizer's error estimate and the $-2\Delta\ln L=1$ error convention.

What would settle it

Repeat the same greedy mass-scan simulation with a non-uniform background (for example, an exponential fall) and with the background normalization and slope fitted rather than fixed; if the observed-to-true significance ratio for a genuinely 3σ signal stays within a few percent of unity, the claimed 13% rate bias and 20% mass-error inflation would not generalize.

Watch

Extended reading notes

Core claim

The core discovery is that the greedy choice of the most significant excess in an unknown-mass scan is a biased estimator: background fluctuations near the true signal are merged with the fitted peak, inflating the measured rate by about 13% when the reported significance is 3σ and degrading the mass error by 10–20%, with non-Gaussian tails beyond 5σ. In the toy model, the relation between the significance expected at the true mass, x, and the mean observed maximum significance, y, is well described by y = (x^n + c^n)^(1/n) with c = 2.04 and n = 3.06, so an observed 3σ corresponds to an underlying 2.66σ and an observed 5σ to an underlying 4.9σ. Inverting this relation, an observed 3σ hint requires roughly 2.4 times the original dataset as additional data (960 inverse femtobarns rather than 720 after an initial 400) to reach a 5σ discovery, and the mass uncertainty remains underestimated until the significance is at least 5σ.

Load-bearing premise

The quantitative 10–13% rate bias, 20% mass-error inflation, and 2.4-times luminosity factor come from a toy model with a flat known background, a Gaussian signal of known width, and no systematic errors, so they carry over to real experiments only if those idealizations are representative.

Editorial extensions

If this is right

  • An observed 3σ bump-hunt excess corresponds to an underlying signal of only about 2.66σ at the true mass, so fitted signal rates from evidence-level peaks are biased high by roughly 13%.
  • The mass of a particle claimed at 3–4σ carries an underestimated uncertainty: the core error is too small by 10–20%, and there are tails with pulls beyond 5, so early mass measurements should be quoted with enlarged errors.
  • If the signal width or detector resolution is not known perfectly, the rate bias grows and can reach 10% even at 5σ, so fitting the width freely makes the overestimate worse.
  • Turning a 3σ hint into a 5σ discovery requires about 960 inverse femtobarns of additional data after an initial 400 inverse femtobarns, not the 720 a naive scaling suggests.
  • The bias is unavoidable whenever the mass is unknown and the reported significance is the maximum over a scan; it is absent only when the signal mass is fixed in advance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same greedy-selection mechanism should apply to any search that optimizes over nuisance parameters—unknown width, unknown line shape, or unknown location in a multidimensional space—so analogous biases should appear in axion-line, dark-matter, and gravitational-wave searches, with magnitudes growing with the number of scanned hypotheses.
  • A practical correction is possible: use the inverse of the heuristic mapping to deflate an observed significance before setting a luminosity target, effectively a trials correction for the signal strength rather than the p-value.
  • For real evidence-level excesses, such as the ~95 GeV diphoton excess, the measured signal rate should drift downward as more data accumulate and the significance should grow more slowly than the square root of luminosity; tracking those two quantities over time would provide a direct empirical test of the bias.
  • Combining multiple channels or datasets at evidence level should inflate each channel's rate and mass error by the channel-specific bias before averaging, otherwise the combined result will inherit a low mass error and a high rate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 8 minor

Summary. This paper studies the selection ('greedy') bias introduced when a bump-hunt search scans over an unknown mass and reports the most significant excess. Using an idealized toy model with a uniform background of known rate and a Gaussian signal of known width, the authors simulate many experiments with injected signals of various strengths, scan over a mass range, and record the maximum fitted signal significance. They find that the mean observed significance as a function of the true (single-mass) significance is well described by y=(x^n+c^n)^(1/n) with c=2.04 and n=3.06 for a scan range mimicking H→γγ. Inverting this curve, they claim an observed 3σ excess corresponds to a true significance of about 2.66σ, i.e., a rate overestimate of ~13%, and that reaching 5σ requires 2.4 times the original dataset (e.g., 960 fb^{-1} vs 720 fb^{-1}). They also quantify the mass-measurement pull in ensembles conditioned on the observed significance, finding that the Gaussian core of the mass pull has width ~1.19 at 3σ and non-Gaussian tails, supporting the claim that mass uncertainties are underestimated by ~20%. The main quantitative rate-bias claims are, however, based on inverting the mean calibration curve rather than computing the conditional distribution of the true rate given an observed significance; this is the major issue.

Significance. The paper addresses a real and often overlooked selection effect: quoting the most significant mass hypothesis in a search with unknown particle mass positively biases the measured signal rate and underestimates mass errors. The toy model is transparent and the fit validation is meticulous (pull distribution with RMS 0.9998, non-closure of only 0.4%). The mass-pull analysis in Section 4 is methodologically sound because it conditions on the observed significance using an explicit prior on the true signal rate. If the rate-bias calibration is corrected to the same standard, the paper will provide a useful quantitative caution for the interpretation of 3σ hints and for luminosity planning. The qualitative conclusion is robust; the quantitative LHC extrapolation is currently not directly supported and should be revised.

major comments (2)
  1. [Section 3, Eq. (2), Fig. 3(left)] The paper fits the heuristic y=(x^n+c^n)^(1/n) to the mean observed significance y as a function of the true significance x, then inverts this relation to state that an observed 3σ excess corresponds to a true significance of 2.66σ and a rate overestimate of 13%. The abstract and Section 5 make claims about an observed 3σ excess, but the quantity implied by those claims is the conditional expectation E[x|y=3] (or a posterior interval), not g^{-1}(3) with g(x)=E[y|x]. Because Fig. 2 shows substantial scatter of y for fixed x and g is nonlinear, these two quantities can differ materially; for the uniform prior on x used in Section 4, the concavity of g^{-1} suggests E[x|y=3] is smaller than 2.66σ, so the rate bias may be larger than 13%. The paper does not compute this conditional quantity anywhere in Sections 3 or 5, despite doing exactly this type of conditioning in Section 4 for the mass-pull distribution. Please re-derive the rate-bias claim by generating an ensemble with a prior on true rates, conditioning on observed significance, and reporting E[x|y] (or the median and a credible interval).
  2. [Section 5, luminosity projection] The statement that a 3σ evidence at 400 fb^{-1} requires 960 fb^{-1} rather than 720 fb^{-1} relies on the same inverted curve. There are two distinct quantities being conflated: the true significance x such that the *median* observed significance is 3 (i.e., g(x)=3), and the conditional expected true significance given an observed 3σ. The forward mapping g(x)=5 is the right target for planning a median 5σ reach, but the starting point "an observed 3σ evidence" is not the same as the x that yields median 3σ. After correcting the calibration as suggested in the previous comment, the luminosity projection should be re-derived; it may increase substantially (e.g., if E[x|y=3]≈2.4σ, the required total luminosity would be about (5/2.4)^2≈4.3 times the original rather than 3.4 times). The current 960/fb number is therefore not directly supported by the analysis.
minor comments (8)
  1. [Section 2, paragraph 2] The phrase "this ensures that so that there is negligible signal density" contains a duplicated "that so that" and should be corrected.
  2. [Section 2, paragraph 3] The word "whuch" is a typo and should read "which".
  3. [Eq. (2)] The notation "n√" is malformed; it should be displayed as an n-th root, e.g., \sqrt[n]{x^n+c^n}.
  4. [Section 3, near Fig. 2] The sentence "The other data show the impact of scanning a range for the highest excess" is ambiguous; it should read "The other curves show".
  5. [Section 5, paragraph 1] The phrase "has the advantage that that they are independent" contains a repeated "that" and should be corrected.
  6. [Abstract] The phrase "the data-doubling period is now measured in years" is colloquial; consider "the time to double the dataset is now measured in years" for clarity.
  7. [Section 4] The prior on true signal rates (uniform between 0 and 8σ) is a strong assumption; the authors test robustness to the prior upper limit and to conditioning on expected significance, but the manuscript should state explicitly that the quoted mass-pull widths are conditional on that prior choice.
  8. [References] Reference [5] is an internal CDF note from 2000; if a more accessible version or DOI exists, it should be cited; otherwise, the URL should be checked for stability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the bias claims are Monte Carlo measurements with injected truth, and Eq. (2) is an empirical calibration curve rather than a target input.

full rationale

The paper's derivation chain is self-contained and driven by toy simulations with fully known injected signal rates, masses, and background levels. The fitted function of Eq. (2), y = (x^n + c^n)^(1/n), is a heuristic parameterization of the simulated mean observed significance as a function of the true expected significance; c and n are fitted to those simulated data, not to the claims being made. The inversion of y = 3 to obtain x = 2.66, and the resulting 13% rate bias, is an interpolation through this calibration curve, and the 960/fb versus 720/fb luminosity projection follows arithmetically from that same curve. Whether one should instead condition on the observed significance when estimating the true rate is a statistical-validity concern about the estimator, not a circularity: the paper does not assume the conclusion it is trying to measure. The width-uncertainty results come from separate toy fits with floating width, and the Section 4 mass-pull analysis explicitly constructs an ensemble conditioned on observed significance. The only prior work on the effect, the CDF note by Dorigo and Schmitt, is cited as independent context obtained late in the project and is not load-bearing for Eq. (2) or the toy results. The paper also explicitly states that the numerical importance depends on the shapes of signal and background distributions, which is a scope caveat rather than a circular move. In sum, the quantitative conclusions are Monte Carlo measurements with known injected truth; no load-bearing step reduces to its own input.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The central quantitative claims rest on two fitted parameters (c and n) in a heuristic function, plus several domain assumptions about the idealization of the experiment. No new physical entities are introduced.

free parameters (2)
  • c = 2.04
    Heuristic parameter in Eq. 2 fitted to toy simulation data; represents the mean observed significance under background-only for the scan range used. It depends on the scanned range.
  • n = 3.06
    Heuristic exponent in Eq. 2 fitted to toy simulation data; controls the transition between noise-dominated and signal-dominated regimes.
assumptions (5)
  • domain assumption The background is uniform with a known rate.
    Stated in Section 2: 'The background is taken to be uniform in mass within the considered range' and the background rate is assumed known. This simplifies the likelihood and is not representative of all real searches.
  • domain assumption The signal shape is Gaussian with known width (for the rate bias part).
    Section 2: the signal is a narrow resonance, observed width is Gaussian and known to the experimenter. Width uncertainty is studied separately in Section 3.
  • domain assumption MIGRAD errors approximate p-values.
    Section 2: they validate with a pull distribution from 10M toys, but this is an approximation that could fail with non-Gaussian likelihoods, especially when the width is fitted.
  • domain assumption The toy model is representative of real LHC bump-hunt searches.
    The paper applies the toy model results to LHC luminosity planning in Section 5. The Discussion concedes that 'numerical importance will depend in detail on the shape of signal and background distributions', but the abstract and conclusions present the numbers as broadly applicable.
  • ad hoc to paper Eq. 2 provides an adequate description of the bias.
    Section 3: the functional form y = (x^n + c^n)^(1/n) is chosen heuristically and fitted to simulation data. It is not derived from first principles and its adequacy is only stated, not quantified with residuals.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Biased rate estimates in bump-hunt searches." pith.science (2026). https://pith.science/paper/7HT2VGMX

@misc{pith2026250601505,
  author       = {Pith},
  title        = {Pith review of: Biased rate estimates in bump-hunt searches},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7HT2VGMX}},
  note         = {Machine review of arXiv:2506.01505}
}
read the original abstract

The cleanest way to discover a new particle is generally the "bump-hunt" methodology: looking for a localised excess in a mass (or related) distribution. However, if the mass of the particle being discovered is not known the procedure of quoting the most significant excess seen is "greedy", and random fluctuations from the background can be merged with the signal. This means that an observed 3{\sigma} evidence probably has the rate over-estimated by of order 10% and the mass uncertainty underestimated by more like 20%. If the width of the particle being discovered is also unknown, or the experimental resolution uncertain, the effect grows larger. In the context of LHC, the data-doubling period is now measured in years. If evidence for a genuine signal is obtained at the 3{\sigma} level, the true signal rate probably corresponds to a lower significance. The additional data required to produce a 5{\sigma} discovery may be hundreds of inverse femtobarn more than might be naively expected.

Figures

Figures reproduced from arXiv: 2506.01505 by the authors.

Figure 1
Figure 1. Left: an example toy experiment with a large signal injected as mass zero. Right: the pull distribution of the [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Left: The mean signal size as a function of the expected with various different scanning ranges. Right: The [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Left: The mean significance found as a function of that expected when the signal is tested only at the mass at [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Left: the pull in fitted mass for experiments with an observed significance 3 [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

11 extracted references · 10 canonical work pages

  1. [1]

    Observation of a new particle in the search for the Standard Model Higgs boson with the ATLAS detector at the LHC

    ATLAS collaboration. Observation of a new particle in the search for the Standard Model Higgs boson with the ATLAS detector at the LHC. Physics Letters B, 716(1):1–29, 2012

  2. [2]

    Observation of a new boson at a mass of 125 GeV with the CMS experiment at the LHC

    CMS Collaboration. Observation of a new boson at a mass of 125 GeV with the CMS experiment at the LHC. Physics Letters B, 716(1):30–61, 2012

  3. [3]

    Procedure for the LHC Higgs boson search combination in Summer 2011

    ATLAS and CMS Collaborations. Procedure for the LHC Higgs boson search combination in Summer 2011. CMS- NOTE-2011-005, ATL-PHYS-PUB-2011-011 https: // cds. cern. ch/ record/ 1379837? ln= en, 8 2011

  4. [4]

    Trial factors for the look elsewhere effect in high energy physics

    Eilam Gross and Ofer Vitells. Trial factors for the look elsewhere effect in high energy physics. The European Physical Journal C, 70(1–2):525–530, October 2010

  5. [5]

    T Dorigo and M. Schmitt. On the significance of the dimuon mass bump and the greedy bump bias. CDF/DOC/TOP/CDFR/5239 http: // www. pd. infn. it/~ dorigo/ cdf5239_ Sign_ mm_ bump_ and_ GBB. ps, 2000

  6. [6]

    Search for a standard model-like Higgs boson in the mass range between 70 and 110 GeV in the diphoton final state in proton-proton collisions at √s=8 and 13 TeV

    CMS Collaboration. Search for a standard model-like Higgs boson in the mass range between 70 and 110 GeV in the diphoton final state in proton-proton collisions at √s=8 and 13 TeV. Physics Letters B, 793:320–347, 2019

  7. [7]

    The infamous 95 GeV excess at LEP: two b or not two b? JHEP, (223), October 2024

    Patrick Janot. The infamous 95 GeV excess at LEP: two b or not two b? JHEP, (223), October 2024

  8. [8]

    root-project/root: v6.18/02 https://doi.org/10.5281/zenodo.3895860, August 2019

    Rene Brun et al. root-project/root: v6.18/02 https://doi.org/10.5281/zenodo.3895860, August 2019

Show all 11 references
  1. [9]

    F. James. MINUIT: Function Minimization and Error Analysis Reference Manual, 1998. CERN Program Library Long Writeups https://cds.cern.ch/record/2296388

  2. [10]

    Measurement of the Higgs boson mass from the H → γγ and H → ZZ ∗ → 4ℓ channels in pp collisions at center-of-mass energies of 7 and 8 TeV with the ATLAS detector

    ATLAS Collaboration. Measurement of the Higgs boson mass from the H → γγ and H → ZZ ∗ → 4ℓ channels in pp collisions at center-of-mass energies of 7 and 8 TeV with the ATLAS detector. Phys. Rev. D, 90:052004, Sep 2014

  3. [11]

    Study of the Mass and Spin-Parity of the Higgs Boson Candidate via Its Decays to Z Boson Pairs

    CMS Collaboration. Study of the Mass and Spin-Parity of the Higgs Boson Candidate via Its Decays to Z Boson Pairs. Phys. Rev. Lett., 110:081803, Feb 2013. 6

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.