{"id":"b8584ba3-b0b6-409a-9d0b-d5cdc947f1d3","arxiv_id":"2506.01505","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In bump-hunt searches with unknown mass, greedy selection of the most significant excess biases the measured rate upward by roughly 10% at 3 sigma and makes mass errors too small by roughly 20%.","lead":"This paper uses simple Monte Carlo simulations to show that when particle physicists search for a new particle without knowing its mass, the standard practice of reporting the most significant bump overestimates the signal rate by about 10% and underestimates the mass uncertainty by about 20%. The authors estimate that reaching a 5 sigma discovery from a 3 sigma hint could require about a third more data than naive extrapolation suggests.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10–13% rate-bias claim rests on inverting a mean calibration curve; conditioning on the observed significance could shift the number materially.","rationale":"The paper's qualitative message—that scanning for an unknown mass selects upward background fluctuations and biases rate and mass estimates—is well supported by its own simulations and is physically sensible. The reader's conditional verdict is appropriate. However, the most load-bearing quantitative step is the inversion of the fitted mean curve. The paper never justifies that the inverse-mean mapping equals the conditional expectation given an observed significance, and the large scatter visible in Fig. 2 makes that equality implausible. Because the headline 13% bias and the §5 luminosity multiplier are direct outputs of this inversion, the quantitative claims are less secure than the qualitative ones. The suggested test would settle whether the inversion introduces a material error, and if so, by how much. This does not change the verdict of CONDITIONAL, since the requested conditioning check is a concrete addition to the reader's existing requirements.","tokens_in":7052,"tokens_out":5936,"duration_ms":68675,"concrete_test":"Run the §2 toy with a prior over the true signal significance (e.g., uniform x∈[0,5]σ, or the flat 0–8σ prior used in §4). For each toy, scan mass and record y (maximum observed significance). Select all toys with y∈[2.75,3.25] and compute the mean and median injected x, plus the 68% interval; compare with g^{-1}(3)=2.66 from Eq. (2). Repeat for y∈[4.75,5.25] to check the 5σ step. If the conditional mean differs from g^{-1}(y) by more than ~0.2σ, the 13% rate bias and the 960/fb luminosity projection need to be recomputed with the conditional calibration.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Abstract's headline rate-bias number is obtained by fitting Eq. (2), y=(x^n+c^n)^{1/n}, to the mean observed significance as a function of the true significance for fixed injected signals, and then inverting: a 3σ observation is mapped to x=2.66σ and a 13% rate bias. But the claim 'an observed 3σ evidence probably has the rate over-estimated' requires the conditional expectation E[x | y=3], not g^{-1}(3). These agree only if the observed y is a deterministic or zero-mean symmetric function of x. In this problem it is neither: Fig. 2 shows large scatter in y for fixed x, and even for x=0 the mean scanned maximum is c=2.04 with a broad upper tail. An observed 3σ excess can therefore frequently arise from a smaller true signal (or no signal) than the inverse-mean curve suggests, and the size of the bias quoted in the Abstract is not directly supported. The distinction is not pedantic: the §5 luminosity projection (960/fb vs 720/fb) is computed from the same inverted curve. Notably, §4 constructs a prior over true rates and conditions on observed significance for the mass-pull analysis, but §3 and §5 never apply that conditioning to the rate calibration.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the selection ('greedy') bias introduced when a bump-hunt search scans over an unknown mass and reports the most significant excess. Using an idealized toy model with a uniform background of known rate and a Gaussian signal of known width, the authors simulate many experiments with injected signals of various strengths, scan over a mass range, and record the maximum fitted signal significance. They find that the mean observed significance as a function of the true (single-mass) significance is well described by y=(x^n+c^n)^(1/n) with c=2.04 and n=3.06 for a scan range mimicking H→γγ. Inverting this curve, they claim an observed 3σ excess corresponds to a true significance of about 2.66σ, i.e., a rate overestimate of ~13%, and that reaching 5σ requires 2.4 times the original dataset (e.g., 960 fb^{-1} vs 720 fb^{-1}). They also quantify the mass-measurement pull in ensembles conditioned on the observed significance, finding that the Gaussian core of the mass pull has width ~1.19 at 3σ and non-Gaussian tails, supporting the claim that mass uncertainties are underestimated by ~20%. The main quantitative rate-bias claims are, however, based on inverting the mean calibration curve rather than computing the conditional distribution of the true rate given an observed significance; this is the major issue.","tokens_in":7300,"tokens_out":13474,"duration_ms":138026,"significance":"The paper addresses a real and often overlooked selection effect: quoting the most significant mass hypothesis in a search with unknown particle mass positively biases the measured signal rate and underestimates mass errors. The toy model is transparent and the fit validation is meticulous (pull distribution with RMS 0.9998, non-closure of only 0.4%). The mass-pull analysis in Section 4 is methodologically sound because it conditions on the observed significance using an explicit prior on the true signal rate. If the rate-bias calibration is corrected to the same standard, the paper will provide a useful quantitative caution for the interpretation of 3σ hints and for luminosity planning. The qualitative conclusion is robust; the quantitative LHC extrapolation is currently not directly supported and should be revised.","major_comments":[{"comment":"The paper fits the heuristic y=(x^n+c^n)^(1/n) to the mean observed significance y as a function of the true significance x, then inverts this relation to state that an observed 3σ excess corresponds to a true significance of 2.66σ and a rate overestimate of 13%. The abstract and Section 5 make claims about an observed 3σ excess, but the quantity implied by those claims is the conditional expectation E[x|y=3] (or a posterior interval), not g^{-1}(3) with g(x)=E[y|x]. Because Fig. 2 shows substantial scatter of y for fixed x and g is nonlinear, these two quantities can differ materially; for the uniform prior on x used in Section 4, the concavity of g^{-1} suggests E[x|y=3] is smaller than 2.66σ, so the rate bias may be larger than 13%. The paper does not compute this conditional quantity anywhere in Sections 3 or 5, despite doing exactly this type of conditioning in Section 4 for the mass-pull distribution. Please re-derive the rate-bias claim by generating an ensemble with a prior on true rates, conditioning on observed significance, and reporting E[x|y] (or the median and a credible interval).","section":"Section 3, Eq. (2), Fig. 3(left)"},{"comment":"The statement that a 3σ evidence at 400 fb^{-1} requires 960 fb^{-1} rather than 720 fb^{-1} relies on the same inverted curve. There are two distinct quantities being conflated: the true significance x such that the *median* observed significance is 3 (i.e., g(x)=3), and the conditional expected true significance given an observed 3σ. The forward mapping g(x)=5 is the right target for planning a median 5σ reach, but the starting point \"an observed 3σ evidence\" is not the same as the x that yields median 3σ. After correcting the calibration as suggested in the previous comment, the luminosity projection should be re-derived; it may increase substantially (e.g., if E[x|y=3]≈2.4σ, the required total luminosity would be about (5/2.4)^2≈4.3 times the original rather than 3.4 times). The current 960/fb number is therefore not directly supported by the analysis.","section":"Section 5, luminosity projection"}],"minor_comments":[{"comment":"The phrase \"this ensures that so that there is negligible signal density\" contains a duplicated \"that so that\" and should be corrected.","section":"Section 2, paragraph 2"},{"comment":"The word \"whuch\" is a typo and should read \"which\".","section":"Section 2, paragraph 3"},{"comment":"The notation \"n√\" is malformed; it should be displayed as an n-th root, e.g., \\sqrt[n]{x^n+c^n}.","section":"Eq. (2)"},{"comment":"The sentence \"The other data show the impact of scanning a range for the highest excess\" is ambiguous; it should read \"The other curves show\".","section":"Section 3, near Fig. 2"},{"comment":"The phrase \"has the advantage that that they are independent\" contains a repeated \"that\" and should be corrected.","section":"Section 5, paragraph 1"},{"comment":"The phrase \"the data-doubling period is now measured in years\" is colloquial; consider \"the time to double the dataset is now measured in years\" for clarity.","section":"Abstract"},{"comment":"The prior on true signal rates (uniform between 0 and 8σ) is a strong assumption; the authors test robustness to the prior upper limit and to conditioning on expected significance, but the manuscript should state explicitly that the quoted mass-pull widths are conditional on that prior choice.","section":"Section 4"},{"comment":"Reference [5] is an internal CDF note from 2000; if a more accessible version or DOI exists, it should be cited; otherwise, the URL should be checked for stability.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is relevant to the journal's readership and the qualitative message is important. The main statistical flaw (inverse regression) is fixable within the scope of the paper, so I recommend major revision rather than rejection. The authors should also ensure that the CDF internal note is adequately credited, although they do cite it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the greedy bump-hunt bias is real, and this paper is a decent quantitative look at it. But the headline \"observed 3σ means rate overestimated by ~10%\" is not actually backed by the simulations as stated. Good news first. The paper ships a clean toy MC, checks pulls, quotes a 0.4% non-closure in the signal extraction, and is unusually candid about the prior CDF note and the idealized setup. The mass-pull section is the strongest part: the conditional ensemble in Section 4 directly addresses the \"given what I saw\" question, and the 10–20% core-width inflation plus tails is a concrete, useful result. The width-uncertainty extension is also sensible and points to a larger effect.\n\nThe soft spot is exactly the one flagged in the stress-test. The 13% rate bias comes from fitting y = (x^n + c^n)^(1/n) to the mean observed significance as a function of the true value, then inverting: y=3 maps to x≈2.7. But the claim in the abstract is about an experiment that has observed 3σ, i.e. E[x|y=3], and that is not g^{-1}(3) unless y is a deterministic or symmetric zero-mean function of x. Fig. 2 shows it is neither – at x=0 the mean scanned maximum is 2.04 with a large upper tail. The same inversion feeds the 960/fb vs 720/fb projection. Since Section 4 already does the conditioning properly for masses, applying the same logic to rates would tighten the paper substantially. Without that, I'd quote the qualitative message but not the 13% or 2.4x numbers.\n\nThe other caveats are proportionate: uniform known background, known width, one binning; the authors say the magnitude depends on details, which is true. The non-closure, left unexplained, is small and probably not load-bearing.\n\nVerdict: worth a serious referee, but the referee should push for a conditional rate-bias estimate and uncertainty bands on c and n. I wouldn't desk-reject it.","headline":"Real effect, honest toy study, but the headline 13% rate bias is a mean-calibration inversion and needs a conditional treatment before I'd trust the number.","tokens_in":7819,"tokens_out":2926,"would_cite":false,"duration_ms":34135,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In a bump-hunt search with unknown mass, reporting the most significant excess overestimates the signal rate by roughly 10% and underestimates the mass uncertainty by 10–20%, so upgrading a 3σ hint to 5σ requires about 960 inverse…","keywords":["bump-hunt","greedy bias","look-elsewhere effect","mass scan","signal rate bias","mass uncertainty","luminosity scaling","resonance search"],"falsifier":"Repeat the same greedy mass-scan simulation with a non-uniform background (for example, an exponential fall) and with the background normalization and slope fitted rather than fixed; if the observed-to-true significance ratio for a genuinely 3σ signal stays within a few percent of unity, the claimed 13% rate bias and 20% mass-error inflation would not generalize.","tokens_in":6841,"feed_emoji":"🔬","tokens_out":12402,"duration_ms":122682,"temperature":0.7,"pith_summary":"The paper establishes that when a resonance search scans over unknown mass and reports the most significant excess, random background fluctuations near the true signal are absorbed into the fitted peak. In a toy model with a uniform, known background and a Gaussian signal of known width, an observed 3σ excess corresponds to a true signal of only about 2.66σ, meaning the rate is overestimated by roughly 13% and the mass uncertainty is underestimated by 10–20%. The paper derives a heuristic mapping between true and observed significance and uses it to show that turning a 3σ hint into a 5σ discovery requires about 960 inverse femtobarns of additional data after an initial 400 inverse femtobarns, not the 720 that a naive scaling suggests. If the signal width or detector resolution is also unknown, the rate bias grows and can reach 10% even at 5σ. The numerical magnitudes depend on the shapes of the signal and background, but the mechanism is universal.","feed_headline":"Bump-hunt searches overstate new-particle rates by ~10%","feed_subtitle":"A 3σ hint is really ~2.7σ, so a 5σ discovery needs ~960 inverse femtobarns, not 720, after the first 400.","key_machinery":"The load-bearing object is the heuristic mapping from true to observed significance, $y=\\left(x^n+c^n\\right)^{1/n}$ with $c=2.04$ and $n=3.06$, calibrated from toy fits over a scan range of $\\pm8\\sigma$ that mimics diphoton Higgs searches. The toy model is a binned negative-log-likelihood fit of a Gaussian signal on a uniform known background (1000 events per bin, signal width $\\sigma=0.5$), with the mass and rate left free and the most significant point across the scanned mass grid reported. The mapping converts an observed 3σ into the underlying 2.66σ used for luminosity extrapolation, and it is the reason the required additional luminosity becomes 960 inverse femtobarns rather than 720. For the mass bias, the machinery is the pull distribution of the fitted mass from the greedy scan, compared against the minimizer's error estimate and the $-2\\Delta\\ln L=1$ error convention.","core_discovery":"The core discovery is that the greedy choice of the most significant excess in an unknown-mass scan is a biased estimator: background fluctuations near the true signal are merged with the fitted peak, inflating the measured rate by about 13% when the reported significance is 3σ and degrading the mass error by 10–20%, with non-Gaussian tails beyond 5σ. In the toy model, the relation between the significance expected at the true mass, x, and the mean observed maximum significance, y, is well described by y = (x^n + c^n)^(1/n) with c = 2.04 and n = 3.06, so an observed 3σ corresponds to an underlying 2.66σ and an observed 5σ to an underlying 4.9σ. Inverting this relation, an observed 3σ hint requires roughly 2.4 times the original dataset as additional data (960 inverse femtobarns rather than 720 after an initial 400) to reach a 5σ discovery, and the mass uncertainty remains underestimated until the significance is at least 5σ.","pith_inferences":["The same greedy-selection mechanism should apply to any search that optimizes over nuisance parameters—unknown width, unknown line shape, or unknown location in a multidimensional space—so analogous biases should appear in axion-line, dark-matter, and gravitational-wave searches, with magnitudes growing with the number of scanned hypotheses.","A practical correction is possible: use the inverse of the heuristic mapping to deflate an observed significance before setting a luminosity target, effectively a trials correction for the signal strength rather than the p-value.","For real evidence-level excesses, such as the ~95 GeV diphoton excess, the measured signal rate should drift downward as more data accumulate and the significance should grow more slowly than the square root of luminosity; tracking those two quantities over time would provide a direct empirical test of the bias.","Combining multiple channels or datasets at evidence level should inflate each channel's rate and mass error by the channel-specific bias before averaging, otherwise the combined result will inherit a low mass error and a high rate."],"forward_implications":["An observed 3σ bump-hunt excess corresponds to an underlying signal of only about 2.66σ at the true mass, so fitted signal rates from evidence-level peaks are biased high by roughly 13%.","The mass of a particle claimed at 3–4σ carries an underestimated uncertainty: the core error is too small by 10–20%, and there are tails with pulls beyond 5, so early mass measurements should be quoted with enlarged errors.","If the signal width or detector resolution is not known perfectly, the rate bias grows and can reach 10% even at 5σ, so fitting the width freely makes the overestimate worse.","Turning a 3σ hint into a 5σ discovery requires about 960 inverse femtobarns of additional data after an initial 400 inverse femtobarns, not the 720 a naive scaling suggests.","The bias is unavoidable whenever the mass is unknown and the reported significance is the maximum over a scan; it is absent only when the signal mass is fixed in advance."],"supporting_citations":[{"why":"The Higgs discovery paper that provides the real bump-hunt example and the H to gamma gamma scan range used to calibrate the toy model.","marker":"[1]"},{"why":"The companion Higgs discovery paper that establishes the same discovery methodology and scan range.","marker":"[2]"},{"why":"Quantifies the look-elsewhere effect and trials factor for the background-only case, which the paper contrasts with the signal-present greedy bias.","marker":"[4]"},{"why":"An earlier note that described the greedy bump bias for a dimuon mass bump, supporting the claim that the effect is known.","marker":"[5]"},{"why":"Reports the 2.9 sigma diphoton excess near 95 GeV that motivates the discussion of how significance should evolve with more data.","marker":"[6]"},{"why":"Discusses the significance scaling of the 95 GeV excess, a specific analysis the paper says overstates the case by ignoring the greedy bias.","marker":"[7]"},{"why":"The fitting software used to generate and fit the toy mass spectra in all the simulations.","marker":"[8]"},{"why":"Supplies the error estimation method (including the -2 Delta ln L = 1 convention) whose coverage is tested in the mass-pull studies.","marker":"[9]"}],"fun_headline_variants":["Bump-hunt overestimates signal rates by ~13%","Greedy bump-hunt bias inflates rates, shrinks mass errors","Unknown-mass scans overstate signal rates and mass precision","Bump-hunt scans bias rate estimates up to 13%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative 10–13% rate bias, 20% mass-error inflation, and 2.4-times luminosity factor come from a toy model with a flat known background, a Gaussian signal of known width, and no systematic errors, so they carry over to real experiments only if those idealizations are representative.","fun_headline_variants_meta":{"raw":{"variants":["Bump-hunt overestimates signal rates by ~13%","Greedy bump-hunt bias inflates rates, shrinks mass errors","Unknown-mass scans overstate signal rates and mass precision","Bump-hunt scans bias rate estimates up to 13%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001077,"raw_usage":{"total_tokens":4514,"prompt_tokens":961,"completion_tokens":3553,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":3479}},"tokens_in":577,"tokens_out":3553,"duration_ms":28296,"temperature":1.0,"reasoning_tokens":3479,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:39:32.073518+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the same greedy mass-scan simulation with a non-uniform background (for example, an exponential fall) and with the background normalization and slope fitted rather than fixed; if the observed-to-true significance ratio for a genuinely 3σ signal stays within a few percent of unity, the claimed 13% rate bias and 20% mass-error inflation would not generalize.","supporting_citations":[{"cited_title":"Observation of a new particle in the search for the Standard Model Higgs boson with the ATLAS detector at the LHC","cited_arxiv_id":null,"evidence_quote":"The Higgs discovery paper that provides the real bump-hunt example and the H to gamma gamma scan range used to calibrate the toy model."},{"cited_title":"Observation of a new boson at a mass of 125 GeV with the CMS experiment at the LHC","cited_arxiv_id":null,"evidence_quote":"The companion Higgs discovery paper that establishes the same discovery methodology and scan range."},{"cited_title":"Trial factors for the look elsewhere effect in high energy physics","cited_arxiv_id":null,"evidence_quote":"Quantifies the look-elsewhere effect and trials factor for the background-only case, which the paper contrasts with the signal-present greedy bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"An earlier note that described the greedy bump bias for a dimuon mass bump, supporting the claim that the effect is known."},{"cited_title":"Search for a standard model-like Higgs boson in the mass range between 70 and 110 GeV in the diphoton final state in proton-proton collisions at √s=8 and 13 TeV","cited_arxiv_id":null,"evidence_quote":"Reports the 2.9 sigma diphoton excess near 95 GeV that motivates the discussion of how significance should evolve with more data."},{"cited_title":"The infamous 95 GeV excess at LEP: two b or not two b? JHEP, (223), October 2024","cited_arxiv_id":null,"evidence_quote":"Discusses the significance scaling of the 95 GeV excess, a specific analysis the paper says overstates the case by ignoring the greedy bias."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the error estimation method (including the -2 Delta ln L = 1 convention) whose coverage is tested in the mass-pull studies."}],"review_version":1}