{"id":"822e450c-dc56-42d2-a26f-c157b042559f","arxiv_id":"2607.29436","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Refined asymptotic formulas for median discovery significance in Poisson counting experiments with background uncertainty, validated with Monte Carlo and improved at low counts via r* corrections.","lead":"This paper derives closed-form approximations for the expected discovery significance of a Poisson counting experiment with an uncertain background constrained by a control measurement, and adds higher-order corrections that improve accuracy at small event counts. It matters because experimental searches for new particles need reliable sensitivity estimates to plan measurements and optimize event selections.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uncertain-background validation relies on an approximate MC reference; exact conditional p-values are needed before the claimed low-yield improvements can be assessed.","rationale":"I read the paper as a careful extension of known asymptotic methods: the derivation of Eq. (20) is coherent, the code is released, and the known-background comparisons are supported by exact Poisson calculations. The primary load-bearing weakness is not the derivation itself but the benchmark used to validate the uncertain-background case. The reader identified the Asimov approximation as the weakest assumption and noted that the MC reference is an approximate profile-construction method. I agree that the Asimov approximation is unquantified, but I think the more specific and decisive concern is the approximate nature of the reference: without an exact or demonstrably unbiased reference for the on/off problem, agreement with the MC points does not validate the claimed improvements. The proposed conditional binomial test provides a concrete, nuisance-free exact reference for the on/off likelihood ratio statistic. This concern does not overturn the paper's positive contributions—the known-background higher-order corrections are credible, and Eq. (20) is a useful reference formula—but it does mean the central claim for the uncertain-background case should be regarded as conditional on further validation. Since the reader's verdict is already CONDITIONAL, my assessment does not change it.","tokens_in":938,"tokens_out":1359,"duration_ms":82371,"concrete_test":"For representative low-yield points (e.g., s=2,b=1,tau=1; s=5,b=1,tau=0.5; s=2,b=1,sigma_b/b=0.25), replace the profile-construction MC reference with an exact conditional reference. Under s=0, n | N=n+m ~ Binomial(N, 1/(1+tau)), a distribution free of b, so for each observed pair (n,m) the exact conditional p-value of q0 is sum over n'=0..N of Binom(N, p)(n') * I[q0(n',N-n') >= q0_obs], with p=1/(1+tau). Generate high-statistics pseudo-data under (s,b), compute these exact p-values, convert to Z, and take the median. Compare this exact median with Eq. (20) and with the q0* Asimov prediction. If the deviations from the exact reference are substantially larger than those against the profile-construction MC, the central low-yield validation claim is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that Eq. (20) and q0* give accurate median discovery significance with background uncertainty, improving on s/sqrt(b+sigma_b^2) at low yields, is validated for the on/off problem only against a Monte Carlo reference built from profile construction/hybrid resampling (Sec. 5, Figs. 2 and 6). For each simulated pair (n,m), the background-only p-value is estimated by generating s=0 pseudo-data with b fixed to the profiled value b-hat0 = (n+m)/(1+tau) (Eq. 16). This plug-in null distribution does not propagate the uncertainty in the nuisance parameter into the reference p-value; it is an approximate method with imperfect coverage, not an exact frequentist calculation. Therefore, agreement between Eq. (20)/q0* and these MC points does not by itself establish that the formulae are accurate, and the claimed 'meaningful improvements at small event yields' in the uncertain-background case could be an artifact of matching a biased reference. The known-background case (Figs. 1, 3, 5) has an exact Poisson reference and is not subject to this criticism, but the central on/off results—the new Eq. (20) and the q0* corrections in Fig. 6—are not independently benchmarked.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper treats discovery significance in Poisson counting experiments, first for known background and then for an uncertain background constrained by a Poisson control measurement (the on/off problem). Using the profile likelihood ratio test statistic q0 and the Asimov data set, it derives Eq. (20), a compact expression for the median discovery significance in the uncertain-background case, and shows that this reduces to s/sqrt(b+sigma_b^2) in the limit of small s/b and small sigma_b^2/b. The second part of the paper applies the Barndorff-Nielsen r* higher-order correction, defining q0* = [max(0,r*(0))]^2, and applies it to observed and median significances, with a continuity correction in the known-background case. The authors report Monte Carlo comparisons and conclude that the new formulas improve accuracy at low event yields, especially relative to the simple s/sqrt(b+sigma_b^2) formula. The algebraic derivation from the profile likelihood to Eq. (20) is internally consistent, and the paper provides code to reproduce the figures.","tokens_in":12492,"tokens_out":6153,"duration_ms":76388,"significance":"If the numerical claims hold, Eq. (20) supplies a stable, citable reference for a formula already used in the literature, and the q0* correction extends the usable range of asymptotic formulae for discovery significance to very small Poisson counts. Strengths of the paper include the explicit derivation of Eq. (20), reproduction of the known-background limiting cases, and the availability of open-source code for all figures. The main qualification is that the uncertain-background validation is performed against a Monte Carlo reference built from profile construction/hybrid resampling, which is itself approximate; this limits the strength of the low-yield improvement claims for the central on/off case until an exact or better-controlled reference is supplied.","major_comments":[{"comment":"The MC reference for the uncertain-background model uses the profile-construction/hybrid-resampling procedure: for each simulated (n,m), the background-only p-value is estimated by generating s=0 pseudo-data with b fixed to the profiled value bhat0=(n+m)/(1+tau) of Eq. (16). This plug-in null distribution does not propagate the nuisance-parameter uncertainty and has no guaranteed frequentist coverage, so agreement between Eq. (20)/q0* and these MC points is not an independent validation. Since the on/off model has an exact ancillary reduction under s=0 — conditional on N=n+m, n is Binomial(N, 1/(1+tau)), free of b — exact or near-exact p-values can be computed and used as a benchmark. Please add such a comparison, or at minimum report the difference between the profile-construction reference and the exact conditional-binomial reference for representative (s,b,tau) values, with MC uncerta","section":"Sec. 5, Figs. 2, 4, 6"},{"comment":"The auxiliary statistic u(s) in Eq. (35), and its s=0 specialization Eq. (36), are central to the q0* correction, but they are stated without derivation; the text defers to Refs. [17-19]. Since the r* correction is the second main result of the paper, please provide a derivation in an appendix or online supplement, and state the regularity conditions under which the r* approximation is expected to hold for this discrete exponential-family model. Also clarify the limiting conventions at the boundaries n=0 or m=0: the text assigns u(0)=0 via sqrt(x) ln x -> 0, but the expressions contain ln(n tau / m), and the behavior of r*(0) at those boundary points should be stated more carefully.","section":"Sec. 6.4, Eq. (35)"},{"comment":"The claims that q0* 'improves' agreement with the MC median and that the higher-order corrections provide 'meaningful improvements at small event yields' are supported only by visual inspection of the figures. Please quantify the improvement, e.g., mean and maximum absolute difference in Z units between the Asimov Z_A and the MC median over the plotted b range, before and after the r*/continuity corrections, ideally split into low-yield and large-yield regions. In addition, report the number of MC toys and the fraction of omitted points at the stated 5-sigma toy-resolution ceiling in Figs. 2 and 6; the omission of ceiling-reaching points can distort the apparent median in small-b regions.","section":"Sec. 6.5, Figs. 5 and 6"}],"minor_comments":[{"comment":"The abstract's phrase 'meaningful improvements at small event yields' should be tempered until a quantitative measure is provided; consider saying 'improvements in the specific cases shown' or citing the numerical residuals.","section":"Abstract and Sec. 1"},{"comment":"The provenance of Eq. (20) in the unpublished 2012 SLAC presentation [2] is acknowledged, but since the paper aims to provide a stable reference, it would be useful to state explicitly which of the subsequent applications [3,4] used this exact formula rather than the limiting s/sqrt(b+sigma_b^2) form.","section":"Sec. 5, Eq. (20)"},{"comment":"For the known-background continuity correction, the paper applies the half-unit shift only to q0* and not to q0, to preserve the standard first-order result. This is a reasonable convention, but it should be flagged prominently in the conclusions, where the improved performance of q0* (cc) is partly attributable to the continuity correction rather than to the r* correction alone.","section":"Sec. 6.2"},{"comment":"The left and right panels use different parameterizations (tau fixed versus sigma_b/b fixed), which is useful, but the reader must compare across rows and columns carefully; adding a small panel label such as 'tau = 0.5, 2' and 'sigma_b/b = 0.25, 1' inside each panel would improve readability.","section":"Figs. 2 and 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful consolidation of a formula that has been floating in the literature since an unpublished 2012 talk, and the algebraic derivation is sound. The main risk is that the uncertain-background validation is performed against an approximate Monte Carlo reference, which weakens the low-yield claims. The authors can likely address this within the manuscript's scope by adding an exact conditional-binomial benchmark for the on/off model and by quantifying the visual claims. I do not see a need for rejection, but the validation section should be strengthened before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, honest methodology paper. The main result is the formal derivation of Eq. (20), the median discovery significance for a Poisson counting experiment with a background constrained by a control measurement. It reduces to s/sqrt(b+sigma_b^2) in the right limit but corrects it at low yields. The application of Barndorff-Nielsen r* corrections to both observed and Asimov significances is new, and the known-background case with the continuity correction tracks the exact Poisson tail well even at b=1. Code is released, which is good for reproducibility.\n\nWhat is genuinely new: Eq. (20) did not appear in print before, only in Cowan's 2012 conference presentation, and the paper gives a clean derivation from the profile likelihood. The r* treatment for these Poisson models, with the auxiliary statistic u(0), is a real extension.\n\nThe soft spots: the validation of the uncertain-background case rests on Monte Carlo comparisons whose reference is profile construction/hybrid resampling with b fixed to the profiled value b-hat0. That plug-in null does not propagate the nuisance uncertainty; it is a standard approximate method but not an exact frequentist calculation. So the agreement in Figs. 2 and 6 shows the formulas match that approximate method, not that they are exact. The known-background comparisons use the exact Poisson tail, so those are solid. The visual-only comparisons without error bars make the claimed improvements hard to quantify. Also, the Asimov approximation to the median remains unchecked analytically at low counts; MC agreement is evidence but not a bound.\n\nNone of this is fatal. The paper is careful to flag the lack of a continuity correction in the 2D case and notes where q0* loses accuracy near the known-background limit. It is a paper about useful approximations, not exactness. But the conclusions could be phrased more cautiously: 'meaningful improvements at small event yields' should be read as 'improvements relative to an approximate MC reference.'\n\nWho should read it: HEP experimentalists planning searches and cut optimization, and statisticians working on asymptotics for discrete models. It deserves a serious referee, with a request to add MC error bars, compare against exact conditional p-values (or at least a different method from Cousins et al.) for the on/off case, and to say something about where the Asimov approximation fails. I would cite it.","headline":"Useful, honest derivation of the on/off median-significance formula with r* corrections; the main caveat is that the uncertain-background MC reference is itself approximate, so the low-yield improvements are not independently grounded.","tokens_in":12944,"tokens_out":2246,"would_cite":true,"duration_ms":26017,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper derives closed-form approximations for the median discovery significance in Poisson counting experiments with uncertain background, and shows they remain accurate at low event yields where the standard s/√(b+σ_b²) formula fails.","keywords":["discovery significance","median sensitivity","Asimov data set","profile likelihood ratio","Poisson counting experiment","background uncertainty","higher-order asymptotics","r* statistic"],"falsifier":"Enumerate exactly the joint Poisson distribution over (n,m) for a few small parameter choices (e.g., b=1, τ=1 with s=2 and s=5), compute the true median of the discovery significance using exact background-only p-values, and compare with Eq. (20) and with the r*-corrected Asimov value; a persistent discrepancy larger than the claimed low-yield accuracy would falsify the central claim. The same check should avoid using profile-construction resampling as the reference, since that is itself approximate.","tokens_in":12099,"feed_emoji":"🔬","tokens_out":8633,"duration_ms":84125,"temperature":0.7,"pith_summary":"The paper sets out to give particle physicists a reliable, closed-form way to predict how significant a discovery will be in a simple counting experiment before the experiment is run, even when the expected background is not known exactly. It derives Eq. (20), the full Asimov-data-set expression for the median significance when the background is constrained by a Poisson control measurement, and shows it reduces to the familiar s/√(b+σ_b²) only in the limit of small signal and small background uncertainty. At small event yields, the simple formula overestimates the median sensitivity, while the full expression tracks Monte Carlo results. The paper then applies a higher-order asymptotic correction (the r* statistic) along with, in the known-background case, a continuity correction, extending accuracy to expected counts as low as one. A sympathetic reader would take away that the paper's compact formulae give quantitatively better projections of experimental sensitivity, with errors small enough to matter for cut optimisation in rare-event searches.","feed_headline":"Median-sensitivity formula outdoes s/√b at low event counts","feed_subtitle":"Derived from profile likelihood with higher-order corrections, it sharpens projections for rare-event searches.","key_machinery":"The central object is the Asimov data set: replacing the actual observed counts n and m by their expectation values s+b and τb inside the profile-likelihood-ratio statistic q0, giving a closed-form estimate of the median significance. The decisive identity is Eq. (20), obtained by eliminating the control-region scale factor τ through σ_b² = b/τ, which reduces to the naive s/√(b+σ_b²) only at leading order. The second mechanism is the r* statistic, a higher-order asymptotic correction to the signed likelihood-ratio root r(0), defined as r*(0) = r(0) + (1/r(0)) ln|u(0)/r(0)|, with u(0) the auxiliary statistic computed from log-likelihood derivatives; it sharpens the Gaussian approximation of t","core_discovery":"The paper establishes that the median discovery significance in a Poisson counting experiment with a Poisson-constrained background is, to first asymptotic order, the quantity in Eq. (20), a closed form involving the expected signal s, expected background b, and inferred background uncertainty σ_b. The familiar s/√(b+σ_b²) is only the leading-order limit of this expression, and the paper shows it overestimates sensitivity when s/b or σ_b²/b are not small. It further establishes that applying the r* higher-order correction at the same Asimov point improves both the observed and the median significance at low event yields, and that in the known-background case combining r* with a continuity co","pith_inferences":["The paper's own conclusions flag that no canonical continuity correction exists for the two-count (n,m) discrete problem; a natural extension would be a mid-p or half-bin correction in some transformed coordinate, and the accuracy gain there is untested.","The Asimov approximation itself is never bounded analytically; at extremely low counts, the step structure of the median significance suggests that a direct saddlepoint or exact enumeration approach could differ from the Asimov point in ways the r* correction cannot repair.","Because Eq. (20) is expressed in terms of σ_b via the control measurement, it could plausibly be adapted to expected exclusion limits (not just discovery) by the same profile-likelihood machinery, although the paper does not pursue that.","If the higher-order corrections hold up in the low-yield regime, they could reduce the reliance on toy-Monte-Carlo sensitivity scans in current searches, since the closed forms are effectively instant and differentiable in the physics parameters."],"forward_implications":["Planned experiments can compute expected discovery sensitivity without toy Monte Carlo, using only s, b, and the control-region scale factor, via a closed formula that is continuous and instantly evaluable.","Projections based on s/√(b+σ_b²) overstate the sensitivity of low-statistics searches; cut choices and exposure decisions based on that formula may be suboptimal in exactly the regimes where experiments are background-limited.","The r*-corrected Asimov significance interpolates the stepwise Monte Carlo median of the significance, giving an analytic curve that can be used for smooth sensitivity scans.","For the known-background case, the continuity-corrected r* significance matches the exact Poisson-tail significance down to an expected background of one event, so it can serve as a fast check of exact calculations.","The same formalism applies directly to the astrophysical on/off problem, where the same Poisson control-measurement structure is standard."],"fun_headline_variants":["Better median significance formula for rare-event searches","New formula sharpens discovery sensitivity estimates","Beyond s/√b: improved median sensitivity for counting experiments","Higher-order corrections boost sensitivity projections at low yields","Median discovery significance: closed form beats old approximation"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the median discovery significance is well approximated by evaluating the test statistic at the single Asimov data point (n=s+b, m=τb); the paper demonstrates this by Monte Carlo agreement (Secs. 4–5) but does not provide an analytic error bound, and Sec. 6.5 acknowledges that the r* correction only improves the significance at that fixed Asimov point, not the Asimov approximation itself. In the uncertain-background case the Monte Carlo referen","fun_headline_variants_meta":{"raw":{"variants":["Better median significance formula for rare-event searches","New formula sharpens discovery sensitivity estimates","Beyond s/√b: improved median sensitivity for counting experiments","Higher-order corrections boost sensitivity projections at low yields","Median discovery significance: closed form beats old approximation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001019,"raw_usage":{"total_tokens":4141,"prompt_tokens":756,"completion_tokens":3385,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":3328}},"tokens_in":500,"tokens_out":3385,"duration_ms":24770,"temperature":1.0,"reasoning_tokens":3328,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T07:07:52.026631+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Enumerate exactly the joint Poisson distribution over (n,m) for a few small parameter choices (e.g., b=1, τ=1 with s=2 and s=5), compute the true median of the discovery significance using exact background-only p-values, and compare with Eq. (20) and with the r*-corrected Asimov value; a persistent discrepancy larger than the claimed low-yield accuracy would falsify the central claim. The same check should avoid using profile-construction resampling as the reference, since that is itself approximate.","supporting_citations":[],"review_version":1}