REVIEW 3 major objections 3 minor 2 cited by
The Benjamini–Hochberg procedure can fail to control the false discovery rate for correlated two-sided Gaussian p-values.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 01:01 UTC pith:J6MN3G5H
load-bearing objection Abstract-only claim of a certified counterexample that BH can exceed nominal FDR for correlated two-sided Gaussians, killing a 20-year conjecture if the interval-arithmetic bound holds. the 3 major comments →
The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
There exists a factor model of correlated two-sided Gaussian p-values such that the Benjamini–Hochberg procedure run at nominal level α = 0.01 has FDR greater than 0.0104 for every sufficiently large number of hypotheses; the inequality is certified by rigorous interval arithmetic and thereby disproves the long-standing conjecture that BH controls FDR for all such tests.
What carries the argument
An explicit Gaussian factor model that generates the joint distribution of the two-sided p-values, together with an interval-arithmetic certificate that rigorously lower-bounds the limiting FDR of BH under that model.
Load-bearing premise
The constructed factor model must lie inside the precise dependence class (correlated two-sided Gaussians) for which the twenty-year conjecture claimed FDR control; otherwise the excess FDR would not constitute a genuine counter-example.
What would settle it
Either a tighter analytic or numerical bound showing that the FDR of BH under the given factor model is at most 0.01 for all large n, or an independent re-verification of the interval-arithmetic certificate that fails to recover the strict lower bound 0.0104.
If this is right
- BH cannot be invoked as a black-box FDR-controlling procedure for arbitrary positive dependence among two-sided Gaussian tests.
- Any theoretical guarantee that previously relied on the disproved conjecture must be re-examined or restricted to narrower dependence classes.
- Practitioners analyzing two-sided Gaussian statistics under factor-type correlation should either verify FDR control by other means or adopt more conservative multiple-testing procedures.
- The same factor-model construction supplies a concrete test case against which any proposed extension of BH can be checked.
Where Pith is reading between the lines
- The counter-example is specific to two-sided tests; the one-sided Gaussian case may still enjoy FDR control under the same dependence, suggesting a sharp distinction that future theory should map.
- Because the excess is modest (only a few percent above nominal), the practical inflation may be small for many data sets, yet the logical failure already forces a revision of textbooks and software documentation that state the conjecture as fact.
- Interval-arithmetic certification of asymptotic FDR bounds could become a standard tool for settling other open dependence questions in multiple testing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript asserts that the Benjamini–Hochberg (BH) procedure need not control the false discovery rate at its nominal level for correlated two-sided Gaussian p-values. The authors construct a factor model under which, at α = 0.01, a rigorous interval-arithmetic certificate establishes FDR > 0.0104 for all sufficiently large numbers of hypotheses, thereby disproving a conjecture widely believed for roughly twenty years. Monte Carlo experiments are reported as consistent with the asymptotic lower bound. The proof is stated to have been obtained by GPT-5.6 Pro and carefully checked by the author. Only the abstract is available for this review; the full construction, certificate, and Monte Carlo design are not present.
Significance. If the claimed certificate is correct and the factor model lies inside the dependence class for which the conjecture asserted FDR control, the result is a high-impact negative finding in multiple-testing theory. BH is foundational, and a rigorous counterexample for two-sided Gaussian tests under factor dependence would sharpen the known sufficient conditions (e.g., PRDS) and guide practice. The use of interval arithmetic to produce a machine-checkable numerical lower bound, together with Monte Carlo consistency checks, is a methodological strength worth crediting once the certificate can be inspected. The AI-assisted provenance is secondary provided the author-checked certificate is independently verifiable.
major comments (3)
- [Abstract (certificate claim)] The central claim rests on a rigorous interval-arithmetic certificate that FDR exceeds 0.0104 at α = 0.01 for all large m. With only the abstract available, the factor loadings, the asymptotic FDR expression, the interval bounds, and the software/implementation of the certificate cannot be examined. Without those details the load-bearing lower bound cannot be verified, so the disproof of the conjecture cannot yet be accepted.
- [Abstract (factor-model claim)] For the construction to refute the twenty-year conjecture it must lie inside the dependence class the conjecture was understood to cover (correlated two-sided Gaussian p-values, typically under positive or factor dependence). The abstract asserts a “factor model” of such p-values but does not state the precise correlation structure or its relation to known sufficient conditions (PRDS, etc.). Explicit membership must be established so the example is not outside the conjecture’s scope.
- [Abstract (proof provenance)] Because the proof was produced by GPT-5.6 Pro, the manuscript must document the author’s verification steps (interval-arithmetic code, bound propagation, and any hand checks) in enough detail for independent reproduction. An author statement that the proof was “carefully checked” is not itself a certificate; the verification trail is load-bearing for a machine-assisted numerical proof.
minor comments (3)
- [Abstract] The abstract should state the precise asymptotic regime (e.g., fixed factor loadings as m → ∞, or any sparsity/null proportion assumptions) so readers can immediately see the scope of the claimed FDR lower bound.
- [Abstract (Monte Carlo sentence)] When the full text is supplied, the Monte Carlo design (number of replications, m values, random-number generation, and how empirical FDR is estimated) should be reported with enough precision to allow independent reproduction of the consistency claim.
- [Abstract] The abstract would benefit from a one-sentence pointer to the known positive results (e.g., BH under PRDS or independence) that the counterexample is intended to sit outside, clarifying the logical gap being closed.
Circularity Check
No circularity: abstract-only counterexample with independent FDR target and interval-arithmetic lower bound
full rationale
Only the abstract is available. It claims an explicit factor-model construction of correlated two-sided Gaussian p-values for which an interval-arithmetic certificate shows FDR > 0.0104 at nominal α = 0.01 for all large m, thereby disproving a long-standing conjecture. The target quantity (FDR) is defined independently of the model parameters; the reported excess is presented as a rigorous lower bound, not a fitted or self-defined quantity. No equations, self-citations, uniqueness theorems, or ansatzes appear in the available text, so none of the six circularity patterns can be exhibited by quotation. The derivation is therefore self-contained against the abstract's own claims; residual risk is verification of the unseen certificate, not circularity. Score 0 is the honest finding.
Axiom & Free-Parameter Ledger
free parameters (2)
- factor-model loadings / correlation strength
- nominal level α = 0.01
axioms (4)
- domain assumption Two-sided p-values arise from correlated Gaussian test statistics under a factor model.
- standard math The Benjamini–Hochberg step-up procedure is applied at fixed level α to the resulting p-values.
- ad hoc to paper Interval arithmetic yields a rigorous lower bound on the asymptotic FDR of the constructed model.
- domain assumption The constructed dependence lies in the class for which the twenty-year conjecture claimed FDR control.
read the original abstract
We show that the Benjamini--Hochberg procedure can fail to control the false discovery rate (FDR) at its nominal level for correlated two-sided Gaussian $p$-values. We construct a factor model for which, at level $\alpha=0.01$, a rigorous interval-arithmetic certificate proves $FDR>0.0104$ for all sufficiently large numbers of hypotheses. This disproves a conjecture widely believed to be true for twenty years. Monte Carlo experiments are consistent with the theoretical result. The proof was obtained by GPT-5.6 Pro and carefully checked by the author.
Forward citations
Cited by 2 Pith papers
-
Mirror and knockoff+ thresholds under dependence
Mirror and knockoff+ thresholds can have FDR far above nominal under dependence, even for uniform PRDS p-values and equicorrelated Gaussian scores.
-
How Much Can Gaussian Dependence Inflate the Benjamini-Hochberg Procedure's FDR?
No universal multiplicative FDR bound holds for Benjamini-Hochberg under Gaussian dependence: worst-case FDR/q diverges like sqrt(log(1/q)), with exact constants in natural cases.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.