Pith. sign in

REVIEW 3 major objections 3 minor 2 cited by

The Benjamini–Hochberg procedure can fail to control the false discovery rate for correlated two-sided Gaussian p-values.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 01:01 UTC pith:J6MN3G5H

load-bearing objection Abstract-only claim of a certified counterexample that BH can exceed nominal FDR for correlated two-sided Gaussians, killing a 20-year conjecture if the interval-arithmetic bound holds. the 3 major comments →

arxiv 2607.12208 v1 pith:J6MN3G5H submitted 2026-07-13 math.ST cs.AIstat.MEstat.TH

The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests

classification math.ST cs.AIstat.MEstat.TH MSC 62J1562H1562F03
keywords Benjamini-Hochbergfalse discovery rateFDR controlcorrelated p-valuestwo-sided Gaussian testsfactor modelinterval arithmeticmultiple testing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper establishes that the Benjamini–Hochberg (BH) procedure does not always control the false discovery rate at its nominal level when the underlying p-values come from correlated two-sided Gaussian tests. The authors exhibit an explicit factor model of dependence for which, at the conventional level α = 0.01, a machine-checked interval-arithmetic argument proves that the FDR stays strictly above 0.0104 once the number of hypotheses is large enough. The construction therefore supplies a concrete counter-example to a conjecture that had been widely regarded as true for two decades. Monte Carlo simulations of the same model are consistent with the certified lower bound. A sympathetic reader cares because BH is the default multiple-testing method in genomics, imaging, and many other high-dimensional settings that routinely produce two-sided Gaussian statistics; if the conjecture fails, those analyses may be reporting more false discoveries than their nominal level suggests.

Core claim

There exists a factor model of correlated two-sided Gaussian p-values such that the Benjamini–Hochberg procedure run at nominal level α = 0.01 has FDR greater than 0.0104 for every sufficiently large number of hypotheses; the inequality is certified by rigorous interval arithmetic and thereby disproves the long-standing conjecture that BH controls FDR for all such tests.

What carries the argument

An explicit Gaussian factor model that generates the joint distribution of the two-sided p-values, together with an interval-arithmetic certificate that rigorously lower-bounds the limiting FDR of BH under that model.

Load-bearing premise

The constructed factor model must lie inside the precise dependence class (correlated two-sided Gaussians) for which the twenty-year conjecture claimed FDR control; otherwise the excess FDR would not constitute a genuine counter-example.

What would settle it

Either a tighter analytic or numerical bound showing that the FDR of BH under the given factor model is at most 0.01 for all large n, or an independent re-verification of the interval-arithmetic certificate that fails to recover the strict lower bound 0.0104.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • BH cannot be invoked as a black-box FDR-controlling procedure for arbitrary positive dependence among two-sided Gaussian tests.
  • Any theoretical guarantee that previously relied on the disproved conjecture must be re-examined or restricted to narrower dependence classes.
  • Practitioners analyzing two-sided Gaussian statistics under factor-type correlation should either verify FDR control by other means or adopt more conservative multiple-testing procedures.
  • The same factor-model construction supplies a concrete test case against which any proposed extension of BH can be checked.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The counter-example is specific to two-sided tests; the one-sided Gaussian case may still enjoy FDR control under the same dependence, suggesting a sharp distinction that future theory should map.
  • Because the excess is modest (only a few percent above nominal), the practical inflation may be small for many data sets, yet the logical failure already forces a revision of textbooks and software documentation that state the conjecture as fact.
  • Interval-arithmetic certification of asymptotic FDR bounds could become a standard tool for settling other open dependence questions in multiple testing.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript asserts that the Benjamini–Hochberg (BH) procedure need not control the false discovery rate at its nominal level for correlated two-sided Gaussian p-values. The authors construct a factor model under which, at α = 0.01, a rigorous interval-arithmetic certificate establishes FDR > 0.0104 for all sufficiently large numbers of hypotheses, thereby disproving a conjecture widely believed for roughly twenty years. Monte Carlo experiments are reported as consistent with the asymptotic lower bound. The proof is stated to have been obtained by GPT-5.6 Pro and carefully checked by the author. Only the abstract is available for this review; the full construction, certificate, and Monte Carlo design are not present.

Significance. If the claimed certificate is correct and the factor model lies inside the dependence class for which the conjecture asserted FDR control, the result is a high-impact negative finding in multiple-testing theory. BH is foundational, and a rigorous counterexample for two-sided Gaussian tests under factor dependence would sharpen the known sufficient conditions (e.g., PRDS) and guide practice. The use of interval arithmetic to produce a machine-checkable numerical lower bound, together with Monte Carlo consistency checks, is a methodological strength worth crediting once the certificate can be inspected. The AI-assisted provenance is secondary provided the author-checked certificate is independently verifiable.

major comments (3)
  1. [Abstract (certificate claim)] The central claim rests on a rigorous interval-arithmetic certificate that FDR exceeds 0.0104 at α = 0.01 for all large m. With only the abstract available, the factor loadings, the asymptotic FDR expression, the interval bounds, and the software/implementation of the certificate cannot be examined. Without those details the load-bearing lower bound cannot be verified, so the disproof of the conjecture cannot yet be accepted.
  2. [Abstract (factor-model claim)] For the construction to refute the twenty-year conjecture it must lie inside the dependence class the conjecture was understood to cover (correlated two-sided Gaussian p-values, typically under positive or factor dependence). The abstract asserts a “factor model” of such p-values but does not state the precise correlation structure or its relation to known sufficient conditions (PRDS, etc.). Explicit membership must be established so the example is not outside the conjecture’s scope.
  3. [Abstract (proof provenance)] Because the proof was produced by GPT-5.6 Pro, the manuscript must document the author’s verification steps (interval-arithmetic code, bound propagation, and any hand checks) in enough detail for independent reproduction. An author statement that the proof was “carefully checked” is not itself a certificate; the verification trail is load-bearing for a machine-assisted numerical proof.
minor comments (3)
  1. [Abstract] The abstract should state the precise asymptotic regime (e.g., fixed factor loadings as m → ∞, or any sparsity/null proportion assumptions) so readers can immediately see the scope of the claimed FDR lower bound.
  2. [Abstract (Monte Carlo sentence)] When the full text is supplied, the Monte Carlo design (number of replications, m values, random-number generation, and how empirical FDR is estimated) should be reported with enough precision to allow independent reproduction of the consistency claim.
  3. [Abstract] The abstract would benefit from a one-sentence pointer to the known positive results (e.g., BH under PRDS or independence) that the counterexample is intended to sit outside, clarifying the logical gap being closed.

Circularity Check

0 steps flagged

No circularity: abstract-only counterexample with independent FDR target and interval-arithmetic lower bound

full rationale

Only the abstract is available. It claims an explicit factor-model construction of correlated two-sided Gaussian p-values for which an interval-arithmetic certificate shows FDR > 0.0104 at nominal α = 0.01 for all large m, thereby disproving a long-standing conjecture. The target quantity (FDR) is defined independently of the model parameters; the reported excess is presented as a rigorous lower bound, not a fitted or self-defined quantity. No equations, self-citations, uniqueness theorems, or ansatzes appear in the available text, so none of the six circularity patterns can be exhibited by quotation. The derivation is therefore self-contained against the abstract's own claims; residual risk is verification of the unseen certificate, not circularity. Score 0 is the honest finding.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

Abstract-only review: free parameters are those needed to define the factor model that produces the counterexample; axioms are the standard Gaussian and BH setup plus the claim that the model sits inside the conjectured class; no new physical entities are invented.

free parameters (2)
  • factor-model loadings / correlation strength
    The abstract’s counterexample is a constructed factor model; the specific loadings that make FDR exceed α are chosen by the authors and are free parameters of the construction (exact values not given in the abstract).
  • nominal level α = 0.01
    The numerical certificate is stated at a fixed α; while α is a user choice, the concrete excess 0.0104 is tied to this choice and is part of the reported construction.
axioms (4)
  • domain assumption Two-sided p-values arise from correlated Gaussian test statistics under a factor model.
    Standard multiple-testing setup; the abstract takes Gaussianity and the factor structure as given for the counterexample.
  • standard math The Benjamini–Hochberg step-up procedure is applied at fixed level α to the resulting p-values.
    Definition of the procedure whose FDR is being evaluated.
  • ad hoc to paper Interval arithmetic yields a rigorous lower bound on the asymptotic FDR of the constructed model.
    The abstract’s certificate rests on this numerical method; its correctness for the particular expressions is paper-specific and not independently verified here.
  • domain assumption The constructed dependence lies in the class for which the twenty-year conjecture claimed FDR control.
    Necessary for the result to be a counterexample rather than an instance outside the conjecture.

pith-pipeline@v1.1.0-grok45 · 6003 in / 2636 out tokens · 32180 ms · 2026-07-15T01:01:27.352110+00:00 · methodology

0 comments
read the original abstract

We show that the Benjamini--Hochberg procedure can fail to control the false discovery rate (FDR) at its nominal level for correlated two-sided Gaussian $p$-values. We construct a factor model for which, at level $\alpha=0.01$, a rigorous interval-arithmetic certificate proves $FDR>0.0104$ for all sufficiently large numbers of hypotheses. This disproves a conjecture widely believed to be true for twenty years. Monte Carlo experiments are consistent with the theoretical result. The proof was obtained by GPT-5.6 Pro and carefully checked by the author.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mirror and knockoff+ thresholds under dependence

    math.ST 2026-07 accept novelty 8.0

    Mirror and knockoff+ thresholds can have FDR far above nominal under dependence, even for uniform PRDS p-values and equicorrelated Gaussian scores.

  2. How Much Can Gaussian Dependence Inflate the Benjamini-Hochberg Procedure's FDR?

    math.ST 2026-07 accept novelty 8.0

    No universal multiplicative FDR bound holds for Benjamini-Hochberg under Gaussian dependence: worst-case FDR/q diverges like sqrt(log(1/q)), with exact constants in natural cases.