Pith. sign in

REVIEW 1 major objections 3 minor 16 references

For correlated Gaussian tests, the Benjamini-Hochberg FDR can exceed any fixed multiple of q, with the ratio growing like sqrt(log(1/q)).

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 00:58 UTC pith:H3LILGVU

load-bearing objection Disproves the long-open Gaussian FDR conjecture with sharp asymptotics; the proof is heavy but the central claims check out. the 1 major comments →

arxiv 2607.14812 v3 pith:H3LILGVU submitted 2026-07-16 math.ST stat.TH

How Much Can Gaussian Dependence Inflate the Benjamini-Hochberg Procedure's FDR?

classification math.ST stat.TH MSC 62F0362F0562H15
keywords Benjamini-Hochberg procedurefalse discovery rateGaussian dependencetwo-sided testsone-sided testsone-common-factor modelworst-case FDRpositive regression dependence
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper asks how far the Benjamini-Hochberg (BH) procedure's false discovery rate can rise above its nominal level q when test statistics are Gaussian with arbitrary correlations. It proves that for both two-sided and one-sided Gaussian tests, the worst-case FDR is at least an explicit function of q that is strictly larger than q for every q and whose ratio to q diverges as q tends to 0. This overturns a long-standing folklore conjecture that BH controls FDR for correlated two-sided Gaussian z-tests, and it rules out any universal multiplicative bound of the form Cq. The lower bounds are achieved by finite Gaussian models with covariance matrices of one-common-factor form, and matching upper bounds in that class show the q sqrt(log(1/q)) order is sharp; for one-sided tests the leading constant is determined exactly.

Core claim

Two-sided: the supremum over N, mean vector, and correlation matrix of the FDR of BH at level q is at least l=(q) = q sqrt(log(1/q))/(2 sqrt(pi)) + c_l q + o(q) with c_l = 0.6492828..., and l=(q) > q for every q in (0,1). One-sided: at least l<=(q) = q sqrt(log(1/q))/sqrt(pi) + q/2 + o(q), also > q. Since these lower bounds divided by q diverge, no finite C can satisfy FDR <= Cq for all Gaussian dependence. Within the one-common-factor class, the paper proves FDR = O(q sqrt(log(1/q))) for two-sided tests and FDR = q sqrt(log(1/q))/sqrt(pi) + O(q) for one-sided tests, matching the one-sided lower-bound constant; the exact two-sided constant is left open.

What carries the argument

The engine is a family of finite Gaussian designs (a common factor plus independent noise) with three population components: nulls with tiny loading, a point-mass primary non-null, and a continuum of 'secondary' non-nulls that shapes the conditional p-value CDF. The analysis centers on the limiting BH cutoff C(z) = inf{c >= u_q : Rbar_z(c) >= 1} for the normalized mixture CDF; the proof works only when the cutoff is a 'regular first crossing'—the CDF stays strictly below 1 before the cutoff, strictly above 1 just after it, and the null contribution is continuous there. A likelihood-ratio envelope M(y) = sup_{a>=0} e^{-a^2/2} cosh(ay), crossed via q M(y_q) = 1, fixes the two-sided secondary-m

Load-bearing premise

The load-bearing technical premise is that the limiting BH cutoff is a 'regular first crossing' for almost every common-factor value—the normalized p-value CDF must stay strictly below 1 up to the cutoff and strictly above 1 just after it, with the null contribution continuous there; the finite-sample FDR limit is obtained by dominated convergence only at such regular points.

What would settle it

Evaluate the normalized mixture CDF Rbar_z(c) on a fine grid in z around the boundary points -1 and -y_q (and around 0 and y_q^+ in the one-sided construction). If for some z in a set of positive measure the curve touches 1 before its first crossing or fails to exceed 1 immediately after, the regular-first-crossing condition fails and the dominated-convergence step cannot be applied; equally, a direct simulation of the paper's finite model at these parameters that clearly violates the stated lower bound would falsify the construction.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • At every q in (0,1), BH can be anticonservative under Gaussian dependence: worst-case FDR is at least l=(q) > q for two-sided tests and l<=(q) > q for one-sided tests.
  • The ratio of worst-case FDR to q diverges as q goes to 0, so no universal multiplicative constant C can bound the FDR uniformly over all Gaussian correlations.
  • Restricting to one-common-factor Gaussian models does not restore FDR control; the order q sqrt(log(1/q)) persists in both two-sided and one-sided settings.
  • For one-sided common-factor tests, the leading constant is exactly 1/sqrt(pi): upper and lower bounds match, so the inflation rate is determined within that class.
  • Dependence direction matters: the two-sided construction works with positive correlations because the absolute-value transformation breaks PRDS, while the one-sided construction requires negative null-to-non-null correlations to violate PRDS.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: in high-throughput applications that report BH at very small q, the paper's numbers (roughly 1.36x inflation at q=0.01 two-sided and 1.83x one-sided for the constructed designs) imply that dependence-agnostic reporting could be materially anticonservative at even smaller q.
  • Editorial inference: the two-sided leading constant is open; the natural testable target is to decide whether the worst-case two-sided class has the larger 1/sqrt(pi) constant or a genuinely smaller one.
  • Editorial inference: the construction's dependence on the Gaussian envelope M(y) suggests the same divergence rate should be checked for heavier-tailed or non-normal symmetric likelihoods; the result may be qualitatively different there.
  • Editorial inference: a practical validation is to simulate the paper's one-common-factor designs with q as low as 10^-4 and compare empirical FDR against q sqrt(log(1/q))/sqrt(pi) to confirm the predicted rate.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 3 minor

Summary. The paper studies the worst-case FDR of the Benjamini–Hochberg procedure for Gaussian z-tests under arbitrary correlation. For two-sided tests it constructs a q-indexed family of finite Gaussian models with FDR at least an explicit ℓ=(q)>q, where ℓ=(q)=q√(log(1/q))/(2√π)+cℓ q+o(q); for one-sided tests, a sign-reversed common-factor construction gives ℓ≤(q)=q√(log(1/q))/√π+q/2+o(q). Together these disprove any universal multiplicative bound Cq. The paper also proves common-factor upper bounds: O(q√log(1/q)) for two-sided tests and q√log(1/q)/√π+O(q) for one-sided tests, the latter matching the one-sided lower constant. The proofs use a random two-level design, a regular-first-crossing analysis of the limiting BH cutoff via DKW-type empirical-CDF convergence, and Gaussian tail expansions; the lower constants are explicit functions of q and the standard normal distribution, with no fitted parameters.

Significance. If correct, this settles a long-standing question: BH does not control FDR under arbitrary Gaussian dependence, and the inflation is worse than multiplicative. The explicit lower bounds are concrete and falsifiable, and the one-sided common-factor constant is proven sharp. The 'regular first crossing' machinery is a useful technical contribution. I checked the load-bearing regularity verifications (Lemmas 2.15, 2.19, 2.20, 3.11–3.13) and the dominated-convergence argument; I did not find a gap. The two-sided upper and lower constants are not matched, but the paper states this. The main fixes needed are presentational, especially notation.

major comments (1)
  1. [§1, Eq. (1), (5), (9), (141); Appendix A] The paper's notation for Φ is internally inconsistent. The text defines Φ and φ=Φ′ as the standard normal CDF and density, but then uses Φ(x) as the upper-tail survival function in p(x)=2Φ(x)=2{1−Φ(x)}, p+(x)=Φ(x)=1−Φ(x), λ(x)=φ(x)/Φ(x), and throughout the tail expansions. In particular, Eq. (5) defines ℓ=(q) with qΦ(1)+... and Φ(yq); using the stated CDF convention, ℓ=(q) can exceed 1 for some q∈(0,1), which is impossible for an FDR lower bound. This notational conflict affects Eqs. (141), (154), (161), (185), Lemma A.2, and many other places. The central proofs appear to be correct when read with Φ as the upper-tail function, and Eq. (9) requires the CDF elsewhere. The manuscript must be revised to introduce, say, Φ̄=1−Φ for the tail and use it consistently. This is a load-bearing notational issue because the main theorem's exact formula and the numerical tables are otherwise uninterpr
minor comments (3)
  1. [§2.7, Proposition 2.2] The constant cℓ=0.6492828... is reported with no indication of the quadrature method or error bound. Since cℓ is not needed for the divergence claim, this is a presentation issue, but a brief note on numerical accuracy would help.
  2. [§4.2, Theorem 4.2] The condition q≤2Φ(1)≈0.3173 should be restated after the Φ notation is fixed. It currently reads oddly because, under the paper's stated CDF convention, 2Φ(1) would be about 1.6826, not 0.3173.
  3. [§1] The quoted phrase 'convincing simutheoretical evidence' should be marked as a quotation error or 'sic'.

Circularity Check

0 steps flagged

No significant circularity: the lower bounds are explicit mathematical constructions, not fitted predictions, and the self-citations are not load-bearing.

full rationale

The paper's central claims are existence lower bounds: for every q and δ it constructs a finite Gaussian model whose BH FDR exceeds an explicit function. The two-sided bound ℓ=(q) is defined in Eq. (5) solely from q, the standard normal distribution, and the envelope M; the one-sided bound ℓ≤(q) in Eq. (141) is likewise an explicit formula. The construction in Table 3/4 chooses null and non-null masses and loadings to force the limiting conditional FDP D*(z) to equal q, qM(-z), or 1 on the relevant z-regimes, with τr = O(r^2) determined by the integral equation (38)-(40) and cr calibrated by (41). This is an adversarial design for a lower-bound proof, not a fitted parameter subsequently renamed as a prediction. The convergence is proved by dominated convergence after verifying regular first crossings via Lemmas 2.15, 2.19, 2.20, and 3.11-3.13; none of these lemmas assumes the target ℓ values. The upper bounds in Theorems 4.2 and 4.4 are independent and match the qualitative rate. Self-citations such as Fithian-Lei and Luo-Lei appear only in the Discussion and are not used to justify the main theorems. No circular step can be exhibited from the paper's equations or citations.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The central claims rest only on standard Gaussian probability, empirical-process inequalities, and proved analytic lemmas. No fitted parameters appear; the numerical constant c_ell is a quadrature value of a derived integral. The design measure nu_{q,r} is a proof device, not an invented entity with independent empirical content.

axioms (4)
  • domain assumption Gaussian observations T ~ N(theta, Sigma) with unit marginal variances
    Problem definition; all results are conditional on this model.
  • standard math Dvoretzky-Kiefer-Wolfowitz inequality
    Used in Lemma 2.6 to obtain uniform empirical-CDF bounds conditional on Z.
  • standard math Gaussian Mills-ratio and tail expansions (Lemma A.2)
    Used throughout for quantile asymptotics and tail-ratio limits; proved in the appendix.
  • standard math Karlin-Rinott absolute-value multinormal MTP2 criterion
    Used in Remark 4.3 to establish PRDN of two-sided null p-values in the common-factor class; not needed for the main lower-bound theorems.

pith-pipeline@v1.3.0-alltime-deepseek · 48921 in / 12801 out tokens · 103748 ms · 2026-08-02T00:58:40.758039+00:00 · methodology

0 comments
read the original abstract

We study the worst-case false discovery rate (FDR) of the Benjamini-Hochberg procedure for both one- and two-sided Gaussian tests when the correlation matrix is otherwise unrestricted. In each setting we construct a $q$-indexed family of finite Gaussian models whose FDR divided by $q$ diverges as $q\downarrow0$, disproving any universal multiplicative FDR bound. For two-sided tests, the supremum over the number of hypotheses, mean vector, and correlation matrix is at least an explicit $\ell_{=}(q)>q$ satisfying \[ \ell_{=}(q)=\frac{q\sqrt{\log(1/q)}}{2\sqrt{\pi}}+c_\ell q+o(q), \qquad c_\ell=0.6492828\ldots. \] For the one-sided hypotheses $H_i:\theta_i\leq0$, a sign-reversed one-common-factor construction gives the stronger explicit lower bound $\ell_{\le}(q)>q$, with \[ \ell_{\le}(q)=\frac{q\sqrt{\log(1/q)}}{\sqrt\pi} +\frac q2+o(q). \] Finally, we prove an $O\{q\sqrt{\log(1/q)}\}$ upper bound for the two-sided {one-common-factor} class and the matching upper bound $q\sqrt{\log(1/q)}/\sqrt\pi+O(q)$ for the one-sided one-common-factor class.

Figures

Figures reproduced from arXiv: 2607.14812 by Lihua Lei.

Figure 1
Figure 1. Figure 1: Numerical evaluation of ℓ=(q) − q and ℓ=(q)/q over the grid q = 0.005k, k = 1, . . . , 199. q 0.01 0.05 0.10 0.20 ℓ=(q) − q 0.0036 0.0129 0.0209 0.0314 ℓ=(q)/q 1.3553 1.2571 1.2094 1.1570 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Numerical evaluation of ℓ≤(q) − q and ℓ≤(q)/q over the grid q = 0.005k, k = 1, . . . , 199 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

16 extracted references · 4 canonical work pages

  1. [1]

    Rina Foygel Barber and Emmanuel J. Candès. Controlling the false discovery rate via knockoffs. The Annals of Statistics, 43(5):2055–2085,

  2. [12]

    Anat Reiner-Benaim

    doi: 10.1214/25-AOS2543. Anat Reiner-Benaim. FDR control by the BH procedure for two-sided correlated tests with implications to gene expression data analysis.Biometrical Journal, 49(1):107–126,

  3. [15]

    Weijie J

    doi: 10.1016/j.jspi.2024.106238. Weijie J. Su. The FDR-linking theorem,

  4. [1981]

    Xiao Li and William Fithian

    doi: 10.1214/aos/1176345583. Xiao Li and William Fithian. Whiteout: When do fixed-X knockoffs fail?,

  5. [1995]

    Yoav Benjamini and Daniel Yekutieli

    doi: 10.1111/j.2517-6161.1995.tb02031.x. Yoav Benjamini and Daniel Yekutieli. The control of the false discovery rate in multiple testing under dependency.The Annals of Statistics, 29(4):1165–1188,

  6. [2001]

    Ziyu Chi, Aaditya Ramdas, and Ruodu Wang

    doi: 10.1214/aos/1013699998. Ziyu Chi, Aaditya Ramdas, and Ruodu Wang. Multiple testing under negative dependence.Bernoulli, 31(2):1230–1255,

  7. [2006]

    Edgar Dobriban

    doi: 10.1239/jap/1143936242. Edgar Dobriban. The Benjamini–Hochberg procedure can fail to control the FDR for correlated two-sided Gaussian tests,

  8. [2007]

    Marine Roux.Inférence de graphes par une procédure de test multiple avec application en Neuroimagerie

    doi: 10.1002/bimj.200510313. Marine Roux.Inférence de graphes par une procédure de test multiple avec application en Neuroimagerie. PhD thesis, Université Grenoble Alpes (ComUE), September

  9. [2010]

    URL https://doi.org/10.1111/j.1467-9868.2010.00746.x

    doi: 10.1111/j.1467-9868.2010.00746.x. URL https://doi.org/10.1111/j.1467-9868.2010.00746.x. Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: A practical and powerful approach to multiple testing.Journal of the Royal Statistical Society: Series B, 57(1):289–300,

  10. [2015]

    Yoav Benjamini

    doi: 10.1214/15-AOS1337. Yoav Benjamini. Discovering the false discovery rate.Journal of the Royal Statistical Society: Series 65 B (Statistical Methodology), 72(4):405–416,

  11. [2018]

    66 A Auxiliary lemmas This appendix collects the properties of the envelope and Gaussian tails used to define and analyze νq,r

    URLhttps://arxiv.org/abs/1812.08965. 66 A Auxiliary lemmas This appendix collects the properties of the envelope and Gaussian tails used to define and analyze νq,r. The following lemma characterizes the likelihood-ratio envelope. Lemma A.1.For0 ≤y≤ 1, M(y) =

  12. [2021]

    Yixiang Luo, William Fithian, and Lihua Lei

    URLhttps: //arxiv.org/abs/2107.06388. Yixiang Luo, William Fithian, and Lihua Lei. Improving knockoffs with conditional calibration. The Annals of Statistics, 53(5):2283–2302,

  13. [2022]

    Yosef Hochberg and Dror Rom

    doi: 10.1214/21-AOS2137. Yosef Hochberg and Dror Rom. Extensions of multiple testing procedures based on simes’ test. Journal of Statistical Planning and Inference, 48(2):141–152,

  14. [2023]

    URLhttps://arxiv.org/abs/2304.05261. Sanat K. Sarkar and Shiyu Zhang. Shifted BH methods for controlling false discovery rate in multiple testing of the means of correlated normals against two-sided alternatives.Journal of Statistical Planning and Inference, 236:106238,

  15. [2025]

    Antonio Colangelo, Alfred Müller, and Marco Scarsini

    doi: 10.3150/24-BEJ1768. Antonio Colangelo, Alfred Müller, and Marco Scarsini. Positive dependence and weak convergence. Journal of Applied Probability, 43(1):48–59,

  16. [2026]

    William Fithian and Lihua Lei

    URLhttps://arxiv.org/abs/2607.12208. William Fithian and Lihua Lei. Conditional calibration for false discovery rate control under dependence.The Annals of Statistics, 50(6):3091–3118,