Pith. sign in

REVIEW 4 minor 24 references

The mirror and knockoff+ thresholds do not control the false discovery rate under dependence unless null signs can be flipped independently.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 19:05 UTC pith:2EGR5GLG

load-bearing objection The central claim holds up: the bare mirror/knockoff+ count threshold is not an FDR guarantee under PRDS, Gaussian equicorrelation, exchangeability, or pairwise uncorrelatedness, and the paper proves it with exact, self-contained counterexamples.

arxiv 2607.17084 v2 pith:2EGR5GLG submitted 2026-07-19 math.ST stat.MEstat.TH

Mirror and knockoff+ thresholds under dependence

classification math.ST stat.MEstat.TH MSC 62F0362H15
keywords false discovery ratemirror statisticsknockoff+PRDSdependencesign flipsmultiple testingexchangeability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper takes the mirror and knockoff+ thresholds—rules that reject when enough scores fall on the discovery side compared with the control side—and asks whether they still control the false discovery rate for dependent test statistics. It establishes that they do not, by constructing families of null distributions that satisfy the usual symmetry and dependence conditions (uniform margins, PRDS, Gaussianity, exchangeability, pairwise uncorrelatedness) yet drive the FDR above its nominal level. The central reason is that these thresholds adapt to the data by comparing only two current tail counts, which can all be elevated together by a shared latent factor. A sympathetic reader should care because these thresholds are used in practice outside the exact knockoff construction; the paper shows that the plus-one adjustment and symmetry alone are not protection. The paper is careful to note that valid knockoff statistics, which have conditional sign flips, are not contradicted.

Core claim

The paper's central claim is that the two-count mirror/knockoff+ rule is not an FDR guarantee when applied to generic dependent scores, even if each null marginal is symmetric or uniform. It proves the claim with exact counterexamples: a full-support PRDS family of uniform p-values whose FDR at q=0.1 is 17.4% and can approach 1/2; standard equicorrelated Gaussian null scores for which liminf_m FDR_m ≥ 1/2 for every fixed ρ>0; and an exchangeable, pairwise-uncorrelated symmetric construction with FDR arbitrarily close to 1. The paper also proves an impossibility result: for q<1/2, no deterministic monotone function of the two current tail counts can repair the threshold over the full-support

What carries the argument

The central objects are the two counting rules: the mirror threshold and the knockoff+ threshold, both of which compare control-side counts (L(u) or N_-(t)) against discovery-side counts (R(u) or N_+(t)). The proof mechanism that destroys validity is a lemma stating that if a block of all-null scores is positive together, the threshold is forced to pass, so the FDR equals the probability of such a block. Each counterexample builds a joint law that makes this block event likely while keeping the desired marginal or dependence properties: a latent Bernoulli mixture for the PRDS example, a common Gaussian factor for the equicorrelation result, and an exchangeable mixture LiZ + σε_i for the near

Load-bearing premise

The conclusions rest on applying the bare two-count threshold to dependent scores that do not have the conditional sign-flip property; if a valid fixed-X or model-X knockoff construction is actually used, the negative results do not apply.

What would settle it

Run the knockoff+ threshold at q=0.1 on m=5000 all-null standard Gaussian equicorrelated scores with ρ=0.05 and repeat 10,000 times; Theorem 1 predicts an empirical FDR near 0.11 and rising with m, so observing the FDR stably below 0.1 would be a direct contradiction.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • For any target level q<1/2, standard Gaussian null scores with a fixed positive equicorrelation ρ will eventually produce FDR at least 1/2 as m grows, regardless of how small ρ is.
  • The mirror threshold can fail within the PRDS family, a positive-dependence condition that is sufficient for some standard step-up procedures but not for this adaptive two-tail rule.
  • Exchangeability and pairwise uncorrelatedness do not imply that the threshold's signs align weakly; FDR can be arbitrarily close to one while every null marginal is the same continuous symmetric distribution.
  • No deterministic monotone rule based only on the two current tail counts can give a distribution-free FDR repair over the full-support PRDS class for q<1/2; power and validity cannot both be achieved without extra calibrated information.
  • The results leave valid knockoff+ theory intact: procedures that produce conditionally independent fair coin flip signs keep the FDR bound.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: the failure mechanism suggests that any two-count adaptive rule will be fragile whenever test statistics share a latent common factor, even one of small variance; applied users should screen for such factors before applying mirror-type thresholds.
  • Inference: the results point to a possible repair direction the paper leaves implicit: using the entire mirror process (all thresholds) or randomized thresholds could bypass the deterministic current-count limitation, at the cost of more complex theory.
  • Inference: the Gaussian equicorrelation theorem plausibly extends to other one-factor models with heavy-tailed factor loadings, where the liminf lower bound may be even closer to one; this is a testable extension.
  • Inference: for practitioners, the paper implies that the 'plus one' in knockoff+ is doing less protective work than is sometimes assumed when exchangeability is only approximate; correction factors from robust knockoff theory may need to be enforced rather than treated as negligible.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper studies the mirror and knockoff+ thresholds — procedures that compare discovery-side counts with control-side counts — when applied to dependent scores or p-values that lack the conditional sign-flip property of valid knockoff statistics. It constructs four negative results: (i) an exactly uniform, full-support PRDS p-value model at nominal q=0.1 with FDR 17.4% for m=11, and FDR arbitrarily close to 1/2 within the same family; (ii) standard Gaussian null scores with any fixed positive equicorrelation ρ have liminf FDR at least 1/2 as m→∞; (iii) exchangeable, pairwise-uncorrelated, marginally symmetric scores can have FDR arbitrarily close to 1; and (iv) for q<1/2, no deterministic monotone threshold based only on the two current tail counts can give a nontrivial distribution-free repair over the full-support PRDS class. The paper carefully distinguishes these failures from valid knockoff theory, which relies on conditional coordinatewise sign flips rather than marginal symmetry, PRDS, exchangeability, or pairwise uncorrelatedness.

Significance. If the results hold, this is a valuable clarification of the limitations of count-comparison FDR methods. The paper provides exact, self-contained counterexamples with explicit formulas, and its scope limitations are stated clearly. The construction of a full-support PRDS example with uniform margins is technically clean and directly challenges the intuition that PRDS suffices for adaptive two-tail comparisons as it does for Benjamini–Hochberg. The Gaussian equicorrelation result is striking: no matter how small ρ>0 is, the FDR lower bound is 1/2 in the limit. The impossibility result for monotone count-only corrections is a useful negative result for attempts to fix the threshold by simple modifications. The paper also correctly credits existing methods (valid knockoffs, data splitting, conditional calibration) that add the additional structure needed for validity. Overall, the manuscript is a substantive theoretical contribution with machine-checkable-style proofs: all derivations are explicit and no parameters are fitted to data.

minor comments (4)
  1. [Section 6, proof of Theorem 1] The notation 'let \Phi = 1 - \Phi' is self-referential and should read 'let \bar\Phi = 1 - \Phi'. This is a typographical issue and does not affect the argument.
  2. [Section 1.1] The displayed FDR calculation writes '0.910' and '0.110'; these are intended as powers 0.9^{10} and 0.1^{10}. Please fix the formatting and similarly in the exact formula (5).
  3. [Section 5] The sentence 'The corresponding scores have symmetric uniform margins and satisfy W d=-W' should read 'W \stackrel{d}{=} -W' for clarity.
  4. [General] There are frequent missing spaces in 'p-values' and 'pvalue' throughout the abstract and introduction; a copyedit pass would improve readability.

Circularity Check

0 steps flagged

No significant circularity: the counterexamples are self-contained existence proofs; only self-citations are contextual.

full rationale

The paper's derivation chain is self-contained. Theorem 1 rests on Mills-ratio tail bounds and a conditional law-of-large-numbers argument; Theorem 2 constructs an explicit finite mixture and verifies exchangeability, zero pairwise covariance, and full-support density; Proposition 3 verifies PRDS by direct monotonicity calculations; Proposition 4 is a dichotomy following from monotonicity of the count-only rule. No parameter is fitted to data and then called a prediction, and no equation equates the target FDR failure with an input by definition. The only self-citations ([9], joint mirror; [24], covariate-adaptive FDR) occur in the related-work survey and are not used as load-bearing support for the paper's negative results. Proposition 2's proof is omitted and cited to Barber and Candès and Candès et al., but that is independent external support and is not what the paper claims to establish. The paper also explicitly limits its scope to the bare mirror/knockoff+ threshold without the conditional sign-flip property, so the apparent scope restriction is stated rather than hidden.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 0 invented entities

The construction parameters (ε, π, a, b, σ, m) are legitimate degrees of freedom for existence proofs, not hidden fits; the paper's contribution is to show that within these families the threshold fails. The axioms are standard probabilistic tools and known external results. No new entities are introduced.

free parameters (6)
  • ε (mixture split in Proposition 3) = 0.1 in the m=11 example; ↓0 for the 1/2 limit
    Controls how much mass the two conditional densities put on each half; chosen small so the all-positive event has probability >q or near 1/2. This is a counterexample tuning parameter, not a data fit.
  • m (number of hypotheses) = 11 in Proposition 3; large in Theorem 2
    Required to satisfy m>1/q in Proposition 3 or to ensure K/m≈π in Theorem 2; chosen to make the block inequality (4) pass.
  • π (mixing probability in Theorem 2) = 0.05 in the numerical study
    Chosen so π/(1−π)<q, ensuring G_m has probability tending to 1; also makes E(L_i)=0, giving pairwise uncorrelated scores.
  • a and b=πa/(1−π) (scale constants in Theorem 2) = a=1, b=1/19 in the numerical study
    a>b gives separation: for z>0 the a-block dominates in magnitude, for z<0 the b-block does; b is determined by π and a.
  • σ (noise scale in Theorem 2) = 0.0005–0.2 in the numerical study; σ↓0 in the proof
    Small noise lets the deterministic block pattern determine signs and magnitudes; the theorem sends σ to 0.
  • ρ (equicorrelation in Theorem 1)
    Universal quantifier in the theorem (any fixed ρ>0); not fitted, but included because the lower bound depends on it.
axioms (5)
  • standard math Mills inequalities x/(1+x^2) φ(x) ≤ \barΦ(x) ≤ φ(x)/x for x>0.
    Used in Theorem 1 to bound the lower-to-upper tail ratio at a fixed threshold and to choose tδ.
  • standard math Law of large numbers and Fatou's lemma.
    Used in Theorem 1 to pass from conditional tail probabilities to rejection probability and in Theorem 2 for the FDR limit as σ↓0.
  • standard math PRDS sufficiency for one-sided Gaussian p-values with nonnegative correlations (Benjamini–Yekutieli, Sarkar).
    Used in Corollary 1 to assert that the Gaussian example is PRDS.
  • standard math Conditional sign-flip sufficiency for knockoff+ FDR control (Barber–Candès, Candès et al.).
    Used as the external benchmark that the counterexamples do not contradict; not used to derive any counterexample.
  • standard math Global-null identity FDR = P(any rejection).
    Simplifies all FDR computations under the global null; standard for a fixed null set.

pith-pipeline@v1.3.0-alltime-deepseek · 12395 in / 27874 out tokens · 256207 ms · 2026-08-01T19:05:15.967004+00:00 · methodology

0 comments
read the original abstract

Many multiple-testing methods compare the two sides of a null distribution to control the false discovery rate (FDR). The mirror and knockoff+ thresholds use large positive scores or small p-values as discoveries and the opposite tail as a control. For valid knockoff statistics, this comparison is justified because, conditional on their magnitudes and the nonnull information, the null signs are independent fair coin flips. We show what can go wrong without this property. First, we construct exactly uniform p-values satisfying positive regression dependence on a subset (PRDS), with a positive joint density on the whole unit cube. At nominal level 10%, an eleven-hypothesis example has FDR 17.4%, and the FDR can approach one half. Second, for standard Gaussian null scores with any fixed positive equicorrelation, the liminf of the FDR is at least one half. Third, we construct exchangeable, pairwise-uncorrelated scores with a common continuous symmetric distribution for which the FDR is arbitrarily close to one; transforming them by their common null distribution gives exactly uniform p-values with the same failure. Finally, for target levels q<1/2, no deterministic monotone threshold based only on the two current tail counts gives a nontrivial distribution-free repair over the full-support PRDS class: it either fails or never rejects. These results do not contradict knockoff theory. They show that adaptively comparing control-side and discovery-side counts is not, by itself, an FDR guarantee. Validity depends on the joint behavior of the null signs, not only on marginal symmetry, Gaussianity, PRDS, exchangeability, or pairwise uncorrelatedness.

Figures

Figures reproduced from arXiv: 2607.17084 by Xianyang Zhang.

Figure 1
Figure 1. Figure 1: Finite-sample behavior of the mirror/knockoff+ threshold under the global null, at [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 3 linked inside Pith

  1. [1]

    Barber, R. F. and Candès, E. J. (2015). Controlling the false discovery rate via knockoffs.Ann. Statist.43, 2055–2085. 16

  2. [2]

    F., Candès, E

    Barber, R. F., Candès, E. J. and Samworth, R. J. (2020). Robust inference with knockoffs.Ann. Statist.48, 1409–1431

  3. [3]

    and Hochberg, Y

    Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing.J. R. Statist. Soc. B57, 289–300

  4. [4]

    and Yekutieli, D

    Benjamini, Y. and Yekutieli, D. (2001). The control of the false discovery rate in multiple testing under dependency.Ann. Statist.29, 1165–1188

  5. [5]

    and Roquain, E

    Blanchard, G. and Roquain, E. (2009). Adaptive false discovery rate control under independence and dependence.J. Mach. Learn. Res.10, 2837–2871

  6. [6]

    J., Fan, Y., Janson, L

    Candès, E. J., Fan, Y., Janson, L. and Lv, J. (2018). Panning for gold: model-X knockoffs for high-dimensional controlled variable selection.J. R. Statist. Soc. B80, 551–577

  7. [7]

    and Liu, J

    Dai, C., Lin, B., Xing, X. and Liu, J. S. (2023a). False discovery rate control via data splitting.J. Am. Statist. Assoc.118, 2503–2520

  8. [8]

    and Liu, J

    Dai, C., Lin, B., Xing, X. and Liu, J. S. (2023b). A scale-free approach for false discovery rate control in generalized linear models.J. Am. Statist. Assoc.118, 1551–1565

  9. [9]

    and Zhang, X

    Deng, L., He, K. and Zhang, X. (2024). Joint mirror procedure: controlling false discovery rate for identifying simultaneous signals.Biometrics80, ujae142

  10. [10]

    Dobriban, E. (2026). The Benjamini–Hochberg procedure can fail to control the FDR for correlated two-sided Gaussian tests.arXiv:2607.12208

  11. [11]

    and Zou, C

    Du, L., Guo, X., Sun, W. and Zou, C. (2023). False discovery rate control under general dependence by symmetrized data aggregation.J. Am. Statist. Assoc.118, 607–621

  12. [12]

    and Roters, M

    Finner, H., Dickhaus, T. and Roters, M. (2007). Dependency and false discovery rate: asymptotics. Ann. Statist.35, 1432–1455

  13. [13]

    and Lei, L

    Fithian, W. and Lei, L. (2022). Conditional calibration for false discovery rate control under dependence.Ann. Statist.50, 3091–3118

  14. [14]

    and Liu, J

    Guo, Y., Lin, B. and Liu, J. S. (2026). PRADAS: prior-assisted data splitting for false discovery rate control.arXiv:2604.19517

  15. [15]

    and Janson, L

    Huang, D. and Janson, L. (2020). Relaxing the assumptions of knockoffs by conditioning.Ann. Statist.48, 3021–3042

  16. [16]

    and Fithian, W

    Lei, L. and Fithian, W. (2018). AdaPT: an interactive procedure for multiple testing with side information.J. R. Statist. Soc. B80, 649–679

  17. [17]

    and Fithian, W

    Lei, L., Ramdas, A. and Fithian, W. (2021). A general interactive framework for false discovery rate control under structural constraints.Biometrika108, 253–267

  18. [18]

    Sarkar, S. K. (2002). Some results on false discovery rate in stepwise multiple testing procedures. Ann. Statist.30, 239–257

  19. [19]

    Storey, J. D. (2002). A direct approach to false discovery rates.J. R. Statist. Soc. B64, 479–498

  20. [20]

    D., Taylor, J

    Storey, J. D., Taylor, J. E. and Siegmund, D. (2004). Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach.J. R. Statist. Soc. B66, 187–205. 17

  21. [21]

    Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference.J. Am. Statist. Assoc.22, 209–212

  22. [22]

    and Jiang, B

    Wu, Y., Yuan, P. and Jiang, B. (2026). Bi-Gaussian mirrors for false discovery rate control. arXiv:2604.24056

  23. [23]

    and Liu, J

    Xing, X., Zhao, Z. and Liu, J. S. (2023). Controlling false discovery rate using Gaussian mirrors.J. Am. Statist. Assoc.118, 222–241

  24. [24]

    and Chen, J

    Zhang, X. and Chen, J. (2022). Covariate adaptive false discovery rate control with applications to omics-wide multiple testing.J. Am. Statist. Assoc.117, 411–427. 18