REVIEW 4 minor 24 references
The mirror and knockoff+ thresholds do not control the false discovery rate under dependence unless null signs can be flipped independently.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:05 UTC pith:2EGR5GLG
load-bearing objection The central claim holds up: the bare mirror/knockoff+ count threshold is not an FDR guarantee under PRDS, Gaussian equicorrelation, exchangeability, or pairwise uncorrelatedness, and the paper proves it with exact, self-contained counterexamples.
Mirror and knockoff+ thresholds under dependence
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the two-count mirror/knockoff+ rule is not an FDR guarantee when applied to generic dependent scores, even if each null marginal is symmetric or uniform. It proves the claim with exact counterexamples: a full-support PRDS family of uniform p-values whose FDR at q=0.1 is 17.4% and can approach 1/2; standard equicorrelated Gaussian null scores for which liminf_m FDR_m ≥ 1/2 for every fixed ρ>0; and an exchangeable, pairwise-uncorrelated symmetric construction with FDR arbitrarily close to 1. The paper also proves an impossibility result: for q<1/2, no deterministic monotone function of the two current tail counts can repair the threshold over the full-support
What carries the argument
The central objects are the two counting rules: the mirror threshold and the knockoff+ threshold, both of which compare control-side counts (L(u) or N_-(t)) against discovery-side counts (R(u) or N_+(t)). The proof mechanism that destroys validity is a lemma stating that if a block of all-null scores is positive together, the threshold is forced to pass, so the FDR equals the probability of such a block. Each counterexample builds a joint law that makes this block event likely while keeping the desired marginal or dependence properties: a latent Bernoulli mixture for the PRDS example, a common Gaussian factor for the equicorrelation result, and an exchangeable mixture LiZ + σε_i for the near
Load-bearing premise
The conclusions rest on applying the bare two-count threshold to dependent scores that do not have the conditional sign-flip property; if a valid fixed-X or model-X knockoff construction is actually used, the negative results do not apply.
What would settle it
Run the knockoff+ threshold at q=0.1 on m=5000 all-null standard Gaussian equicorrelated scores with ρ=0.05 and repeat 10,000 times; Theorem 1 predicts an empirical FDR near 0.11 and rising with m, so observing the FDR stably below 0.1 would be a direct contradiction.
If this is right
- For any target level q<1/2, standard Gaussian null scores with a fixed positive equicorrelation ρ will eventually produce FDR at least 1/2 as m grows, regardless of how small ρ is.
- The mirror threshold can fail within the PRDS family, a positive-dependence condition that is sufficient for some standard step-up procedures but not for this adaptive two-tail rule.
- Exchangeability and pairwise uncorrelatedness do not imply that the threshold's signs align weakly; FDR can be arbitrarily close to one while every null marginal is the same continuous symmetric distribution.
- No deterministic monotone rule based only on the two current tail counts can give a distribution-free FDR repair over the full-support PRDS class for q<1/2; power and validity cannot both be achieved without extra calibrated information.
- The results leave valid knockoff+ theory intact: procedures that produce conditionally independent fair coin flip signs keep the FDR bound.
Where Pith is reading between the lines
- Inference: the failure mechanism suggests that any two-count adaptive rule will be fragile whenever test statistics share a latent common factor, even one of small variance; applied users should screen for such factors before applying mirror-type thresholds.
- Inference: the results point to a possible repair direction the paper leaves implicit: using the entire mirror process (all thresholds) or randomized thresholds could bypass the deterministic current-count limitation, at the cost of more complex theory.
- Inference: the Gaussian equicorrelation theorem plausibly extends to other one-factor models with heavy-tailed factor loadings, where the liminf lower bound may be even closer to one; this is a testable extension.
- Inference: for practitioners, the paper implies that the 'plus one' in knockoff+ is doing less protective work than is sometimes assumed when exchangeability is only approximate; correction factors from robust knockoff theory may need to be enforced rather than treated as negligible.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the mirror and knockoff+ thresholds — procedures that compare discovery-side counts with control-side counts — when applied to dependent scores or p-values that lack the conditional sign-flip property of valid knockoff statistics. It constructs four negative results: (i) an exactly uniform, full-support PRDS p-value model at nominal q=0.1 with FDR 17.4% for m=11, and FDR arbitrarily close to 1/2 within the same family; (ii) standard Gaussian null scores with any fixed positive equicorrelation ρ have liminf FDR at least 1/2 as m→∞; (iii) exchangeable, pairwise-uncorrelated, marginally symmetric scores can have FDR arbitrarily close to 1; and (iv) for q<1/2, no deterministic monotone threshold based only on the two current tail counts can give a nontrivial distribution-free repair over the full-support PRDS class. The paper carefully distinguishes these failures from valid knockoff theory, which relies on conditional coordinatewise sign flips rather than marginal symmetry, PRDS, exchangeability, or pairwise uncorrelatedness.
Significance. If the results hold, this is a valuable clarification of the limitations of count-comparison FDR methods. The paper provides exact, self-contained counterexamples with explicit formulas, and its scope limitations are stated clearly. The construction of a full-support PRDS example with uniform margins is technically clean and directly challenges the intuition that PRDS suffices for adaptive two-tail comparisons as it does for Benjamini–Hochberg. The Gaussian equicorrelation result is striking: no matter how small ρ>0 is, the FDR lower bound is 1/2 in the limit. The impossibility result for monotone count-only corrections is a useful negative result for attempts to fix the threshold by simple modifications. The paper also correctly credits existing methods (valid knockoffs, data splitting, conditional calibration) that add the additional structure needed for validity. Overall, the manuscript is a substantive theoretical contribution with machine-checkable-style proofs: all derivations are explicit and no parameters are fitted to data.
minor comments (4)
- [Section 6, proof of Theorem 1] The notation 'let \Phi = 1 - \Phi' is self-referential and should read 'let \bar\Phi = 1 - \Phi'. This is a typographical issue and does not affect the argument.
- [Section 1.1] The displayed FDR calculation writes '0.910' and '0.110'; these are intended as powers 0.9^{10} and 0.1^{10}. Please fix the formatting and similarly in the exact formula (5).
- [Section 5] The sentence 'The corresponding scores have symmetric uniform margins and satisfy W d=-W' should read 'W \stackrel{d}{=} -W' for clarity.
- [General] There are frequent missing spaces in 'p-values' and 'pvalue' throughout the abstract and introduction; a copyedit pass would improve readability.
Circularity Check
No significant circularity: the counterexamples are self-contained existence proofs; only self-citations are contextual.
full rationale
The paper's derivation chain is self-contained. Theorem 1 rests on Mills-ratio tail bounds and a conditional law-of-large-numbers argument; Theorem 2 constructs an explicit finite mixture and verifies exchangeability, zero pairwise covariance, and full-support density; Proposition 3 verifies PRDS by direct monotonicity calculations; Proposition 4 is a dichotomy following from monotonicity of the count-only rule. No parameter is fitted to data and then called a prediction, and no equation equates the target FDR failure with an input by definition. The only self-citations ([9], joint mirror; [24], covariate-adaptive FDR) occur in the related-work survey and are not used as load-bearing support for the paper's negative results. Proposition 2's proof is omitted and cited to Barber and Candès and Candès et al., but that is independent external support and is not what the paper claims to establish. The paper also explicitly limits its scope to the bare mirror/knockoff+ threshold without the conditional sign-flip property, so the apparent scope restriction is stated rather than hidden.
Axiom & Free-Parameter Ledger
free parameters (6)
- ε (mixture split in Proposition 3) =
0.1 in the m=11 example; ↓0 for the 1/2 limit
- m (number of hypotheses) =
11 in Proposition 3; large in Theorem 2
- π (mixing probability in Theorem 2) =
0.05 in the numerical study
- a and b=πa/(1−π) (scale constants in Theorem 2) =
a=1, b=1/19 in the numerical study
- σ (noise scale in Theorem 2) =
0.0005–0.2 in the numerical study; σ↓0 in the proof
- ρ (equicorrelation in Theorem 1)
axioms (5)
- standard math Mills inequalities x/(1+x^2) φ(x) ≤ \barΦ(x) ≤ φ(x)/x for x>0.
- standard math Law of large numbers and Fatou's lemma.
- standard math PRDS sufficiency for one-sided Gaussian p-values with nonnegative correlations (Benjamini–Yekutieli, Sarkar).
- standard math Conditional sign-flip sufficiency for knockoff+ FDR control (Barber–Candès, Candès et al.).
- standard math Global-null identity FDR = P(any rejection).
read the original abstract
Many multiple-testing methods compare the two sides of a null distribution to control the false discovery rate (FDR). The mirror and knockoff+ thresholds use large positive scores or small p-values as discoveries and the opposite tail as a control. For valid knockoff statistics, this comparison is justified because, conditional on their magnitudes and the nonnull information, the null signs are independent fair coin flips. We show what can go wrong without this property. First, we construct exactly uniform p-values satisfying positive regression dependence on a subset (PRDS), with a positive joint density on the whole unit cube. At nominal level 10%, an eleven-hypothesis example has FDR 17.4%, and the FDR can approach one half. Second, for standard Gaussian null scores with any fixed positive equicorrelation, the liminf of the FDR is at least one half. Third, we construct exchangeable, pairwise-uncorrelated scores with a common continuous symmetric distribution for which the FDR is arbitrarily close to one; transforming them by their common null distribution gives exactly uniform p-values with the same failure. Finally, for target levels q<1/2, no deterministic monotone threshold based only on the two current tail counts gives a nontrivial distribution-free repair over the full-support PRDS class: it either fails or never rejects. These results do not contradict knockoff theory. They show that adaptively comparing control-side and discovery-side counts is not, by itself, an FDR guarantee. Validity depends on the joint behavior of the null signs, not only on marginal symmetry, Gaussianity, PRDS, exchangeability, or pairwise uncorrelatedness.
Figures
Reference graph
Works this paper leans on
-
[1]
Barber, R. F. and Candès, E. J. (2015). Controlling the false discovery rate via knockoffs.Ann. Statist.43, 2055–2085. 16
2015
-
[2]
F., Candès, E
Barber, R. F., Candès, E. J. and Samworth, R. J. (2020). Robust inference with knockoffs.Ann. Statist.48, 1409–1431
2020
-
[3]
and Hochberg, Y
Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing.J. R. Statist. Soc. B57, 289–300
1995
-
[4]
and Yekutieli, D
Benjamini, Y. and Yekutieli, D. (2001). The control of the false discovery rate in multiple testing under dependency.Ann. Statist.29, 1165–1188
2001
-
[5]
and Roquain, E
Blanchard, G. and Roquain, E. (2009). Adaptive false discovery rate control under independence and dependence.J. Mach. Learn. Res.10, 2837–2871
2009
-
[6]
J., Fan, Y., Janson, L
Candès, E. J., Fan, Y., Janson, L. and Lv, J. (2018). Panning for gold: model-X knockoffs for high-dimensional controlled variable selection.J. R. Statist. Soc. B80, 551–577
2018
-
[7]
and Liu, J
Dai, C., Lin, B., Xing, X. and Liu, J. S. (2023a). False discovery rate control via data splitting.J. Am. Statist. Assoc.118, 2503–2520
-
[8]
and Liu, J
Dai, C., Lin, B., Xing, X. and Liu, J. S. (2023b). A scale-free approach for false discovery rate control in generalized linear models.J. Am. Statist. Assoc.118, 1551–1565
-
[9]
and Zhang, X
Deng, L., He, K. and Zhang, X. (2024). Joint mirror procedure: controlling false discovery rate for identifying simultaneous signals.Biometrics80, ujae142
2024
-
[10]
Dobriban, E. (2026). The Benjamini–Hochberg procedure can fail to control the FDR for correlated two-sided Gaussian tests.arXiv:2607.12208
Pith/arXiv arXiv 2026
-
[11]
and Zou, C
Du, L., Guo, X., Sun, W. and Zou, C. (2023). False discovery rate control under general dependence by symmetrized data aggregation.J. Am. Statist. Assoc.118, 607–621
2023
-
[12]
and Roters, M
Finner, H., Dickhaus, T. and Roters, M. (2007). Dependency and false discovery rate: asymptotics. Ann. Statist.35, 1432–1455
2007
-
[13]
and Lei, L
Fithian, W. and Lei, L. (2022). Conditional calibration for false discovery rate control under dependence.Ann. Statist.50, 3091–3118
2022
-
[14]
Guo, Y., Lin, B. and Liu, J. S. (2026). PRADAS: prior-assisted data splitting for false discovery rate control.arXiv:2604.19517
Pith/arXiv arXiv 2026
-
[15]
and Janson, L
Huang, D. and Janson, L. (2020). Relaxing the assumptions of knockoffs by conditioning.Ann. Statist.48, 3021–3042
2020
-
[16]
and Fithian, W
Lei, L. and Fithian, W. (2018). AdaPT: an interactive procedure for multiple testing with side information.J. R. Statist. Soc. B80, 649–679
2018
-
[17]
and Fithian, W
Lei, L., Ramdas, A. and Fithian, W. (2021). A general interactive framework for false discovery rate control under structural constraints.Biometrika108, 253–267
2021
-
[18]
Sarkar, S. K. (2002). Some results on false discovery rate in stepwise multiple testing procedures. Ann. Statist.30, 239–257
2002
-
[19]
Storey, J. D. (2002). A direct approach to false discovery rates.J. R. Statist. Soc. B64, 479–498
2002
-
[20]
D., Taylor, J
Storey, J. D., Taylor, J. E. and Siegmund, D. (2004). Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach.J. R. Statist. Soc. B66, 187–205. 17
2004
-
[21]
Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference.J. Am. Statist. Assoc.22, 209–212
1927
-
[22]
Wu, Y., Yuan, P. and Jiang, B. (2026). Bi-Gaussian mirrors for false discovery rate control. arXiv:2604.24056
Pith/arXiv arXiv 2026
-
[23]
and Liu, J
Xing, X., Zhao, Z. and Liu, J. S. (2023). Controlling false discovery rate using Gaussian mirrors.J. Am. Statist. Assoc.118, 222–241
2023
-
[24]
and Chen, J
Zhang, X. and Chen, J. (2022). Covariate adaptive false discovery rate control with applications to omics-wide multiple testing.J. Am. Statist. Assoc.117, 411–427. 18
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.