REVIEW 3 major objections 8 minor 28 references
One randomization step unifies three families of FDR tests
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-09 16:12 UTC pith:HYJVAGUX
load-bearing objection Genuine unification of BH, conformal, and competition tests, but the broadest competition class is left as a conjecture the 3 major comments →
The Randomized BH Procedure: A Generalized Framework Encompassing Conformal and Competition Tests
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The randomized BH-valid condition replaces per-hypothesis distributional requirements with an average requirement over a uniformly random null hypothesis. Under this condition, the standard BH procedure controls FDR at level απ₀. The three embedding theorems show that classical BH-valid p-variables embed via stochastic domination (Theorem 13), conformal p-values embed because exchangeability implies identical marginals which satisfy the randomized PRDS (Theorem 12), and competition statistics embed after reformulation as weighted p-values on a random subset (Theorem 14). The unifying mechanism is that all three methods use null hypotheses as mutual evidence: conformal tests use calibration数据
What carries the argument
Definition 9 (Randomized BH-valid): For a uniform random null index J, require (1) P(P_J ≤ t) ≤ t (randomized p-variable) and (2) P(P_J^(·) ∈ C | P_J ≤ t) nondecreasing in t for nondecreasing sets C (randomized PRDS). Theorem 11 proves FDR ≤ απ₀ under these conditions. The framework connects to the three special cases through: stochastic domination (Corollary 1, Theorem 13), exchangeability-induced identical marginals (Lemma 2, Theorem 12), and competition-statistic reformulation as weighted BH on random subsets (Theorem 6, Theorem 14).
Load-bearing premise
The randomized PRDS condition requires that, after conditioning on a uniformly random null p-value being small, the remaining null p-values' order statistics are stochastically nondecreasing. This is an average condition that can fail when different null hypotheses have different marginal distributions, because the uniform index is no longer uniform after conditioning on being small—the conditioning reweights toward hypotheses with more mass near zero.
What would settle it
Construct null p-variables with heterogeneous marginal distributions where conditioning on P_J ≤ t shifts probability mass toward a subset of nulls whose remaining coordinates violate monotonicity of the order statistics. The paper's own example in Supplementary S6.1 (independent compound p-variables with P₁ ~ Unif(1.5/m, 1) and P₂ = 1/m) demonstrates exactly this failure: the randomized PRDS condition is violated, and FDR exceeds the nominal level.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the Randomized BH procedure, a generalized FDR-control framework that introduces randomization into the Benjamini-Hochberg procedure. The central contribution is a set of embedding theorems (Theorems 12-14) showing that the classical BH procedure, conformal tests, and competition tests all arise as special cases of this framework under a 'randomized BH-valid' condition (Definition 9), which is shown to be strictly weaker than classical BH-validity. Two applications are developed: a theoretical analysis of compound p-values explaining their FDR inflation, and a practical approach for integrating distributed multiple-testing problems with global FDR control. The theoretical development is careful, with detailed proofs in the supplementary materials.
Significance. The unification of conformal and competition methods under a common BH-based framework is a conceptually valuable contribution that clarifies the relationship between two independently developed families of null-distribution-free methods. The relaxation of exchangeability for conformal tests (Theorem 10) and the ability to integrate heterogeneous testing problems (Algorithms 1-2, Theorem 18) are practically significant. The analysis of compound p-values (Section 4.1, Proposition 1) provides a clean explanation for known FDR inflation phenomena. The framework is built on standard, well-established results (BH FDR control, weighted BH, competition statistics), and the embedding derivations are genuine rather than circular.
major comments (3)
- The unification claim for competition tests is incomplete. Theorem 14 rigorously embeds only 'competition statistics' (Definition 6, equality case: P{L_j=1|W,L_{-j}} = r/(1+r)). The more general 'sub-competition statistics' (Definition 6, inequality case) are not rigorously embedded. Supplementary S5.2 explicitly states: 'However, we do not give a rigorous proof for this claim, and we leave it as a heuristic conjecture.' Since key methods such as the Knockoff filter can rely on sub-competition statistics (the inequality form), the paper's claim to 'encompass competition tests' is overstated relative to what is proven. The authors should either (a) complete the proof for sub-competition statistics, (b) clearly state in the main text (not just the supplement) that the embedding covers only the equality case, or (c) restrict the unification claim accordingly. As written, the abstract and §3
- The relationship between the randomized PRDS condition (Definition 9) and existing dependence concepts needs sharper articulation. Definition 9 requires P(P_J^(·) ∈ C | P_J ≤ t) to be nondecreasing in t for a uniformly random null index J. The paper shows this holds under exchangeability (Theorem 12) and classical PRDS (Theorem 13), but the condition's practical verifiability for methods not already covered by these cases is unclear. For compound p-values (Section 4.1), the condition fails (Supplementary S6.1), which is a key insight. However, for the integration applications (Section 4.2), the paper assumes randomized p-variables satisfy PRDS (Theorem 17) without clarifying when this holds beyond the conformal and competition special cases already embedded. The authors should clarify what new practical scenarios the framework enables that are not already covered by the embedding theore.
- Theorem 19 introduces a correction function c(r) to handle arbitrary inter-group dependence in integrated competition tests, with c(1) = 0.5201. The derivation (Supplementary S7.4) bounds the FDR via a supremum over binomial tail probabilities, but the tightness of this bound and the resulting power loss are not assessed. Since c(1) ≈ 0.52 implies roughly halving the effective signal, the practical utility of this result is questionable. The paper should either provide simulation evidence for the multi-group case under dependence (current simulations in §5.2 only cover independent groups) or discuss the conservativeness of c(r) more explicitly.
minor comments (8)
- Section 3.1: The phrase 'approximate equivalence' in the section title is somewhat misleading. Theorems 5 and 6 establish exact equivalences between reformulated procedures, not approximations. The 'approximate' refers to the informal connection between conformal and competition methods, but the theorems themselves are exact. Clarifying this distinction would help readers.
- Definition 9: The notation P_J^(·) for order statistics of the vector P with the J-th component removed is introduced somewhat abruptly. Adding an explicit definition (e.g., 'P_J^(·) denotes the order statistics of P_{-J}') at the point of first use would improve readability.
- Theorem 19: The correction function c(r) is defined implicitly via an expectation involving a binomial supremum. Providing a closed-form approximation or a table of values for common r (beyond r=1) would help practitioners assess how conservative the bound is.
- Section 5.3: The comparison with e-value-based integration methods (MHTE and CPE) is somewhat one-sided. The paper notes that e-value transformations 'frequently fail to produce valid test results' but does not explore whether alternative e-value calibration strategies might address the identified tightness problem. A more balanced discussion or a reference to ongoing work on e-value integration would strengthen the section.
- Supplementary S5.2: The heuristic argument for embedding sub-competition statistics is difficult to follow. The construction of Z and Z' and the decomposition Z* = Z - Z ∘ Z' could benefit from more intuitive explanation of why this decomposition is expected to preserve the sub-competition property.
- Figure 2 (Section 5.2): The probability density plots of p-values transformed from competition statistics are mentioned but the figure itself appears to be referenced without sufficient detail in the caption. Adding axis labels and describing what the 'different groups' represent would make the figure self-contained.
- The paper uses 'randomized p-variables' (Definition 9, condition 1) and 'compound p-variables' (Definition 10) in closely related contexts. The relationship between these two concepts (noted briefly in §4.1: 'if P is a collection of compound p-variables... then Q = (P_1/π_0, ...) satisfies the condition of randomized p-variables') could be stated more prominently, as it connects the paper's framework to the recent compound p-value literature.
- Section 3.3, Corollary 1: The dominated PRDS condition is introduced and immediately applied to conformal tests with contaminated calibration sets. A brief remark on how this result relates to existing work on conformal robustness (e.g., Bates et al. 2023) would provide helpful context.
Circularity Check
No significant circularity: the framework is built on standard FDR control results and the embedding theorems are genuine derivations, not definitional restatements
full rationale
The paper's central claim is that classical BH, conformal tests, and competition tests are all special cases of the Randomized BH framework. Walking the derivation chain: (1) The randomized BH-valid condition (Definition 9) is defined independently—it requires a uniformly random null index J to satisfy a p-variable condition and a randomized PRDS condition. This is not defined in terms of the conclusion (FDR control). (2) Theorem 11 proves FDR ≤ απ₀ from the randomized BH-valid condition using a standard step-by-step argument (decomposing FDR, applying sub-uniformity, telescoping via the PRDS monotonicity on nested sets C_k). The proof does not assume FDR control to prove FDR control. (3) The embedding theorems (Theorems 12-14) show that existing procedures satisfy the randomized BH-valid condition. Theorem 12 shows identical marginals + PRDS implies randomized BH-valid by averaging over H₀—this is a genuine derivation, not a renaming. Theorem 13 constructs a dominating variable Q via probability integral transform and shows dominated PRDS—a real construction. Theorem 14 shows competition statistics yield randomized BH-valid p-variables via the conditional uniformity of V₊(0) (Lemma S4). (4) The integration algorithms (Algorithms 1-2) derive FDR control from Theorem 17, which itself follows from the randomized BH framework's properties. No step reduces to its own inputs by construction. The self-citations (e.g., to Ignatiadis et al. 2023 for weighted BH, Barber & Samworth 2025 for compound p-values) are to independent results that do not assume the present paper's conclusions. The one caveat noted by the skeptic—that sub-competition statistics embedding is conjectural (Supplementary S5.2)—is a gap in coverage, not a circularity: the paper explicitly acknowledges it as a heuristic conjecture rather than claiming a proven result. This is a correctness/completeness concern, not a circular derivation.
Axiom & Free-Parameter Ledger
free parameters (2)
- Pre-screening threshold τ =
0.5 (in simulations)
- Correction function c(r) =
c(1)=0.5201
axioms (5)
- standard math FDR control for the BH procedure under PRDS (Theorem 1)
- standard math FDR control for the eBH procedure with compound e-values (Theorem 2)
- standard math FDR control for the competition procedure (Theorem 3)
- standard math FDR control for weighted BH (Theorem 7)
- standard math FDR-Linking theorem
invented entities (2)
-
Randomized BH-valid p-variables
independent evidence
-
Dominated PRDS
independent evidence
read the original abstract
We propose the Randomized BH procedure, a generalized framework that introduces randomization into the Benjamini-Hochberg (BH) procedure for false discovery rate (FDR) control. We show that the classical BH procedure as well as two recently developed families of null-distribution-free methods, competition tests and conformal tests, share a common theoretical foundation and arise as special cases of this framework. The Randomized BH relaxes the classical conditions required for FDR control, frees conformal tests from the strict exchangeability assumption while allowing calibration sets to be constructed via pre-screening, and enables competition tests to be analyzed through the well-developed theory of p-values. We illustrate the framework through two applications. The first is a deeper analysis of compound p-values through the lens of randomization. The second is a general approach for integrating distributed multiple testing problems while maintaining global FDR control. Numerical simulations confirm the validity of our theoretical results.
Figures
Reference graph
Works this paper leans on
-
[1]
Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry , author =. Nature Methods , volume =. 2007 , publisher =
work page 2007
-
[2]
Multiple hypothesis testing methods for large-scale peptide identification in computational proteomics , author=. 2013 , school=
work page 2013
-
[3]
A theoretical foundation of the target-decoy search strategy for false discovery rate control in proteomics , author=. arXiv preprint arXiv:1501.00537 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[4]
Acta Mathematicae Applicatae Sinica, English Series , volume=
Null-free false discovery rate control using decoy permutations , author=. Acta Mathematicae Applicatae Sinica, English Series , volume=. 2022 , publisher=
work page 2022
-
[5]
Journal of the Royal statistical society: series B (Methodological) , volume=
Controlling the false discovery rate: a practical and powerful approach to multiple testing , author=. Journal of the Royal statistical society: series B (Methodological) , volume=. 1995 , publisher=
work page 1995
-
[6]
Benjamini, Yoav and Yekutieli, Daniel , journal =
-
[7]
The Annals of statistics , pages=
Controlling the false discovery rate via knockoffs , author=. The Annals of statistics , pages=. 2015 , publisher=
work page 2015
-
[8]
A Power and Prediction Analysis for Knockoffs with Lasso Statistics
A power and prediction analysis for knockoffs with lasso statistics , author=. arXiv preprint arXiv:1712.06465 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[9]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Panning for gold:‘model-X’knockoffs for high dimensional controlled variable selection , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2018 , publisher=
work page 2018
-
[10]
Journal of the American Statistical Association , volume=
False discovery rate control via data splitting , author=. Journal of the American Statistical Association , volume=. 2023 , publisher=
work page 2023
-
[11]
Journal of the American Statistical Association , volume=
False discovery rate control under general dependence by symmetrized data aggregation , author=. Journal of the American Statistical Association , volume=. 2023 , publisher=
work page 2023
-
[12]
Vovk, Vladimir and Shafer, Glenn and Gammerman, Alexander , month =. 2005 , doi =
work page 2005
-
[13]
Wang, Ruodu and Ramdas, Aaditya , journal =
-
[14]
Bates, Stephen and Candès, Emmanuel and Lei, Lihua and Romano, Yaniv and Sesia, Matteo , journal =. 2023 , doi =
work page 2023
- [15]
-
[16]
Ignatiadis, Nikolaos and Wang, Ruodu and Ramdas, Aaditya , journal =
-
[17]
False Discovery Rate Adjustments for Average Significance Level Controlling Tests , author=. 2025 , eprint=
work page 2025
-
[18]
False discovery rate control with compound p-values , author=. 2025 , eprint=
work page 2025
-
[19]
Asymptotic and compound e-values: multiple testing and empirical Bayes , author=. 2025 , eprint=
work page 2025
-
[20]
Ren, Zhimei and Barber, Rina Foygel , journal =
-
[21]
Block, Henry W. and Savits, Thomas H. and Wang, Jie , journal =. 2008 , doi =
work page 2008
- [22]
- [23]
-
[24]
E-values: Calibration, combination and applications , volume=
Vovk, Vladimir and Wang, Ruodu , year=. E-values: Calibration, combination and applications , volume=. The Annals of Statistics , publisher=. doi:10.1214/20-aos2020 , number=
-
[25]
Rina Foygel Barber and Emmanuel J. Cand. Bernoulli , number =. 2024 , doi =
work page 2024
- [26]
- [27]
- [28]
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.