Pith. sign in

REVIEW 3 major objections 8 minor 28 references

One randomization step unifies three families of FDR tests

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 16:12 UTC pith:HYJVAGUX

load-bearing objection Genuine unification of BH, conformal, and competition tests, but the broadest competition class is left as a conjecture the 3 major comments →

arxiv 2607.07245 v1 pith:HYJVAGUX submitted 2026-07-08 math.ST stat.TH

The Randomized BH Procedure: A Generalized Framework Encompassing Conformal and Competition Tests

classification math.ST stat.TH
keywords testsframeworkprocedurecompetitionconformalcontrolrandomizedclassical
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper proposes the Randomized BH procedure, which introduces a uniform random draw over null hypotheses into the Benjamini-Hochberg (BH) step-up procedure for false discovery rate control. The central claim is that three previously separate testing frameworks—classical BH (which needs known null distributions), conformal tests (which need exchangeability), and competition tests (which need paired pseudo-variables)—are all special cases of this single randomized framework. The key technical object is the 'randomized BH-valid' condition: instead of requiring every individual null p-value to satisfy distributional and dependence constraints, one requires only that a uniformly chosen null index satisfies them on average. This is strictly weaker than the classical condition. The paper proves the embedding theorems (Theorems 12–14) showing that classical BH-valid p-variables, exchangeable conformal p-values, and competition statistics all satisfy this relaxed condition. As applications, the framework explains why compound p-values can lose FDR control (the average condition breaks when marginals are heterogeneous) and provides algorithms for merging multiple independent testing problems into one global BH procedure while controlling FDR at the combined level.

Core claim

The randomized BH-valid condition replaces per-hypothesis distributional requirements with an average requirement over a uniformly random null hypothesis. Under this condition, the standard BH procedure controls FDR at level απ₀. The three embedding theorems show that classical BH-valid p-variables embed via stochastic domination (Theorem 13), conformal p-values embed because exchangeability implies identical marginals which satisfy the randomized PRDS (Theorem 12), and competition statistics embed after reformulation as weighted p-values on a random subset (Theorem 14). The unifying mechanism is that all three methods use null hypotheses as mutual evidence: conformal tests use calibration数据

What carries the argument

Definition 9 (Randomized BH-valid): For a uniform random null index J, require (1) P(P_J ≤ t) ≤ t (randomized p-variable) and (2) P(P_J^(·) ∈ C | P_J ≤ t) nondecreasing in t for nondecreasing sets C (randomized PRDS). Theorem 11 proves FDR ≤ απ₀ under these conditions. The framework connects to the three special cases through: stochastic domination (Corollary 1, Theorem 13), exchangeability-induced identical marginals (Lemma 2, Theorem 12), and competition-statistic reformulation as weighted BH on random subsets (Theorem 6, Theorem 14).

Load-bearing premise

The randomized PRDS condition requires that, after conditioning on a uniformly random null p-value being small, the remaining null p-values' order statistics are stochastically nondecreasing. This is an average condition that can fail when different null hypotheses have different marginal distributions, because the uniform index is no longer uniform after conditioning on being small—the conditioning reweights toward hypotheses with more mass near zero.

What would settle it

Construct null p-variables with heterogeneous marginal distributions where conditioning on P_J ≤ t shifts probability mass toward a subset of nulls whose remaining coordinates violate monotonicity of the order statistics. The paper's own example in Supplementary S6.1 (independent compound p-variables with P₁ ~ Unif(1.5/m, 1) and P₂ = 1/m) demonstrates exactly this failure: the randomized PRDS condition is violated, and FDR exceeds the nominal level.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The paper proposes the Randomized BH procedure, a generalized FDR-control framework that introduces randomization into the Benjamini-Hochberg procedure. The central contribution is a set of embedding theorems (Theorems 12-14) showing that the classical BH procedure, conformal tests, and competition tests all arise as special cases of this framework under a 'randomized BH-valid' condition (Definition 9), which is shown to be strictly weaker than classical BH-validity. Two applications are developed: a theoretical analysis of compound p-values explaining their FDR inflation, and a practical approach for integrating distributed multiple-testing problems with global FDR control. The theoretical development is careful, with detailed proofs in the supplementary materials.

Significance. The unification of conformal and competition methods under a common BH-based framework is a conceptually valuable contribution that clarifies the relationship between two independently developed families of null-distribution-free methods. The relaxation of exchangeability for conformal tests (Theorem 10) and the ability to integrate heterogeneous testing problems (Algorithms 1-2, Theorem 18) are practically significant. The analysis of compound p-values (Section 4.1, Proposition 1) provides a clean explanation for known FDR inflation phenomena. The framework is built on standard, well-established results (BH FDR control, weighted BH, competition statistics), and the embedding derivations are genuine rather than circular.

major comments (3)
  1. The unification claim for competition tests is incomplete. Theorem 14 rigorously embeds only 'competition statistics' (Definition 6, equality case: P{L_j=1|W,L_{-j}} = r/(1+r)). The more general 'sub-competition statistics' (Definition 6, inequality case) are not rigorously embedded. Supplementary S5.2 explicitly states: 'However, we do not give a rigorous proof for this claim, and we leave it as a heuristic conjecture.' Since key methods such as the Knockoff filter can rely on sub-competition statistics (the inequality form), the paper's claim to 'encompass competition tests' is overstated relative to what is proven. The authors should either (a) complete the proof for sub-competition statistics, (b) clearly state in the main text (not just the supplement) that the embedding covers only the equality case, or (c) restrict the unification claim accordingly. As written, the abstract and §3
  2. The relationship between the randomized PRDS condition (Definition 9) and existing dependence concepts needs sharper articulation. Definition 9 requires P(P_J^(·) ∈ C | P_J ≤ t) to be nondecreasing in t for a uniformly random null index J. The paper shows this holds under exchangeability (Theorem 12) and classical PRDS (Theorem 13), but the condition's practical verifiability for methods not already covered by these cases is unclear. For compound p-values (Section 4.1), the condition fails (Supplementary S6.1), which is a key insight. However, for the integration applications (Section 4.2), the paper assumes randomized p-variables satisfy PRDS (Theorem 17) without clarifying when this holds beyond the conformal and competition special cases already embedded. The authors should clarify what new practical scenarios the framework enables that are not already covered by the embedding theore.
  3. Theorem 19 introduces a correction function c(r) to handle arbitrary inter-group dependence in integrated competition tests, with c(1) = 0.5201. The derivation (Supplementary S7.4) bounds the FDR via a supremum over binomial tail probabilities, but the tightness of this bound and the resulting power loss are not assessed. Since c(1) ≈ 0.52 implies roughly halving the effective signal, the practical utility of this result is questionable. The paper should either provide simulation evidence for the multi-group case under dependence (current simulations in §5.2 only cover independent groups) or discuss the conservativeness of c(r) more explicitly.
minor comments (8)
  1. Section 3.1: The phrase 'approximate equivalence' in the section title is somewhat misleading. Theorems 5 and 6 establish exact equivalences between reformulated procedures, not approximations. The 'approximate' refers to the informal connection between conformal and competition methods, but the theorems themselves are exact. Clarifying this distinction would help readers.
  2. Definition 9: The notation P_J^(·) for order statistics of the vector P with the J-th component removed is introduced somewhat abruptly. Adding an explicit definition (e.g., 'P_J^(·) denotes the order statistics of P_{-J}') at the point of first use would improve readability.
  3. Theorem 19: The correction function c(r) is defined implicitly via an expectation involving a binomial supremum. Providing a closed-form approximation or a table of values for common r (beyond r=1) would help practitioners assess how conservative the bound is.
  4. Section 5.3: The comparison with e-value-based integration methods (MHTE and CPE) is somewhat one-sided. The paper notes that e-value transformations 'frequently fail to produce valid test results' but does not explore whether alternative e-value calibration strategies might address the identified tightness problem. A more balanced discussion or a reference to ongoing work on e-value integration would strengthen the section.
  5. Supplementary S5.2: The heuristic argument for embedding sub-competition statistics is difficult to follow. The construction of Z and Z' and the decomposition Z* = Z - Z ∘ Z' could benefit from more intuitive explanation of why this decomposition is expected to preserve the sub-competition property.
  6. Figure 2 (Section 5.2): The probability density plots of p-values transformed from competition statistics are mentioned but the figure itself appears to be referenced without sufficient detail in the caption. Adding axis labels and describing what the 'different groups' represent would make the figure self-contained.
  7. The paper uses 'randomized p-variables' (Definition 9, condition 1) and 'compound p-variables' (Definition 10) in closely related contexts. The relationship between these two concepts (noted briefly in §4.1: 'if P is a collection of compound p-variables... then Q = (P_1/π_0, ...) satisfies the condition of randomized p-variables') could be stated more prominently, as it connects the paper's framework to the recent compound p-value literature.
  8. Section 3.3, Corollary 1: The dominated PRDS condition is introduced and immediately applied to conformal tests with contaminated calibration sets. A brief remark on how this result relates to existing work on conformal robustness (e.g., Bates et al. 2023) would provide helpful context.

Circularity Check

0 steps flagged

No significant circularity: the framework is built on standard FDR control results and the embedding theorems are genuine derivations, not definitional restatements

full rationale

The paper's central claim is that classical BH, conformal tests, and competition tests are all special cases of the Randomized BH framework. Walking the derivation chain: (1) The randomized BH-valid condition (Definition 9) is defined independently—it requires a uniformly random null index J to satisfy a p-variable condition and a randomized PRDS condition. This is not defined in terms of the conclusion (FDR control). (2) Theorem 11 proves FDR ≤ απ₀ from the randomized BH-valid condition using a standard step-by-step argument (decomposing FDR, applying sub-uniformity, telescoping via the PRDS monotonicity on nested sets C_k). The proof does not assume FDR control to prove FDR control. (3) The embedding theorems (Theorems 12-14) show that existing procedures satisfy the randomized BH-valid condition. Theorem 12 shows identical marginals + PRDS implies randomized BH-valid by averaging over H₀—this is a genuine derivation, not a renaming. Theorem 13 constructs a dominating variable Q via probability integral transform and shows dominated PRDS—a real construction. Theorem 14 shows competition statistics yield randomized BH-valid p-variables via the conditional uniformity of V₊(0) (Lemma S4). (4) The integration algorithms (Algorithms 1-2) derive FDR control from Theorem 17, which itself follows from the randomized BH framework's properties. No step reduces to its own inputs by construction. The self-citations (e.g., to Ignatiadis et al. 2023 for weighted BH, Barber & Samworth 2025 for compound p-values) are to independent results that do not assume the present paper's conclusions. The one caveat noted by the skeptic—that sub-competition statistics embedding is conjectural (Supplementary S5.2)—is a gap in coverage, not a circularity: the paper explicitly acknowledges it as a heuristic conjecture rather than claiming a proven result. This is a correctness/completeness concern, not a circular derivation.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 2 invented entities

The framework introduces two key definitions (randomized BH-valid, dominated PRDS) that are the load-bearing concepts. The free parameters are simulation-specific or derived from theoretical bounds. The axioms are standard results from the multiple testing literature.

free parameters (2)
  • Pre-screening threshold τ = 0.5 (in simulations)
    Used in Section 5.1 to construct calibration sets via pre-screening; chosen for simulation, not theoretically derived.
  • Correction function c(r) = c(1)=0.5201
    Introduced in Theorem 19 for FDR control under arbitrary inter-group dependence; derived as the inverse of an expectation involving a supremum over binomial distributions.
axioms (5)
  • standard math FDR control for the BH procedure under PRDS (Theorem 1)
    Classical result from Benjamini & Yekutieli (2001); used as the foundation for generalization.
  • standard math FDR control for the eBH procedure with compound e-values (Theorem 2)
    From Wang & Ramdas (2022); used for comparison.
  • standard math FDR control for the competition procedure (Theorem 3)
    Prior result from the competition testing literature; the paper provides alternative proofs.
  • standard math FDR control for weighted BH (Theorem 7)
    From Ignatiadis et al. (2023); used as a building block for the generalized framework.
  • standard math FDR-Linking theorem
    From Su (2018); invoked in the proof of Theorem 16 for negative dependence.
invented entities (2)
  • Randomized BH-valid p-variables independent evidence
    purpose: Central definition replacing classical BH-validity with average conditions over a random null index
    The definition is falsifiable: if the randomized PRDS condition fails (as shown for compound p-values), FDR control is lost. The embedding theorems provide independent verification that existing procedures satisfy this condition.
  • Dominated PRDS independent evidence
    purpose: Relaxation allowing test statistics to be stochastically dominated by p-variables
    Corollary 1 provides a concrete falsifiable condition; the conformal BH with contaminated calibration sets is shown to satisfy it.

pith-pipeline@v1.1.0-glm · 51218 in / 2392 out tokens · 330242 ms · 2026-07-09T16:12:00.414765+00:00 · methodology

0 comments
read the original abstract

We propose the Randomized BH procedure, a generalized framework that introduces randomization into the Benjamini-Hochberg (BH) procedure for false discovery rate (FDR) control. We show that the classical BH procedure as well as two recently developed families of null-distribution-free methods, competition tests and conformal tests, share a common theoretical foundation and arise as special cases of this framework. The Randomized BH relaxes the classical conditions required for FDR control, frees conformal tests from the strict exchangeability assumption while allowing calibration sets to be constructed via pre-screening, and enables competition tests to be analyzed through the well-developed theory of p-values. We illustrate the framework through two applications. The first is a deeper analysis of compound p-values through the lens of randomization. The second is a general approach for integrating distributed multiple testing problems while maintaining global FDR control. Numerical simulations confirm the validity of our theoretical results.

Figures

Figures reproduced from arXiv: 2607.07245 by Mingzhou Deng, Yan Fu.

Figure 1
Figure 1. Figure 1: The boxplots of FDR and Power under different target control levels [PITH_FULL_IMAGE:figures/full_fig_p020_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Probability densities of the p-values transformed from competition statistics under different groups. [PITH_FULL_IMAGE:figures/full_fig_p021_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Integration test of two competition procedures. [PITH_FULL_IMAGE:figures/full_fig_p022_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Integration test of randomized BH procedure and competition procedure. [PITH_FULL_IMAGE:figures/full_fig_p022_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Integration test and e-value-based integration. [PITH_FULL_IMAGE:figures/full_fig_p023_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

28 extracted references · 28 canonical work pages · 2 internal anchors

  1. [1]

    Nature Methods , volume =

    Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry , author =. Nature Methods , volume =. 2007 , publisher =

  2. [2]

    2013 , school=

    Multiple hypothesis testing methods for large-scale peptide identification in computational proteomics , author=. 2013 , school=

  3. [3]

    A theoretical foundation of the target-decoy search strategy for false discovery rate control in proteomics

    A theoretical foundation of the target-decoy search strategy for false discovery rate control in proteomics , author=. arXiv preprint arXiv:1501.00537 , year=

  4. [4]

    Acta Mathematicae Applicatae Sinica, English Series , volume=

    Null-free false discovery rate control using decoy permutations , author=. Acta Mathematicae Applicatae Sinica, English Series , volume=. 2022 , publisher=

  5. [5]

    Journal of the Royal statistical society: series B (Methodological) , volume=

    Controlling the false discovery rate: a practical and powerful approach to multiple testing , author=. Journal of the Royal statistical society: series B (Methodological) , volume=. 1995 , publisher=

  6. [6]

    Benjamini, Yoav and Yekutieli, Daniel , journal =

  7. [7]

    The Annals of statistics , pages=

    Controlling the false discovery rate via knockoffs , author=. The Annals of statistics , pages=. 2015 , publisher=

  8. [8]

    A Power and Prediction Analysis for Knockoffs with Lasso Statistics

    A power and prediction analysis for knockoffs with lasso statistics , author=. arXiv preprint arXiv:1712.06465 , year=

  9. [9]

    Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=

    Panning for gold:‘model-X’knockoffs for high dimensional controlled variable selection , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2018 , publisher=

  10. [10]

    Journal of the American Statistical Association , volume=

    False discovery rate control via data splitting , author=. Journal of the American Statistical Association , volume=. 2023 , publisher=

  11. [11]

    Journal of the American Statistical Association , volume=

    False discovery rate control under general dependence by symmetrized data aggregation , author=. Journal of the American Statistical Association , volume=. 2023 , publisher=

  12. [12]

    2005 , doi =

    Vovk, Vladimir and Shafer, Glenn and Gammerman, Alexander , month =. 2005 , doi =

  13. [13]

    Wang, Ruodu and Ramdas, Aaditya , journal =

  14. [14]

    2023 , doi =

    Bates, Stephen and Candès, Emmanuel and Lei, Lihua and Romano, Yaniv and Sesia, Matteo , journal =. 2023 , doi =

  15. [15]

    2025 , doi =

    Ramdas, Aaditya and Wang, Ruodu , journal =. 2025 , doi =

  16. [16]

    Ignatiadis, Nikolaos and Wang, Ruodu and Ramdas, Aaditya , journal =

  17. [17]

    2025 , eprint=

    False Discovery Rate Adjustments for Average Significance Level Controlling Tests , author=. 2025 , eprint=

  18. [18]

    2025 , eprint=

    False discovery rate control with compound p-values , author=. 2025 , eprint=

  19. [19]

    2025 , eprint=

    Asymptotic and compound e-values: multiple testing and empirical Bayes , author=. 2025 , eprint=

  20. [20]

    Ren, Zhimei and Barber, Rina Foygel , journal =

  21. [21]

    and Savits, Thomas H

    Block, Henry W. and Savits, Thomas H. and Wang, Jie , journal =. 2008 , doi =

  22. [22]

    2024 , eprint=

    Multiple testing under negative dependence , author=. 2024 , eprint=

  23. [23]

    2018 , eprint=

    The FDR-Linking Theorem , author=. 2018 , eprint=

  24. [24]

    E-values: Calibration, combination and applications , volume=

    Vovk, Vladimir and Wang, Ruodu , year=. E-values: Calibration, combination and applications , volume=. The Annals of Statistics , publisher=. doi:10.1214/20-aos2020 , number=

  25. [25]

    Rina Foygel Barber and Emmanuel J. Cand. Bernoulli , number =. 2024 , doi =

  26. [26]

    , journal =

    Storey, John D. , journal =. 2002 , doi =

  27. [27]

    1997 , doi =

    Benjamini, Yoav and Hochberg, Yosef , journal =. 1997 , doi =

  28. [28]

    , journal =

    Sarkar, Sanat K. , journal =. 2002 , doi =