{"id":"df9ccc0a-5f47-4148-91bf-f712b281203b","arxiv_id":"2607.07245","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"The classical BH, conformal, and competition testing procedures are shown to be special cases of a generalized Randomized BH framework that relaxes exchangeability and enables integration of distributed tests.","lead":"The paper unifies three multiple-testing methods—classical BH, conformal tests, and competition tests—under a single 'Randomized BH' framework by introducing randomization into the BH procedure. This matters because it relaxes restrictive assumptions (like exchangeability) and provides a principled way to combine heterogeneous testing results.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Sub-competition embedding is explicitly conjectural, creating a gap in the unification claim for the broadest class of competition methods.","rationale":"The reader correctly identifies the randomized PRDS condition as load-bearing and notes the sub-competition gap. I partially agree: the PRDS concern is real but the paper itself is transparent about it (Section 4.1, S6.1). The more acute issue is the sub-competition embedding gap, which the reader mentions but underweights. The unification claim is rigorously proven for classical BH (Theorem 13), conformal tests (Theorem 12), and exact competition statistics (Theorem 14). The sub-competition case is left as conjecture (S5.2), meaning the framework does not fully encompass all competition-based methods as claimed. This is a genuine gap but it is localized: the core theoretical machinery is sound, the embedding theorems for the three named special cases (BH, conformal, exact competition) are rigorously proven, and the applications (compound p-values, integrated tests) follow from the proven cases. The CONDITIONAL verdict is appropriate because the headline claim of encompassing 'competition tests' broadly is not fully substantiated, though the framework's theoretical contributions are substantial and largely valid. A reader should be aware that methods relying on sub-competition statistics (inequality form) are not yet rigorously covered.","tokens_in":48478,"tokens_out":563,"duration_ms":634063,"concrete_test":"Construct a concrete instance of sub-competition statistics where P{L_j=1|W,L_{-j}} < r/(1+r) strictly for some null j (not equality), and verify whether the randomized BH-valid condition (Definition 9) holds. If it fails, the unification claim does not extend to sub-competition statistics and the headline should be qualified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that competition tests are a special case of the Randomized BH framework (Theorem 14). However, Theorem 14 only covers 'competition statistics' where P{L_j=1|W,L_{-j}} = r/(1+r) exactly. The more general 'sub-competition statistics' (Definition 6), where this holds as an inequality, are NOT rigorously embedded. Supplementary S5.2 explicitly states: 'However, we do not give a rigorous proof for this claim, and we leave it as a heuristic conjecture.' Since key methods like the Knockoff filter rely on sub-competition statistics (the inequality form), the unification is incomplete for precisely the methods that motivate the competition framework. The paper's claim to 'encompass competition tests' is overstated relative to what is proven.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The paper proposes the Randomized BH procedure, a generalized FDR-control framework that introduces randomization into the Benjamini-Hochberg procedure. The central contribution is a set of embedding theorems (Theorems 12-14) showing that the classical BH procedure, conformal tests, and competition tests all arise as special cases of this framework under a 'randomized BH-valid' condition (Definition 9), which is shown to be strictly weaker than classical BH-validity. Two applications are developed: a theoretical analysis of compound p-values explaining their FDR inflation, and a practical approach for integrating distributed multiple-testing problems with global FDR control. The theoretical development is careful, with detailed proofs in the supplementary materials.","tokens_in":49123,"tokens_out":1626,"duration_ms":287094,"significance":"The unification of conformal and competition methods under a common BH-based framework is a conceptually valuable contribution that clarifies the relationship between two independently developed families of null-distribution-free methods. The relaxation of exchangeability for conformal tests (Theorem 10) and the ability to integrate heterogeneous testing problems (Algorithms 1-2, Theorem 18) are practically significant. The analysis of compound p-values (Section 4.1, Proposition 1) provides a clean explanation for known FDR inflation phenomena. The framework is built on standard, well-established results (BH FDR control, weighted BH, competition statistics), and the embedding derivations are genuine rather than circular.","major_comments":[{"comment":"The unification claim for competition tests is incomplete. Theorem 14 rigorously embeds only 'competition statistics' (Definition 6, equality case: P{L_j=1|W,L_{-j}} = r/(1+r)). The more general 'sub-competition statistics' (Definition 6, inequality case) are not rigorously embedded. Supplementary S5.2 explicitly states: 'However, we do not give a rigorous proof for this claim, and we leave it as a heuristic conjecture.' Since key methods such as the Knockoff filter can rely on sub-competition statistics (the inequality form), the paper's claim to 'encompass competition tests' is overstated relative to what is proven. The authors should either (a) complete the proof for sub-competition statistics, (b) clearly state in the main text (not just the supplement) that the embedding covers only the equality case, or (c) restrict the unification claim accordingly. As written, the abstract and §3","section":null},{"comment":"The relationship between the randomized PRDS condition (Definition 9) and existing dependence concepts needs sharper articulation. Definition 9 requires P(P_J^(·) ∈ C | P_J ≤ t) to be nondecreasing in t for a uniformly random null index J. The paper shows this holds under exchangeability (Theorem 12) and classical PRDS (Theorem 13), but the condition's practical verifiability for methods not already covered by these cases is unclear. For compound p-values (Section 4.1), the condition fails (Supplementary S6.1), which is a key insight. However, for the integration applications (Section 4.2), the paper assumes randomized p-variables satisfy PRDS (Theorem 17) without clarifying when this holds beyond the conformal and competition special cases already embedded. The authors should clarify what new practical scenarios the framework enables that are not already covered by the embedding theore.","section":null},{"comment":"Theorem 19 introduces a correction function c(r) to handle arbitrary inter-group dependence in integrated competition tests, with c(1) = 0.5201. The derivation (Supplementary S7.4) bounds the FDR via a supremum over binomial tail probabilities, but the tightness of this bound and the resulting power loss are not assessed. Since c(1) ≈ 0.52 implies roughly halving the effective signal, the practical utility of this result is questionable. The paper should either provide simulation evidence for the multi-group case under dependence (current simulations in §5.2 only cover independent groups) or discuss the conservativeness of c(r) more explicitly.","section":null}],"minor_comments":[{"comment":"Section 3.1: The phrase 'approximate equivalence' in the section title is somewhat misleading. Theorems 5 and 6 establish exact equivalences between reformulated procedures, not approximations. The 'approximate' refers to the informal connection between conformal and competition methods, but the theorems themselves are exact. Clarifying this distinction would help readers.","section":null},{"comment":"Definition 9: The notation P_J^(·) for order statistics of the vector P with the J-th component removed is introduced somewhat abruptly. Adding an explicit definition (e.g., 'P_J^(·) denotes the order statistics of P_{-J}') at the point of first use would improve readability.","section":null},{"comment":"Theorem 19: The correction function c(r) is defined implicitly via an expectation involving a binomial supremum. Providing a closed-form approximation or a table of values for common r (beyond r=1) would help practitioners assess how conservative the bound is.","section":null},{"comment":"Section 5.3: The comparison with e-value-based integration methods (MHTE and CPE) is somewhat one-sided. The paper notes that e-value transformations 'frequently fail to produce valid test results' but does not explore whether alternative e-value calibration strategies might address the identified tightness problem. A more balanced discussion or a reference to ongoing work on e-value integration would strengthen the section.","section":null},{"comment":"Supplementary S5.2: The heuristic argument for embedding sub-competition statistics is difficult to follow. The construction of Z and Z' and the decomposition Z* = Z - Z ∘ Z' could benefit from more intuitive explanation of why this decomposition is expected to preserve the sub-competition property.","section":null},{"comment":"Figure 2 (Section 5.2): The probability density plots of p-values transformed from competition statistics are mentioned but the figure itself appears to be referenced without sufficient detail in the caption. Adding axis labels and describing what the 'different groups' represent would make the figure self-contained.","section":null},{"comment":"The paper uses 'randomized p-variables' (Definition 9, condition 1) and 'compound p-variables' (Definition 10) in closely related contexts. The relationship between these two concepts (noted briefly in §4.1: 'if P is a collection of compound p-variables... then Q = (P_1/π_0, ...) satisfies the condition of randomized p-variables') could be stated more prominently, as it connects the paper's framework to the recent compound p-value literature.","section":null},{"comment":"Section 3.3, Corollary 1: The dominated PRDS condition is introduced and immediately applied to conformal tests with contaminated calibration sets. A brief remark on how this result relates to existing work on conformal robustness (e.g., Bates et al. 2023) would provide helpful context.","section":null}],"recommendation":"major_revision","confidential_remarks":"The core framework is sound and the embedding theorems for the equality case are rigorous. The main issue is the gap for sub-competition statistics (S5.2), which the authors themselves acknowledge as conjectural. Since the paper's central selling point is unification, and the unification is incomplete for the broadest class of competition methods, major revision is warranted. The authors may be able to close this gap or, alternatively, reframe the contribution to accurately reflect the scope of what is proven. The integration application is interesting but somewhat orthogonal to the main unification claim; if the sub-competition gap cannot be closed, the paper could potentially be repositioned around the conformal relaxation and integration results, with the competition embedding presented as a partial result. I note that the paper is dense and would benefit from careful editing for clarity, but the theoretical content is substantial."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper builds a single randomized BH framework and shows that classical BH, conformal tests, and competition tests all embed as special cases. That unification is real and non-trivial. The embedding theorems (Theorems 12–14) are carefully constructed, and the randomized BH-valid condition (Definition 9) is a genuine weakening of classical BH-validity—it replaces per-index requirements with average conditions over the null set, which is the right level of abstraction for this problem. The two applications—compound p-value analysis and integration of distributed tests—are well-motivated and the integration algorithms (Algorithms 1–2) come with FDR guarantees under independence. The compound p-value analysis (Section 4.1, Supplementary S6.1) is honest about where the framework breaks down, which I appreciate. The three proof methods for competition FDR control in Supplementary S2 are a nice pedagogical touch. The stress-test concern about sub-competition statistics lands. Theorem 14 covers competition statistics where P{L_j=1|W,L_{-j}} = r/(1+r) holds with equality. The more general sub-competition case (inequality form, Definition 6) is explicitly left as a heuristic conjecture in Supplementary S5.2—the authors say so directly. Since knockoff filters rely on the inequality form, the unification is incomplete for exactly the methods that motivate the competition framework. The abstract's claim to 'encompass competition tests' is slightly overstated relative to what is proven. That said, the equality case is the cleaner theoretical setting, and the conjecture is plausible. This is a gap in completeness, not a flaw in what is proven. The randomized PRDS condition is load-bearing, and the paper is upfront that it fails for compound p-values with heterogeneous marginals. The simulations are Gaussian-only and there is no code link, which limits reproducibility assessment. The correction function c(r) in Theorem 19 is defined implicitly via a supremum over binomial expectations—workable but not closed-form. Overall: the core unification is sound for the cases with rigorous proofs. The sub-competition gap should be acknowledged more prominently in the main text rather than buried in the supplement. This deserves a serious referee—it is a substantial theoretical contribution with a real but addressable gap.","headline":"Genuine unification of BH, conformal, and competition tests, but the broadest competition class is left as a conjecture","tokens_in":49031,"tokens_out":548,"would_cite":true,"duration_ms":146783,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"One randomization step unifies three families of FDR tests","keywords":[],"falsifier":"Construct null p-variables with heterogeneous marginal distributions where conditioning on P_J ≤ t shifts probability mass toward a subset of nulls whose remaining coordinates violate monotonicity of the order statistics. The paper's own example in Supplementary S6.1 (independent compound p-variables with P₁ ~ Unif(1.5/m, 1) and P₂ = 1/m) demonstrates exactly this failure: the randomized PRDS condition is violated, and FDR exceeds the nominal level.","tokens_in":48680,"feed_emoji":"🎲","tokens_out":807,"duration_ms":161147,"temperature":0.7,"pith_summary":"The paper proposes the Randomized BH procedure, which introduces a uniform random draw over null hypotheses into the Benjamini-Hochberg (BH) step-up procedure for false discovery rate control. The central claim is that three previously separate testing frameworks—classical BH (which needs known null distributions), conformal tests (which need exchangeability), and competition tests (which need paired pseudo-variables)—are all special cases of this single randomized framework. The key technical object is the 'randomized BH-valid' condition: instead of requiring every individual null p-value to satisfy distributional and dependence constraints, one requires only that a uniformly chosen null index satisfies them on average. This is strictly weaker than the classical condition. The paper proves the embedding theorems (Theorems 12–14) showing that classical BH-valid p-variables, exchangeable conformal p-values, and competition statistics all satisfy this relaxed condition. As applications, the framework explains why compound p-values can lose FDR control (the average condition breaks when marginals are heterogeneous) and provides algorithms for merging multiple independent testing problems into one global BH procedure while controlling FDR at the combined level.","feed_headline":"One randomization step unifies three families of FDR tests","feed_subtitle":"Classical BH, conformal, and competition methods all reduce to a single randomized procedure with weaker assumptions","key_machinery":"Definition 9 (Randomized BH-valid): For a uniform random null index J, require (1) P(P_J ≤ t) ≤ t (randomized p-variable) and (2) P(P_J^(·) ∈ C | P_J ≤ t) nondecreasing in t for nondecreasing sets C (randomized PRDS). Theorem 11 proves FDR ≤ απ₀ under these conditions. The framework connects to the three special cases through: stochastic domination (Corollary 1, Theorem 13), exchangeability-induced identical marginals (Lemma 2, Theorem 12), and competition-statistic reformulation as weighted BH on random subsets (Theorem 6, Theorem 14).","core_discovery":"The randomized BH-valid condition replaces per-hypothesis distributional requirements with an average requirement over a uniformly random null hypothesis. Under this condition, the standard BH procedure controls FDR at level απ₀. The three embedding theorems show that classical BH-valid p-variables embed via stochastic domination (Theorem 13), conformal p-values embed because exchangeability implies identical marginals which satisfy the randomized PRDS (Theorem 12), and competition statistics embed after reformulation as weighted p-values on a random subset (Theorem 14). The unifying mechanism is that all three methods use null hypotheses as mutual evidence: conformal tests use calibration数据","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["A single randomized BH framework subsumes classical BH, conformal, and competition tests","Randomization turns three distinct FDR methods into special cases of one BH procedure","BH, conformal, and competition tests all embed in a single randomized FDR framework","One randomized BH procedure covers classical, conformal, and competition testing","Relaxing BH assumptions via randomization unifies three null-distribution-free methods"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The randomized PRDS condition requires that, after conditioning on a uniformly random null p-value being small, the remaining null p-values' order statistics are stochastically nondecreasing. This is an average condition that can fail when different null hypotheses have different marginal distributions, because the uniform index is no longer uniform after conditioning on being small—the conditioning reweights toward hypotheses with more mass near zero.","fun_headline_variants_meta":{"raw":{"variants":["A single randomized BH framework subsumes classical BH, conformal, and competition tests","Randomization turns three distinct FDR methods into special cases of one BH procedure","BH, conformal, and competition tests all embed in a single randomized FDR framework","One randomized BH procedure covers classical, conformal, and competition testing","Relaxing BH assumptions via randomization unifies three null-distribution-free methods","Randomized BH replaces per-hypothesis conditions with an average requirement over nulls","Conformal tests relax exchangeability within a unified randomized BH framework","A randomized BH condition weakens assumptions while unifying three FDR test families","Three FDR testing traditions reduce to one randomized BH procedure","Compound p-values and distributed testing via a randomized BH lens"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":2370,"prompt_tokens":480,"completion_tokens":1890,"prompt_tokens_details":null},"tokens_in":480,"tokens_out":1890,"duration_ms":86556,"temperature":1.0,"reasoning_tokens":1436,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T16:12:00.414765+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Construct null p-variables with heterogeneous marginal distributions where conditioning on P_J ≤ t shifts probability mass toward a subset of nulls whose remaining coordinates violate monotonicity of the order statistics. The paper's own example in Supplementary S6.1 (independent compound p-variables with P₁ ~ Unif(1.5/m, 1) and P₂ = 1/m) demonstrates exactly this failure: the randomized PRDS condition is violated, and FDR exceeds the nominal level.","supporting_citations":[],"review_version":1}