{"id":"f47d4ecd-dad8-44c7-bb66-cbf5e839a7f8","arxiv_id":"2507.21465","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper derives finite-sample bounds on the false discovery rate of the Benjamini-Hochberg procedure for compound p-values: a constant-factor bound under independence, a near-optimal quadratic bound under the global null, and an O(log m) bound under positive dependence.","lead":"This paper proves that the Benjamini-Hochberg method, a standard tool for controlling errors across many simultaneous tests, still keeps the false discovery rate under control when individual test statistics are only valid on average across all tests, losing at most a constant factor. It also shows that if the tests are positively correlated, this safety guarantee degrades by a logarithmic factor, and it gives practical examples where the relaxed validity is useful.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 14's unproved numerical limit is the weakest load-bearing step: Theorem 1's advertised constant 1.93α depends on the assertion lim c_ℓ ≤ 1.9227.","rationale":"The reader's weakest-assumption assessment matches my own: the proof of Theorem 1 is complete up to the numerical assertion in Lemma 14. The qualitative O(α) result has independent support—Proposition 2 shows a constant-factor inflation is necessary, and Theorem 4 uses a different argument for the global null—so even a failure of the specific constant would not overturn the paper's main qualitative contribution; it would force a revision of the advertised 1.93. I checked the reduction to Lemmas 13 and 14 and found no further flaw: the use of positive association for independent variables, the Bernoulli-supremum bounds, and the Hoeffding applications all appear valid under the stated conditions. The lower-bound constructions in Propositions 2, 5, and 6 also satisfy the compound-p-value condition at the relevant thresholds and piecewise for intermediate values. The false identity in Lemma 24 is a genuine but isolated error: it is not used in the proof of Theorem 4, where only the second identity is invoked. The PRDS lower bound in Proposition 6 is plausible and does not change my assessment. Overall, the conditional verdict is appropriate and no verdict adjustment is needed.","tokens_in":32100,"tokens_out":32373,"duration_ms":347212,"concrete_test":"Write an independent script that evaluates the recurrence in Lemma 14 exactly as defined, using Lemma 23 to compute the supremum, with high-precision or interval arithmetic for ℓ = 1, . . . , 10^5, and either (a) confirms limsup c_ℓ ≤ 1.9227 with a rigorous tail bound, or (b) finds the first ℓ with c_ℓ > 1.9227. If (a) holds, the conditional concern is resolved; if (b) occurs, Theorem 1's constant must be revised or the proof modified. As a separate check, recompute the two identities in Lemma 24 and correct the first one, confirming that only the second identity is needed for Theorem 4.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is Theorem 1, whose proof funnels through Lemma 14. After reducing the FDP to a conditional expectation, the proof bounds each q̄_i by B_i(αk_i) and invokes Lemma 14. In that lemma, a nonlinear recurrence c_ℓ = c_{ℓ−1} + sup_t [P(Pois(t) ≥ ℓ−1) − c_{ℓ−1} P(Pois(t) ≥ ℓ)] is introduced, and the proof states without any supporting computation, code, interval bound, or tail estimate: 'We can verify numerically that lim_{\\ell\\to\\infty} c_\\ell \\le 1.9227.' If that numerical limit is wrong or not reproducible, the stated FDR ≤ 1.93α is not established by the given argument, even though a different constant might still yield O(α). This is not a stylistic quibble: the constant is the advertised quantitative content of Theorem 1, and the recurrence is nonlinear enough that ad hoc numerical verification is genuinely load-bearing. A separate, non-load-bearing blemish is that the first identity in Lemma 24 is numerically false (for t = 0.5 the right-hand side is 0 while the left-hand side already exceeds 0.393 from the k = 1 term); only the second identity is used in Theorem 4, so this is easily corrected and does not affect the main theorem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies false discovery rate control of the Benjamini–Hochberg (BH) procedure when the input p-values are compound p-values, i.e., they satisfy average superuniformity over the true nulls. Under independence, Theorem 1 gives the central bound FDR ≤ 1.93α, Theorem 3 extends this to approximate compound p-values, and Theorem 4 gives the sharper global-null bound FDR ≤ α + 2α². These upper bounds are complemented by lower bounds showing that a constant-factor inflation is unavoidable under independence (Proposition 2, FDR ≥ 7α/6) and that the global-null excess is at least α²/4 (Proposition 5). Proposition 6 shows that, under PRDS, the O(log m) upper bound of Armstrong cannot be improved beyond a constant factor. The paper also presents several motivating constructions of compound p-values (empirical Bayes, decreasing densities, permutation tests, Monte Carlo p-values, data alignment, Gaussian means) and an application to Upworthy headline data. Proofs are deferred to appendices and are largely detailed and self-contained.","tokens_in":32376,"tokens_out":9032,"duration_ms":104123,"significance":"If the main results are correct, this is a substantial contribution: it shows that independence, rather than positive dependence, is the assumption that rescues BH from the O(log m) inflation seen under arbitrary dependence for compound p-values, and it pins down the worst-case constant-factor inflation up to the factor 7/6 versus 1.93. The lower-bound constructions are explicit and checkable, and the examples and data analysis demonstrate practical relevance. The paper is also careful to benchmark against the existing results of Armstrong, Benjamini–Yekutieli, and Guo–Rao. The main weakness is that the advertised constant 1.93 in Theorem 1 rests on an unverified numerical claim in Lemma 14; until a reproducible numerical verification or a rigorous analytic bound is supplied, the exact constant is not established, although the qualitative O(α) conclusion is very likely sound.","major_comments":[{"comment":"The proof of Lemma 14 asserts 'We can verify numerically that lim_{\\ell\\to\\infty} c_\\ell \\le 1.9227' and provides no computation, code, certified interval bound, or tail estimate. This numerical assertion is the only step that converts the recursive inequality (23) into the constant 1.93 in Theorem 1, so it is load-bearing: without a reproducible verification (or an analytic bound replacing the numerical limit), the advertised constant is not proved, even though a qualitative O(α) bound might survive by a different route.","section":"Appendix A.2, Lemma 14"},{"comment":"The formal definition of the lower-bound construction says 'p1 ∼ Unif(1.5/m, 1)', but the subsequent compound-p-value check and the FDR calculation use the two-point distribution P(p1 = 1.5/m) = P(p1 = 1) = 1/2. Under the uniform distribution, the stated value of \\(\\sum_{i\\in H_0} P(p_i \\le 1.5/m)\\) would be 1 rather than 1.5, so the proof as written does not match the construction. Please correct the displayed definition to the two-point distribution that is actually used in the calculation.","section":"Appendix B.1, Proposition 2"}],"minor_comments":[{"comment":"The first identity in Lemma 24 is numerically false: for t = 0.5, the right-hand side equals 0 while the left-hand side is at least P(Pois(0.5) ≥ 1) > 0.393. The second identity is the one used in Theorem 4, so the main result is not affected, but the lemma should be corrected or the first identity removed.","section":"Appendix F.4, Lemma 24"},{"comment":"In the compound-p-value verification, the displayed sum is concluded to equal αℓ, but the computation actually yields α′ℓ. Since α′ ≤ α, the inequality still verifies the compound-p-value condition; please fix the displayed equality.","section":"Appendix B.3, Proposition 6"},{"comment":"The claim that the construction satisfies the PRDS condition (9) is asserted without proof. A short verification would be helpful, since this is a central lower-bound result and the conditioning event changes in a non-obvious way when t crosses the atom of p_i.","section":"Appendix B.3, Proposition 6"},{"comment":"The notation is confusing because U1(s1) is used both for the event {P_m j=1 1{U_j ≤ s_{1,j}} ≥ 1} and for the random variable U1. Please introduce distinct notation for the events, for instance E_i(s_i).","section":"Appendix A.4, Lemma 15"},{"comment":"The captions refer to the 'Uworthy' data set; the correct name is 'Upworthy'. There is also a duplicated article 'the' on page 7 ('estimate the the false discovery proportion') that should be corrected.","section":"Figures 1 and 2"}],"recommendation":"major_revision","confidential_remarks":"The numerical verification in Lemma 14 should be made available in a reproducible form (e.g., a short script or a rigorously certified interval computation) before publication; the reported constant 1.9227 is otherwise not checkable. I have no concerns about novelty or scope; the paper is a good fit and the central arguments are sound apart from the identified load-bearing numerical step and the typographical inconsistency in Proposition 2. I would be comfortable recommending acceptance after these points are resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core results here are real and new. For compound p-values under independence, Theorem 1's FDR ≤ 1.93α plus Proposition 2's 7/6α lower bound settles the qualitative picture: independence gives a constant-factor inflation, not the O(log m) blow-up seen under arbitrary dependence. The global-null bounds (α + 2α^2 upper, α + α^2/4 lower) are nice, and the PRDS lower bound is a useful negative result showing that positive dependence does not rescue BH in this setting. The examples are genuinely illustrative and the empirical section is appropriately treated as a demonstration, not a load-bearing claim.\n\nThe proofs are largely careful and self-contained. The leave-one-out reformulation, the use of Lemma 13 (positive association), and the Poisson tail comparisons all work. I also appreciate the honest benchmarking against Armstrong, Ignatiadis–Sen, and Ignatiadis et al.\n\nThe soft spots are minor relative to the overall structure, but they are real. Lemma 14 contains the line \"We can verify numerically that lim c_ℓ ≤ 1.9227\" with no code, no interval computation, no formal bound. That numerical limit is exactly what the 1.93α constant rests on. The recurrence is nonlinear enough that I would not call this a cosmetic issue; it is a gap in the proof as written. A reproducible script or a short interval-arithmetic argument would close it. The qualitative O(α) conclusion would likely survive, so this is fixable rather than fatal.\n\nAlso, the first identity in Lemma 24 is numerically false as stated (check t=0.5); only the second identity is used in Theorem 4, so the main global-null result stands. Still, the authors should correct it.\n\nWhere does this leave us? The paper is a serious contribution to multiple testing theory. I would send it to a knowledgeable referee, not desk reject it, but I would flag Lemma 14 and ask for a complete proof or a reproducible verification of the constant. If that step gets fixed, the paper is publishable in a strong statistics journal. Even if the constant turns out to be a bit larger, the qualitative independence result and the PRDS lower bound are valuable on their own.","headline":"Genuinely new results on FDR for compound p-values, with one unproved numerical step that should be fixed before the advertised constant is taken at face value.","tokens_in":32911,"tokens_out":1270,"would_cite":true,"duration_ms":17168,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62J15"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that independent compound p-values keep the BH procedure's false discovery rate at most $1.93\\alpha$, and that the constant cannot drop below $\\frac{7}{6}$.","keywords":["false discovery rate","Benjamini-Hochberg procedure","compound p-values","multiple testing","independence","positive dependence","global null","FDR control"],"falsifier":"Compute the sequence $c_\\ell$ from Lemma 14 to a large index, say $10^6$ or $10^7$ terms, and compare its limiting value with $1.9227$; if the limit exceeds $1.93$, the constant in Theorem 1 is false, while reproducing the limit confirms the lemma's numerical claim.","tokens_in":31888,"feed_emoji":"📊","tokens_out":10508,"duration_ms":105862,"temperature":0.7,"pith_summary":"Compound p-values relax the usual guarantee for p-values: instead of each true-null value being superuniform, only the average over the true nulls is required to be controlled. This paper asks what happens when the Benjamini–Hochberg (BH) procedure is run on such values. Its central result is that, under independence, the BH procedure keeps the false discovery rate at most $1.93\\alpha$, where $\\alpha$ is the nominal level, in contrast to the $O(\\log m)$ inflation that is unavoidable under arbitrary dependence. The paper also constructs worst-case distributions forcing FDR at least $\\frac{7}{6}\\alpha$, proves a sharper $\\alpha + 2\\alpha^2$ bound under the global null, and shows that positive dependence still permits the logarithmic blow-up. Several examples show when compound p-values arise naturally and can be more powerful than ordinary p-values.","feed_headline":"FDR below 1.93 alpha for independent compound p-values","feed_subtitle":"Independence replaces the log-factor FDR blow-up with a small constant, making compound p-values usable in large multiple-testing screens.","key_machinery":"The engine of the proof is a leave-one-out reformulation of the BH procedure (Lemma 12): for each possible number $i$ of rejected nulls, $k_i$ is the number of rejections the procedure would make if exactly $i$ true nulls were rejected, and $I$ is the actual number of rejected nulls, so the false discovery proportion is $I/(1\\vee k_I)$. Conditional on the non-null p-values, independence is used through a positive-correlation inequality for independent variables (Lemma 13) to bound the conditional probabilities, and through $B_i(t)$, the worst-case probability that a sum of independent Bernoulli variables with mean at most $t$ reaches $i$. The remaining task is a deterministic optimization problem, solved by the recursively defined sequence $c_1=1$, $c_2=1.5$, and $c_\\ell = c_{\\ell-1} + P\\{\\operatorname{Pois}((\\ell-1)/c_{\\ell-1}) \\ge \\ell-1\\} - c_{\\ell-1}P\\{\\operatorname{Pois}((\\ell-1)/c_{\\ell-1}) \\ge \\ell\\}$, whose limit is verified numerically to be at most $1.9227$. That constant is what converts the tail bounds into the $1.93\\alpha$ FDR guarantee.","core_discovery":"The central claim is that independence protects the BH procedure from the worst effects of compound p-values. Theorem 1 states that for independent compound p-values, the FDR of the BH procedure at level $\\alpha$ is at most $1.93\\alpha$ for every $m$ and every set of true nulls; Proposition 2 shows the constant is genuinely needed, with a distribution whose FDR is exactly $\\frac{7}{6}\\alpha$. Under the global null (all hypotheses true) the bound improves to $\\alpha + 2\\alpha^2$, with a matching lower bound of $\\alpha + \\alpha^2/4$. Under the PRDS positive-dependence condition, however, the paper constructs compound p-values whose FDR is at least $\\frac{3}{8}\\min\\{\\alpha h_m, 1\\}$, so the $O(\\log m)$ inflation cannot be avoided in that setting. The same techniques extend to approximate compound p-values, and the paper gives examples where compound p-values arise from empirical Bayes, decreasing-density, permutation, and Monte Carlo testing settings.","pith_inferences":["A natural next step, left open by the paper, is to optimize the deterministic problem in Lemma 14 directly; the true worst-case constant may well be closer to $\\frac{7}{6}$ than to $1.93$.","The conditional-independence structure in the permutation and Monte Carlo examples suggests a transferable recipe: if a pooled statistic is exchangeable after conditioning on empirical distributions, the resulting p-values become independent compound p-values, so the $1.93\\alpha$ bound applies in many designs beyond those listed.","The contrast with the e-value analogue of BH, which does allow level-$\\alpha$ FDR control even under dependence, indicates that the logarithmic penalty for compound p-values may be intrinsic to p-value aggregation rather than to the average-validity idea itself."],"forward_implications":["For any number $m$ of hypotheses, independent compound p-values can be fed into the BH procedure with only a constant-factor FDR inflation, so the harmonic-factor correction required under arbitrary dependence is unnecessary.","The worst-case constant lies between $\\frac{7}{6}$ and $1.93$, so exact FDR control at level $\\alpha$ is impossible for this procedure without extra assumptions on the marginal distributions.","Under the global null the excess FDR is $O(\\alpha^2)$, meaning that for small $\\alpha$ the BH procedure is nearly an exact level-$\\alpha$ test even though individual p-values are only valid on average.","The PRDS condition, which restores level-$\\alpha$ control for ordinary p-values, does not rescue compound p-values: FDR can still grow like $\\log m$, so independence is what is doing the real work.","Approximate compound p-values inherit the same type of guarantee, with the $1.93\\alpha$ bound inflated by a factor $1+\\epsilon$ when the approximation error is of the multiplicative type."],"supporting_citations":[{"why":"Defines the BH procedure and proves FDR control at level alpha for independent p-values; the benchmark that compound p-values are compared against.","marker":"Benjamini and Hochberg, 1995"},{"why":"Introduces the PRDS condition and the FDR <= alpha h_m bound for arbitrary dependence, and supplies the leave-one-out framework used in the proofs.","marker":"Benjamini and Yekutieli, 2001"},{"why":"Introduces compound p-values and establishes the FDR <= alpha h_m bound that the paper improves under independence.","marker":"Armstrong, 2022"},{"why":"Gives the construction showing alpha h_m is tight for arbitrary dependence, which is adapted in the PRDS lower bound.","marker":"Guo and Rao, 2008"},{"why":"Provides the Poisson tail bound for sums of independent Bernoulli variables that controls B_i(t) in Lemma 14 and the global-null bound.","marker":"Hoeffding, 1956"},{"why":"Develops the leave-one-out characterization of BH that underlies Lemma 12 and Theorem 20.","marker":"Ferreira and Zwinderman, 2006"}],"fun_headline_variants":["Independence keeps BH FDR below 1.93α for compound p-values","Independent compound p-values: BH FDR at most 1.93α","FDR inflated by log m under dependence, but independence caps at 1.93α","Compound p-values: independence ensures FDR ≤ 1.93α, unlike O(log m)"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument for the $1.93\\alpha$ bound relies on Lemma 14, whose final step states without a written derivation that the recursively defined sequence $c_\\ell$ has limit at most $1.9227$; if that numerical verification is wrong or unreproducible, the stated constant is not established by the proof as written.","fun_headline_variants_meta":{"raw":{"variants":["Independence keeps BH FDR below 1.93α for compound p-values","Independent compound p-values: BH FDR at most 1.93α","FDR inflated by log m under dependence, but independence caps at 1.93α","Compound p-values: independence ensures FDR ≤ 1.93α, unlike O(log m)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000628,"raw_usage":{"total_tokens":2908,"prompt_tokens":955,"completion_tokens":1953,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":1862}},"tokens_in":571,"tokens_out":1953,"duration_ms":16288,"temperature":1.0,"reasoning_tokens":1862,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:43:21.154519+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the sequence $c_\\ell$ from Lemma 14 to a large index, say $10^6$ or $10^7$ terms, and compare its limiting value with $1.9227$; if the limit exceeds $1.93$, the constant in Theorem 1 is false, while reproducing the limit confirms the lemma's numerical claim.","supporting_citations":[{"cited_title":"24 where the next-to-last step holds since hL ≥ hm/2","cited_arxiv_id":null,"evidence_quote":"Gives the construction showing alpha h_m is tight for arbitrary dependence, which is adapted in the PRDS lower bound."},{"cited_title":"So far, then, we have proved that ℓX i=1 i ti · ¯qi ℓY j=i+1 (1 − ¯qj) ≤ cℓ−1 + ℓ tℓ − cℓ−1 + · P {Pois(tℓ) ≥ ℓ}","cited_arxiv_id":null,"evidence_quote":"Provides the Poisson tail bound for sums of independent Bernoulli variables that controls B_i(t) in Lemma 14 and the global-null bound."}],"review_version":1}