REVIEW 3 major objections 4 minor 3 references
A grouped, selectively weighted false discovery rate procedure
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper proves that a two-stage, selectively weighted FDR procedure controls the false discovery rate when only some groups contain real effects, provided groups are fixed beforehand and null proportions are conservatively estimated.
desk verdict Useful refinement of grouped FDR control with solid theory for fixed partitions, but the real-data application breaks the paper's own assumptions; the gap is fixable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing identity is the oracle sGBH weight $v_j = \pi_{j0}(1-\tilde\pi_0)/(1-\pi_{j0})$ applied only to p-values in selected groups, together with the identity $(1-\tilde\pi_0)m_S = (1-\pi_0)m$, which makes the oracle sGBH and oracle GBH coincide under sparse configurations. The other central ingredient is a reciprocally conservative estimator of the proportion of true nulls, meaning $E[1/\hat\pi_{j0,k}] \le 1/\pi_{j0}$ when a true-null p-value is set to zero; this inequality is what lets the proof bound the FDR non-asymptotically. A KS uniformity test serves as the selection mechanism that recovers the set of interesting groups.
What would settle it
Simulate independent uniform p-values under the global null, form groups by the paper's application thresholds (p > 0.7, 0.15 < p < 0.7, p < 0.15), run the plug-in adaptive sGBH at level 0.05, and average the false discovery proportion; if it materially exceeds 0.05, the data-dependent grouping breaks the super-uniformity assumption that Theorem 2 requires.
Extended reading notes
Core claim
A grouped, selectively weighted procedure can preserve FDR control while testing only a subset of hypotheses. Under a sparse configuration, where some groups contain no false nulls, sGBH estimates the set of interesting groups, weights p-values only inside those groups, and applies BH to the weighted p-values. Theorem 2 shows that if each groupwise null proportion estimator is non-increasing and reciprocally conservative, or consistent, and if the group-selection step consistently recovers the interesting set, then the plug-in adaptive sGBH is non-asymptotically conservative up to an explicit additive error, and asymptotically has FDR no larger than the nominal level. The paper also shows that the oracle versions of sGBH and GBH make exactly the same rejections under a nontrivial sparse configuration, so the practical power gain comes from how the adaptive versions estimate the interesting set and the null proportions.
Load-bearing premise
The FDR guarantee assumes the group partition is fixed before the p-values are seen, but the paper's own application builds the groups by thresholding the same p-values used for testing, which makes true-null p-values in a group follow truncated uniforms that are not super-uniform.
Editorial extensions
If this is right
- With a fixed group partition, the plug-in adaptive sGBH controls FDR non-asymptotically up to an additive term that vanishes when the interesting set and the groupwise null proportions are consistently estimated.
- Under a nontrivial sparse configuration, the oracle sGBH and oracle GBH reject exactly the same hypotheses, so any power advantage in practice must come from estimation of the interesting set and proportions rather than from the oracle weighting scheme.
- The generic adaptive sGBH, using the weights of Nandi and Sarkar (2018), inherits non-asymptotic FDR control under independence and is nearly insensitive to its tuning parameter in simulations.
- In the prostate cancer application, the plug-in adaptive sGBH reports 759 differentially expressed genes at FDR level 0.05, compared with 487 for the plug-in adaptive GBH.
- As the total number of hypotheses grows, the power gap between the adaptive sGBH and the adaptive GBH shrinks, because the estimates of the interesting set and null proportions converge to their oracle values.
Reading between the lines
- A natural repair for the data-dependent grouping in the application is sample splitting: build the groups on one part of the data and test on the other, which would restore the theorem's requirement that the partition be fixed before seeing the tested p-values.
- The reciprocal-conservativeness mechanism is not tied to groups specifically; any two-stage procedure that selects hypotheses and then reweights their p-values must account for the selection step, and the same inequality could be adapted to selection based on effect-size estimates rather than uniformity tests.
- When groups are formed by p-value bins, true-null p-values in a bin follow truncated uniforms that are not super-uniform, so FDR inflation under the global null is expected; quantifying that inflation would provide a practical calibration for bin-based grouping.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a two-stage grouped FDR procedure (sGBH) for testing structured hypotheses. In the first stage, groups of hypotheses judged to contain signals (the "interesting" groups) are selected; in the second stage, only p-values in those groups are weighted by estimated groupwise null proportions and subjected to BH. The authors prove that the oracle sGBH coincides with the oracle GBH under sparse configurations (Theorem 1), give an FDR upper bound for the plug-in adaptive sGBH under conditions of reciprocal conservativeness and consistency (Theorem 2), provide a KS-based consistency lemma for selecting interesting groups (Lemma 1), and report simulations and a prostate cancer data application suggesting power gains over adaptive GBH.
Significance. The theoretical results are presented with explicit conditions rather than assuming the conclusion; Theorem 2's non-asymptotic bound honestly includes the costs of estimating S and the null proportions, and the proof strategy via reciprocal conservativeness is sound and generalizes earlier work. If the FDR guarantee held for the demonstrated application, the method would give practitioners a useful alternative to GBH when most groups are null. The potential power improvement is supported by simulations. However, the practical demonstration in Section 5 is not covered by the stated theory, and the simulation setting for the plug-in version with Jin's estimator is not fully covered either; these gaps substantially reduce the strength of the empirical claims.
major comments (3)
- [Section 5 (Application)] The application constructs the groups H1, H2, H3 by thresholding the same p-values that are later tested (p>0.7, 0.15<p<0.7, p<0.15), with H2 and H3 randomly subsampled. Section 3 explicitly assumes the partition {G_j} is known before seeing the p-values, and Section 2 assumes every null p-value is super-uniform. For a true null p-value in H2, the conditional distribution given 0.15<p<0.7 is F(t)=(t-0.15)/0.55, so F(0.6)=0.818>0.6; it is not super-uniform. Therefore the key inequality used in the proof of Theorem 2 (Appendix A.3, around Eq. (26)) does not hold for the procedure as run, and Proposition 1 also does not apply. The reported comparison of 759 vs 487 rejections is thus not backed by an FDR guarantee. The caution at the end of Section 5 acknowledges this, but the section still presents the improvement without a valid theoretical basis. Please either use a pre-specified partition based on external information, extend the theory to threshold-based data-dependent partitions, or reclassify the application as exploratory without FDR claims.
- [Section 4 (Simulation study)] The plug-in adaptive sGBH in the main simulations uses Jin's estimator to estimate groupwise null proportions, but Theorem 2's non-asymptotic bound requires each \hat{\pi}_{j0} to be non-increasing and reciprocally conservative; no such property is proved for Jin's estimator. The asymptotic part of Theorem 2 requires PRDS, and Section 4.1 notes that two-sided p-values "may not" satisfy PRDS. Consequently the simulation results with Jin's estimator and two-sided p-values are not covered by either arm of Theorem 2. The text in Section 4.2 stating that the adaptive sGBH "is conservative" is therefore an empirical observation, not a consequence of the proved theory; please state this distinction explicitly, or verify the required conditions in the simulations.
- [Section 3.1 and Section 5] The application uses Simes test to select \hat{S}, and the generic adaptive sGBH simulations of Section 4 use Simes at various \xi. Theorem 2's consistency condition for \hat{S} is supplied by Lemma 1 for the KS test only; the paper does not prove that Simes-based selection is consistent under the sparse configurations considered. If Simes is to be used in the main procedure, a consistency result (or a simulation-based demonstration that selection errors vanish) is needed to bring the application under Theorem 2.
minor comments (4)
- [Definition 2] The inequality \hat{\pi}^{\dagger}_{0,k} \le \hat{\pi}^{\dagger}_0 stated immediately after (5) is reversed for a non-increasing estimator: setting a p-value to 0 should not decrease the estimate. The correct inequality is \ge, and the proof of Theorem 2 actually relies on this corrected direction.
- [Lemma 1] Condition (8) is a limit of random quantities; please state the mode of convergence (in probability, almost surely) or replace it by a high-level assumption that the KS test is consistent, as the proof uses it to conclude Pr(\hat{S}=S)\to 1.
- [Section 5] The random sampling of 1500 genes from p-value intervals makes the partition random as well as threshold-based; even if the threshold-based grouping were pre-specified, the random subsampling adds another layer of data dependence that is not present in the theoretical setup.
- [References and supplementary material] There are small typographical errors: "mutliple" in the Benjamini and Yekutieli reference, and "there exits a constant" in Theorem 4 of the supplementary material.
Circularity Check
No circularity: sGBH's FDR theorems are conditional on explicit estimator and group-selection assumptions; the Section 5 gap is a stated limitation, not a circular derivation.
full rationale
The central claim, Theorem 2, is a conditional non-asymptotic FDR bound whose assumptions are explicit: independence or PRDS, a non-increasing and reciprocally conservative or consistent null-proportion estimator, and consistency of the estimated interesting-group set S. The proof in Appendix A.3 uses reciprocal conservativeness as an assumed property of the estimator and explicitly adds the selection-error terms Pr(hat S != S) and Pr(hat_pi0,S > check_pi0) to the bound. This is not a fitted input being renamed a prediction; it is a trade-off statement. The consistency results imported from prior work by the same authors (Chen 2018; Nandi and Sarkar 2018; Chen et al. 2017) are mathematical theorems with stated distributional and sparsity conditions, so they constitute independent support rather than circular self-citation. No uniqueness theorem or ansatz is smuggled in by citation. The Section 5 application constructs groups by thresholding the same p-values, which may violate the super-uniformity needed by Theorem 2, and the paper itself cautions: "the FDRs of the adaptive GBH and sGBH may exceed the specified normal level 0.05 due to a potential violation of the assumptions that ensure their non-asymptotic conservativeness." This is an honest limitation about assumption coverage, not a circular reduction: the reported procedure is not equivalent to its inputs by construction. Overall, the derivation chain is self-contained relative to its explicitly stated conditions, and no circular step is exhibited.
Assumptions & free parameters
free parameters (6)
- lambda (tuning parameter for generic weights and Storey estimator) =
0.25, 0.5, 0.75 (simulations); 0.5 (application figures)
- beta (KS test Type I error) =
0.025 (simulations); 0.1 (application check)
- xi (Simes test Type I error) =
0.05, 0.1, 0.2 (simulations); 0.1 (application)
- gamma in Jin's estimator =
0.5 (simulations); (delta - 0.0001)/2 with delta as PCS index (application)
- p-value thresholds for group construction in application =
0.7 and 0.15
- group sizes in application =
1374, 1500, 1500
assumptions (7)
- domain assumption Each null p-value is super-uniform and min p_i > 0 almost surely (Section 2, first paragraph).
- domain assumption The group partition {G_j} is known and fixed a priori (Section 3, first paragraph).
- domain assumption Sparse configuration: pi_j0 = 1 for j outside S (Definition 1).
- domain assumption P-values are mutually independent (Theorem 2) or satisfy PRDS (Theorem 2 Part II, Proposition 1).
- domain assumption Null proportion estimators are non-increasing and reciprocally conservative (Definition 2, Theorem 2).
- domain assumption Signal separation condition (8) in Lemma 1: the empirical process of false-null p-values in each interesting group diverges from uniform at rate sqrt(n_j)(1-pi_j0).
- domain assumption Jin's estimator consistency requires PCS with n^{-2}||S||_1 = O(n^{-delta}) and condition (12) on minimum signal size (Section 3.2).
Cite this review
Pith. "Pith review of A grouped, selectively weighted false discovery rate procedure." pith.science (2026). https://pith.science/paper/OK4IQWUT
@misc{pith2026190805319,
author = {Pith},
title = {Pith review of: A grouped, selectively weighted false discovery rate procedure},
year = {2026},
howpublished = {\url{https://pith.science/paper/OK4IQWUT}},
note = {Machine review of arXiv:1908.05319}
}
read the original abstract
False discovery rate (FDR) control in structured hypotheses testing is an important topic in simultaneous inference. Most existing methods that aim to utilize group structure among hypotheses either employ the groupwise mixture model or weight all p-values or hypotheses. Thus, their powers can be improved when the groupwise mixture model is inappropriate or when most groups contain only true null hypotheses. Motivated by this, we propose a grouped, selectively weighted FDR procedure, which we refer to as "sGBH". Specifically, without employing the groupwise mixture model, sGBH identifies groups of hypotheses of interest, weights p-values in each such group only, and tests only the selected hypotheses using the weighted p-values. The sGBH subsumes a standard grouped, weighted FDR procedure which we refer to as "GBH". We provide simple conditions to ensure the conservativeness of sGBH, together with empirical evidence on its much improved power over GBH. The new procedure is applied to a gene expression study.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Basu, P., Cai, T. T., Das, K. and Sun, W. (2018). Weighted false discovery rate control in large-scale multiple testing, J. Amer. Statist. Assoc. 113(523): 1172–1183. Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing, J. R. Statist. Soc. Ser. B 57(1): 289–300. Benjamini, Y. a...
work page Pith review arXiv 2018
-
[2]
LetCm =⋂ j∈S{ˆπj0 =πj0} andDm be the complement ofCm. Then, ˆαm≤ ∑ j∈S ∑ k∈Sj0 E [ 1 R1 { pjk πj0 (1− ˜π0) 1−πj0 ≤ Rα m }] + Pr (Dm) + Pr ( ˆS⁄=S ) ≤α + Pr (Dm) + Pr ( ˆS⁄=S ) , where the second inequality follows from the conservativeness of the oracle sGBH under PRDS. So, the claim holds. A.4 Proof of Lemma 1 For eachj∈{ 1,...,l }, let Fj be the empiric...
work page 1950
-
[3]
(33) If in addition ˆπ0,S consistently estimates ˜π0, then lim supm→∞ ˜αm≤ α
such that Pr (ˆπ0,S≤ ˇπ0)> 0 and ˜αm≤ α m 1 1− ˇπ0 ∑ j∈S nj (1−πj0) + α m ∑ j/∈S njπj0 + Pr (ˆπ0,S > ˇπ0). (33) If in addition ˆπ0,S consistently estimates ˜π0, then lim supm→∞ ˜αm≤ α. On the other hand, if{pi}m i=1 have the property of PRDS and ˆπj0 is consistent for πj0 uniformly in j∈ S (with necessarily being non-increasing or reciprocally conservativ...
work page 2017
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.