Pith. sign in

REVIEW 3 major objections 4 minor 3 references

A grouped, selectively weighted false discovery rate procedure

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper proves that a two-stage, selectively weighted FDR procedure controls the false discovery rate when only some groups contain real effects, provided groups are fixed beforehand and null proportions are conservatively estimated.

desk verdict Useful refinement of grouped FDR control with solid theory for fixed partitions, but the real-data application breaks the paper's own assumptions; the gap is fixable. read the letter →

arxiv 1908.05319 v1 pith:OK4IQWUT submitted 2019-08-14 stat.ME

classification stat.ME MSC 62J1562F0362F05
keywords falsediscoveryrategroupedhypothesesweightedp-valuesadaptiveproceduresparseconfigurationproportionoftruenullsmultipletestingreciprocalconservativeness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces sGBH, a two-stage procedure for false discovery rate control when hypotheses are grouped and most groups contain no real effects. Rather than weighting every p-value, as grouped BH procedures do, sGBH first screens groups with a uniformity test, then weights and tests only the p-values in the groups it declares interesting. The paper proves that, when the group partition is fixed before the p-values are seen and groupwise null proportions are estimated by reciprocally conservative estimators, the FDR is bounded by the nominal level plus the probability of mis-selecting groups or mis-estimating proportions, and tends asymptotically to no more than the nominal level. The practical payoff is power: in simulations and a prostate cancer gene-expression analysis, sGBH reports more discoveries than standard grouped BH at the same nominal FDR.

What carries the argument

The load-bearing identity is the oracle sGBH weight $v_j = \pi_{j0}(1-\tilde\pi_0)/(1-\pi_{j0})$ applied only to p-values in selected groups, together with the identity $(1-\tilde\pi_0)m_S = (1-\pi_0)m$, which makes the oracle sGBH and oracle GBH coincide under sparse configurations. The other central ingredient is a reciprocally conservative estimator of the proportion of true nulls, meaning $E[1/\hat\pi_{j0,k}] \le 1/\pi_{j0}$ when a true-null p-value is set to zero; this inequality is what lets the proof bound the FDR non-asymptotically. A KS uniformity test serves as the selection mechanism that recovers the set of interesting groups.

What would settle it

Simulate independent uniform p-values under the global null, form groups by the paper's application thresholds (p > 0.7, 0.15 < p < 0.7, p < 0.15), run the plug-in adaptive sGBH at level 0.05, and average the false discovery proportion; if it materially exceeds 0.05, the data-dependent grouping breaks the super-uniformity assumption that Theorem 2 requires.

Watch

Extended reading notes

Core claim

A grouped, selectively weighted procedure can preserve FDR control while testing only a subset of hypotheses. Under a sparse configuration, where some groups contain no false nulls, sGBH estimates the set of interesting groups, weights p-values only inside those groups, and applies BH to the weighted p-values. Theorem 2 shows that if each groupwise null proportion estimator is non-increasing and reciprocally conservative, or consistent, and if the group-selection step consistently recovers the interesting set, then the plug-in adaptive sGBH is non-asymptotically conservative up to an explicit additive error, and asymptotically has FDR no larger than the nominal level. The paper also shows that the oracle versions of sGBH and GBH make exactly the same rejections under a nontrivial sparse configuration, so the practical power gain comes from how the adaptive versions estimate the interesting set and the null proportions.

Load-bearing premise

The FDR guarantee assumes the group partition is fixed before the p-values are seen, but the paper's own application builds the groups by thresholding the same p-values used for testing, which makes true-null p-values in a group follow truncated uniforms that are not super-uniform.

Editorial extensions

If this is right

  • With a fixed group partition, the plug-in adaptive sGBH controls FDR non-asymptotically up to an additive term that vanishes when the interesting set and the groupwise null proportions are consistently estimated.
  • Under a nontrivial sparse configuration, the oracle sGBH and oracle GBH reject exactly the same hypotheses, so any power advantage in practice must come from estimation of the interesting set and proportions rather than from the oracle weighting scheme.
  • The generic adaptive sGBH, using the weights of Nandi and Sarkar (2018), inherits non-asymptotic FDR control under independence and is nearly insensitive to its tuning parameter in simulations.
  • In the prostate cancer application, the plug-in adaptive sGBH reports 759 differentially expressed genes at FDR level 0.05, compared with 487 for the plug-in adaptive GBH.
  • As the total number of hypotheses grows, the power gap between the adaptive sGBH and the adaptive GBH shrinks, because the estimates of the interesting set and null proportions converge to their oracle values.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural repair for the data-dependent grouping in the application is sample splitting: build the groups on one part of the data and test on the other, which would restore the theorem's requirement that the partition be fixed before seeing the tested p-values.
  • The reciprocal-conservativeness mechanism is not tied to groups specifically; any two-stage procedure that selects hypotheses and then reweights their p-values must account for the selection step, and the same inequality could be adapted to selection based on effect-size estimates rather than uniformity tests.
  • When groups are formed by p-value bins, true-null p-values in a bin follow truncated uniforms that are not super-uniform, so FDR inflation under the global null is expected; quantifying that inflation would provide a practical calibration for bin-based grouping.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper proposes a two-stage grouped FDR procedure (sGBH) for testing structured hypotheses. In the first stage, groups of hypotheses judged to contain signals (the "interesting" groups) are selected; in the second stage, only p-values in those groups are weighted by estimated groupwise null proportions and subjected to BH. The authors prove that the oracle sGBH coincides with the oracle GBH under sparse configurations (Theorem 1), give an FDR upper bound for the plug-in adaptive sGBH under conditions of reciprocal conservativeness and consistency (Theorem 2), provide a KS-based consistency lemma for selecting interesting groups (Lemma 1), and report simulations and a prostate cancer data application suggesting power gains over adaptive GBH.

Significance. The theoretical results are presented with explicit conditions rather than assuming the conclusion; Theorem 2's non-asymptotic bound honestly includes the costs of estimating S and the null proportions, and the proof strategy via reciprocal conservativeness is sound and generalizes earlier work. If the FDR guarantee held for the demonstrated application, the method would give practitioners a useful alternative to GBH when most groups are null. The potential power improvement is supported by simulations. However, the practical demonstration in Section 5 is not covered by the stated theory, and the simulation setting for the plug-in version with Jin's estimator is not fully covered either; these gaps substantially reduce the strength of the empirical claims.

major comments (3)
  1. [Section 5 (Application)] The application constructs the groups H1, H2, H3 by thresholding the same p-values that are later tested (p>0.7, 0.15<p<0.7, p<0.15), with H2 and H3 randomly subsampled. Section 3 explicitly assumes the partition {G_j} is known before seeing the p-values, and Section 2 assumes every null p-value is super-uniform. For a true null p-value in H2, the conditional distribution given 0.15<p<0.7 is F(t)=(t-0.15)/0.55, so F(0.6)=0.818>0.6; it is not super-uniform. Therefore the key inequality used in the proof of Theorem 2 (Appendix A.3, around Eq. (26)) does not hold for the procedure as run, and Proposition 1 also does not apply. The reported comparison of 759 vs 487 rejections is thus not backed by an FDR guarantee. The caution at the end of Section 5 acknowledges this, but the section still presents the improvement without a valid theoretical basis. Please either use a pre-specified partition based on external information, extend the theory to threshold-based data-dependent partitions, or reclassify the application as exploratory without FDR claims.
  2. [Section 4 (Simulation study)] The plug-in adaptive sGBH in the main simulations uses Jin's estimator to estimate groupwise null proportions, but Theorem 2's non-asymptotic bound requires each \hat{\pi}_{j0} to be non-increasing and reciprocally conservative; no such property is proved for Jin's estimator. The asymptotic part of Theorem 2 requires PRDS, and Section 4.1 notes that two-sided p-values "may not" satisfy PRDS. Consequently the simulation results with Jin's estimator and two-sided p-values are not covered by either arm of Theorem 2. The text in Section 4.2 stating that the adaptive sGBH "is conservative" is therefore an empirical observation, not a consequence of the proved theory; please state this distinction explicitly, or verify the required conditions in the simulations.
  3. [Section 3.1 and Section 5] The application uses Simes test to select \hat{S}, and the generic adaptive sGBH simulations of Section 4 use Simes at various \xi. Theorem 2's consistency condition for \hat{S} is supplied by Lemma 1 for the KS test only; the paper does not prove that Simes-based selection is consistent under the sparse configurations considered. If Simes is to be used in the main procedure, a consistency result (or a simulation-based demonstration that selection errors vanish) is needed to bring the application under Theorem 2.
minor comments (4)
  1. [Definition 2] The inequality \hat{\pi}^{\dagger}_{0,k} \le \hat{\pi}^{\dagger}_0 stated immediately after (5) is reversed for a non-increasing estimator: setting a p-value to 0 should not decrease the estimate. The correct inequality is \ge, and the proof of Theorem 2 actually relies on this corrected direction.
  2. [Lemma 1] Condition (8) is a limit of random quantities; please state the mode of convergence (in probability, almost surely) or replace it by a high-level assumption that the KS test is consistent, as the proof uses it to conclude Pr(\hat{S}=S)\to 1.
  3. [Section 5] The random sampling of 1500 genes from p-value intervals makes the partition random as well as threshold-based; even if the threshold-based grouping were pre-specified, the random subsampling adds another layer of data dependence that is not present in the theoretical setup.
  4. [References and supplementary material] There are small typographical errors: "mutliple" in the Benjamini and Yekutieli reference, and "there exits a constant" in Theorem 4 of the supplementary material.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: sGBH's FDR theorems are conditional on explicit estimator and group-selection assumptions; the Section 5 gap is a stated limitation, not a circular derivation.

full rationale

The central claim, Theorem 2, is a conditional non-asymptotic FDR bound whose assumptions are explicit: independence or PRDS, a non-increasing and reciprocally conservative or consistent null-proportion estimator, and consistency of the estimated interesting-group set S. The proof in Appendix A.3 uses reciprocal conservativeness as an assumed property of the estimator and explicitly adds the selection-error terms Pr(hat S != S) and Pr(hat_pi0,S > check_pi0) to the bound. This is not a fitted input being renamed a prediction; it is a trade-off statement. The consistency results imported from prior work by the same authors (Chen 2018; Nandi and Sarkar 2018; Chen et al. 2017) are mathematical theorems with stated distributional and sparsity conditions, so they constitute independent support rather than circular self-citation. No uniqueness theorem or ansatz is smuggled in by citation. The Section 5 application constructs groups by thresholding the same p-values, which may violate the super-uniformity needed by Theorem 2, and the paper itself cautions: "the FDRs of the adaptive GBH and sGBH may exceed the specified normal level 0.05 due to a potential violation of the assumptions that ensure their non-asymptotic conservativeness." This is an honest limitation about assumption coverage, not a circular reduction: the reported procedure is not equivalent to its inputs by construction. Overall, the derivation chain is self-contained relative to its explicitly stated conditions, and no circular step is exhibited.

Assumptions & free parameters 6 free parameters · 7 assumptions · 0 invented entities

The procedure's theoretical guarantees rest on standard multiple-testing assumptions (super-uniform p-values, dependency conditions) and on the availability of estimators with specific properties (reciprocal conservativeness, consistency). The application adds free grouping parameters that are not covered by the theory. No new entities are introduced.

free parameters (6)
  • lambda (tuning parameter for generic weights and Storey estimator) = 0.25, 0.5, 0.75 (simulations); 0.5 (application figures)
    User-chosen; FDR and power vary little across values (Figure 5).
  • beta (KS test Type I error) = 0.025 (simulations); 0.1 (application check)
    Hand-selected; controls the probability of misestimating the set of interesting groups.
  • xi (Simes test Type I error) = 0.05, 0.1, 0.2 (simulations); 0.1 (application)
    Hand-selected; smaller xi reduces false inclusion of null groups.
  • gamma in Jin's estimator = 0.5 (simulations); (delta - 0.0001)/2 with delta as PCS index (application)
    Must lie in (0, delta/2]; trades off consistency rate against variance.
  • p-value thresholds for group construction in application = 0.7 and 0.15
    Data-dependent thresholds used to partition genes into H1/H2/H3; not part of the theory, which assumes a fixed partition.
  • group sizes in application = 1374, 1500, 1500
    The number of genes in each group after thresholding and random sampling; arbitrary choices affecting results.
assumptions (7)
  • domain assumption Each null p-value is super-uniform and min p_i > 0 almost surely (Section 2, first paragraph).
    Required for the BH and weighting arguments; the application's truncated p-values violate super-uniformity.
  • domain assumption The group partition {G_j} is known and fixed a priori (Section 3, first paragraph).
    The theory conditions on the partition; the application groups genes by p-value thresholds, which is data-dependent.
  • domain assumption Sparse configuration: pi_j0 = 1 for j outside S (Definition 1).
    Central setting for the main theorems and the claimed power gain.
  • domain assumption P-values are mutually independent (Theorem 2) or satisfy PRDS (Theorem 2 Part II, Proposition 1).
    Standard dependency conditions for FDR control.
  • domain assumption Null proportion estimators are non-increasing and reciprocally conservative (Definition 2, Theorem 2).
    The non-asymptotic FDR bound depends on these properties; Storey's estimator has them.
  • domain assumption Signal separation condition (8) in Lemma 1: the empirical process of false-null p-values in each interesting group diverges from uniform at rate sqrt(n_j)(1-pi_j0).
    Needed for KS test to consistently estimate S; fails in weak-signal settings.
  • domain assumption Jin's estimator consistency requires PCS with n^{-2}||S||_1 = O(n^{-delta}) and condition (12) on minimum signal size (Section 3.2).
    Assumptions imported from Chen (2018) to guarantee consistency of the plug-in weights.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A grouped, selectively weighted false discovery rate procedure." pith.science (2026). https://pith.science/paper/OK4IQWUT

@misc{pith2026190805319,
  author       = {Pith},
  title        = {Pith review of: A grouped, selectively weighted false discovery rate procedure},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OK4IQWUT}},
  note         = {Machine review of arXiv:1908.05319}
}
read the original abstract

False discovery rate (FDR) control in structured hypotheses testing is an important topic in simultaneous inference. Most existing methods that aim to utilize group structure among hypotheses either employ the groupwise mixture model or weight all p-values or hypotheses. Thus, their powers can be improved when the groupwise mixture model is inappropriate or when most groups contain only true null hypotheses. Motivated by this, we propose a grouped, selectively weighted FDR procedure, which we refer to as "sGBH". Specifically, without employing the groupwise mixture model, sGBH identifies groups of hypotheses of interest, weights p-values in each such group only, and tests only the selected hypotheses using the weighted p-values. The sGBH subsumes a standard grouped, weighted FDR procedure which we refer to as "GBH". We provide simple conditions to ensure the conservativeness of sGBH, together with empirical evidence on its much improved power over GBH. The new procedure is applied to a gene expression study.

Figures

Figures reproduced from arXiv: 1908.05319 by the authors.

Figure 1
Figure 1. A schematic comparison between GBH and sGBH. sGBH selects groups of hypotheses of interest and weights only pvalues in each such groupwhereas GBH does not select groups of interest and weights only p-values in each such group, whereas GBH does not select groups [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FDRs and powers of the quasi-adaptive sGBH (coded as “sGBH” and as circle in [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. FDRs and powers of the plug-in adaptive sGBH (“sGBH”) and GBH (“GBH”) based [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: FDRs and powers of the generic adaptive sGBH (“sGBH”) and GBH (“GBH”) based [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: FDR and power of the generic adaptive sGBH (“sGBH”) based on two-sided p-values [PITH_FULL_IMAGE:figures/full_fig_p018_5.png]
Figure 6
Figure 6. Figure 6: FDRs and powers of the generic adaptive sGBH (“sGBH”) and GBH (“GBH”) based [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: FDRs and powers of the generic adaptive sGBH (“sGBH”) and GBH (“GBH”) based [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

3 extracted references · 3 canonical work pages

  1. [1]

    Consistent estimation of the proportion of false nulls and FDR for adaptive multiple testing Normal means under weak dependence

    Basu, P., Cai, T. T., Das, K. and Sun, W. (2018). Weighted false discovery rate control in large-scale multiple testing, J. Amer. Statist. Assoc. 113(523): 1172–1183. Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing, J. R. Statist. Soc. Ser. B 57(1): 289–300. Benjamini, Y. a...

  2. [2]

    LetCm =⋂ j∈S{ˆπj0 =πj0} andDm be the complement ofCm. Then, ˆαm≤ ∑ j∈S ∑ k∈Sj0 E [ 1 R1 { pjk πj0 (1− ˜π0) 1−πj0 ≤ Rα m }] + Pr (Dm) + Pr ( ˆS⁄=S ) ≤α + Pr (Dm) + Pr ( ˆS⁄=S ) , where the second inequality follows from the conservativeness of the oracle sGBH under PRDS. So, the claim holds. A.4 Proof of Lemma 1 For eachj∈{ 1,...,l }, let Fj be the empiric...

  3. [3]

    (33) If in addition ˆπ0,S consistently estimates ˜π0, then lim supm→∞ ˜αm≤ α

    such that Pr (ˆπ0,S≤ ˇπ0)> 0 and ˜αm≤ α m 1 1− ˇπ0 ∑ j∈S nj (1−πj0) + α m ∑ j/∈S njπj0 + Pr (ˆπ0,S > ˇπ0). (33) If in addition ˆπ0,S consistently estimates ˜π0, then lim supm→∞ ˜αm≤ α. On the other hand, if{pi}m i=1 have the property of PRDS and ˆπj0 is consistent for πj0 uniformly in j∈ S (with necessarily being non-increasing or reciprocally conservativ...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.