Pith. sign in

REVIEW 3 major objections 24 references

The GBS step-down procedure is sharply asymptotically minimax for sparse Gaussian multiple testing

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-08 19:30 UTC pith:UUO2V545

load-bearing objection Solid theory paper that closes the sharp-minimax gap for classical GBS under Abraham et al. (2024) benchmarks; the main technical risk is uniformity of the shifted-order-statistic errors at the detection boundary. the 3 major comments →

arxiv 2607.05893 v1 pith:UUO2V545 submitted 2026-07-07 math.ST stat.TH

Sharp Asymptotic Minimaxity of the Gavrilov-Benjamini-Sarkar Step-Down Testing Procedure in Sparse Gaussian Sequence Models

classification math.ST stat.TH MSC 62H1562C2062G10
keywords multiple testingstep-down procedureGavrilov-Benjamini-Sarkarsparse Gaussian sequence modelasymptotic minimaxityHamming lossfalse discovery proportionbeta-min condition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper shows that the classical Gavrilov–Benjamini–Sarkar (GBS) step-down multiple testing procedure matches the best possible asymptotic risk in sparse Gaussian sequence models. In these models one observes independent noisy coordinates and wants to decide which coordinates carry a signal, under sparsity. Abraham and coauthors recently identified sharp asymptotic minimax benchmarks for this problem under two standard loss measures—normalized Hamming loss and the combined false-discovery-plus-false-non-discovery loss—and under both the classical beta-min signal classes and more general heterogeneous signal-strength classes. Several other procedures were already known to attain those benchmarks; the GBS step-down rule had not been shown to do so. The paper proves that GBS is sharply asymptotically minimax under both losses on the beta-min classes, that it attains the sharp FDP+FNP benchmark on the heterogeneous classes, and that it achieves a matching sharp asymptotic upper bound under Hamming loss on those same heterogeneous classes. The argument rests on a shifted order-statistic inequality for the GBS critical values together with deterministic and empirical signal-crossing comparisons for ordered alternative p-values, and deliberately avoids reducing the step-down rule to a single implicit threshold. That makes the result specific to the geometry of a genuine step-down procedure whose rejection set depends on the whole ordered sequence of p-values.

Core claim

In sparse Gaussian sequence models, the Gavrilov–Benjamini–Sarkar step-down procedure attains the sharp asymptotic minimax risk under both normalized Hamming loss and combined FDP+FNP loss over classical beta-min classes, and attains the corresponding sharp asymptotic minimax benchmark under FDP+FNP (and a sharp asymptotic upper bound under Hamming) over the heterogeneous signal-strength classes of Abraham et al. (2024).

What carries the argument

A shifted order-statistic inequality for the GBS critical values, combined with deterministic and empirical signal-crossing arguments that compare ordered alternative p-values to those critical values. This machinery works with the full ordered sequence of p-values and does not localize the rejection rule to a single threshold.

Load-bearing premise

The analysis is tied to independent unit-variance Gaussian noise and to the specific asymptotic sparsity and signal-strength geometry of the Abraham et al. (2024) parameter classes; outside that model and regime the sharp-rate identification and the GBS matching arguments need not transfer.

What would settle it

In a sparse Gaussian sequence simulation at the beta-min or heterogeneous boundary regimes of Abraham et al., compute the asymptotic normalized Hamming risk and the FDP+FNP risk of GBS; if either risk stays strictly above the known minimax benchmark by a positive constant, the sharp-minimax claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The manuscript studies the classical Gavrilov–Benjamini–Sarkar (GBS) step-down multiple testing procedure in the sparse Gaussian sequence model and compares it to the sharp asymptotic minimax benchmarks of Abraham et al. (2024). The main claims are that GBS is sharply asymptotically minimax under normalized Hamming loss and under the combined FDP+FNP loss on the classical beta-min classes, that it attains a sharp asymptotic upper bound for Hamming risk on the heterogeneous signal-strength classes of Abraham et al., and that it attains the corresponding sharp FDP+FNP benchmark on those heterogeneous classes. The technical development rests on a shifted order-statistic inequality for the GBS critical values, together with deterministic and empirical signal-crossing arguments for ordered alternative p-values, deliberately avoiding single-threshold localization used for BH and empirical-Bayes ℓ-value analyses.

Significance. If the sharp-constant claims hold, the paper closes a genuine gap: sharp asymptotic minimaxity theory for sparse multiple testing has been available for BH-type and empirical-Bayes procedures, but not for a classical genuine step-down rule whose rejection set depends on the entire ordered p-value path. Establishing that GBS matches the Abraham et al. (2024) benchmarks under both Hamming and FDP+FNP losses on beta-min classes, and the FDP+FNP benchmark on heterogeneous classes, is a substantial contribution to the multiple-testing literature. The geometric approach via a shifted order-statistic inequality is of independent technical interest as a template for other step-down procedures.

major comments (3)
  1. The central sharp-constant upper bounds rest on the shifted order-statistic inequality for GBS critical values controlling all additive/multiplicative errors uniformly down to the beta-min detection boundary a_n = sqrt(2 log(n/s)) + o(1). Because GBS is path-dependent, a local discrepancy near the extreme-order-statistic quantile can propagate into the final rejection set. The manuscript must make explicit, for each of the Hamming and FDP+FNP upper bounds on the beta-min classes, that every error term introduced by the shift is o(of the leading Abraham et al. risk term) uniformly over the full class, including the critical regime. If uniformity is only established away from the boundary, the claims reduce from sharp asymptotic minimaxity to rate-optimality and should be restated accordingly.
  2. For the heterogeneous signal-strength classes the paper claims a sharp asymptotic upper bound under Hamming loss and attainment of the sharp FDP+FNP benchmark. The heterogeneous geometry is strictly richer than beta-min; the signal-crossing arguments must therefore be checked against the worst-case configurations that attain the Abraham et al. lower bounds (not only against typical or average configurations). The manuscript should identify which configurations are used to match the lower-bound constants and confirm that the GBS path does not leave a residual gap of constant order under those configurations.
  3. The comparison with Abraham et al. (2024) is load-bearing for the word “sharp.” The manuscript should state, for each loss and each parameter class, the precise form of the Abraham et al. benchmark being matched (including any logarithmic or lower-order terms that are retained or discarded) and verify that the GBS risk equals (1+o(1)) times that benchmark rather than merely O(of the benchmark). Any place where only an upper bound of the correct order is obtained should be labeled as such, as already done for Hamming risk on the heterogeneous classes.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for a careful and constructive report. The three major comments concern (i) uniform control of shift-induced error terms down to the beta-min boundary, (ii) verification of signal-crossing arguments against the worst-case heterogeneous configurations that attain the Abraham et al. lower bounds, and (iii) explicit matching of the precise Abraham et al. benchmarks (including which lower-order terms are retained). We agree that the manuscript should make these points more transparent. In the revision we will strengthen the statements and proofs so that every claimed sharp constant is visibly (1+o(1)) times the corresponding Abraham et al. benchmark, uniformly over the full parameter classes, and we will label any result that is only order-sharp as such. Below we respond point by point.

read point-by-point responses
  1. Referee: The central sharp-constant upper bounds rest on the shifted order-statistic inequality for GBS critical values controlling all additive/multiplicative errors uniformly down to the beta-min detection boundary a_n = sqrt(2 log(n/s)) + o(1). Because GBS is path-dependent, a local discrepancy near the extreme-order-statistic quantile can propagate into the final rejection set. The manuscript must make explicit, for each of the Hamming and FDP+FNP upper bounds on the beta-min classes, that every error term introduced by the shift is o(of the leading Abraham et al. risk term) uniformly over the full class, including the critical regime. If uniformity is only established away from the boundary, the claims reduce from sharp asymptotic minimaxity to rate-optimality and should be restated accordingly.

    Authors: We agree that uniformity of all shift-induced errors down to the boundary a_n = sqrt(2 log(n/s))+o(1) is essential for the sharp-constant claims, and that path-dependence makes this non-automatic. The shifted order-statistic inequality is already stated and applied so that the additive and multiplicative discrepancies it produces are o(1) relative to the Abraham et al. leading risk terms, uniformly over the full beta-min classes (including the critical regime). The same holds for the deterministic and empirical signal-crossing arguments that control the ordered alternative p-values. We will, however, make this uniformity fully explicit in the revised manuscript: for each of the Hamming and FDP+FNP upper bounds on the beta-min classes we will isolate every error term generated by the shift, verify that it is o(of the corresponding Abraham et al. risk) uniformly over the whole class, and record the resulting (1+o(1)) statements in the theorems and proofs. If any intermediate estimate were only valid away from the boundary, we would restate the claim as rate-optimality; our analysis does not require such a restriction, and the revision will make that clear. revision: yes

  2. Referee: For the heterogeneous signal-strength classes the paper claims a sharp asymptotic upper bound under Hamming loss and attainment of the sharp FDP+FNP benchmark. The heterogeneous geometry is strictly richer than beta-min; the signal-crossing arguments must therefore be checked against the worst-case configurations that attain the Abraham et al. lower bounds (not only against typical or average configurations). The manuscript should identify which configurations are used to match the lower-bound constants and confirm that the GBS path does not leave a residual gap of constant order under those configurations.

    Authors: We agree that the heterogeneous geometry is richer than beta-min and that the signal-crossing arguments must be validated on the configurations that attain the Abraham et al. lower bounds, not merely on typical signals. In the present analysis the upper bounds are obtained by comparing the ordered alternative p-values to the GBS critical curve via deterministic and empirical crossing lemmas that hold pathwise for every configuration in the heterogeneous classes; the worst-case risk is then read off by maximizing the resulting expression over those classes, which recovers the Abraham et al. constants for FDP+FNP and the stated sharp upper bound for Hamming. We will revise the manuscript to identify explicitly the configurations (or sequences of configurations) that attain the Abraham et al. lower-bound constants, to verify the crossing lemmas on those configurations, and to confirm that the GBS rejection path leaves no residual gap of constant order. For Hamming risk on the heterogeneous classes we already claim only a sharp asymptotic upper bound (not full minimaxity); that distinction will be retained and emphasized. revision: yes

  3. Referee: The comparison with Abraham et al. (2024) is load-bearing for the word “sharp.” The manuscript should state, for each loss and each parameter class, the precise form of the Abraham et al. benchmark being matched (including any logarithmic or lower-order terms that are retained or discarded) and verify that the GBS risk equals (1+o(1)) times that benchmark rather than merely O(of the benchmark). Any place where only an upper bound of the correct order is obtained should be labeled as such, as already done for Hamming risk on the heterogeneous classes.

    Authors: We agree that the word “sharp” requires an explicit, term-by-term comparison with the Abraham et al. (2024) benchmarks. In the revision we will, for each loss (normalized Hamming; FDP+FNP) and each parameter class (beta-min; heterogeneous), state the precise form of the Abraham et al. benchmark that we match, including which logarithmic or lower-order terms are retained or discarded. We will then verify that the GBS risk equals (1+o(1)) times that benchmark (rather than merely O of it) wherever we claim sharp asymptotic minimaxity or attainment of the sharp benchmark—namely, both losses on beta-min classes, and FDP+FNP on the heterogeneous classes. Places where only an upper bound of the correct order (or a sharp upper bound that is not matched by a GBS-specific lower bound) is obtained will be labeled as such; this is already the case for Hamming risk on the heterogeneous classes, and that labeling will be kept and made more prominent in the statements of the theorems. revision: yes

Circularity Check

0 steps flagged

No significant circularity: GBS is shown to attain external Abraham et al. (2024) minimax benchmarks via a new geometric analysis of the step-down path, not by defining rates in terms of GBS or by fitting.

full rationale

The paper’s central claims are that the classical Gavrilov–Benjamini–Sarkar (GBS) step-down procedure attains the sharp asymptotic minimax benchmarks previously established by Abraham et al. (2024) for normalized Hamming loss and FDP+FNP loss over classical beta-min classes, and attains the corresponding FDP+FNP benchmark (with a matching upper bound for Hamming) over the heterogeneous signal-strength classes of the same reference. Those benchmarks are external: they are defined and proved for the sparse Gaussian sequence model by Abraham et al., independent of any particular procedure and in particular independent of GBS. GBS itself is a classical, parameter-free step-down rule whose critical values are fixed by the ordered p-value sequence; the paper does not fit free parameters to data and then relabel the fit as a prediction. The technical engine is a shifted order-statistic inequality for the GBS critical values together with deterministic/empirical signal-crossing arguments; these are presented as intrinsic to the geometry of the GBS path and as a replacement for single-threshold localization used for BH or empirical-Bayes ℓ-value methods. That device is a proof technique, not a redefinition of the risk or of the benchmark. There is no self-definitional loop (X defined via Y and then used to “derive” Y), no fitted-input-called-prediction, no load-bearing uniqueness theorem imported solely from the same author, and no renaming of a known empirical pattern as a new first-principles result. Self-citation, if any, is not load-bearing for the minimax identification: the target rates come from Abraham et al. (2024). The analysis is therefore self-contained asymptotic theory against external benchmarks; residual concerns about uniformity of error terms down to the detection boundary are correctness/technical risks, not circularity. Score 0 is the honest finding.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

Pure asymptotic decision-theory paper. Free parameters are none in the fitting sense; the model uses standard Gaussian noise and sparsity/signal classes from Abraham et al. Axioms are standard probability/order-statistic facts plus the domain setup of sparse Gaussian sequences and the GBS critical-value recursion. No invented particles or mediators. The main technical novelty is a shifted order-statistic inequality treated as a derived lemma, not an axiom.

axioms (3)
  • domain assumption Observations are independent N(μ_i, 1) in a sparse Gaussian sequence model with sparsity and signal classes as in Abraham et al. (2024).
    Standard sparse multiple-testing setup; the sharp rates and matching arguments are stated for this model and do not automatically extend outside it.
  • domain assumption GBS critical values are defined by the classical step-down recursion on ordered p-values (Gavrilov-Benjamini-Sarkar).
    The procedure under study; its geometry is the object of the shifted order-statistic inequality.
  • standard math Standard asymptotic order-statistic and extreme-value behavior of Gaussian p-values under null and alternative.
    Background probabilistic tools used for signal-crossing and risk asymptotics.

pith-pipeline@v0.9.1-grok · 6380 in / 2596 out tokens · 39261 ms · 2026-07-08T19:30:49.541615+00:00 · methodology

0 comments
read the original abstract

We investigate the sharp asymptotic minimaxity of the classical Gavrilov-Benjamini-Sarkar (GBS) step-down multiple testing procedure in sparse Gaussian sequence models. Abraham et al. (2024) recently established sharp asymptotic minimax benchmarks for sparse multiple testing under both the classical beta-min framework and more general heterogeneous signal-strength classes. Although several multiple testing procedures attain these benchmarks, corresponding results for the GBS procedure have remained unavailable. We prove that the GBS procedure is sharply asymptotically minimax under both normalized Hamming loss and the combined $\mathrm{FDP}+\mathrm{FNP}$ loss over the classical beta-min parameter classes. We further establish a sharp asymptotic upper bound for the normalized Hamming risk over the heterogeneous signal-strength classes and show that the GBS procedure attains the corresponding sharp asymptotic minimax benchmark under the combined $\mathrm{FDP}+\mathrm{FNP}$ loss. Our analysis is based on a shifted order-statistic inequality for the GBS critical values together with deterministic and empirical signal-crossing arguments for ordered alternative $p$-values. Unlike existing analyses of Benjamini-Hochberg procedures and empirical Bayes $\ell$-value methods, the proposed approach is intrinsic to the geometry of the GBS step-down procedure and avoids localization of a single implicit rejection threshold. Consequently, our results extend sharp asymptotic minimaxity theory to a genuinely step-down multiple testing procedure whose rejection rule depends on the entire ordered sequence of $p$-values.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 24 canonical work pages

  1. [1]

    Abraham, K., Castillo, I., and Roquain, E. (2024). Sharp multiple testing boundary for sparse sequences.The Annals of Statistics, 52(4):1564–1591. 1, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 17, 18, 33, 35, 41

  2. [2]

    L., and Johnstone, I

    Abramovich, F., Benjamini, Y., Donoho, D. L., and Johnstone, I. M. (2006). Adapting to unknown sparsity by controlling the false discovery rate.The Annals of Statistics, 34(2):584–653. 3

  3. [3]

    and Chen, S

    Arias-Castro, E. and Chen, S. (2017). Distribution-free multiple testing.Electronic Journal of Statistics, 11(1):1983–2001. 3, 4

  4. [4]

    and Nurushev, N

    Belitser, E. and Nurushev, N. (2020). Needles and straw in a haystack: Robust confidence for possibly sparse sequences.The Annals of Statistics, 48(2):956–982. 4

  5. [5]

    and Hochberg, Y

    Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing.Journal of the Royal Statistical Society: Series B, 57(1):289–300. 2

  6. [6]

    M., and Yekutieli, D

    Benjamini, Y., Krieger, A. M., and Yekutieli, D. (2006). Adaptive linear step-up procedures that control the false discovery rate.Biometrika, 93(3):491–507. 2

  7. [7]

    and Yekutieli, D

    Benjamini, Y. and Yekutieli, D. (2001). The control of the false discovery rate in multiple testing under dependency.The Annals of Statistics, 29(4):1165–1188. 2

  8. [8]

    and Roquain, E

    Blanchard, G. and Roquain, E. (2009). Adaptive fdr control under independence and de- pendence.Journal of Machine Learning Research, 10:2837–2871. 2

  9. [9]

    Bogdan, M., Chakrabarti, A., Frommlet, F., and Ghosh, J. K. (2011). Asymptotic bayes- optimality under sparsity of some multiple testing procedures.The Annals of Statistics, 39(3):1551–1579. 3, 5

  10. [10]

    and Ghosh, J

    Datta, J. and Ghosh, J. K. (2013). Asymptotic properties of bayes risk for the horseshoe prior.Bayesian Analysis, 8(1):111–132. 3

  11. [11]

    Fromont, M., Lerasle, M., and Reynaud-Bouret, P. (2016). Family-wise separation rates for multiple testing.The Annals of Statistics, 44(6):2533–2563. 3

  12. [12]

    Gavrilov, Y., Benjamini, Y., and Sarkar, S. K. (2009). An adaptive step-down procedure with proven fdr control under independence.The Annals of Statistics, 37(2):619–629. 1, 2, 3, 7, 10, 11, 18

  13. [13]

    and Chakrabarti, A

    Ghosh, P. and Chakrabarti, A. (2017). Asymptotic optimality of one-group shrinkage priors in sparse high-dimensional problems.Bayesian Analysis, 12(4):1133–1161. 3, 5

  14. [14]

    Ghosh, P., Tang, X., Ghosh, M., and Chakrabarti, A. (2016). Asymptotic properties of bayes risk of a general class of shrinkage priors in multiple hypothesis testing under sparsity. Bayesian Analysis, 11(3):753–796. 3, 5

  15. [15]

    Guo, W., He, L., and Sarkar, S. K. (2014). Further results on controlling the false discovery proportion.The Annals of Statistics, 42(3):1070–1101. 2

  16. [16]

    and Roquain, E

    Neuvial, P. and Roquain, E. (2012). On false discovery rate thresholding for classification under sparsity.The Annals of Statistics, 40(5):2572–2600. 2

  17. [17]

    and Chakrabarti, A

    Paul, S. and Chakrabarti, A. (2025). Posterior contraction rate and asymptotic bayes opti- mality for one group global-local shrinkage priors in sparse normal means problem.Annals of the Institute of Statistical Mathematics, 77:1–33. 3, 5

  18. [18]

    Paul, S., Ghosh, P., and Chakrabarti, A. (2025). Sharp asymptotic minimaxity for multiple testing using one-group shrinkage priors. arXiv:2505.16428. 5

  19. [19]

    I., and Wainwright, M

    Rabinovich, M., Ramdas, A., Jordan, M. I., and Wainwright, M. J. (2020). Optimal rates and trade-offs in multiple testing.Statistica Sinica, 30(2):741–762. 4

  20. [20]

    Sarkar, S. K. (2006). False discovery and false nondiscovery rates in single-step multiple testing procedures.The Annals of Statistics, 34:394–415. 2

  21. [21]

    Sarkar, S. K. (2008). On methods controlling the false discovery rate.Sankhy¯ a: The Indian Journal of Statistics, 70-A(2):135–168. 2

  22. [22]

    Storey, J. D. (2002). A direct approach to false discovery rates.Journal of the Royal Statistical Society: Series B, 64(3):479–498. 2

  23. [23]

    D., Taylor, J

    Storey, J. D., Taylor, J. E., and Siegmund, D. (2004). Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: A unified approach.Journal of the Royal Statistical Society: Series B, 66(1):187–205. 2

  24. [24]

    and Cai, T

    Sun, W. and Cai, T. T. (2009). Large-scale multiple testing under dependence.Journal of the Royal Statistical Society: Series B, 71(2):393–424. 2 44