REVIEW 3 major objections 24 references
The GBS step-down procedure is sharply asymptotically minimax for sparse Gaussian multiple testing
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-08 19:30 UTC pith:UUO2V545
load-bearing objection Solid theory paper that closes the sharp-minimax gap for classical GBS under Abraham et al. (2024) benchmarks; the main technical risk is uniformity of the shifted-order-statistic errors at the detection boundary. the 3 major comments →
Sharp Asymptotic Minimaxity of the Gavrilov-Benjamini-Sarkar Step-Down Testing Procedure in Sparse Gaussian Sequence Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In sparse Gaussian sequence models, the Gavrilov–Benjamini–Sarkar step-down procedure attains the sharp asymptotic minimax risk under both normalized Hamming loss and combined FDP+FNP loss over classical beta-min classes, and attains the corresponding sharp asymptotic minimax benchmark under FDP+FNP (and a sharp asymptotic upper bound under Hamming) over the heterogeneous signal-strength classes of Abraham et al. (2024).
What carries the argument
A shifted order-statistic inequality for the GBS critical values, combined with deterministic and empirical signal-crossing arguments that compare ordered alternative p-values to those critical values. This machinery works with the full ordered sequence of p-values and does not localize the rejection rule to a single threshold.
Load-bearing premise
The analysis is tied to independent unit-variance Gaussian noise and to the specific asymptotic sparsity and signal-strength geometry of the Abraham et al. (2024) parameter classes; outside that model and regime the sharp-rate identification and the GBS matching arguments need not transfer.
What would settle it
In a sparse Gaussian sequence simulation at the beta-min or heterogeneous boundary regimes of Abraham et al., compute the asymptotic normalized Hamming risk and the FDP+FNP risk of GBS; if either risk stays strictly above the known minimax benchmark by a positive constant, the sharp-minimax claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies the classical Gavrilov–Benjamini–Sarkar (GBS) step-down multiple testing procedure in the sparse Gaussian sequence model and compares it to the sharp asymptotic minimax benchmarks of Abraham et al. (2024). The main claims are that GBS is sharply asymptotically minimax under normalized Hamming loss and under the combined FDP+FNP loss on the classical beta-min classes, that it attains a sharp asymptotic upper bound for Hamming risk on the heterogeneous signal-strength classes of Abraham et al., and that it attains the corresponding sharp FDP+FNP benchmark on those heterogeneous classes. The technical development rests on a shifted order-statistic inequality for the GBS critical values, together with deterministic and empirical signal-crossing arguments for ordered alternative p-values, deliberately avoiding single-threshold localization used for BH and empirical-Bayes ℓ-value analyses.
Significance. If the sharp-constant claims hold, the paper closes a genuine gap: sharp asymptotic minimaxity theory for sparse multiple testing has been available for BH-type and empirical-Bayes procedures, but not for a classical genuine step-down rule whose rejection set depends on the entire ordered p-value path. Establishing that GBS matches the Abraham et al. (2024) benchmarks under both Hamming and FDP+FNP losses on beta-min classes, and the FDP+FNP benchmark on heterogeneous classes, is a substantial contribution to the multiple-testing literature. The geometric approach via a shifted order-statistic inequality is of independent technical interest as a template for other step-down procedures.
major comments (3)
- The central sharp-constant upper bounds rest on the shifted order-statistic inequality for GBS critical values controlling all additive/multiplicative errors uniformly down to the beta-min detection boundary a_n = sqrt(2 log(n/s)) + o(1). Because GBS is path-dependent, a local discrepancy near the extreme-order-statistic quantile can propagate into the final rejection set. The manuscript must make explicit, for each of the Hamming and FDP+FNP upper bounds on the beta-min classes, that every error term introduced by the shift is o(of the leading Abraham et al. risk term) uniformly over the full class, including the critical regime. If uniformity is only established away from the boundary, the claims reduce from sharp asymptotic minimaxity to rate-optimality and should be restated accordingly.
- For the heterogeneous signal-strength classes the paper claims a sharp asymptotic upper bound under Hamming loss and attainment of the sharp FDP+FNP benchmark. The heterogeneous geometry is strictly richer than beta-min; the signal-crossing arguments must therefore be checked against the worst-case configurations that attain the Abraham et al. lower bounds (not only against typical or average configurations). The manuscript should identify which configurations are used to match the lower-bound constants and confirm that the GBS path does not leave a residual gap of constant order under those configurations.
- The comparison with Abraham et al. (2024) is load-bearing for the word “sharp.” The manuscript should state, for each loss and each parameter class, the precise form of the Abraham et al. benchmark being matched (including any logarithmic or lower-order terms that are retained or discarded) and verify that the GBS risk equals (1+o(1)) times that benchmark rather than merely O(of the benchmark). Any place where only an upper bound of the correct order is obtained should be labeled as such, as already done for Hamming risk on the heterogeneous classes.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive report. The three major comments concern (i) uniform control of shift-induced error terms down to the beta-min boundary, (ii) verification of signal-crossing arguments against the worst-case heterogeneous configurations that attain the Abraham et al. lower bounds, and (iii) explicit matching of the precise Abraham et al. benchmarks (including which lower-order terms are retained). We agree that the manuscript should make these points more transparent. In the revision we will strengthen the statements and proofs so that every claimed sharp constant is visibly (1+o(1)) times the corresponding Abraham et al. benchmark, uniformly over the full parameter classes, and we will label any result that is only order-sharp as such. Below we respond point by point.
read point-by-point responses
-
Referee: The central sharp-constant upper bounds rest on the shifted order-statistic inequality for GBS critical values controlling all additive/multiplicative errors uniformly down to the beta-min detection boundary a_n = sqrt(2 log(n/s)) + o(1). Because GBS is path-dependent, a local discrepancy near the extreme-order-statistic quantile can propagate into the final rejection set. The manuscript must make explicit, for each of the Hamming and FDP+FNP upper bounds on the beta-min classes, that every error term introduced by the shift is o(of the leading Abraham et al. risk term) uniformly over the full class, including the critical regime. If uniformity is only established away from the boundary, the claims reduce from sharp asymptotic minimaxity to rate-optimality and should be restated accordingly.
Authors: We agree that uniformity of all shift-induced errors down to the boundary a_n = sqrt(2 log(n/s))+o(1) is essential for the sharp-constant claims, and that path-dependence makes this non-automatic. The shifted order-statistic inequality is already stated and applied so that the additive and multiplicative discrepancies it produces are o(1) relative to the Abraham et al. leading risk terms, uniformly over the full beta-min classes (including the critical regime). The same holds for the deterministic and empirical signal-crossing arguments that control the ordered alternative p-values. We will, however, make this uniformity fully explicit in the revised manuscript: for each of the Hamming and FDP+FNP upper bounds on the beta-min classes we will isolate every error term generated by the shift, verify that it is o(of the corresponding Abraham et al. risk) uniformly over the whole class, and record the resulting (1+o(1)) statements in the theorems and proofs. If any intermediate estimate were only valid away from the boundary, we would restate the claim as rate-optimality; our analysis does not require such a restriction, and the revision will make that clear. revision: yes
-
Referee: For the heterogeneous signal-strength classes the paper claims a sharp asymptotic upper bound under Hamming loss and attainment of the sharp FDP+FNP benchmark. The heterogeneous geometry is strictly richer than beta-min; the signal-crossing arguments must therefore be checked against the worst-case configurations that attain the Abraham et al. lower bounds (not only against typical or average configurations). The manuscript should identify which configurations are used to match the lower-bound constants and confirm that the GBS path does not leave a residual gap of constant order under those configurations.
Authors: We agree that the heterogeneous geometry is richer than beta-min and that the signal-crossing arguments must be validated on the configurations that attain the Abraham et al. lower bounds, not merely on typical signals. In the present analysis the upper bounds are obtained by comparing the ordered alternative p-values to the GBS critical curve via deterministic and empirical crossing lemmas that hold pathwise for every configuration in the heterogeneous classes; the worst-case risk is then read off by maximizing the resulting expression over those classes, which recovers the Abraham et al. constants for FDP+FNP and the stated sharp upper bound for Hamming. We will revise the manuscript to identify explicitly the configurations (or sequences of configurations) that attain the Abraham et al. lower-bound constants, to verify the crossing lemmas on those configurations, and to confirm that the GBS rejection path leaves no residual gap of constant order. For Hamming risk on the heterogeneous classes we already claim only a sharp asymptotic upper bound (not full minimaxity); that distinction will be retained and emphasized. revision: yes
-
Referee: The comparison with Abraham et al. (2024) is load-bearing for the word “sharp.” The manuscript should state, for each loss and each parameter class, the precise form of the Abraham et al. benchmark being matched (including any logarithmic or lower-order terms that are retained or discarded) and verify that the GBS risk equals (1+o(1)) times that benchmark rather than merely O(of the benchmark). Any place where only an upper bound of the correct order is obtained should be labeled as such, as already done for Hamming risk on the heterogeneous classes.
Authors: We agree that the word “sharp” requires an explicit, term-by-term comparison with the Abraham et al. (2024) benchmarks. In the revision we will, for each loss (normalized Hamming; FDP+FNP) and each parameter class (beta-min; heterogeneous), state the precise form of the Abraham et al. benchmark that we match, including which logarithmic or lower-order terms are retained or discarded. We will then verify that the GBS risk equals (1+o(1)) times that benchmark (rather than merely O of it) wherever we claim sharp asymptotic minimaxity or attainment of the sharp benchmark—namely, both losses on beta-min classes, and FDP+FNP on the heterogeneous classes. Places where only an upper bound of the correct order (or a sharp upper bound that is not matched by a GBS-specific lower bound) is obtained will be labeled as such; this is already the case for Hamming risk on the heterogeneous classes, and that labeling will be kept and made more prominent in the statements of the theorems. revision: yes
Circularity Check
No significant circularity: GBS is shown to attain external Abraham et al. (2024) minimax benchmarks via a new geometric analysis of the step-down path, not by defining rates in terms of GBS or by fitting.
full rationale
The paper’s central claims are that the classical Gavrilov–Benjamini–Sarkar (GBS) step-down procedure attains the sharp asymptotic minimax benchmarks previously established by Abraham et al. (2024) for normalized Hamming loss and FDP+FNP loss over classical beta-min classes, and attains the corresponding FDP+FNP benchmark (with a matching upper bound for Hamming) over the heterogeneous signal-strength classes of the same reference. Those benchmarks are external: they are defined and proved for the sparse Gaussian sequence model by Abraham et al., independent of any particular procedure and in particular independent of GBS. GBS itself is a classical, parameter-free step-down rule whose critical values are fixed by the ordered p-value sequence; the paper does not fit free parameters to data and then relabel the fit as a prediction. The technical engine is a shifted order-statistic inequality for the GBS critical values together with deterministic/empirical signal-crossing arguments; these are presented as intrinsic to the geometry of the GBS path and as a replacement for single-threshold localization used for BH or empirical-Bayes ℓ-value methods. That device is a proof technique, not a redefinition of the risk or of the benchmark. There is no self-definitional loop (X defined via Y and then used to “derive” Y), no fitted-input-called-prediction, no load-bearing uniqueness theorem imported solely from the same author, and no renaming of a known empirical pattern as a new first-principles result. Self-citation, if any, is not load-bearing for the minimax identification: the target rates come from Abraham et al. (2024). The analysis is therefore self-contained asymptotic theory against external benchmarks; residual concerns about uniformity of error terms down to the detection boundary are correctness/technical risks, not circularity. Score 0 is the honest finding.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Observations are independent N(μ_i, 1) in a sparse Gaussian sequence model with sparsity and signal classes as in Abraham et al. (2024).
- domain assumption GBS critical values are defined by the classical step-down recursion on ordered p-values (Gavrilov-Benjamini-Sarkar).
- standard math Standard asymptotic order-statistic and extreme-value behavior of Gaussian p-values under null and alternative.
read the original abstract
We investigate the sharp asymptotic minimaxity of the classical Gavrilov-Benjamini-Sarkar (GBS) step-down multiple testing procedure in sparse Gaussian sequence models. Abraham et al. (2024) recently established sharp asymptotic minimax benchmarks for sparse multiple testing under both the classical beta-min framework and more general heterogeneous signal-strength classes. Although several multiple testing procedures attain these benchmarks, corresponding results for the GBS procedure have remained unavailable. We prove that the GBS procedure is sharply asymptotically minimax under both normalized Hamming loss and the combined $\mathrm{FDP}+\mathrm{FNP}$ loss over the classical beta-min parameter classes. We further establish a sharp asymptotic upper bound for the normalized Hamming risk over the heterogeneous signal-strength classes and show that the GBS procedure attains the corresponding sharp asymptotic minimax benchmark under the combined $\mathrm{FDP}+\mathrm{FNP}$ loss. Our analysis is based on a shifted order-statistic inequality for the GBS critical values together with deterministic and empirical signal-crossing arguments for ordered alternative $p$-values. Unlike existing analyses of Benjamini-Hochberg procedures and empirical Bayes $\ell$-value methods, the proposed approach is intrinsic to the geometry of the GBS step-down procedure and avoids localization of a single implicit rejection threshold. Consequently, our results extend sharp asymptotic minimaxity theory to a genuinely step-down multiple testing procedure whose rejection rule depends on the entire ordered sequence of $p$-values.
Reference graph
Works this paper leans on
-
[1]
Abraham, K., Castillo, I., and Roquain, E. (2024). Sharp multiple testing boundary for sparse sequences.The Annals of Statistics, 52(4):1564–1591. 1, 4, 5, 6, 7, 8, 9, 10, 12, 13, 14, 15, 16, 17, 18, 33, 35, 41
work page 2024
-
[2]
Abramovich, F., Benjamini, Y., Donoho, D. L., and Johnstone, I. M. (2006). Adapting to unknown sparsity by controlling the false discovery rate.The Annals of Statistics, 34(2):584–653. 3
work page 2006
-
[3]
Arias-Castro, E. and Chen, S. (2017). Distribution-free multiple testing.Electronic Journal of Statistics, 11(1):1983–2001. 3, 4
work page 2017
-
[4]
Belitser, E. and Nurushev, N. (2020). Needles and straw in a haystack: Robust confidence for possibly sparse sequences.The Annals of Statistics, 48(2):956–982. 4
work page 2020
-
[5]
Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing.Journal of the Royal Statistical Society: Series B, 57(1):289–300. 2
work page 1995
-
[6]
Benjamini, Y., Krieger, A. M., and Yekutieli, D. (2006). Adaptive linear step-up procedures that control the false discovery rate.Biometrika, 93(3):491–507. 2
work page 2006
-
[7]
Benjamini, Y. and Yekutieli, D. (2001). The control of the false discovery rate in multiple testing under dependency.The Annals of Statistics, 29(4):1165–1188. 2
work page 2001
-
[8]
Blanchard, G. and Roquain, E. (2009). Adaptive fdr control under independence and de- pendence.Journal of Machine Learning Research, 10:2837–2871. 2
work page 2009
-
[9]
Bogdan, M., Chakrabarti, A., Frommlet, F., and Ghosh, J. K. (2011). Asymptotic bayes- optimality under sparsity of some multiple testing procedures.The Annals of Statistics, 39(3):1551–1579. 3, 5
work page 2011
-
[10]
Datta, J. and Ghosh, J. K. (2013). Asymptotic properties of bayes risk for the horseshoe prior.Bayesian Analysis, 8(1):111–132. 3
work page 2013
-
[11]
Fromont, M., Lerasle, M., and Reynaud-Bouret, P. (2016). Family-wise separation rates for multiple testing.The Annals of Statistics, 44(6):2533–2563. 3
work page 2016
-
[12]
Gavrilov, Y., Benjamini, Y., and Sarkar, S. K. (2009). An adaptive step-down procedure with proven fdr control under independence.The Annals of Statistics, 37(2):619–629. 1, 2, 3, 7, 10, 11, 18
work page 2009
-
[13]
Ghosh, P. and Chakrabarti, A. (2017). Asymptotic optimality of one-group shrinkage priors in sparse high-dimensional problems.Bayesian Analysis, 12(4):1133–1161. 3, 5
work page 2017
-
[14]
Ghosh, P., Tang, X., Ghosh, M., and Chakrabarti, A. (2016). Asymptotic properties of bayes risk of a general class of shrinkage priors in multiple hypothesis testing under sparsity. Bayesian Analysis, 11(3):753–796. 3, 5
work page 2016
-
[15]
Guo, W., He, L., and Sarkar, S. K. (2014). Further results on controlling the false discovery proportion.The Annals of Statistics, 42(3):1070–1101. 2
work page 2014
-
[16]
Neuvial, P. and Roquain, E. (2012). On false discovery rate thresholding for classification under sparsity.The Annals of Statistics, 40(5):2572–2600. 2
work page 2012
-
[17]
Paul, S. and Chakrabarti, A. (2025). Posterior contraction rate and asymptotic bayes opti- mality for one group global-local shrinkage priors in sparse normal means problem.Annals of the Institute of Statistical Mathematics, 77:1–33. 3, 5
work page 2025
- [18]
-
[19]
Rabinovich, M., Ramdas, A., Jordan, M. I., and Wainwright, M. J. (2020). Optimal rates and trade-offs in multiple testing.Statistica Sinica, 30(2):741–762. 4
work page 2020
-
[20]
Sarkar, S. K. (2006). False discovery and false nondiscovery rates in single-step multiple testing procedures.The Annals of Statistics, 34:394–415. 2
work page 2006
-
[21]
Sarkar, S. K. (2008). On methods controlling the false discovery rate.Sankhy¯ a: The Indian Journal of Statistics, 70-A(2):135–168. 2
work page 2008
-
[22]
Storey, J. D. (2002). A direct approach to false discovery rates.Journal of the Royal Statistical Society: Series B, 64(3):479–498. 2
work page 2002
-
[23]
Storey, J. D., Taylor, J. E., and Siegmund, D. (2004). Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: A unified approach.Journal of the Royal Statistical Society: Series B, 66(1):187–205. 2
work page 2004
-
[24]
Sun, W. and Cai, T. T. (2009). Large-scale multiple testing under dependence.Journal of the Royal Statistical Society: Series B, 71(2):393–424. 2 44
work page 2009
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.