REVIEW 3 major objections 4 minor 41 references
Adversarially robust multiple testing in high dimensions
T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Quantile-winsorized step-down multiple testing controls the familywise error rate in finite samples under adversarial contamination, with bounds approaching the nominal level even as the dimension grows super-polynomially under only slightl
desk verdict Two-sample winsorized Gaussian approximation and bootstrap-critical-value control are the real contributions; the one-sample FWER bounds are repackaged from unproved companion results, so the headline moment/dimension claim is only as good as those imports. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is coordinate-wise quantile winsorization: every coordinate of each contaminated vector is clipped at data-dependent order-statistic thresholds chosen slightly beyond the assumed contamination fraction, and the same clipping is used to build both a location estimator and a covariance estimator. A Gaussian approximation inequality over hyperrectangles transfers the distribution of the centered, self-normalized winsorized mean vector to a Gaussian with the true covariance (or correlation) matrix up to an explicit rate; a covariance concentration bound extends this to the self-normalized statistic. The step-down algorithm then compares test statistics against critical values that are
What would settle it
Simulate one-sample data with exactly m=2.5 moments, choose d≈exp(n^{0.2}) and η≈n^{-0.4} so the stated regime holds, and compute the sup-over-hyperrectangles distance in the imported Gaussian approximation (Theorem A.2). If that distance does not vanish at the claimed rate, or if the empirical FWER of Algorithm 3 with a large B exceeds α plus a constant times the paper's error terms, the central claim collapses. A more targeted check is the covariance concentration bound: verify whether max_{j,k}|Σ̂_{j,k}-Σ_{j,k}| exceeds C·d_n with probability no greater than 24/n in the one-sample case.
Extended reading notes
Core claim
The paper's central claim is that step-down multiple testing procedures based on quantile-winsorized statistics achieve strong familywise error rate control under adversarial contamination. For a target level α, the probability that the procedure's reported set excludes at least one genuinely null coordinate is bounded by α plus explicit finite-sample error terms; for the bootstrap implementation the additional term is B^{-1}. The bounds are uniform over distributions whose coordinates have only m>2 moments and whose variances are bounded away from zero and above, and they converge to α as n grows in regimes where √(n log d) η^{1−1/m}→0 and log(d)/n^{(m−2)/(5m−2)}→0. In the two-sample case t
Load-bearing premise
The whole edifice rests on the imported one-sample Gaussian approximation inequalities for quantile-winsorized means and their stated rates; if those inequalities are wrong, or secretly need stronger moment or eigenvalue conditions, the finite-sample FWER bounds and the super-polynomial-dimension convergence claims do not follow.
Editorial extensions
If this is right
- Under Assumption 2.2 the one-sample algorithms' FWER upper bounds converge to α, so the step-down procedures become asymptotically exact while d grows faster than any polynomial and the number of contaminated observations can diverge.
- Because only m>2 moments are required, the same guarantees apply to heavy-tailed data without sub-Gaussian or sub-exponential tail conditions.
- The two-sample Gaussian approximation and covariance concentration inequalities imply finite-sample size control for robust global tests of equality of high-dimensional means, including a feasible bootstrap version.
- The bootstrap algorithms (3 and 6) explicitly charge B^{-1} for estimating critical values, so a user can make the cost as small as desired by increasing B; the tabulated values suggest B = 5,000 or 10,000 makes the extra term negligible.
- All six procedures provide strong control—the error is bounded no matter which set of coordinates is truly null—not merely weak control under the global null.
Reading between the lines
- The step-down construction could be repurposed to build simultaneous confidence intervals for the coordinates of a high-dimensional mean, since the FWER bound is uniform over the unknown set of true nulls; the paper does not pursue this.
- The two-sample approximation is assembled by conditioning on one sample and reusing the one-sample result, which suggests a recursive template for k-sample or stratified robust tests, but the paper stops at two samples.
- The finite-sample bounds are upper bounds; whether practical FWER is much smaller than α and how much power the winsorization costs near m=2, or with contamination at the boundary η≈1/2, are quantitative questions the paper does not answer and that simulation could settle.
- The tuning parameters λ are user-chosen; the paper inherits practical defaults from numerical experience, so a systematic sensitivity analysis of the λ choices would be a natural follow-up.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes step-down multiple testing procedures for coordinate-wise null hypotheses about high-dimensional means under adversarial contamination, using quantile-winsorized means and covariance estimators. In the one-sample case, finite-sample FWER bounds are derived from Gaussian approximation inequalities for winsorized means that are restated (without proof) from companion papers; in the two-sample case, new Gaussian approximation inequalities and covariance concentration bounds are developed. Theorems 2.1, 2.2, 2.4, 3.6, 3.7, and 3.8 provide finite-sample upper bounds of the form α + C(A + B [+ C]) that converge to α under Assumptions 2.2 and 3.3. Algorithms 3 and 6 explicitly account for the error from replacing exact Gaussian critical values by bootstrap order statistics, adding a B^{-1} term. The paper is a preliminary version with no numerical experiments.
Significance. If the imported one-sample Gaussian approximation results are correct under the stated moment conditions, the paper makes a useful contribution: it extends robust mean testing to multiple testing with strong FWER control, gives explicit finite-sample bounds that incorporate the cost of bootstrap-based critical values, and provides two-sample Gaussian approximation results that may be of independent interest. The treatment of data-dependent critical values in Propositions B.1 and C.1 is a genuine technical step. However, the advertised dimension-growth claim appears overstated relative to the paper's own rates, and the central one-sample inequalities are not proved or made available in the manuscript, so the significance is conditional on external results.
major comments (3)
- [Appendix A, Theorems A.1–A.2] These one-sample Gaussian approximation inequalities are restated from Kock and Preinerstorfer (2025a) without proof. Every FWER bound in the paper — Theorems 2.1, 2.2, 2.4 and, through Theorem 3.4, the two-sample Theorems 3.6–3.8 — relies on them. The rates A_X, B_X in Eq. (20) are the mechanism by which the bounds converge under Assumption 2.2. Since the companion papers are also preprints, the reader cannot verify that the stated rates require only m_X > 2 moments and the eigenvalue lower bound b_1. Please include proofs of Theorems A.1 and A.2 in an appendix, or at minimum state explicitly and verify that no additional conditions (e.g., finite fourth moments, bounded density, stronger eigenvalue separation) are needed.
- [Assumption 2.2 / Eq. (20)] The abstract and Section 2.3.2 claim the procedures allow the number of hypotheses to grow 'exponentially with sample size.' But Assumption 2.2 requires log(d)/n^{(m_X-2)/(5m_X-2)} -> 0, and the rates in Eq. (20) are consistent with this. Since (m_X-2)/(5m_X-2) < 1/5 for every m_X, this permits at most log d = o(n^{1/5}), i.e., d = exp(o(n^{1/5})), which is super-polynomial but not exponential in the usual sense. The headline claim should be corrected to 'super-polynomially' or the dimension condition should be strengthened to actually allow log d = o(n) if that is the intended claim.
- [Assumption 3.2 / Theorem 3.3] The two-sample results require the contaminated samples (X~_1,...,X~_nX) and (Y~_1,...,Y~_nY) to be independent. Under the stated adversarial contamination model, a single adversary who inspects both original samples before contaminating them may produce contaminated samples that are dependent, even when the original samples are independent. Independence of the original samples does not imply independence of the contaminated samples. Please clarify the threat model: is contamination assumed to be applied independently to each sample? If so, state this explicitly and discuss whether it is standard for two-sample adversarial robustness; if not, the two-sample theorems do not cover the stated model.
minor comments (4)
- [Abstract / Section 2.3.2] As noted above, 'exponentially with sample size' should be replaced by a precise statement such as 'super-polynomially' or 'faster than any polynomial rate'.
- [General] The manuscript reports no simulations or code. Given that Algorithms 3 and 6 introduce a specific bootstrap order statistic m(B,α,d), a small simulation study or at least a reproducible numerical example would help calibrate B and demonstrate finite-sample FWER behavior. This is not required for the theory but would improve the paper's usefulness.
- [Proof of Theorem 3.3 / Eq. (C.1)] There is a stray 's' before the bound on |[II]|; the line should read '|[II]| ≤ ...' instead of 's |[II]| ≤ ...'. This is typographical but should be fixed.
- [Eq. (30)] The notation 'w_XY' for sqrt(n_X n_Y/(n_X+n_Y)) is introduced without comment. It is clear from context, but a brief definition would avoid confusion with a product w_X w_Y.
Circularity Check
No significant circularity: the FWER bounds are derived from explicitly stated imported Gaussian approximation theorems, with no fitted-parameter or definitional circularity.
full rationale
The derivation chain is conditional rather than circular. The one-sample FWER bounds (Theorems 2.1, 2.2, 2.4) take the imported Gaussian approximation inequalities (Theorems A.1 and A.2, restated from Kock & Preinerstorfer 2025a) as explicit premises and then apply the standard Romano–Wolf step-down argument, augmented for the feasible bootstrap procedure by the in-paper elementary Lemma 2.3 to account for estimated critical values. The imported theorems are not defined in terms of the FWER; they bound the Kolmogorov distance over hyperrectangles between the law of centered winsorized means and a Gaussian, with rates A_X, B_X depending on n, d, η, m but not on the target FWER bound. No parameter in the paper is fitted to data and then renamed a prediction: the winsorization and bootstrap tuning parameters are user-specified, and α is fixed. The two-sample results are genuine extensions: Theorem 3.3 derives a two-sample Gaussian approximation from the one-sample Theorem A.1 via a conditioning argument, and Theorem 3.4 adds a covariance concentration step (Proposition 3.1) and a perturbation lemma. The paper is transparent that A.1/A.2 and Lemma A.4 come from companion papers by the same authors, so the contribution is conditional on those results; a hidden extra condition in the companion papers would invalidate the advertised claims, but that is a verification/correctness risk, not circularity. There is no uniqueness theorem imported from the authors, no ansatz smuggled in via citation, and no renaming of a known result as an organizing principle. Hence no circular step can be exhibited by quoting an equation that reduces to its own input.
Assumptions & free parameters
free parameters (3)
- λ1,X, λ'1,X (one-sample winsorization contamination factors) =
not fitted; recommendation λ1,X=λ'1,X≈1.01 (Remark 2.1)
- λ2,X, λ'2,X (one-sample winsorization log-term factors) =
not fitted; recommendation λ2,X=0.1, λ'2,X=0.07 (Remark 2.1)
- λ1,Y, λ'1,Y, λ2,Y, λ'2,Y (two-sample analogues) =
same recommendations as one-sample (Section 3.2)
assumptions (6)
- domain assumption One-sample Gaussian approximation inequalities for winsorized means (Theorems A.1/A.2, restated from Kock & Preinerstorfer 2025a) hold with stated rates under Assumption 2.1.
- domain assumption Contamination model: adversary can replace at most η_X n_X (resp. η_Y n_Y) observations arbitrarily, and η is known, non-random, <1/2.
- domain assumption Moment bounds: E|X|^m <∞ for m>2, coordinate variances bounded away from zero, m-th moments bounded above (Assumptions 2.1/3.1).
- domain assumption Independence of the two samples (Assumption 3.2).
- domain assumption Asymptotic regimes Assumptions 2.2/3.3: sqrt(n log d) η^{1-1/m}→0 and log(d)/n^{(m-2)/(5m-2)}→0.
- standard math Khatri-Šidák inequality and monotonicity of maxima quantiles (14), (17), (18), (19).
Cite this review
Pith. "Pith review of Adversarially robust multiple testing in high dimensions." pith.science (2026). https://pith.science/paper/GVECYUI2
@misc{pith2026260729507,
author = {Pith},
title = {Pith review of: Adversarially robust multiple testing in high dimensions},
year = {2026},
howpublished = {\url{https://pith.science/paper/GVECYUI2}},
note = {Machine review of arXiv:2607.29507}
}
read the original abstract
Robust multiple testing procedures for assessing equality restrictions on the coordinates of high-dimensional mean vectors are proposed. Our procedures are based on quantile-winsorization techniques, approximately control the familywise error rate (strongly) under adversarial contamination, and allow the number of hypotheses to grow exponentially with sample size, despite requiring only slightly more than two moments. Technically, our one-sample results build on recent Gaussian approximation inequalities for the distribution of high-dimensional quantile-winsorized means, whereas our two-sample results are based on extensions thereof, which we develop here and which could be of some interest in their own right.
Reference graph
Works this paper leans on
-
[1]
Bai, Z. and H. Saranadasa (1996): Effect of high dimension: by an example of a two sample problem, Statistica Sinica, 311--329
1996
-
[2]
Belloni, A., V. Chernozhukov, D. Chetverikov, C. Hansen, and K. Kato (2018): High-dimensional econometrics and regularized GMM, arXiv preprint arXiv:1806.01888
arXiv 2018
-
[3]
Bhatt, S., G. Fang, P. Li, and G. Samorodnitsky (2022): Minimax m-estimation under adversarial contamination, in International Conference on Machine Learning, PMLR, 1906--1924
2022
-
[4]
Liu, and Y
Cai, T., W. Liu, and Y. Xia (2014): Two-sample test of high dimensional means under dependence, Journal of the Royal Statistical Society Series B: Statistical Methodology, 76, 349--372
2014
-
[5]
Diakonikolas, and R
Cheng, Y., I. Diakonikolas, and R. Ge (2019): High-dimensional robust mean estimation in nearly-linear time, in Proceedings of the thirtieth annual ACM-SIAM symposium on discrete algorithms, SIAM, 2755--2771
2019
-
[6]
Chetverikov, and K
Chernozhukov, V., D. Chetverikov, and K. Kato (2013): Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors, The Annals of Statistics, 2786--2819
2013
-
[7]
--- -.1pt --- -.1pt --- (2017): Central limit theorems and bootstrap in high dimensions, Annals of Probability, 45, 2309--2352
2017
-
[8]
Dalalyan, A. S. and A. Minasyan (2022): All-in-one robust estimator of the gaussian mean, Annals of Statistics, 50, 1193--1219
2022
Show all 41 references
-
[9]
Dempster, A. P. (1958): A high dimensional two sample significance test, The Annals of Mathematical Statistics, 995--1010
1958
-
[10]
Depersin, J. and G. Lecu \'e (2022): Robust sub-Gaussian estimation of a mean vector in nearly linear time, Annals of Statistics, 50, 511--536
2022
-
[11]
Kamath, D
Diakonikolas, I., G. Kamath, D. Kane, J. Li, A. Moitra, and A. Stewart (2019): Robust estimators in high-dimensions without the computational intractability, SIAM Journal on Computing, 48, 742--864
2019
-
[12]
Dudoit, S., J. P. Shaffer, and J. C. Boldrick (2003): Multiple hypothesis testing in microarray experiments, Statistical Science, 18, 71--103
2003
-
[13]
Gin \'e , E. and R. Nickl (2016): Mathematical foundations of infinite-dimensional statistical models, Cambridge University Press
2016
-
[14]
Li, and F
Hopkins, S., J. Li, and F. Zhang (2020): Robust and heavy-tailed mean estimation made simple, via regret minimization, Advances in Neural Information Processing Systems, 33, 11902--11912
2020
-
[15]
(1931): The Generalization of Student's Ratio, Ann
Hotelling, H. (1931): The Generalization of Student's Ratio, Ann. Math. Statist., 2, 360--378
1931
-
[16]
Huang, Y., C. Li, R. Li, and S. Yang (2022): An overview of tests on high-dimensional means, Journal of Multivariate Analysis, 188, 104813
2022
-
[17]
Jiang, Y., X. Wang, C. Wen, Y. Jiang, and H. Zhang (2024): Nonparametric two-sample tests of high dimensional mean vectors via random integration, Journal of the American Statistical Association, 119, 701--714
2024
-
[18]
Khatri, C. G. (1967): On certain inequalities for normal distributions and their applications to simultaneous confidence bounds, The Annals of Mathematical Statistics, 1853--1867
1967
-
[19]
Kock, A. B. and D. Preinerstorfer (2023): Consistency of p-norm based tests in high dimensions: Characterization, monotonicity, domination, Bernoulli, 29, 2544--2573
2023
-
[20]
--- -.1pt --- -.1pt --- (2024): A remark on moment-dependent phase transitions in high-dimensional Gaussian approximations, Statistics & Probability Letters, 211, 110149
2024
-
[21]
--- -.1pt --- -.1pt --- (2025 a ): High-dimensional Gaussian and bootstrap approximations for robust means, arXiv preprint 2504.08435
2025 arXiv
-
[22]
--- -.1pt --- -.1pt --- (2025 b ): Winsorized mean estimation with heavy tails and adversarial contamination, arXiv preprint arXiv:2504.08482
2025 arXiv
-
[23]
--- -.1pt --- -.1pt --- (2026 a ): Enhanced power enhancements for testing many moment equalities: Beyond the 2 -and -norm, to appear in Journal of the American Statistical Association
2026
-
[24]
--- -.1pt --- -.1pt --- (2026 b ): Robustness for free: asymptotic size and power of max-tests in high dimensions, arXiv preprint 2601.14013
2026 arXiv
-
[25]
Lai, K. A., A. B. Rao, and S. Vempala (2016): Agnostic estimation of mean and covariance, in 2016 IEEE 57th Annual Symposium on Foundations of Computer Science (FOCS), IEEE, 665--674
2016
-
[26]
Liu, M. and M. E. Lopes (2024): Robust max statistics for high-dimensional inference, arXiv preprint arXiv:2409.16683
2024
-
[27]
Lugosi, G. and S. Mendelson (2021): Robust multivariate mean estimation: The optimality of trimmed mean , Annals of Statistics, 49, 393--410
2021
-
[28]
Minasyan, A. and N. Zhivotovskiy (2023): Statistically optimal robust mean and covariance estimation for anisotropic gaussians, arXiv preprint arXiv:2301.09024
2023 arXiv
-
[29]
(2023): Efficient median of means estimator, in The Thirty Sixth Annual Conference on Learning Theory, PMLR, 5925--5933
Minsker, S. (2023): Efficient median of means estimator, in The Thirty Sixth Annual Conference on Learning Theory, PMLR, 5925--5933
2023
-
[30]
Minsker, S. and M. Ndaoud (2021): Robust and efficient mean estimation: an approach based on the properties of self-normalized sums, Electronic Journal of Statistics, 15, 6036--6070
2021
-
[31]
Orenstein, and Z
Oliveira, R., P. Orenstein, and Z. Rico (2025): Finite-sample properties of the trimmed mean, arXiv preprint arXiv:2501.03694
2025 arXiv
-
[32]
Qiu, J., S. X. Chen, and Q.-M. Shao (2025): Self-normalized Cram \'e r type moderate deviation theorem for Gaussian approximation, The Annals of Statistics, 53, 1319--1346
2025
-
[33]
(2024): Robust high-dimensional Gaussian and bootstrap approximations for trimmed sample means, arXiv preprint arXiv:2410.22085
Resende, L. (2024): Robust high-dimensional Gaussian and bootstrap approximations for trimmed sample means, arXiv preprint arXiv:2410.22085
2024 arXiv
-
[34]
Romano, J. P. and M. Wolf (2005): Exact and approximate stepdown methods for multiple hypothesis testing, Journal of the American Statistical Association, 100, 94--108
2005
-
[35]
(1967): Rectangular confidence regions for the means of multivariate normal distributions, Journal of the American statistical association, 62, 626--633
S id \'a k, Z. (1967): Rectangular confidence regions for the means of multivariate normal distributions, Journal of the American statistical association, 62, 626--633
1967
-
[36]
Peng, and R
Wang, L., B. Peng, and R. Li (2015): A high-dimensional nonparametric multivariate test for mean vector, Journal of the American Statistical Association, 110, 1658--1669
2015
-
[37]
Xu, G., L. Lin, P. Wei, and W. Pan (2016): An adaptive two-sample test for high-dimensional means, Biometrika, 103, 609--624
2016
-
[38]
Xue, K. and F. Yao (2020): Distribution and correlation-free two-sample test of high-dimensional means, The Annals of Statistics, 48, 1304--1328
2020
-
[39]
Zheng, and R
Yang, S., S. Zheng, and R. Li (2024): A new test for high-dimensional two-sample mean problems with consideration of correlation structure, Annals of Statistics, 52, 2217--2240
2024
-
[40]
Zhang, D. and W. B. Wu (2017): Gaussian approximation for high dimensional time series , Annals of Statistics, 45, 1895--1919
2017
-
[41]
Zhang, J.-T., J. Guo, B. Zhou, and M.-Y. Cheng (2020): A simple two-sample test in high dimensions based on L^2 -norm, Journal of the American Statistical Association, 115, 1011--1027
2020
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.