Pith. sign in

REVIEW 3 major objections 4 minor 41 references

Adversarially robust multiple testing in high dimensions

T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Quantile-winsorized step-down multiple testing controls the familywise error rate in finite samples under adversarial contamination, with bounds approaching the nominal level even as the dimension grows super-polynomially under only slightl

desk verdict Two-sample winsorized Gaussian approximation and bootstrap-critical-value control are the real contributions; the one-sample FWER bounds are repackaged from unproved companion results, so the headline moment/dimension claim is only as good as those imports. read the letter →

arxiv 2607.29507 v1 pith:GVECYUI2 submitted 2026-07-31 math.ST stat.TH

classification math.STstat.TH MSC 62H1562G3562G2062E17
keywords adversarialcontaminationmultipletestingfamilywiseerrorratewinsorizedmeanshigh-dimensionalGaussianapproximationtwo-samplemeanbootstrapcriticalvaluesheavy-taileddistributions
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Adversarial contamination usually destroys tests built on coordinate averages; the paper claims that quantile-winsorized means and covariance estimators restore multiple testing guarantees. It builds step-down procedures whose probability of missing at least one true null hypothesis is at most the nominal level α plus explicit finite-sample terms A, B, C (and 1/B when bootstrap quantiles are used). These terms go to zero in regimes where the number of hypotheses d can grow faster than any polynomial in the sample size and the number of contaminated observations can diverge, while the data only need slightly more than two moments. The same construction is extended to two-sample equality-of-means testing, including a new Gaussian approximation for winsorized two-sample statistics and covariance concentration inequalities. A practitioner who trusts the imported one-sample Gaussian approximation gets robust, feasible algorithms with strong familywise error control, not just a heuristic.

What carries the argument

The engine is coordinate-wise quantile winsorization: every coordinate of each contaminated vector is clipped at data-dependent order-statistic thresholds chosen slightly beyond the assumed contamination fraction, and the same clipping is used to build both a location estimator and a covariance estimator. A Gaussian approximation inequality over hyperrectangles transfers the distribution of the centered, self-normalized winsorized mean vector to a Gaussian with the true covariance (or correlation) matrix up to an explicit rate; a covariance concentration bound extends this to the self-normalized statistic. The step-down algorithm then compares test statistics against critical values that are

What would settle it

Simulate one-sample data with exactly m=2.5 moments, choose d≈exp(n^{0.2}) and η≈n^{-0.4} so the stated regime holds, and compute the sup-over-hyperrectangles distance in the imported Gaussian approximation (Theorem A.2). If that distance does not vanish at the claimed rate, or if the empirical FWER of Algorithm 3 with a large B exceeds α plus a constant times the paper's error terms, the central claim collapses. A more targeted check is the covariance concentration bound: verify whether max_{j,k}|Σ̂_{j,k}-Σ_{j,k}| exceeds C·d_n with probability no greater than 24/n in the one-sample case.

Watch

Extended reading notes

Core claim

The paper's central claim is that step-down multiple testing procedures based on quantile-winsorized statistics achieve strong familywise error rate control under adversarial contamination. For a target level α, the probability that the procedure's reported set excludes at least one genuinely null coordinate is bounded by α plus explicit finite-sample error terms; for the bootstrap implementation the additional term is B^{-1}. The bounds are uniform over distributions whose coordinates have only m>2 moments and whose variances are bounded away from zero and above, and they converge to α as n grows in regimes where √(n log d) η^{1−1/m}→0 and log(d)/n^{(m−2)/(5m−2)}→0. In the two-sample case t

Load-bearing premise

The whole edifice rests on the imported one-sample Gaussian approximation inequalities for quantile-winsorized means and their stated rates; if those inequalities are wrong, or secretly need stronger moment or eigenvalue conditions, the finite-sample FWER bounds and the super-polynomial-dimension convergence claims do not follow.

Editorial extensions

If this is right

  • Under Assumption 2.2 the one-sample algorithms' FWER upper bounds converge to α, so the step-down procedures become asymptotically exact while d grows faster than any polynomial and the number of contaminated observations can diverge.
  • Because only m>2 moments are required, the same guarantees apply to heavy-tailed data without sub-Gaussian or sub-exponential tail conditions.
  • The two-sample Gaussian approximation and covariance concentration inequalities imply finite-sample size control for robust global tests of equality of high-dimensional means, including a feasible bootstrap version.
  • The bootstrap algorithms (3 and 6) explicitly charge B^{-1} for estimating critical values, so a user can make the cost as small as desired by increasing B; the tabulated values suggest B = 5,000 or 10,000 makes the extra term negligible.
  • All six procedures provide strong control—the error is bounded no matter which set of coordinates is truly null—not merely weak control under the global null.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The step-down construction could be repurposed to build simultaneous confidence intervals for the coordinates of a high-dimensional mean, since the FWER bound is uniform over the unknown set of true nulls; the paper does not pursue this.
  • The two-sample approximation is assembled by conditioning on one sample and reusing the one-sample result, which suggests a recursive template for k-sample or stratified robust tests, but the paper stops at two samples.
  • The finite-sample bounds are upper bounds; whether practical FWER is much smaller than α and how much power the winsorization costs near m=2, or with contamination at the boundary η≈1/2, are quantitative questions the paper does not answer and that simulation could settle.
  • The tuning parameters λ are user-chosen; the paper inherits practical defaults from numerical experience, so a systematic sensitivity analysis of the λ choices would be a natural follow-up.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes step-down multiple testing procedures for coordinate-wise null hypotheses about high-dimensional means under adversarial contamination, using quantile-winsorized means and covariance estimators. In the one-sample case, finite-sample FWER bounds are derived from Gaussian approximation inequalities for winsorized means that are restated (without proof) from companion papers; in the two-sample case, new Gaussian approximation inequalities and covariance concentration bounds are developed. Theorems 2.1, 2.2, 2.4, 3.6, 3.7, and 3.8 provide finite-sample upper bounds of the form α + C(A + B [+ C]) that converge to α under Assumptions 2.2 and 3.3. Algorithms 3 and 6 explicitly account for the error from replacing exact Gaussian critical values by bootstrap order statistics, adding a B^{-1} term. The paper is a preliminary version with no numerical experiments.

Significance. If the imported one-sample Gaussian approximation results are correct under the stated moment conditions, the paper makes a useful contribution: it extends robust mean testing to multiple testing with strong FWER control, gives explicit finite-sample bounds that incorporate the cost of bootstrap-based critical values, and provides two-sample Gaussian approximation results that may be of independent interest. The treatment of data-dependent critical values in Propositions B.1 and C.1 is a genuine technical step. However, the advertised dimension-growth claim appears overstated relative to the paper's own rates, and the central one-sample inequalities are not proved or made available in the manuscript, so the significance is conditional on external results.

major comments (3)
  1. [Appendix A, Theorems A.1–A.2] These one-sample Gaussian approximation inequalities are restated from Kock and Preinerstorfer (2025a) without proof. Every FWER bound in the paper — Theorems 2.1, 2.2, 2.4 and, through Theorem 3.4, the two-sample Theorems 3.6–3.8 — relies on them. The rates A_X, B_X in Eq. (20) are the mechanism by which the bounds converge under Assumption 2.2. Since the companion papers are also preprints, the reader cannot verify that the stated rates require only m_X > 2 moments and the eigenvalue lower bound b_1. Please include proofs of Theorems A.1 and A.2 in an appendix, or at minimum state explicitly and verify that no additional conditions (e.g., finite fourth moments, bounded density, stronger eigenvalue separation) are needed.
  2. [Assumption 2.2 / Eq. (20)] The abstract and Section 2.3.2 claim the procedures allow the number of hypotheses to grow 'exponentially with sample size.' But Assumption 2.2 requires log(d)/n^{(m_X-2)/(5m_X-2)} -> 0, and the rates in Eq. (20) are consistent with this. Since (m_X-2)/(5m_X-2) < 1/5 for every m_X, this permits at most log d = o(n^{1/5}), i.e., d = exp(o(n^{1/5})), which is super-polynomial but not exponential in the usual sense. The headline claim should be corrected to 'super-polynomially' or the dimension condition should be strengthened to actually allow log d = o(n) if that is the intended claim.
  3. [Assumption 3.2 / Theorem 3.3] The two-sample results require the contaminated samples (X~_1,...,X~_nX) and (Y~_1,...,Y~_nY) to be independent. Under the stated adversarial contamination model, a single adversary who inspects both original samples before contaminating them may produce contaminated samples that are dependent, even when the original samples are independent. Independence of the original samples does not imply independence of the contaminated samples. Please clarify the threat model: is contamination assumed to be applied independently to each sample? If so, state this explicitly and discuss whether it is standard for two-sample adversarial robustness; if not, the two-sample theorems do not cover the stated model.
minor comments (4)
  1. [Abstract / Section 2.3.2] As noted above, 'exponentially with sample size' should be replaced by a precise statement such as 'super-polynomially' or 'faster than any polynomial rate'.
  2. [General] The manuscript reports no simulations or code. Given that Algorithms 3 and 6 introduce a specific bootstrap order statistic m(B,α,d), a small simulation study or at least a reproducible numerical example would help calibrate B and demonstrate finite-sample FWER behavior. This is not required for the theory but would improve the paper's usefulness.
  3. [Proof of Theorem 3.3 / Eq. (C.1)] There is a stray 's' before the bound on |[II]|; the line should read '|[II]| ≤ ...' instead of 's |[II]| ≤ ...'. This is typographical but should be fixed.
  4. [Eq. (30)] The notation 'w_XY' for sqrt(n_X n_Y/(n_X+n_Y)) is introduced without comment. It is clear from context, but a brief definition would avoid confusion with a product w_X w_Y.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the FWER bounds are derived from explicitly stated imported Gaussian approximation theorems, with no fitted-parameter or definitional circularity.

full rationale

The derivation chain is conditional rather than circular. The one-sample FWER bounds (Theorems 2.1, 2.2, 2.4) take the imported Gaussian approximation inequalities (Theorems A.1 and A.2, restated from Kock & Preinerstorfer 2025a) as explicit premises and then apply the standard Romano–Wolf step-down argument, augmented for the feasible bootstrap procedure by the in-paper elementary Lemma 2.3 to account for estimated critical values. The imported theorems are not defined in terms of the FWER; they bound the Kolmogorov distance over hyperrectangles between the law of centered winsorized means and a Gaussian, with rates A_X, B_X depending on n, d, η, m but not on the target FWER bound. No parameter in the paper is fitted to data and then renamed a prediction: the winsorization and bootstrap tuning parameters are user-specified, and α is fixed. The two-sample results are genuine extensions: Theorem 3.3 derives a two-sample Gaussian approximation from the one-sample Theorem A.1 via a conditioning argument, and Theorem 3.4 adds a covariance concentration step (Proposition 3.1) and a perturbation lemma. The paper is transparent that A.1/A.2 and Lemma A.4 come from companion papers by the same authors, so the contribution is conditional on those results; a hidden extra condition in the companion papers would invalidate the advertised claims, but that is a verification/correctness risk, not circularity. There is no uniqueness theorem imported from the authors, no ansatz smuggled in via citation, and no renaming of a known result as an organizing principle. Hence no circular step can be exhibited by quoting an equation that reduces to its own input.

Assumptions & free parameters 3 free parameters · 6 assumptions · 0 invented entities

No hidden fitted constants: the tuning parameters λ are explicitly user-chosen with stated recommendations, and the FWER level α is fixed. The central results are conditional on the imported one-sample Gaussian approximation theorems and the modeling assumptions; there are no invented physical or statistical entities.

free parameters (3)
  • λ1,X, λ'1,X (one-sample winsorization contamination factors) = not fitted; recommendation λ1,X=λ'1,X≈1.01 (Remark 2.1)
    Chosen by hand, following simulations in Kock & Preinerstorfer (2025a); they set how much beyond η must be winsorized.
  • λ2,X, λ'2,X (one-sample winsorization log-term factors) = not fitted; recommendation λ2,X=0.1, λ'2,X=0.07 (Remark 2.1)
    Chosen by hand; technical artifacts from concentration inequalities; affect ε and ε' in (3).
  • λ1,Y, λ'1,Y, λ2,Y, λ'2,Y (two-sample analogues) = same recommendations as one-sample (Section 3.2)
    Same as above for the Y-sample.
assumptions (6)
  • domain assumption One-sample Gaussian approximation inequalities for winsorized means (Theorems A.1/A.2, restated from Kock & Preinerstorfer 2025a) hold with stated rates under Assumption 2.1.
    Invoked as the basis for all one-sample results (Section 2.3.2) and the two-sample extension (Theorem 3.3). Not proved in this paper.
  • domain assumption Contamination model: adversary can replace at most η_X n_X (resp. η_Y n_Y) observations arbitrarily, and η is known, non-random, <1/2.
    Equations (2) and (26); the winsorization fraction ε is built from η. If η is not a valid bound, the procedure's guarantee fails.
  • domain assumption Moment bounds: E|X|^m <∞ for m>2, coordinate variances bounded away from zero, m-th moments bounded above (Assumptions 2.1/3.1).
    Used throughout to control Gaussian approximation errors; the 'slightly more than two moments' claim is exactly this.
  • domain assumption Independence of the two samples (Assumption 3.2).
    Theorem 3.3's conditioning argument requires independence between X- and Y-samples; two-sample FWER bounds in Section 3 rely on it.
  • domain assumption Asymptotic regimes Assumptions 2.2/3.3: sqrt(n log d) η^{1-1/m}→0 and log(d)/n^{(m-2)/(5m-2)}→0.
    Needed for the finite-sample upper bounds to converge to α. The first condition forces the contamination fraction to vanish asymptotically.
  • standard math Khatri-Šidák inequality and monotonicity of maxima quantiles (14), (17), (18), (19).
    Used to compare critical values and to justify worst-case independence critical values.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adversarially robust multiple testing in high dimensions." pith.science (2026). https://pith.science/paper/GVECYUI2

@misc{pith2026260729507,
  author       = {Pith},
  title        = {Pith review of: Adversarially robust multiple testing in high dimensions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GVECYUI2}},
  note         = {Machine review of arXiv:2607.29507}
}
read the original abstract

Robust multiple testing procedures for assessing equality restrictions on the coordinates of high-dimensional mean vectors are proposed. Our procedures are based on quantile-winsorization techniques, approximately control the familywise error rate (strongly) under adversarial contamination, and allow the number of hypotheses to grow exponentially with sample size, despite requiring only slightly more than two moments. Technically, our one-sample results build on recent Gaussian approximation inequalities for the distribution of high-dimensional quantile-winsorized means, whereas our two-sample results are based on extensions thereof, which we develop here and which could be of some interest in their own right.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

41 extracted references · 7 linked inside Pith

  1. [1]

    Bai, Z. and H. Saranadasa (1996): Effect of high dimension: by an example of a two sample problem, Statistica Sinica, 311--329

  2. [2]

    Chernozhukov, D

    Belloni, A., V. Chernozhukov, D. Chetverikov, C. Hansen, and K. Kato (2018): High-dimensional econometrics and regularized GMM, arXiv preprint arXiv:1806.01888

  3. [3]

    Bhatt, S., G. Fang, P. Li, and G. Samorodnitsky (2022): Minimax m-estimation under adversarial contamination, in International Conference on Machine Learning, PMLR, 1906--1924

  4. [4]

    Liu, and Y

    Cai, T., W. Liu, and Y. Xia (2014): Two-sample test of high dimensional means under dependence, Journal of the Royal Statistical Society Series B: Statistical Methodology, 76, 349--372

  5. [5]

    Diakonikolas, and R

    Cheng, Y., I. Diakonikolas, and R. Ge (2019): High-dimensional robust mean estimation in nearly-linear time, in Proceedings of the thirtieth annual ACM-SIAM symposium on discrete algorithms, SIAM, 2755--2771

  6. [6]

    Chetverikov, and K

    Chernozhukov, V., D. Chetverikov, and K. Kato (2013): Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors, The Annals of Statistics, 2786--2819

  7. [7]

    --- -.1pt --- -.1pt --- (2017): Central limit theorems and bootstrap in high dimensions, Annals of Probability, 45, 2309--2352

  8. [8]

    Dalalyan, A. S. and A. Minasyan (2022): All-in-one robust estimator of the gaussian mean, Annals of Statistics, 50, 1193--1219

Show all 41 references
  1. [9]

    Dempster, A. P. (1958): A high dimensional two sample significance test, The Annals of Mathematical Statistics, 995--1010

  2. [10]

    Depersin, J. and G. Lecu \'e (2022): Robust sub-Gaussian estimation of a mean vector in nearly linear time, Annals of Statistics, 50, 511--536

  3. [11]

    Kamath, D

    Diakonikolas, I., G. Kamath, D. Kane, J. Li, A. Moitra, and A. Stewart (2019): Robust estimators in high-dimensions without the computational intractability, SIAM Journal on Computing, 48, 742--864

  4. [12]

    Dudoit, S., J. P. Shaffer, and J. C. Boldrick (2003): Multiple hypothesis testing in microarray experiments, Statistical Science, 18, 71--103

  5. [13]

    Gin \'e , E. and R. Nickl (2016): Mathematical foundations of infinite-dimensional statistical models, Cambridge University Press

  6. [14]

    Li, and F

    Hopkins, S., J. Li, and F. Zhang (2020): Robust and heavy-tailed mean estimation made simple, via regret minimization, Advances in Neural Information Processing Systems, 33, 11902--11912

  7. [15]

    (1931): The Generalization of Student's Ratio, Ann

    Hotelling, H. (1931): The Generalization of Student's Ratio, Ann. Math. Statist., 2, 360--378

  8. [16]

    Huang, Y., C. Li, R. Li, and S. Yang (2022): An overview of tests on high-dimensional means, Journal of Multivariate Analysis, 188, 104813

  9. [17]

    Jiang, Y., X. Wang, C. Wen, Y. Jiang, and H. Zhang (2024): Nonparametric two-sample tests of high dimensional mean vectors via random integration, Journal of the American Statistical Association, 119, 701--714

  10. [18]

    Khatri, C. G. (1967): On certain inequalities for normal distributions and their applications to simultaneous confidence bounds, The Annals of Mathematical Statistics, 1853--1867

  11. [19]

    Kock, A. B. and D. Preinerstorfer (2023): Consistency of p-norm based tests in high dimensions: Characterization, monotonicity, domination, Bernoulli, 29, 2544--2573

  12. [20]

    --- -.1pt --- -.1pt --- (2024): A remark on moment-dependent phase transitions in high-dimensional Gaussian approximations, Statistics & Probability Letters, 211, 110149

  13. [21]

    --- -.1pt --- -.1pt --- (2025 a ): High-dimensional Gaussian and bootstrap approximations for robust means, arXiv preprint 2504.08435

  14. [22]

    --- -.1pt --- -.1pt --- (2025 b ): Winsorized mean estimation with heavy tails and adversarial contamination, arXiv preprint arXiv:2504.08482

  15. [23]

    --- -.1pt --- -.1pt --- (2026 a ): Enhanced power enhancements for testing many moment equalities: Beyond the 2 -and -norm, to appear in Journal of the American Statistical Association

  16. [24]

    --- -.1pt --- -.1pt --- (2026 b ): Robustness for free: asymptotic size and power of max-tests in high dimensions, arXiv preprint 2601.14013

  17. [25]

    Lai, K. A., A. B. Rao, and S. Vempala (2016): Agnostic estimation of mean and covariance, in 2016 IEEE 57th Annual Symposium on Foundations of Computer Science (FOCS), IEEE, 665--674

  18. [26]

    Liu, M. and M. E. Lopes (2024): Robust max statistics for high-dimensional inference, arXiv preprint arXiv:2409.16683

  19. [27]

    Lugosi, G. and S. Mendelson (2021): Robust multivariate mean estimation: The optimality of trimmed mean , Annals of Statistics, 49, 393--410

  20. [28]

    Minasyan, A. and N. Zhivotovskiy (2023): Statistically optimal robust mean and covariance estimation for anisotropic gaussians, arXiv preprint arXiv:2301.09024

  21. [29]

    (2023): Efficient median of means estimator, in The Thirty Sixth Annual Conference on Learning Theory, PMLR, 5925--5933

    Minsker, S. (2023): Efficient median of means estimator, in The Thirty Sixth Annual Conference on Learning Theory, PMLR, 5925--5933

  22. [30]

    Minsker, S. and M. Ndaoud (2021): Robust and efficient mean estimation: an approach based on the properties of self-normalized sums, Electronic Journal of Statistics, 15, 6036--6070

  23. [31]

    Orenstein, and Z

    Oliveira, R., P. Orenstein, and Z. Rico (2025): Finite-sample properties of the trimmed mean, arXiv preprint arXiv:2501.03694

  24. [32]

    Qiu, J., S. X. Chen, and Q.-M. Shao (2025): Self-normalized Cram \'e r type moderate deviation theorem for Gaussian approximation, The Annals of Statistics, 53, 1319--1346

  25. [33]

    (2024): Robust high-dimensional Gaussian and bootstrap approximations for trimmed sample means, arXiv preprint arXiv:2410.22085

    Resende, L. (2024): Robust high-dimensional Gaussian and bootstrap approximations for trimmed sample means, arXiv preprint arXiv:2410.22085

  26. [34]

    Romano, J. P. and M. Wolf (2005): Exact and approximate stepdown methods for multiple hypothesis testing, Journal of the American Statistical Association, 100, 94--108

  27. [35]

    (1967): Rectangular confidence regions for the means of multivariate normal distributions, Journal of the American statistical association, 62, 626--633

    S id \'a k, Z. (1967): Rectangular confidence regions for the means of multivariate normal distributions, Journal of the American statistical association, 62, 626--633

  28. [36]

    Peng, and R

    Wang, L., B. Peng, and R. Li (2015): A high-dimensional nonparametric multivariate test for mean vector, Journal of the American Statistical Association, 110, 1658--1669

  29. [37]

    Xu, G., L. Lin, P. Wei, and W. Pan (2016): An adaptive two-sample test for high-dimensional means, Biometrika, 103, 609--624

  30. [38]

    Xue, K. and F. Yao (2020): Distribution and correlation-free two-sample test of high-dimensional means, The Annals of Statistics, 48, 1304--1328

  31. [39]

    Zheng, and R

    Yang, S., S. Zheng, and R. Li (2024): A new test for high-dimensional two-sample mean problems with consideration of correlation structure, Annals of Statistics, 52, 2217--2240

  32. [40]

    Zhang, D. and W. B. Wu (2017): Gaussian approximation for high dimensional time series , Annals of Statistics, 45, 1895--1919

  33. [41]

    Zhang, J.-T., J. Guo, B. Zhou, and M.-Y. Cheng (2020): A simple two-sample test in high dimensions based on L^2 -norm, Journal of the American Statistical Association, 115, 1011--1027

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.