REVIEW 3 major objections 6 minor 92 references
Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation
T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read For penalized logistic regression in the proportional regime, Gaussian universality and the convex Gaussian min-max theorem hold even for dependent data, and the asymptotic risk depends only on the data's mean and covariance — with data…
desk verdict The dependent CGMT and universality results are real and important; the data augmentation application is not yet proved for the unconstrained estimator because the S_p restriction is left unresolved. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery has three load-bearing parts. First, the block-dependent universality theorem (Theorem 2) uses a Lindeberg-style interpolation between the original data and a matched-covariance Gaussian surrogate, with Assumption 5 requiring joint Gaussian approximation of $k$-tuples of projections $X_{i+r}^{\top}\beta_r$ for $r \le k$, weighted on the sphere, instead of the single projection needed in the independent case. Second, the dependent CGMT (Theorem 5) handles a Gaussian min-max problem whose design matrix $H$ has covariance $\mathrm{Cov}[H_{ji},H_{j'i'}] = \sum_{l=1}^{M} \Sigma^{(l)}_{jj'}\tilde{\Sigma}^{(l)}_{ii'}$, a low-rank sum of factored row and column covariances; for data augmentation $M = 2$ suffices because the covariance of augmented copies is determined by $\mathrm{Var}[\phi_1(Z_1)]$ and $\mathrm{Cov}[\phi_1(Z_1), \phi_2(Z_1)]$. Third, Assumption 7 — that the Gaussian training risk is sharply minimized on the thin shell $|(\beta^{\top}\Sigma_{\text{new}}\beta)^{1/2} - \bar{\chi}| \le \epsilon$ — is what converts training-risk universality into test-risk universality, and the paper verifies it through the CGMT so long as the minimizer-maximizers of the deterministic auxiliary problem lie in the interior of its domain. The final output is the system of ten equations (EQs) whose solution pins down the test risk through the scalar $\bar{\chi}$.
What would settle it
Run the paper's own comparisons at block sizes and aspect ratios beyond the simulated range: if the excess test risk of penalized logistic regression on, say, shifted-gamma covariates differs from the matched-covariance Gaussian surrogate by more than the $\sqrt{k}$ universality bound the authors derive, the claim that dependence enters only through the covariance fails. For the augmentation conclusion, check whether increasing the number of 80%-permutation augmentations eventually lowers the test risk at large $k$; if it does, the qualitative claim that partial invariance is as good as no augmentation would be overturned.
Extended reading notes
Core claim
The central claim is that the asymptotic risk of high-dimensional penalized logistic regression is universal across data distributions once the first two moments are fixed, even when the observations are dependent. Theorems 2 and 3 show that the minimum training risk and the test risk of the estimator on block-dependent, $m$-dependent, or suitably $\beta$-mixing data converge to those of a Gaussian dataset with the same covariance; a key consequence stated in the paper is that uncorrelated dependent data give the same asymptotic risk as independent data. Because the Gaussian surrogate may still have correlated rows and columns, the paper proves a dependent CGMT (Theorem 5) under an assumption that the design covariance factorizes as a sum of Kronecker products of row and column covariance matrices, and uses it to verify the sharp-shell condition that upgrades training-risk universality to test-risk universality. For data augmentation, the estimator trained on $k$ transformed copies of each observation is shown to have a test risk characterized by a deterministic system of ten scalar equations (Theorem 12); the numerical solution shows that full permutation of exchangeable coordinates reduces the test risk substantially, whereas $r_{\text{perm}} = 0.8$ permutations, sign flipping, and cropping without full knowledge of the zero coordinates yield risks within error margins of no augmentation.
Load-bearing premise
The load-bearing premise is that the constrained minimizer on the set $S_p$ is the true minimizer and that, in the Gaussian surrogate, the training risk is uniquely minimized on a thin shell of constant $(\beta^{\top}\Sigma_{\text{new}}\beta)^{1/2}$; the paper verifies the shell condition only when the auxiliary deterministic optimization has interior solutions, and a general $S_p$ bound for all data-augmentation schemes is explicitly left to future work.
Editorial extensions
If this is right
- Uncorrelated but dependent data inherit all previously derived independent-data results for logistic regression, since the asymptotic risk is governed by mean and covariance alone.
- Correlated dependent data can be studied by replacing the data with a Gaussian model of the same covariance and applying the dependent CGMT, avoiding random-matrix theory.
- Data augmentation under the full permutation group of exchangeable coordinate blocks improves the test risk as more augmentations are used.
- Augmenting only a large fraction of the coordinates ($r_{\text{perm}} = 0.8$) leaves the test risk statistically unchanged relative to no augmentation.
- Sign flipping and cropping that do not know the exact zero coordinates of the signal give the same test risk as no augmentation.
Reading between the lines
- If the mean-variance reduction is as general as the proof suggests, the same universality pipeline should transfer to other classifiers whose loss depends on the data through one-dimensional projections, such as SVMs and other generalized linear models — the paper notes this extension is direct.
- The $r_{\text{perm}} = 0.8$ plateau suggests a testable design rule for practitioners: augmentation groups that cover only part of the true symmetry may be wasting compute, and symmetry-learning methods should beat hand-chosen partial augmentations even though the paper does not run that comparison.
- The simulations' observation that training trajectories need different learning rates for t-distributed versus uniform data, even though global minima are universal, indicates that risk universality will not extend to optimization dynamics; proving trajectory non-universality would be a natural sequel.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends two workhorses of high-dimensional statistics — Gaussian universality and the convex Gaussian min-max theorem (CGMT) — from independent observations to dependent ones. For penalized logistic regression with labels generated by a logistic link, it proves universality of the training and test risks (matching to a Gaussian surrogate with the same covariance) under block dependence (Theorem 2), m-dependence, and a restricted class of β-mixing processes (Theorem 3). The CGMT part (Theorem 5/13) handles Gaussian design matrices whose covariance is a sum of Kronecker products of row- and column-side matrices (Assumption 10), with comparison probabilities off by a factor 2^M. The tools are applied to data augmentation schemes — random permutations, sign flipping, and cropping — culminating in a ten-equation fixed-point system (EQs) claimed to characterize the asymptotic test risk (Theorem 12). Simulations across coordinate distributions (Gaussian, uniform, gamma, exponential, t₃) confirm the predicted risk values and illustrate that full permutation invariance helps while partial structure knowledge does not.
Significance. If the main theorems hold, this is a substantial contribution: it removes the independence-of-rows assumption in a principled way for high-dimensional logistic regression, and Theorem 5 is a model-independent tool — the 2^M comparison bounds (Lemmas 45–46) and the low-rank Kronecker-factor covariance structure (Assumption 10) are likely to be reused beyond this paper. The proof work is real: the within-block Lindeberg control (Lemmas 24–28) is a genuine extension of the Montanari–Saeed template, and the deterministic ten-equation system (EQs) is anchored by Lemma 52, which recovers exactly Salehi et al.'s six equations in the isotropic no-augmentation limit. The authors ship reproducible code (footnote in Section 6), give detailed simulation settings (Appendix C), and are unusually candid about what is not proved (the S_p membership in Appendix D.1 and the non-universality of training trajectories in Section 9). The main caveat is that the data-augmentation application's headline claims are conditional on the deferred ℓ∞ bound and on unverified geometric properties of the (DO) limit; these gaps are load-bearing for the practical, unconstrained estimator.
major comments (3)
- [Section 2 (Eq. 4), Section D.2, Appendix D.1] The paper proves universality for the S_p-constrained minimizer (Theorem 15, for any S̃ ⊆ S_p), yet Theorem 2(5) and Theorem 3(7) are stated with the unconstrained min_β R̂_n(β;·), and the practical estimator (3) is the unconstrained minimizer. Section D.2 asserts that Theorem 15 implies (5) 'by setting S̃ = S_p', but this step requires the ℓ∞ membership \hatβ ∈ S_p, i.e. ‖\hatβ‖_∞ ≤ L p^{1/2−r}. Appendix D.1 explicitly leaves this to future work and explains that the standard leave-one-out singular-value route (σ_min(XX^T) ≥ p/C) is 'nearly impossible' when blocks contain identical rows — precisely the within-block dependence created by data augmentation. Until this membership is established for the DA estimators, the abstract's claim to 'establish the impact of data augmentation' overstates what is proved: Theorem 12 applies to the S_p-restricted (OO), not to the unconstrained estimator (3) that is fitted in the simulations. The authors should either supply the ℓ∞ bound for the DA schemes they study, or restate the main theorems and the DA conclusions in explicitly constrained form.
- [Section B.1, Theorem 12 and its proof] Test-risk universality (Theorem 2(6)) requires Assumption 7, and the proof of Theorem 12 verifies Assumption 7 by invoking two premises: that the minimizer-maximizers of (DO) lie in the interior of the domain of optimization, and that restricting the β-domain to |(β^T Σ_new β)^{1/2} − χ̄| > ε changes the (DO) value by Θ(ε²). The first is stated only as an assumption of the theorem, and the second is asserted without derivation. These premises are exactly what makes the Gaussian training risk sharply minimized on the thin shell; no verification for the permutation, sign-flip, or crop schemes is provided. The conclusion |R_test(\hatβ(X,X^Φ)) − R̄_test(χ̄²)| → 0 is therefore conditional on an unverified geometric property of the deterministic limit, and this should be stated as an open condition rather than presented as a completed verification.
- [Section D.5, Eqs. (16)–(19) and the final display of the proof of Theorem 15] The Lindeberg interpolation yields an error bound of order √k γ + √k δ + α^{−1} log(1/δ) + τ, so the result holds for each fixed block size k and contains no control uniform in k. In the data-augmentation application, k (the number of augmentations) is a free parameter that the simulations vary (Figures 1 and 2), and the text draws conclusions 'as more and more augmentations are used'. The theorems and the (EQs) derivation cover the regime where k is held fixed as p → ∞; the k → ∞ regime is outside the current proof. This limitation should be stated explicitly wherever the DA conclusions are drawn, since the figures' x-axis is exactly the parameter that the theory does not let grow.
minor comments (6)
- [Abstract] The phrase 'assumption that significantly limit its applicability' has a subject-verb agreement error and should read 'significantly limits'.
- [Section B.1, Theorem 12] The theorem statement says the test risk is characterized by the parameter set (r, θ, σ, τ), while the system (EQs) solves for ten parameters (α, σ_1, σ_2, τ_1, τ_2, ν_1, ν_2, r_1, r_2, θ); the notation should be aligned between the statement and the equations.
- [Section B.1, Theorem 12 proof] The word 'estbalished' appears in the proof and should be corrected to 'established'.
- [Section 6, Eq. (11) and Section 1.1, Eq. (3)] Eq. (11) restricts the minimization to S_p while the original problem (3) is unconstrained; this distinction is only explained via Remark 1 and Appendix B.3, and it should be stated in the main text that all DA theorems concern the S_p-constrained problem, with the unconstrained version open (see Major Comment 1).
- [Section 2 and throughout] The symbol 'B' is used as a substitute for '≔' throughout the manuscript; it is nonstandard and should be defined at first use or replaced by standard notation.
- [Section 6, Figure 1] The curves for r_perm = 0.8 and no augmentation appear to overlap, and one of the paper's central findings is that partial permutations are no better than no augmentation; the caption should report the trial counts (50 versus 200) and the error-bar convention so that the 'within error margins' claim is checkable.
Circularity Check
No significant circularity: the dependent-CGMT-to-EQ derivation is self-contained and externally anchored by recovery of Salehi et al. (2019); the S_p and interior-minimizer caveats are proof gaps, not circular reductions.
full rationale
The paper's derivation chain is self-contained rather than circular. Section 5 proves the dependent CGMT (Theorem 5) from Gordon's comparison inequality plus a sign-symmetrization argument, and Remark 8 explicitly recovers the standard CGMT and the block-diagonal CGMT of Dhifallah and Lu (2021) as special cases. Sections M and N reduce the data-augmentation logistic problem (OO) through (GO), (PO), (AO), (SO) to the deterministic optimization (DO); the ten parameters in (EQs) are self-consistent fixed-point solutions of the reduced min-max problem, not constants fitted to simulations. The predicted test risk R_bar_test(chi_bar^2) is a function of the DO solution, which is derived from the model, and the figures run genuine gradient descent on non-Gaussian data against Gaussian surrogates, an independent numerical check. External anchoring is explicit: the paper verifies that 'in the isotropic case with no augmentation, it recovers exactly the characterizing equation by Salehi et al. (2019)' (Section 6). The self-citations are either honestly positioned as prior work whose conditions are 'hard to verify and this paper does not cover overparameterized logistic regression' (Section 7), or are peer-reviewed technical lemmas that the paper also states itself (Appendix K.3 reproduces the smoothing construction). The load-bearing caveats, namely the S_p restriction deferred in Appendix D.1 ('To extend this into a proof covering all DA schemes will be left to future work') and Assumption 7's verification via the unproved interior-minimizer regularity of (DO) in Theorem 12, are proof gaps or conditional claims rather than reductions of predictions to their inputs. No step was found in which a fitted quantity is renamed as a prediction or in which a self-citation is the sole justification for a central premise.
Assumptions & free parameters
free parameters (3)
- r_perm, r_flip, r_crop (fractions of coordinates augmented) =
0.8 and 1.0 for permutations; 0.2 for flip and crop in figures
- k (number of augmentations per observation) =
k = 11 and k = 30 in simulations
- S_p shape constants L and r
assumptions (9)
- domain assumption Assumption 1/8: dependence classes, block dependence with block size k, m-dependence, or β-mixing with summable coefficients
- domain assumption Assumption 2/12: logistic link labels, y_i driven by X_i^T beta* minus Logistic noise, or by a block-averaged linear form
- domain assumption Assumption 3: centered sub-Gaussian covariates with sup_i ||X_i||_{psi2} <= K_X / sqrt(n)
- domain assumption Assumptions 5/9: uniform joint Gaussian approximation of (X_{i+r}^T beta_r) over S_p^k and theta in S^{k-1} (and for every fixed d in the mixing case)
- domain assumption Assumption 7: the Gaussian training risk is minimized on a thin Sigma_new-shell, with chi_bar and chi_* bounded away from zero
- domain assumption Assumption 10: low-rank covariance factorization Cov[H_ji, H_j'i'] = sum_l Sigma^{(l)}_{jj'} Sigma_tilde^{(l)}_{ii'}
- ad hoc to paper Assumption 11: Sigma_* = (Sigma^dagger)^{1/2} Cov[phi_1(Z_1), Z_1] (Sigma_o^dagger)^{1/2} and Sigma_*^2 = Sigma_*
- standard math Yu (1994) coupling/embedding for β-mixing blocks
- standard math Gordon's Gaussian comparison inequality (Ledoux-Talagrand form)
invented entities (1)
-
Low-rank factor decomposition (Sigma^{(l)}, Sigma_tilde^{(l)}) of the correlated-Gaussian data matrix covariance
Cite this review
Pith. "Pith review of Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation." pith.science (2026). https://pith.science/paper/TYMOLYNF
@misc{pith2026250215752,
author = {Pith},
title = {Pith review of: Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/TYMOLYNF}},
note = {Machine review of arXiv:2502.15752}
}
abstract
Over the last decade, a wave of research has characterized the exact asymptotic risk of many high-dimensional models in the proportional regime. Two foundational results have driven this progress: Gaussian universality, which shows that the asymptotic risk of estimators trained on non-Gaussian and Gaussian data is equivalent, and the convex Gaussian min-max theorem (CGMT), which characterizes the risk under Gaussian settings. However, these results rely on the assumption that the data consists of independent random vectors--an assumption that significantly limits its applicability to many practical setups. In this paper, we address this limitation by generalizing both results to the dependent setting. More precisely, we prove that Gaussian universality still holds for high-dimensional logistic regression under block dependence, $m$-dependence and special cases of mixing, and establish a novel CGMT framework that accommodates for correlation across both the covariates and observations. Using these results, we establish the impact of data augmentation, a widespread practice in deep learning, on the asymptotic risk.
Figures
Reference graph
Works this paper leans on
-
[1]
Universality in learning from linear measurements
Ehsan Abbasi, Fariborz Salehi, and Babak Hassibi. Universality in learning from linear measurements. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[2]
A novel G aussian min-max theorem and its applications
Danil Akhtiamov, David Bosch, Reza Ghane, K Nithin Varma, and Babak Hassibi. A novel G aussian min-max theorem and its applications. arXiv preprint arXiv:2402.07356, 2024 a
arXiv 2024
-
[3]
Regularized linear regression for binary classification
Danil Akhtiamov, Reza Ghane, and Babak Hassibi. Regularized linear regression for binary classification. In 2024 IEEE International Symposium on Information Theory (ISIT), pages 202--207. IEEE, 2024 b
2024
-
[4]
Multiple fourier series and fourier integrals
Sh A Alimov, RR Ashurov, and AK Pulatov. Multiple fourier series and fourier integrals. Commutative Harmonic Analysis IV: Harmonic Analysis in IR n, pages 1--95, 1992
1992
-
[5]
Liviu Aolaritei, Soroosh Shafieezadeh-Abadeh, and Florian D \"o rfler. The performance of W asserstein distributionally robust M -estimators in high dimensions. arXiv preprint arXiv:2206.13269, 2022
work page Pith review arXiv 2022
-
[6]
Limit theorems for distributions invariant under groups of transformations
Morgane Austern and Peter Orbanz. Limit theorems for distributions invariant under groups of transformations. The Annals of Statistics, 50 0 (4): 0 1960--1991, 2022
1960
-
[7]
Regularized estimation in sparse high-dimensional time series models
Sumanta Basu and George Michailidis. Regularized estimation in sparse high-dimensional time series models. 2015
2015
-
[8]
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences, 116 0 (32): 0 15849--15854, 2019
2019
Show all 92 references
-
[9]
Learning invariances in neural networks from training data
Gregory Benton, Marc Finzi, Pavel Izmailov, and Andrew G Wilson. Learning invariances in neural networks from training data. Advances in neural information processing systems, 33: 0 17605--17616, 2020
2020
-
[10]
Sur l'extension du th \'e or \`e me limite du calcul des probabilit \'e s aux sommes de quantit \'e s d \'e pendantes
Serge Bernstein. Sur l'extension du th \'e or \`e me limite du calcul des probabilit \'e s aux sommes de quantit \'e s d \'e pendantes. Mathematische Annalen, 97: 0 1--59, 1927
1927
-
[11]
Probability and measure
Patrick Billingsley. Probability and measure. John Wiley & Sons, 1995
1995
-
[12]
Logistic regression for dependent binary observations
George Ebow Bonney. Logistic regression for dependent binary observations. Biometrics, pages 951--973, 1987
1987
-
[13]
Basic properties of strong mixing conditions
Richard C Bradley. Basic properties of strong mixing conditions. a survey and some open questions. Probability Surveys, 2: 0 107--114, 2005
2005
-
[14]
Simple technical trading rules and the stochastic properties of stock returns
William Brock, Josef Lakonishok, and Blake LeBaron. Simple technical trading rules and the stochastic properties of stock returns. The Journal of finance, 47 0 (5): 0 1731--1764, 1992
1992
-
[15]
Distributional and lq norm inequalities for polynomials over convex bodies in rn
Anthony Carbery and James Wright. Distributional and lq norm inequalities for polynomials over convex bodies in rn. Mathematical research letters, 8 0 (3): 0 233--248, 2001
2001
-
[16]
Concentration inequalities with exchangeable pairs
Sourav Chatterjee. Concentration inequalities with exchangeable pairs. Stanford University, 2005
2005
-
[17]
A group-theoretic framework for data augmentation
Shuxiao Chen, Edgar Dobriban, and Jane H Lee. A group-theoretic framework for data augmentation. Journal of Machine Learning Research, 21 0 (245): 0 1--71, 2020
2020
-
[18]
Universality of approximate message passing algorithms
Wei-Kuo Chen and Wai-Kit Lam. Universality of approximate message passing algorithms . Electronic Journal of Probability, 26: 0 1 -- 44, 2021
2021
-
[19]
Statistics for spatial data
Noel Cressie. Statistics for spatial data. John Wiley & Sons, 1993
1993
-
[20]
Time series analysis, volume 286
Jonathan D Cryer. Time series analysis, volume 286. Duxbury Press Boston, 1986
1986
-
[21]
Universality laws for G aussian mixtures in generalized linear models
Yatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro, and Lenka Zdeborov \'a . Universality laws for G aussian mixtures in generalized linear models. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[22]
A central limit theorem for globally nonstationary near-epoch dependent functions of mixing processes
James Davidson. A central limit theorem for globally nonstationary near-epoch dependent functions of mixing processes. Econometric theory, 8 0 (3): 0 313--329, 1992
1992
-
[23]
A model of double descent for high-dimensional binary linear classification
Zeyu Deng, Abla Kammoun, and Christos Thrampoulidis. A model of double descent for high-dimensional binary linear classification. Information and Inference: A Journal of the IMA, 11 0 (2): 0 435--495, 2022
2022
-
[24]
A note on empirical processes of strong-mixing sequences
Chandrakant M Deo. A note on empirical processes of strong-mixing sequences. The Annals of Probability, pages 870--875, 1973
1973
-
[25]
On the inherent regularization effects of noise injection during training
Oussama Dhifallah and Yue Lu. On the inherent regularization effects of noise injection during training. In International Conference on Machine Learning, pages 2665--2675. PMLR, 2021
2021
-
[26]
Message-passing algorithms for compressed sensing
David L Donoho, Arian Maleki, and Andrea Montanari. Message-passing algorithms for compressed sensing. Proceedings of the National Academy of Sciences, 106 0 (45): 0 18914--18919, 2009
2009
-
[27]
Lu, and Subhabrata Sen
Rishabh Dudeja, Yue M. Lu, and Subhabrata Sen. Universality of approximate message passing with semirandom matrices. The Annals of Probability, 51 0 (5): 0 1616--1683, 2023
2023
-
[28]
Handbook of spatial statistics
Alan E Gelfand, Peter Diggle, Peter Guttorp, and Montserrat Fuentes. Handbook of spatial statistics. CRC press, 2010
2010
-
[29]
Gaussian universality of perceptrons with random labels
Federica Gerace, Florent Krzakala, Bruno Loureiro, Ludovic Stephan, and Lenka Zdeborov \'a . Gaussian universality of perceptrons with random labels. Physical Review E, 109 0 (3): 0 034305, 2024
2024
-
[30]
Some inequalities for G aussian processes and applications
Yehoram Gordon. Some inequalities for G aussian processes and applications. Israel Journal of Mathematics, 50: 0 265--289, 1985
1985
-
[31]
High dimensional and banded vector autoregressions, 2016
Shaojun Guo, Yazhen Wang, and Qiwei Yao. High dimensional and banded vector autoregressions, 2016. URL https://arxiv.org/abs/1502.07831
2016 arXiv
-
[32]
Universality of regularized regression estimators in high dimensions
Qiyang Han and Yandi Shen. Universality of regularized regression estimators in high dimensions. The Annals of Statistics, 51 0 (4): 0 1799--1823, 2023
2023
-
[33]
Data augmentation as stochastic optimization
Boris Hanin and Yi Sun. Data augmentation as stochastic optimization. 2020
2020
-
[34]
Analysis of dichotomous response data from certain toxicological experiments
JK Haseman and LL Kupper. Analysis of dichotomous response data from certain toxicological experiments. Biometrics, pages 281--293, 1979
1979
-
[35]
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani. Surprises in high-dimensional ridgeless least squares interpolation. The Annals of Statistics, 50 0 (2): 0 949, 2022
2022
-
[36]
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu. Universality laws for high-dimensional learning with random features. IEEE Transactions on Information Theory, 69 0 (3): 0 1932--1964, 2022
1932
-
[37]
Data augmentation in the underparameterized and overparameterized regimes
Kevin Han Huang, Peter Orbanz, and Morgane Austern. Data augmentation in the underparameterized and overparameterized regimes. arXiv preprint arXiv:2202.09134, 2022
2022
-
[38]
A high-dimensional convergence theorem for u-statistics with applications to kernel-based testing
Kevin Han Huang, Xing Liu, Andrew Duncan, and Axel Gandy. A high-dimensional convergence theorem for u-statistics with applications to kernel-based testing. In The Thirty Sixth Annual Conference on Learning Theory, pages 3827--3918. PMLR, 2023
2023
-
[39]
Independent and stationary sequences of random variables
I Ibragimov. Independent and stationary sequences of random variables. Wolters, Noordhoff Pub., 1975
1975
-
[40]
The theory of approximation, volume 11
Dunham Jackson. The theory of approximation, volume 11. American Mathematical Soc., 1930
1930
-
[41]
Precise statistical analysis of classification accuracies for adversarial training
Adel Javanmard and Mahdi Soltanolkotabi. Precise statistical analysis of classification accuracies for adversarial training. The Annals of Statistics, 50 0 (4): 0 2127--2156, 2022
2022
-
[42]
Asymptotic behavior of unregularized and ridge-regularized high-dimensional robust regression estimators : rigorous results, 2013
Noureddine El Karoui. Asymptotic behavior of unregularized and ridge-regularized high-dimensional robust regression estimators : rigorous results, 2013
2013
-
[43]
Label-imbalanced and group-sensitive classification under overparameterization
Ganesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, and Christos Thrampoulidis. Label-imbalanced and group-sensitive classification under overparameterization. Advances in Neural Information Processing Systems, 34: 0 18970--18983, 2021
2021
-
[44]
Applications of the lindeberg principle in communications and statistical learning
Satish Babu Korada and Andrea Montanari. Applications of the lindeberg principle in communications and statistical learning. IEEE transactions on information theory, 57 0 (4): 0 2440--2450, 2011
2011
-
[45]
Universality in block dependent linear models with applications to nonparametric regression
Samriddha Lahiry and Pragya Sur. Universality in block dependent linear models with applications to nonparametric regression. IEEE Transactions on Information Theory, 70 0 (12): 0 8975--9000, 2024
2024
-
[46]
Probability in Banach Spaces: isoperimetry and processes
Michel Ledoux and Michel Talagrand. Probability in Banach Spaces: isoperimetry and processes. Springer Science & Business Media, 1991
1991
-
[47]
The good, the bad and the ugly sides of data augmentation: An implicit spectral regularization perspective
Chi-Heng Lin, Chiraag Kaushik, Eva L Dyer, and Vidya Muthukumar. The good, the bad and the ugly sides of data augmentation: An implicit spectral regularization perspective. Journal of Machine Learning Research, 25 0 (91): 0 1--85, 2024
2024
-
[48]
On the benefits of invariance in neural networks
Clare Lyle, Mark van der Wilk, Marta Kwiatkowska, Yarin Gal, and Benjamin Bloem-Reddy. On the benefits of invariance in neural networks. In Conference on Neural Information Processing Systems: Workshop on Machine Learning with Guarantees, 2019
2019
-
[49]
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari. The generalization error of random features regression: Precise asymptotics and the double descent curve. Communications on Pure and Applied Mathematics, 75 0 (4): 0 667--766, 2022
2022
-
[50]
The role of regularization in classification of high-dimensional noisy G aussian mixture
Francesca Mignacco, Florent Krzakala, Yue Lu, Pierfrancesco Urbani, and Lenka Zdeborova. The role of regularization in classification of high-dimensional noisy G aussian mixture. In International conference on machine learning, pages 6874--6883. PMLR, 2020
2020
-
[51]
Universality of the elastic net error
Andrea Montanari and Phan-Minh Nguyen. Universality of the elastic net error. In 2017 IEEE International Symposium on Information Theory (ISIT), pages 2338--2342. IEEE, 2017
2017
-
[52]
Universality of empirical risk minimization
Andrea Montanari and Basil N Saeed. Universality of empirical risk minimization. In Conference on Learning Theory, pages 4310--4312. PMLR, 2022
2022
-
[53]
Universality of max-margin classifiers
Andrea Montanari, Feng Ruan, Basil Saeed, and Youngtak Sohn. Universality of max-margin classifiers. arXiv preprint arXiv:2310.00176, 2023
2023 arXiv
-
[54]
High dimensional logistic regression under network dependence
Somabha Mukherjee, Ziang Niu, Sagnik Halder, Bhaswar B Bhattacharya, and George Michailidis. High dimensional logistic regression under network dependence. arXiv preprint arXiv:2110.03200, 2021
2021 arXiv
-
[55]
Least squares regression with markovian data: Fundamental limits and algorithms
Dheeraj Nagaraj, Xian Wu, Guy Bresler, Prateek Jain, and Praneeth Netrapalli. Least squares regression with markovian data: Fundamental limits and algorithms. Advances in neural information processing systems, 33: 0 16666--16676, 2020
2020
-
[56]
Nicholson, Ines Wilms, Jacob Bien, and David S
William B. Nicholson, Ines Wilms, Jacob Bien, and David S. Matteson. High dimensional forecasting via interpretable vector autoregression. Journal of Machine Learning Research, 21 0 (166): 0 1--52, 2020. URL http://jmlr.org/papers/v21/19-777.html
2020
-
[57]
From naive mean field theory to the tap equations
Manfred Opper, Ole Winther, et al. From naive mean field theory to the tap equations. Advanced mean field methods: theory and practice, pages 7--20, 2001
2001
-
[58]
Universality laws for randomized dimension reduction, with applications
Samet Oymak and Joel A Tropp. Universality laws for randomized dimension reduction, with applications. Information and Inference: A Journal of the IMA, 7 0 (3): 0 337--446, 2018
2018
-
[59]
Joost D. J. Plate, Rutger R. van de Leur, Luke P.H. Leenen, Falco Hietbrink, Linda M. Peelen, and Marinus J. C. Eijkemans. Incorporating repeated measurements into prediction models in the critical care setting: a framework, systematic review and meta-analysis. BMC Medical Res...
2019
-
[60]
Correlated binary regression with covariates specific to each binary observation
Ross L Prentice. Correlated binary regression with covariates specific to each binary observation. Biometrics, pages 1033--1048, 1988
1988
-
[61]
Locally dependent latent class models with covariates: an application to under-age drinking in the usa
Beth A Reboussin, Edward H Ip, and Mark Wolfson. Locally dependent latent class models with covariates: an application to under-age drinking in the usa. Journal of the Royal Statistical Society Series A: Statistics in Society, 171 0 (4): 0 877--897, 2008
2008
-
[62]
Linear models in statistics
Alvin C Rencher and G Bruce Schaalje. Linear models in statistics. John Wiley & Sons, 2008
2008
-
[63]
Tyrrell Rockafellar
R. Tyrrell Rockafellar. Convex analysis. Princeton Mathematical Series. Princeton University Press, Princeton, N. J., 1970
1970
-
[64]
Fundamentals of Stein’s method
Nathan Ross. Fundamentals of Stein’s method . Probability Surveys, 8: 0 210 -- 293, 2011
2011
-
[65]
Hanson-wright inequality and sub-gaussian concentration, 2013
Mark Rudelson and Roman Vershynin. Hanson-wright inequality and sub-gaussian concentration, 2013
2013
-
[66]
The impact of regularization on high-dimensional logistic regression
Fariborz Salehi, Ehsan Abbasi, and Babak Hassibi. The impact of regularization on high-dimensional logistic regression. Advances in Neural Information Processing Systems, 32, 2019
2019
-
[67]
Local dependence in random graph models: characterization, properties and statistical inference
Michael Schweinberger and Mark S Handcock. Local dependence in random graph models: characterization, properties and statistical inference. Journal of the Royal Statistical Society Series B: Statistical Methodology, 77 0 (3): 0 647--676, 2015
2015
-
[68]
Advanced Data Analysis from an Elementary Point of View
Cosma Rohilla Shalizi. Advanced Data Analysis from an Elementary Point of View. 2019
2019
-
[69]
Shorten and T
C. Shorten and T. M Khoshgoftaar. A survey on image data augmentation for deep learning. Journal of Big Data, 6 0 (1): 0 1--48, 2019
2019
-
[70]
Text data augmentation for deep learning
Connor Shorten, Taghi M Khoshgoftaar, and Borko Furht. Text data augmentation for deep learning. Journal of big Data, 8 0 (1): 0 101, 2021
2021
-
[71]
A framework to characterize performance of lasso algorithms
Mihailo Stojnic. A framework to characterize performance of lasso algorithms. arXiv preprint arXiv:1303.7291, 2013 a
2013 arXiv
-
[72]
Upper-bounding _1 -optimization weak thresholds
Mihailo Stojnic. Upper-bounding _1 -optimization weak thresholds. arXiv preprint arXiv:1303.7289, 2013 b
2013 arXiv
-
[73]
A modern maximum-likelihood theory for high-dimensional logistic regression
Pragya Sur and Emmanuel J Cand \`e s. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences, 116 0 (29): 0 14516--14525, 2019
2019
-
[74]
The impact of multi-optimizers and data augmentation on tensorflow convolutional neural network performance
Arwa Mohammed Taqi, Ahmed Awad, Fadwa Al-Azzo, and Mariofanna Milanova. The impact of multi-optimizers and data augmentation on tensorflow convolutional neural network performance. In Proc. of IEEE MIPR, pages 140--145, 2018
2018
-
[75]
Recovering structured signals in high dimensions via non-smooth convex optimization: Precise performance analysis
Christos Thrampoulidis. Recovering structured signals in high dimensions via non-smooth convex optimization: Precise performance analysis. PhD thesis, California Institute of Technology, 2016
2016
-
[76]
The G aussian min-max theorem in the presence of convexity
Christos Thrampoulidis, Samet Oymak, and Babak Hassibi. The G aussian min-max theorem in the presence of convexity. arXiv preprint arXiv:1408.4837, 2014
2014 arXiv
-
[77]
Regularized linear regression: A precise analysis of the estimation error
Christos Thrampoulidis, Samet Oymak, and Babak Hassibi. Regularized linear regression: A precise analysis of the estimation error. In Conference on Learning Theory, pages 1683--1709. PMLR, 2015
2015
-
[78]
Precise error analysis of regularized m -estimators in high dimensions
Christos Thrampoulidis, Ehsan Abbasi, and Babak Hassibi. Precise error analysis of regularized m -estimators in high dimensions. IEEE Transactions on Information Theory, 64 0 (8): 0 5592--5628, 2018
2018
-
[79]
Analysis of financial time series
Ruey S Tsay. Analysis of financial time series. John wiley & sons, 2005
2005
-
[80]
Some mixing properties of time series models
D Tuan and T Lanh. Some mixing properties of time series models. Stochastic processes and their applications, 19: 0 297--303, 1985
1985
-
[81]
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press, 2018
2018
-
[82]
An overview on data augmentation for machine learning
Svetlana Volkova. An overview on data augmentation for machine learning. In Arthur Gibadullin, editor, Digital and Information Technologies in Economics and Management, pages 143--154, Cham, 2024. Springer Nature Switzerland. ISBN 978-3-031-55349-3
2024
-
[83]
Multivariate geostatistics: an introduction with applications
Hans Wackernagel. Multivariate geostatistics: an introduction with applications. Springer Science & Business Media, 2003
2003
-
[84]
Universality of approximate message passing algorithms and tensor networks
Tianhao Wang, Xinyi Zhong, and Zhou Fan. Universality of approximate message passing algorithms and tensor networks. The Annals of Applied Probability, 34 0 (4): 0 3943--3994, 2024
2024
-
[85]
Regularized estimation in high dimensional time series under mixing conditions
Kam Chung Wong, Ambuj Tewari, and Zifan Li. Regularized estimation in high dimensional time series under mixing conditions. stat, 1050: 0 12, 2016
2016
-
[86]
Lasso guarantees for -mixing heavy-tailed time series
Kam Chung Wong, Zifan Li, and Ambuj Tewari. Lasso guarantees for -mixing heavy-tailed time series. The Annals of Statistics, 48 0 (2): 0 1124--1142, 2020
2020
-
[87]
On the use of repeated measurements in regression analysis with dichotomous responses
Margaret Wu and James H Ware. On the use of repeated measurements in regression analysis with dichotomous responses. Biometrics, pages 513--521, 1979
1979
-
[88]
Generative adversarial symmetry discovery
Jianke Yang, Robin Walters, Nima Dehmamy, and Rose Yu. Generative adversarial symmetry discovery. In International Conference on Machine Learning, pages 39488--39508. PMLR, 2023
2023
-
[89]
Rates of convergence for empirical processes of stationary mixing sequences
Bin Yu. Rates of convergence for empirical processes of stationary mixing sequences. The Annals of Probability, pages 94--116, 1994
1994
-
[90]
Learning local dependence in ordered data
Guo Yu and Jacob Bien. Learning local dependence in ordered data. Journal of Machine Learning Research, 18 0 (42): 0 1--60, 2017
2017
-
[91]
Generalized estimating equation models for correlated data: A review with applications
Christopher JW Zorn. Generalized estimating equation models for correlated data: A review with applications. American Journal of Political Science, pages 470--490, 2001
2001
-
[92]
The generalization performance of erm algorithm with strongly mixing observations
Bin Zou, Luoqing Li, and Zongben Xu. The generalization performance of erm algorithm with strongly mixing observations. Machine learning, 75 0 (3): 0 275--295, 2009
2009
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.