Pith. sign in

REVIEW 3 major objections 6 minor 92 references

Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation

T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read For penalized logistic regression in the proportional regime, Gaussian universality and the convex Gaussian min-max theorem hold even for dependent data, and the asymptotic risk depends only on the data's mean and covariance — with data…

desk verdict The dependent CGMT and universality results are real and important; the data augmentation application is not yet proved for the unconstrained estimator because the S_p restriction is left unresolved. read the letter →

arxiv 2502.15752 v4 pith:TYMOLYNF submitted 2025-02-10 math.ST stat.MLstat.TH

classification math.STstat.MLstat.TH MSC 62F1262J1260F05
keywords high-dimensionallogisticregressionGaussianuniversalityconvexmin-maxtheoremblockdependencem-dependencebeta-mixingdataaugmentationproportionalasymptotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the two standard tools of exact high-dimensional statistics — Gaussian universality and the convex Gaussian min-max theorem (CGMT) — remain valid when the data are dependent rather than independent. For penalized logistic regression in the proportional regime ($p/n \to \kappa$), it proves that the asymptotic training and test risks under block dependence, $m$-dependence, and certain $\beta$-mixing processes coincide with the risks of a Gaussian model with matching mean and covariance (Theorems 2 and 3). It then builds a new CGMT that tolerates correlated rows and columns as long as the covariance has a low-rank factored structure (Theorem 5), and applies both tools to data augmentation, producing a ten-equation system that characterizes the test risk exactly (Theorem 12). The concrete payoff: full random permutation of exchangeable coordinates markedly lowers the test risk, while an 80% permutation leaves it statistically unchanged from no augmentation. If the paper is right, dependence enters high-dimensional risk analysis only through the first two moments, and practitioners receive a quantitative warning that partial invariance knowledge buys almost nothing.

What carries the argument

The machinery has three load-bearing parts. First, the block-dependent universality theorem (Theorem 2) uses a Lindeberg-style interpolation between the original data and a matched-covariance Gaussian surrogate, with Assumption 5 requiring joint Gaussian approximation of $k$-tuples of projections $X_{i+r}^{\top}\beta_r$ for $r \le k$, weighted on the sphere, instead of the single projection needed in the independent case. Second, the dependent CGMT (Theorem 5) handles a Gaussian min-max problem whose design matrix $H$ has covariance $\mathrm{Cov}[H_{ji},H_{j'i'}] = \sum_{l=1}^{M} \Sigma^{(l)}_{jj'}\tilde{\Sigma}^{(l)}_{ii'}$, a low-rank sum of factored row and column covariances; for data augmentation $M = 2$ suffices because the covariance of augmented copies is determined by $\mathrm{Var}[\phi_1(Z_1)]$ and $\mathrm{Cov}[\phi_1(Z_1), \phi_2(Z_1)]$. Third, Assumption 7 — that the Gaussian training risk is sharply minimized on the thin shell $|(\beta^{\top}\Sigma_{\text{new}}\beta)^{1/2} - \bar{\chi}| \le \epsilon$ — is what converts training-risk universality into test-risk universality, and the paper verifies it through the CGMT so long as the minimizer-maximizers of the deterministic auxiliary problem lie in the interior of its domain. The final output is the system of ten equations (EQs) whose solution pins down the test risk through the scalar $\bar{\chi}$.

What would settle it

Run the paper's own comparisons at block sizes and aspect ratios beyond the simulated range: if the excess test risk of penalized logistic regression on, say, shifted-gamma covariates differs from the matched-covariance Gaussian surrogate by more than the $\sqrt{k}$ universality bound the authors derive, the claim that dependence enters only through the covariance fails. For the augmentation conclusion, check whether increasing the number of 80%-permutation augmentations eventually lowers the test risk at large $k$; if it does, the qualitative claim that partial invariance is as good as no augmentation would be overturned.

Watch

Extended reading notes

Core claim

The central claim is that the asymptotic risk of high-dimensional penalized logistic regression is universal across data distributions once the first two moments are fixed, even when the observations are dependent. Theorems 2 and 3 show that the minimum training risk and the test risk of the estimator on block-dependent, $m$-dependent, or suitably $\beta$-mixing data converge to those of a Gaussian dataset with the same covariance; a key consequence stated in the paper is that uncorrelated dependent data give the same asymptotic risk as independent data. Because the Gaussian surrogate may still have correlated rows and columns, the paper proves a dependent CGMT (Theorem 5) under an assumption that the design covariance factorizes as a sum of Kronecker products of row and column covariance matrices, and uses it to verify the sharp-shell condition that upgrades training-risk universality to test-risk universality. For data augmentation, the estimator trained on $k$ transformed copies of each observation is shown to have a test risk characterized by a deterministic system of ten scalar equations (Theorem 12); the numerical solution shows that full permutation of exchangeable coordinates reduces the test risk substantially, whereas $r_{\text{perm}} = 0.8$ permutations, sign flipping, and cropping without full knowledge of the zero coordinates yield risks within error margins of no augmentation.

Load-bearing premise

The load-bearing premise is that the constrained minimizer on the set $S_p$ is the true minimizer and that, in the Gaussian surrogate, the training risk is uniquely minimized on a thin shell of constant $(\beta^{\top}\Sigma_{\text{new}}\beta)^{1/2}$; the paper verifies the shell condition only when the auxiliary deterministic optimization has interior solutions, and a general $S_p$ bound for all data-augmentation schemes is explicitly left to future work.

Editorial extensions

If this is right

  • Uncorrelated but dependent data inherit all previously derived independent-data results for logistic regression, since the asymptotic risk is governed by mean and covariance alone.
  • Correlated dependent data can be studied by replacing the data with a Gaussian model of the same covariance and applying the dependent CGMT, avoiding random-matrix theory.
  • Data augmentation under the full permutation group of exchangeable coordinate blocks improves the test risk as more augmentations are used.
  • Augmenting only a large fraction of the coordinates ($r_{\text{perm}} = 0.8$) leaves the test risk statistically unchanged relative to no augmentation.
  • Sign flipping and cropping that do not know the exact zero coordinates of the signal give the same test risk as no augmentation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the mean-variance reduction is as general as the proof suggests, the same universality pipeline should transfer to other classifiers whose loss depends on the data through one-dimensional projections, such as SVMs and other generalized linear models — the paper notes this extension is direct.
  • The $r_{\text{perm}} = 0.8$ plateau suggests a testable design rule for practitioners: augmentation groups that cover only part of the true symmetry may be wasting compute, and symmetry-learning methods should beat hand-chosen partial augmentations even though the paper does not run that comparison.
  • The simulations' observation that training trajectories need different learning rates for t-distributed versus uniform data, even though global minima are universal, indicates that risk universality will not extend to optimization dynamics; proving trajectory non-universality would be a natural sequel.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper extends two workhorses of high-dimensional statistics — Gaussian universality and the convex Gaussian min-max theorem (CGMT) — from independent observations to dependent ones. For penalized logistic regression with labels generated by a logistic link, it proves universality of the training and test risks (matching to a Gaussian surrogate with the same covariance) under block dependence (Theorem 2), m-dependence, and a restricted class of β-mixing processes (Theorem 3). The CGMT part (Theorem 5/13) handles Gaussian design matrices whose covariance is a sum of Kronecker products of row- and column-side matrices (Assumption 10), with comparison probabilities off by a factor 2^M. The tools are applied to data augmentation schemes — random permutations, sign flipping, and cropping — culminating in a ten-equation fixed-point system (EQs) claimed to characterize the asymptotic test risk (Theorem 12). Simulations across coordinate distributions (Gaussian, uniform, gamma, exponential, t₃) confirm the predicted risk values and illustrate that full permutation invariance helps while partial structure knowledge does not.

Significance. If the main theorems hold, this is a substantial contribution: it removes the independence-of-rows assumption in a principled way for high-dimensional logistic regression, and Theorem 5 is a model-independent tool — the 2^M comparison bounds (Lemmas 45–46) and the low-rank Kronecker-factor covariance structure (Assumption 10) are likely to be reused beyond this paper. The proof work is real: the within-block Lindeberg control (Lemmas 24–28) is a genuine extension of the Montanari–Saeed template, and the deterministic ten-equation system (EQs) is anchored by Lemma 52, which recovers exactly Salehi et al.'s six equations in the isotropic no-augmentation limit. The authors ship reproducible code (footnote in Section 6), give detailed simulation settings (Appendix C), and are unusually candid about what is not proved (the S_p membership in Appendix D.1 and the non-universality of training trajectories in Section 9). The main caveat is that the data-augmentation application's headline claims are conditional on the deferred ℓ∞ bound and on unverified geometric properties of the (DO) limit; these gaps are load-bearing for the practical, unconstrained estimator.

major comments (3)
  1. [Section 2 (Eq. 4), Section D.2, Appendix D.1] The paper proves universality for the S_p-constrained minimizer (Theorem 15, for any S̃ ⊆ S_p), yet Theorem 2(5) and Theorem 3(7) are stated with the unconstrained min_β R̂_n(β;·), and the practical estimator (3) is the unconstrained minimizer. Section D.2 asserts that Theorem 15 implies (5) 'by setting S̃ = S_p', but this step requires the ℓ∞ membership \hatβ ∈ S_p, i.e. ‖\hatβ‖_∞ ≤ L p^{1/2−r}. Appendix D.1 explicitly leaves this to future work and explains that the standard leave-one-out singular-value route (σ_min(XX^T) ≥ p/C) is 'nearly impossible' when blocks contain identical rows — precisely the within-block dependence created by data augmentation. Until this membership is established for the DA estimators, the abstract's claim to 'establish the impact of data augmentation' overstates what is proved: Theorem 12 applies to the S_p-restricted (OO), not to the unconstrained estimator (3) that is fitted in the simulations. The authors should either supply the ℓ∞ bound for the DA schemes they study, or restate the main theorems and the DA conclusions in explicitly constrained form.
  2. [Section B.1, Theorem 12 and its proof] Test-risk universality (Theorem 2(6)) requires Assumption 7, and the proof of Theorem 12 verifies Assumption 7 by invoking two premises: that the minimizer-maximizers of (DO) lie in the interior of the domain of optimization, and that restricting the β-domain to |(β^T Σ_new β)^{1/2} − χ̄| > ε changes the (DO) value by Θ(ε²). The first is stated only as an assumption of the theorem, and the second is asserted without derivation. These premises are exactly what makes the Gaussian training risk sharply minimized on the thin shell; no verification for the permutation, sign-flip, or crop schemes is provided. The conclusion |R_test(\hatβ(X,X^Φ)) − R̄_test(χ̄²)| → 0 is therefore conditional on an unverified geometric property of the deterministic limit, and this should be stated as an open condition rather than presented as a completed verification.
  3. [Section D.5, Eqs. (16)–(19) and the final display of the proof of Theorem 15] The Lindeberg interpolation yields an error bound of order √k γ + √k δ + α^{−1} log(1/δ) + τ, so the result holds for each fixed block size k and contains no control uniform in k. In the data-augmentation application, k (the number of augmentations) is a free parameter that the simulations vary (Figures 1 and 2), and the text draws conclusions 'as more and more augmentations are used'. The theorems and the (EQs) derivation cover the regime where k is held fixed as p → ∞; the k → ∞ regime is outside the current proof. This limitation should be stated explicitly wherever the DA conclusions are drawn, since the figures' x-axis is exactly the parameter that the theory does not let grow.
minor comments (6)
  1. [Abstract] The phrase 'assumption that significantly limit its applicability' has a subject-verb agreement error and should read 'significantly limits'.
  2. [Section B.1, Theorem 12] The theorem statement says the test risk is characterized by the parameter set (r, θ, σ, τ), while the system (EQs) solves for ten parameters (α, σ_1, σ_2, τ_1, τ_2, ν_1, ν_2, r_1, r_2, θ); the notation should be aligned between the statement and the equations.
  3. [Section B.1, Theorem 12 proof] The word 'estbalished' appears in the proof and should be corrected to 'established'.
  4. [Section 6, Eq. (11) and Section 1.1, Eq. (3)] Eq. (11) restricts the minimization to S_p while the original problem (3) is unconstrained; this distinction is only explained via Remark 1 and Appendix B.3, and it should be stated in the main text that all DA theorems concern the S_p-constrained problem, with the unconstrained version open (see Major Comment 1).
  5. [Section 2 and throughout] The symbol 'B' is used as a substitute for '≔' throughout the manuscript; it is nonstandard and should be defined at first use or replaced by standard notation.
  6. [Section 6, Figure 1] The curves for r_perm = 0.8 and no augmentation appear to overlap, and one of the paper's central findings is that partial permutations are no better than no augmentation; the caption should report the trial counts (50 versus 200) and the error-bar convention so that the 'within error margins' claim is checkable.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the dependent-CGMT-to-EQ derivation is self-contained and externally anchored by recovery of Salehi et al. (2019); the S_p and interior-minimizer caveats are proof gaps, not circular reductions.

full rationale

The paper's derivation chain is self-contained rather than circular. Section 5 proves the dependent CGMT (Theorem 5) from Gordon's comparison inequality plus a sign-symmetrization argument, and Remark 8 explicitly recovers the standard CGMT and the block-diagonal CGMT of Dhifallah and Lu (2021) as special cases. Sections M and N reduce the data-augmentation logistic problem (OO) through (GO), (PO), (AO), (SO) to the deterministic optimization (DO); the ten parameters in (EQs) are self-consistent fixed-point solutions of the reduced min-max problem, not constants fitted to simulations. The predicted test risk R_bar_test(chi_bar^2) is a function of the DO solution, which is derived from the model, and the figures run genuine gradient descent on non-Gaussian data against Gaussian surrogates, an independent numerical check. External anchoring is explicit: the paper verifies that 'in the isotropic case with no augmentation, it recovers exactly the characterizing equation by Salehi et al. (2019)' (Section 6). The self-citations are either honestly positioned as prior work whose conditions are 'hard to verify and this paper does not cover overparameterized logistic regression' (Section 7), or are peer-reviewed technical lemmas that the paper also states itself (Appendix K.3 reproduces the smoothing construction). The load-bearing caveats, namely the S_p restriction deferred in Appendix D.1 ('To extend this into a proof covering all DA schemes will be left to future work') and Assumption 7's verification via the unproved interior-minimizer regularity of (DO) in Theorem 12, are proof gaps or conditional claims rather than reductions of predictions to their inputs. No step was found in which a fitted quantity is renamed as a prediction or in which a self-citation is the sole justification for a central premise.

Assumptions & free parameters 3 free parameters · 9 assumptions · 1 invented entities

The central claims rest on structural assumptions rather than on fitted constants: the dependence classes (Assumptions 1/8), the logistic model (Assumptions 2/12), sub-Gaussianity (Assumption 3), the joint Gaussian-approximation condition (Assumptions 5/9), the optimizer-concentration condition (Assumption 7), and the low-rank covariance factorization (Assumption 10). Assumption 11(ii), the idempotency of the augmentation cross-covariance, is the one clearly ad hoc condition, imposed to make the CGMT algebra tractable; the paper admits it is more restrictive than necessary. Classical tools (Gordon's inequality, Yu's embedding) are standard. The parameters of the (EQs) are solutions of a fixed-point system, not free fits; the genuinely free inputs are the user-chosen augmentation fractions, the augmentation count k, and the S_p shape constants L and r. No new physical or algorithmic entities are postulated.

free parameters (3)
  • r_perm, r_flip, r_crop (fractions of coordinates augmented) = 0.8 and 1.0 for permutations; 0.2 for flip and crop in figures
    User-chosen proportions in Section 6. The qualitative conclusion flips at r_perm = 0.8 versus 1.0, so the headline DA message is sensitive to the user's choice, though the constant is not fitted to data.
  • k (number of augmentations per observation) = k = 11 and k = 30 in simulations
    Block size in the universality theorems and replication count in the DA model. The error bounds carry sqrt(k) factors and the theorems are per-fixed-k; cross-k comparisons in Figure 1 are simulation-level.
  • S_p shape constants L and r
    Equation (4) defines S_p = {beta: ||beta||_2 <= L sqrt(p), ||beta||_infinity <= L p^{(1-r)/2}} with fixed L > 0 and r in (0, 1/8). The infinity-norm part is the constraint the paper cannot yet lift for dependent data (Appendix D.1).
assumptions (9)
  • domain assumption Assumption 1/8: dependence classes, block dependence with block size k, m-dependence, or β-mixing with summable coefficients
    Defines the dependent-data regimes the theorems cover. Substituted for the independence assumption in prior universality work.
  • domain assumption Assumption 2/12: logistic link labels, y_i driven by X_i^T beta* minus Logistic noise, or by a block-averaged linear form
    The logistic-regression model. The generalized form (Assumption 12) allows labels to depend on the whole block, which the authors show is needed for data augmentation.
  • domain assumption Assumption 3: centered sub-Gaussian covariates with sup_i ||X_i||_{psi2} <= K_X / sqrt(n)
    Tail condition used throughout the Lindeberg proofs. The paper's own simulations with exponential and t_3 data show universality empirically and it conjectures sub-Gaussianity is not necessary (Section 9, Appendix C).
  • domain assumption Assumptions 5/9: uniform joint Gaussian approximation of (X_{i+r}^T beta_r) over S_p^k and theta in S^{k-1} (and for every fixed d in the mixing case)
    The load-bearing CLT-type condition, strictly stronger than the pointwise normality in Montanari and Saeed (2022), as the paper admits in Section 3.1. Verified for specific DA schemes in Appendix J under local-dependence conditions |N_i| x |N_tilde_i| = o(n^{r/2}).
  • domain assumption Assumption 7: the Gaussian training risk is minimized on a thin Sigma_new-shell, with chi_bar and chi_* bounded away from zero
    Gates test-risk universality in Theorems 2 and 3. The paper routes its verification through the dependent CGMT and the unverified interior-solution condition for the deterministic optimization (Theorem 12).
  • domain assumption Assumption 10: low-rank covariance factorization Cov[H_ji, H_j'i'] = sum_l Sigma^{(l)}_{jj'} Sigma_tilde^{(l)}_{ii'}
    The structural assumption defining the dependent CGMT's scope. It is natural for data augmentation with M = 2 (Lemma 14), but it restricts the CGMT to Gaussian matrices with factorized dependence.
  • ad hoc to paper Assumption 11: Sigma_* = (Sigma^dagger)^{1/2} Cov[phi_1(Z_1), Z_1] (Sigma_o^dagger)^{1/2} and Sigma_*^2 = Sigma_*
    Idempotency forces the augmentation cross-covariance to be a projection, imposed to make the CGMT algebra close and the (EQs) tractable. The paper states it is 'more restrictive than necessary' (Section B.1) and that cropping requires a more complicated system.
  • standard math Yu (1994) coupling/embedding for β-mixing blocks
    Used in Section 4 and Lemma 37-38 to compare dependent blocks with independent blocks in the mixing proof.
  • standard math Gordon's Gaussian comparison inequality (Ledoux-Talagrand form)
    The engine of the dependent CGMT proof (Lemma 45), used to compare the primary and auxiliary Gaussian processes with the low-rank covariance structure.
invented entities (1)
  • Low-rank factor decomposition (Sigma^{(l)}, Sigma_tilde^{(l)}) of the correlated-Gaussian data matrix covariance
    purpose: Enables the dependent CGMT by expressing the covariance of a correlated Gaussian matrix as a sum of M factorized row-and-column covariances (Assumption 10)
    This is a structural assumption on the data distribution, not an empirically detected object. For data augmentation it is shown to hold with M = 2 (Lemma 14), but it is not independently evidenced outside the paper's own verification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation." pith.science (2026). https://pith.science/paper/TYMOLYNF

@misc{pith2026250215752,
  author       = {Pith},
  title        = {Pith review of: Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TYMOLYNF}},
  note         = {Machine review of arXiv:2502.15752}
}
abstract

Over the last decade, a wave of research has characterized the exact asymptotic risk of many high-dimensional models in the proportional regime. Two foundational results have driven this progress: Gaussian universality, which shows that the asymptotic risk of estimators trained on non-Gaussian and Gaussian data is equivalent, and the convex Gaussian min-max theorem (CGMT), which characterizes the risk under Gaussian settings. However, these results rely on the assumption that the data consists of independent random vectors--an assumption that significantly limits its applicability to many practical setups. In this paper, we address this limitation by generalizing both results to the dependent setting. More precisely, we prove that Gaussian universality still holds for high-dimensional logistic regression under block dependence, $m$-dependence and special cases of mixing, and establish a novel CGMT framework that accommodates for correlation across both the covariates and observations. Using these results, we establish the impact of data augmentation, a widespread practice in deep learning, on the asymptotic risk.

Figures

Figures reproduced from arXiv: 2502.15752 by the authors.

Figure 1
Figure 1. Universality of risks of a logistic regressor, trained with different number and amount of random permutations. See Section 6 and Section C for the detailed setup. We also consider data that satisfy m-dependence and β-mixing; see Section 4 for the precise defini￾tions. To relate the labels to their covariates, we assume there is a true signal β ∗ ∈ R p such that P (yi = 1 | X) = σ [PITH_FULL_IMAGE:figures/full_fig_… view at source ↗
Figure 2
Figure 2. Test risks under random cropping and sign flipping. Left. Same setup as Section 6. Right. Signal ratio ρ ∗ = 0.2 and the bottom ⌈s0(1 − ρ ∗ )p⌉ coordinates are known to be zero. zero. The positions of the non-zero coordinates are unknown in general, and ρ ∗ may be known or unknown. This motivates the use of random sign flipping to shrink the estimate βˆ at locations where the entries of β ∗ may be zero. We fix some … view at source ↗
Figure 3
Figure 3. Universality of training risks under cropping and sign flipping. The left and the right plots are the training risk analogues of the left and the right plots of [PITH_FULL_IMAGE:figures/full_fig_p027_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Initial training loss curves for the random permutation setup in [PITH_FULL_IMAGE:figures/full_fig_p028_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

92 extracted references · 66 canonical work pages

  1. [1]

    Universality in learning from linear measurements

    Ehsan Abbasi, Fariborz Salehi, and Babak Hassibi. Universality in learning from linear measurements. Advances in Neural Information Processing Systems, 32, 2019

  2. [2]

    A novel G aussian min-max theorem and its applications

    Danil Akhtiamov, David Bosch, Reza Ghane, K Nithin Varma, and Babak Hassibi. A novel G aussian min-max theorem and its applications. arXiv preprint arXiv:2402.07356, 2024 a

  3. [3]

    Regularized linear regression for binary classification

    Danil Akhtiamov, Reza Ghane, and Babak Hassibi. Regularized linear regression for binary classification. In 2024 IEEE International Symposium on Information Theory (ISIT), pages 202--207. IEEE, 2024 b

  4. [4]

    Multiple fourier series and fourier integrals

    Sh A Alimov, RR Ashurov, and AK Pulatov. Multiple fourier series and fourier integrals. Commutative Harmonic Analysis IV: Harmonic Analysis in IR n, pages 1--95, 1992

  5. [5]

    Wasserstein Distributionally Robust Estimation in High Dimensions: Performance Analysis and Optimal Hyperparameter Tuning

    Liviu Aolaritei, Soroosh Shafieezadeh-Abadeh, and Florian D \"o rfler. The performance of W asserstein distributionally robust M -estimators in high dimensions. arXiv preprint arXiv:2206.13269, 2022

  6. [6]

    Limit theorems for distributions invariant under groups of transformations

    Morgane Austern and Peter Orbanz. Limit theorems for distributions invariant under groups of transformations. The Annals of Statistics, 50 0 (4): 0 1960--1991, 2022

  7. [7]

    Regularized estimation in sparse high-dimensional time series models

    Sumanta Basu and George Michailidis. Regularized estimation in sparse high-dimensional time series models. 2015

  8. [8]

    Reconciling modern machine-learning practice and the classical bias–variance trade-off

    Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences, 116 0 (32): 0 15849--15854, 2019

Show all 92 references
  1. [9]

    Learning invariances in neural networks from training data

    Gregory Benton, Marc Finzi, Pavel Izmailov, and Andrew G Wilson. Learning invariances in neural networks from training data. Advances in neural information processing systems, 33: 0 17605--17616, 2020

  2. [10]

    Sur l'extension du th \'e or \`e me limite du calcul des probabilit \'e s aux sommes de quantit \'e s d \'e pendantes

    Serge Bernstein. Sur l'extension du th \'e or \`e me limite du calcul des probabilit \'e s aux sommes de quantit \'e s d \'e pendantes. Mathematische Annalen, 97: 0 1--59, 1927

  3. [11]

    Probability and measure

    Patrick Billingsley. Probability and measure. John Wiley & Sons, 1995

  4. [12]

    Logistic regression for dependent binary observations

    George Ebow Bonney. Logistic regression for dependent binary observations. Biometrics, pages 951--973, 1987

  5. [13]

    Basic properties of strong mixing conditions

    Richard C Bradley. Basic properties of strong mixing conditions. a survey and some open questions. Probability Surveys, 2: 0 107--114, 2005

  6. [14]

    Simple technical trading rules and the stochastic properties of stock returns

    William Brock, Josef Lakonishok, and Blake LeBaron. Simple technical trading rules and the stochastic properties of stock returns. The Journal of finance, 47 0 (5): 0 1731--1764, 1992

  7. [15]

    Distributional and lq norm inequalities for polynomials over convex bodies in rn

    Anthony Carbery and James Wright. Distributional and lq norm inequalities for polynomials over convex bodies in rn. Mathematical research letters, 8 0 (3): 0 233--248, 2001

  8. [16]

    Concentration inequalities with exchangeable pairs

    Sourav Chatterjee. Concentration inequalities with exchangeable pairs. Stanford University, 2005

  9. [17]

    A group-theoretic framework for data augmentation

    Shuxiao Chen, Edgar Dobriban, and Jane H Lee. A group-theoretic framework for data augmentation. Journal of Machine Learning Research, 21 0 (245): 0 1--71, 2020

  10. [18]

    Universality of approximate message passing algorithms

    Wei-Kuo Chen and Wai-Kit Lam. Universality of approximate message passing algorithms . Electronic Journal of Probability, 26: 0 1 -- 44, 2021

  11. [19]

    Statistics for spatial data

    Noel Cressie. Statistics for spatial data. John Wiley & Sons, 1993

  12. [20]

    Time series analysis, volume 286

    Jonathan D Cryer. Time series analysis, volume 286. Duxbury Press Boston, 1986

  13. [21]

    Universality laws for G aussian mixtures in generalized linear models

    Yatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro, and Lenka Zdeborov \'a . Universality laws for G aussian mixtures in generalized linear models. Advances in Neural Information Processing Systems, 36, 2024

  14. [22]

    A central limit theorem for globally nonstationary near-epoch dependent functions of mixing processes

    James Davidson. A central limit theorem for globally nonstationary near-epoch dependent functions of mixing processes. Econometric theory, 8 0 (3): 0 313--329, 1992

  15. [23]

    A model of double descent for high-dimensional binary linear classification

    Zeyu Deng, Abla Kammoun, and Christos Thrampoulidis. A model of double descent for high-dimensional binary linear classification. Information and Inference: A Journal of the IMA, 11 0 (2): 0 435--495, 2022

  16. [24]

    A note on empirical processes of strong-mixing sequences

    Chandrakant M Deo. A note on empirical processes of strong-mixing sequences. The Annals of Probability, pages 870--875, 1973

  17. [25]

    On the inherent regularization effects of noise injection during training

    Oussama Dhifallah and Yue Lu. On the inherent regularization effects of noise injection during training. In International Conference on Machine Learning, pages 2665--2675. PMLR, 2021

  18. [26]

    Message-passing algorithms for compressed sensing

    David L Donoho, Arian Maleki, and Andrea Montanari. Message-passing algorithms for compressed sensing. Proceedings of the National Academy of Sciences, 106 0 (45): 0 18914--18919, 2009

  19. [27]

    Lu, and Subhabrata Sen

    Rishabh Dudeja, Yue M. Lu, and Subhabrata Sen. Universality of approximate message passing with semirandom matrices. The Annals of Probability, 51 0 (5): 0 1616--1683, 2023

  20. [28]

    Handbook of spatial statistics

    Alan E Gelfand, Peter Diggle, Peter Guttorp, and Montserrat Fuentes. Handbook of spatial statistics. CRC press, 2010

  21. [29]

    Gaussian universality of perceptrons with random labels

    Federica Gerace, Florent Krzakala, Bruno Loureiro, Ludovic Stephan, and Lenka Zdeborov \'a . Gaussian universality of perceptrons with random labels. Physical Review E, 109 0 (3): 0 034305, 2024

  22. [30]

    Some inequalities for G aussian processes and applications

    Yehoram Gordon. Some inequalities for G aussian processes and applications. Israel Journal of Mathematics, 50: 0 265--289, 1985

  23. [31]

    High dimensional and banded vector autoregressions, 2016

    Shaojun Guo, Yazhen Wang, and Qiwei Yao. High dimensional and banded vector autoregressions, 2016. URL https://arxiv.org/abs/1502.07831

  24. [32]

    Universality of regularized regression estimators in high dimensions

    Qiyang Han and Yandi Shen. Universality of regularized regression estimators in high dimensions. The Annals of Statistics, 51 0 (4): 0 1799--1823, 2023

  25. [33]

    Data augmentation as stochastic optimization

    Boris Hanin and Yi Sun. Data augmentation as stochastic optimization. 2020

  26. [34]

    Analysis of dichotomous response data from certain toxicological experiments

    JK Haseman and LL Kupper. Analysis of dichotomous response data from certain toxicological experiments. Biometrics, pages 281--293, 1979

  27. [35]

    Surprises in high-dimensional ridgeless least squares interpolation

    Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani. Surprises in high-dimensional ridgeless least squares interpolation. The Annals of Statistics, 50 0 (2): 0 949, 2022

  28. [36]

    Universality laws for high-dimensional learning with random features

    Hong Hu and Yue M Lu. Universality laws for high-dimensional learning with random features. IEEE Transactions on Information Theory, 69 0 (3): 0 1932--1964, 2022

  29. [37]

    Data augmentation in the underparameterized and overparameterized regimes

    Kevin Han Huang, Peter Orbanz, and Morgane Austern. Data augmentation in the underparameterized and overparameterized regimes. arXiv preprint arXiv:2202.09134, 2022

  30. [38]

    A high-dimensional convergence theorem for u-statistics with applications to kernel-based testing

    Kevin Han Huang, Xing Liu, Andrew Duncan, and Axel Gandy. A high-dimensional convergence theorem for u-statistics with applications to kernel-based testing. In The Thirty Sixth Annual Conference on Learning Theory, pages 3827--3918. PMLR, 2023

  31. [39]

    Independent and stationary sequences of random variables

    I Ibragimov. Independent and stationary sequences of random variables. Wolters, Noordhoff Pub., 1975

  32. [40]

    The theory of approximation, volume 11

    Dunham Jackson. The theory of approximation, volume 11. American Mathematical Soc., 1930

  33. [41]

    Precise statistical analysis of classification accuracies for adversarial training

    Adel Javanmard and Mahdi Soltanolkotabi. Precise statistical analysis of classification accuracies for adversarial training. The Annals of Statistics, 50 0 (4): 0 2127--2156, 2022

  34. [42]

    Asymptotic behavior of unregularized and ridge-regularized high-dimensional robust regression estimators : rigorous results, 2013

    Noureddine El Karoui. Asymptotic behavior of unregularized and ridge-regularized high-dimensional robust regression estimators : rigorous results, 2013

  35. [43]

    Label-imbalanced and group-sensitive classification under overparameterization

    Ganesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, and Christos Thrampoulidis. Label-imbalanced and group-sensitive classification under overparameterization. Advances in Neural Information Processing Systems, 34: 0 18970--18983, 2021

  36. [44]

    Applications of the lindeberg principle in communications and statistical learning

    Satish Babu Korada and Andrea Montanari. Applications of the lindeberg principle in communications and statistical learning. IEEE transactions on information theory, 57 0 (4): 0 2440--2450, 2011

  37. [45]

    Universality in block dependent linear models with applications to nonparametric regression

    Samriddha Lahiry and Pragya Sur. Universality in block dependent linear models with applications to nonparametric regression. IEEE Transactions on Information Theory, 70 0 (12): 0 8975--9000, 2024

  38. [46]

    Probability in Banach Spaces: isoperimetry and processes

    Michel Ledoux and Michel Talagrand. Probability in Banach Spaces: isoperimetry and processes. Springer Science & Business Media, 1991

  39. [47]

    The good, the bad and the ugly sides of data augmentation: An implicit spectral regularization perspective

    Chi-Heng Lin, Chiraag Kaushik, Eva L Dyer, and Vidya Muthukumar. The good, the bad and the ugly sides of data augmentation: An implicit spectral regularization perspective. Journal of Machine Learning Research, 25 0 (91): 0 1--85, 2024

  40. [48]

    On the benefits of invariance in neural networks

    Clare Lyle, Mark van der Wilk, Marta Kwiatkowska, Yarin Gal, and Benjamin Bloem-Reddy. On the benefits of invariance in neural networks. In Conference on Neural Information Processing Systems: Workshop on Machine Learning with Guarantees, 2019

  41. [49]

    The generalization error of random features regression: Precise asymptotics and the double descent curve

    Song Mei and Andrea Montanari. The generalization error of random features regression: Precise asymptotics and the double descent curve. Communications on Pure and Applied Mathematics, 75 0 (4): 0 667--766, 2022

  42. [50]

    The role of regularization in classification of high-dimensional noisy G aussian mixture

    Francesca Mignacco, Florent Krzakala, Yue Lu, Pierfrancesco Urbani, and Lenka Zdeborova. The role of regularization in classification of high-dimensional noisy G aussian mixture. In International conference on machine learning, pages 6874--6883. PMLR, 2020

  43. [51]

    Universality of the elastic net error

    Andrea Montanari and Phan-Minh Nguyen. Universality of the elastic net error. In 2017 IEEE International Symposium on Information Theory (ISIT), pages 2338--2342. IEEE, 2017

  44. [52]

    Universality of empirical risk minimization

    Andrea Montanari and Basil N Saeed. Universality of empirical risk minimization. In Conference on Learning Theory, pages 4310--4312. PMLR, 2022

  45. [53]

    Universality of max-margin classifiers

    Andrea Montanari, Feng Ruan, Basil Saeed, and Youngtak Sohn. Universality of max-margin classifiers. arXiv preprint arXiv:2310.00176, 2023

  46. [54]

    High dimensional logistic regression under network dependence

    Somabha Mukherjee, Ziang Niu, Sagnik Halder, Bhaswar B Bhattacharya, and George Michailidis. High dimensional logistic regression under network dependence. arXiv preprint arXiv:2110.03200, 2021

  47. [55]

    Least squares regression with markovian data: Fundamental limits and algorithms

    Dheeraj Nagaraj, Xian Wu, Guy Bresler, Prateek Jain, and Praneeth Netrapalli. Least squares regression with markovian data: Fundamental limits and algorithms. Advances in neural information processing systems, 33: 0 16666--16676, 2020

  48. [56]

    Nicholson, Ines Wilms, Jacob Bien, and David S

    William B. Nicholson, Ines Wilms, Jacob Bien, and David S. Matteson. High dimensional forecasting via interpretable vector autoregression. Journal of Machine Learning Research, 21 0 (166): 0 1--52, 2020. URL http://jmlr.org/papers/v21/19-777.html

  49. [57]

    From naive mean field theory to the tap equations

    Manfred Opper, Ole Winther, et al. From naive mean field theory to the tap equations. Advanced mean field methods: theory and practice, pages 7--20, 2001

  50. [58]

    Universality laws for randomized dimension reduction, with applications

    Samet Oymak and Joel A Tropp. Universality laws for randomized dimension reduction, with applications. Information and Inference: A Journal of the IMA, 7 0 (3): 0 337--446, 2018

  51. [59]

    Joost D. J. Plate, Rutger R. van de Leur, Luke P.H. Leenen, Falco Hietbrink, Linda M. Peelen, and Marinus J. C. Eijkemans. Incorporating repeated measurements into prediction models in the critical care setting: a framework, systematic review and meta-analysis. BMC Medical Res...

  52. [60]

    Correlated binary regression with covariates specific to each binary observation

    Ross L Prentice. Correlated binary regression with covariates specific to each binary observation. Biometrics, pages 1033--1048, 1988

  53. [61]

    Locally dependent latent class models with covariates: an application to under-age drinking in the usa

    Beth A Reboussin, Edward H Ip, and Mark Wolfson. Locally dependent latent class models with covariates: an application to under-age drinking in the usa. Journal of the Royal Statistical Society Series A: Statistics in Society, 171 0 (4): 0 877--897, 2008

  54. [62]

    Linear models in statistics

    Alvin C Rencher and G Bruce Schaalje. Linear models in statistics. John Wiley & Sons, 2008

  55. [63]

    Tyrrell Rockafellar

    R. Tyrrell Rockafellar. Convex analysis. Princeton Mathematical Series. Princeton University Press, Princeton, N. J., 1970

  56. [64]

    Fundamentals of Stein’s method

    Nathan Ross. Fundamentals of Stein’s method . Probability Surveys, 8: 0 210 -- 293, 2011

  57. [65]

    Hanson-wright inequality and sub-gaussian concentration, 2013

    Mark Rudelson and Roman Vershynin. Hanson-wright inequality and sub-gaussian concentration, 2013

  58. [66]

    The impact of regularization on high-dimensional logistic regression

    Fariborz Salehi, Ehsan Abbasi, and Babak Hassibi. The impact of regularization on high-dimensional logistic regression. Advances in Neural Information Processing Systems, 32, 2019

  59. [67]

    Local dependence in random graph models: characterization, properties and statistical inference

    Michael Schweinberger and Mark S Handcock. Local dependence in random graph models: characterization, properties and statistical inference. Journal of the Royal Statistical Society Series B: Statistical Methodology, 77 0 (3): 0 647--676, 2015

  60. [68]

    Advanced Data Analysis from an Elementary Point of View

    Cosma Rohilla Shalizi. Advanced Data Analysis from an Elementary Point of View. 2019

  61. [69]

    Shorten and T

    C. Shorten and T. M Khoshgoftaar. A survey on image data augmentation for deep learning. Journal of Big Data, 6 0 (1): 0 1--48, 2019

  62. [70]

    Text data augmentation for deep learning

    Connor Shorten, Taghi M Khoshgoftaar, and Borko Furht. Text data augmentation for deep learning. Journal of big Data, 8 0 (1): 0 101, 2021

  63. [71]

    A framework to characterize performance of lasso algorithms

    Mihailo Stojnic. A framework to characterize performance of lasso algorithms. arXiv preprint arXiv:1303.7291, 2013 a

  64. [72]

    Upper-bounding _1 -optimization weak thresholds

    Mihailo Stojnic. Upper-bounding _1 -optimization weak thresholds. arXiv preprint arXiv:1303.7289, 2013 b

  65. [73]

    A modern maximum-likelihood theory for high-dimensional logistic regression

    Pragya Sur and Emmanuel J Cand \`e s. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences, 116 0 (29): 0 14516--14525, 2019

  66. [74]

    The impact of multi-optimizers and data augmentation on tensorflow convolutional neural network performance

    Arwa Mohammed Taqi, Ahmed Awad, Fadwa Al-Azzo, and Mariofanna Milanova. The impact of multi-optimizers and data augmentation on tensorflow convolutional neural network performance. In Proc. of IEEE MIPR, pages 140--145, 2018

  67. [75]

    Recovering structured signals in high dimensions via non-smooth convex optimization: Precise performance analysis

    Christos Thrampoulidis. Recovering structured signals in high dimensions via non-smooth convex optimization: Precise performance analysis. PhD thesis, California Institute of Technology, 2016

  68. [76]

    The G aussian min-max theorem in the presence of convexity

    Christos Thrampoulidis, Samet Oymak, and Babak Hassibi. The G aussian min-max theorem in the presence of convexity. arXiv preprint arXiv:1408.4837, 2014

  69. [77]

    Regularized linear regression: A precise analysis of the estimation error

    Christos Thrampoulidis, Samet Oymak, and Babak Hassibi. Regularized linear regression: A precise analysis of the estimation error. In Conference on Learning Theory, pages 1683--1709. PMLR, 2015

  70. [78]

    Precise error analysis of regularized m -estimators in high dimensions

    Christos Thrampoulidis, Ehsan Abbasi, and Babak Hassibi. Precise error analysis of regularized m -estimators in high dimensions. IEEE Transactions on Information Theory, 64 0 (8): 0 5592--5628, 2018

  71. [79]

    Analysis of financial time series

    Ruey S Tsay. Analysis of financial time series. John wiley & sons, 2005

  72. [80]

    Some mixing properties of time series models

    D Tuan and T Lanh. Some mixing properties of time series models. Stochastic processes and their applications, 19: 0 297--303, 1985

  73. [81]

    High-Dimensional Probability: An Introduction with Applications in Data Science

    Roman Vershynin. High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press, 2018

  74. [82]

    An overview on data augmentation for machine learning

    Svetlana Volkova. An overview on data augmentation for machine learning. In Arthur Gibadullin, editor, Digital and Information Technologies in Economics and Management, pages 143--154, Cham, 2024. Springer Nature Switzerland. ISBN 978-3-031-55349-3

  75. [83]

    Multivariate geostatistics: an introduction with applications

    Hans Wackernagel. Multivariate geostatistics: an introduction with applications. Springer Science & Business Media, 2003

  76. [84]

    Universality of approximate message passing algorithms and tensor networks

    Tianhao Wang, Xinyi Zhong, and Zhou Fan. Universality of approximate message passing algorithms and tensor networks. The Annals of Applied Probability, 34 0 (4): 0 3943--3994, 2024

  77. [85]

    Regularized estimation in high dimensional time series under mixing conditions

    Kam Chung Wong, Ambuj Tewari, and Zifan Li. Regularized estimation in high dimensional time series under mixing conditions. stat, 1050: 0 12, 2016

  78. [86]

    Lasso guarantees for -mixing heavy-tailed time series

    Kam Chung Wong, Zifan Li, and Ambuj Tewari. Lasso guarantees for -mixing heavy-tailed time series. The Annals of Statistics, 48 0 (2): 0 1124--1142, 2020

  79. [87]

    On the use of repeated measurements in regression analysis with dichotomous responses

    Margaret Wu and James H Ware. On the use of repeated measurements in regression analysis with dichotomous responses. Biometrics, pages 513--521, 1979

  80. [88]

    Generative adversarial symmetry discovery

    Jianke Yang, Robin Walters, Nima Dehmamy, and Rose Yu. Generative adversarial symmetry discovery. In International Conference on Machine Learning, pages 39488--39508. PMLR, 2023

  81. [89]

    Rates of convergence for empirical processes of stationary mixing sequences

    Bin Yu. Rates of convergence for empirical processes of stationary mixing sequences. The Annals of Probability, pages 94--116, 1994

  82. [90]

    Learning local dependence in ordered data

    Guo Yu and Jacob Bien. Learning local dependence in ordered data. Journal of Machine Learning Research, 18 0 (42): 0 1--60, 2017

  83. [91]

    Generalized estimating equation models for correlated data: A review with applications

    Christopher JW Zorn. Generalized estimating equation models for correlated data: A review with applications. American Journal of Political Science, pages 470--490, 2001

  84. [92]

    The generalization performance of erm algorithm with strongly mixing observations

    Bin Zou, Luoqing Li, and Zongben Xu. The generalization performance of erm algorithm with strongly mixing observations. Machine learning, 75 0 (3): 0 275--295, 2009

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.