Pith. sign in

REVIEW 2 major objections 3 minor 4 cited by

Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery

T0 review · 2 major / 3 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Under Gaussian design, spectral estimators for multi-index models have exactly characterized top eigenvalues and eigenvector overlaps, and the minimal sample ratio for weak recovery is attained by an explicit preprocessing function.

desk verdict Sharp spectral asymptotics for correlated multi-index models, but the optimality theorem silently assumes y has a density and needs a corrected statement. read the letter →

arxiv 2502.01583 v2 pith:TCZQENFJ submitted 2025-02-03 stat.ML cs.ITcs.LGmath.ITmath.PRmath.STstat.TH

classification stat.MLcs.ITcs.LGmath.ITmath.PRmath.STstat.TH MSC 60B2062H25
keywords multi-indexmodelsspectralestimatorsweakrecoveryphasetransitionrandommatrixtheoryGaussiandesignoptimalpreprocessinghigh-dimensionalasymptotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper analyzes spectral estimators for multi-index models $y_i = q(\langle a_i,w_1^*\rangle,\ldots,\langle a_i,w_p^*\rangle,\varepsilon_i)$ with Gaussian design, where the estimator forms $D_n = \frac{1}{n}A^\top Z A$ with $z_i = T(y_i)$ and returns the top-$p$ eigenvectors. It establishes that as $n,d\to\infty$ with $n/d\to\delta$, the top-$p$ eigenvalues converge almost surely to the explicit values $\zeta_\delta(\alpha_i)$, where $\alpha_i$ solve the $p$-dimensional equation $\det(\zeta_\delta(\alpha)I - R_\infty(\alpha))=0$. Eigenvalues that emerge from the bulk correspond to eigenvectors with non-vanishing overlap with the signal subspace, so weak recovery happens exactly when such an outlier appears. The paper then optimizes over the preprocessing function $T$ and proves that the smallest sample ratio at which any spectral estimator can weakly recover the subspace is $\delta_c$, given in closed form in (4.9), with a matching optimal preprocessing $T^*_\delta$ in (4.10). This matters because spectral methods are widely used warm-starts, and the result turns their design from a heuristic choice into a provably optimal one.

What carries the argument

The load-bearing object is the $p\times p$ matrix $R_\infty(\alpha)=\mathbb{E}[\alpha s s^\top z/(\alpha-z)]$ together with the scalar function $\zeta_\delta(\lambda)=\psi_\delta(\max\{\bar\lambda_\delta,\lambda\})$, where $\psi_\delta(\lambda)=\lambda(1/\delta+\mathbb{E}[z/(\lambda-z)])$ and $\bar\lambda_\delta$ is the minimizer of $\psi_\delta$. Eigenvalue locations are the solutions of $\det(\zeta_\delta(\alpha)I-R_\infty(\alpha))=0$; the proof compresses the $(d-p)$-dimensional spectral problem to this $p$-dimensional equation through a rank-$p$ perturbation formula, and the eigenvector overlaps are expressed through $\zeta_\delta'(\alpha)$ and $R_\infty'(\alpha)$. The fixed-point functions $\tilde L_i(\mu)$ of (5.2)-(5.3), whose limits are obtained from low-rank-perturbation asymptotics, carry the convergence argument, and the optimality proof uses H\"older's inequality to show no $T$ can beat the threshold in (4.9) while $T^*_\delta$ attains it.

What would settle it

Run the model $y=\mathbf{1}\{s_1s_2+\varepsilon>0\}$ with $p=2$, a bounded $T$, and $n/d=\delta$. The expressions in (4.9)-(4.10) require $p(y|s)$ as a Lebesgue density, which does not exist for this atomic output, so the claimed optimal threshold cannot be evaluated directly; any discretization that makes it computable will produce a number that can be compared with the simulated outlier onset, settling whether the hidden regularity assumption is essential to the theorem.

Watch

Extended reading notes

Core claim

The paper claims that for any bounded preprocessing function $T$ with $P(z=0)<1$, the spectral matrix $D_n$ has a precisely describable spectrum: under the stated assumptions, the top $p$ eigenvalues converge almost surely to $\zeta_\delta(\alpha_1)\ge\cdots\ge\zeta_\delta(\alpha_j)>\zeta_\delta(\bar\lambda_\delta)$, where $\alpha_i$ are the solutions of the secular equation, and the remaining $p-j$ eigenvalues collapse to the bulk edge $\zeta_\delta(\bar\lambda_\delta)$. Whenever an eigenvalue is an outlier, i.e. $\alpha_i>\bar\lambda_\delta$, the corresponding eigenvector subspace has asymptotically non-vanishing squared overlap with the signal subspace, and if the eigenvalue stays in the bulk the overlap vanishes. For the weak-recovery threshold $\delta_c$, defined as the infimum over all admissible $T$ of the smallest $\delta$ at which the top-$p$ eigenvectors achieve non-vanishing overlap, the paper proves $\delta_c$ equals the closed-form expression (4.9), and the preprocessing $T^*_\delta$ in (4.10) achieves weak recovery for every $\delta>\delta_c$. This simultaneously provides the first exact asymptotic characterization of spectral methods in general multi-index models and an optimality certificate for one particular choice of $T$.

Load-bearing premise

The main optimality result assumes the response $y$ admits a conditional density $p(y|s)$ with respect to Lebesgue measure, even though the model assumptions do not guarantee this for discrete or mixed outputs.

Editorial extensions

If this is right

  • For any admissible $T$, the top-$p$ eigenvalues and the phase transition are determined by a $p\times p$ equation, so the outlier onset of a spectral estimator is directly computable without simulation.
  • Weak recovery by the top-$p$ eigenvectors occurs exactly when some $\alpha_i>\bar\lambda_\delta$; inside the bulk all overlaps vanish, identifying $\delta_c(T)$ with the appearance of the first outlier.
  • The optimal preprocessing $T^*_\delta$ provably attains the minimal sample ratio over all bounded preprocessing functions, which implies that existing heuristic choices for $T$ are suboptimal in general multi-index models.
  • The threshold formula generalizes the single-index and independent-mixture results to arbitrarily correlated signals and requires only the link function $q$ and the signal covariance $\Sigma$, which can be estimated on a separate sample.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If (4.9) is taken as the fundamental threshold, its integrand resembles a Fisher-information-type quantity for the direction $u$; one could test whether $\delta_c$ also lower-bounds every estimator, spectral or not, in models where computational and statistical thresholds coincide.
  • The hidden absolute-continuity assumption on $y$ suggests a concrete extension: for discrete or mixed outputs, the variational problem should be reformulated with probability mass functions, and the resulting threshold may differ from (4.9).
  • The optimal $T^*_\delta$ depends on the conditional density $p(y|s)$; in practice this density must be estimated, and a finite-sample guarantee for the plug-in spectral estimator (how many samples suffice to stay above threshold) is left implicit by the paper.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper analyzes spectral estimators for multi-index models in the proportional regime n/d -> δ with fixed signal dimension p. For a preprocessing function T, the spectral matrix is D_n = (1/n) A^T Z A with z_i = T(y_i). The main results are: Theorem 4.1 locates the top-p eigenvalues of D_n and exhibits a BBP-type phase transition, with outliers converging to ζδ(α_i) where α_i solve det(ζδ(α)I - R∞(α)) = 0; Theorem 4.2 characterizes the overlaps of the corresponding eigenvectors with the signal subspace; Theorem 4.3 gives the optimal weak-recovery threshold δc over all bounded preprocessing functions and constructs an optimal T* in (4.10). The proofs are based on an equivalent spectral characterization of D_n through the functions L_i, L_{i,j} and on results from Bai-Yao random matrix theory adapted from the single-index case. The paper also proves an equivalence with the threshold of Troiani et al. under simultaneous diagonalizability of the conditional second-moment matrices.

Significance. If the results are correct, this is a substantial advance: it provides the first precise asymptotic characterization of spectral estimators for general multi-index models with correlated signals, it identifies the optimal preprocessing function for weak recovery, and it resolves the conjecture of the parallel work DDM+25 in a wide class of cases. The derivation is self-contained: the threshold is obtained by optimizing a variational bound over preprocessing functions, the optimal T* is constructed by saturating Hölder's inequality, and no fitted parameters enter the asymptotic formulas. The appendices contain detailed proofs, and the numerical experiments in Figures 1 and 2 show good agreement with the predicted overlaps. The main caveat is a hidden regularity assumption in Theorem 4.3: the statement and proof use the conditional density p(y|s) and Lebesgue integrals over y, while the model assumptions allow y to be discrete or mixed. This is a load-bearing gap in the optimality claim, but it is local and can likely be fixed by adding an explicit absolute-continuity condition or by reformulating the result for general output distributions.

major comments (2)
  1. [Theorem 4.3 and Appendix E (Eqs. (4.9), (4.10), (E.4), (E.5))] The statement and proof of the optimality result assume that y admits a conditional density p(y|s) with respect to Lebesgue measure and that the integrals over dy are Lebesgue integrals. Assumptions (A1)-(A5) do not imply this: the link function q and the noise ε are only required to satisfy moment conditions, so y may be discrete or mixed (for example, a classification label or a quantized response). In such cases the conditional density p(y|s) is not defined, the ratio appearing in (E.5) is not meaningful, and the proposed optimal preprocessing T* in (4.10) cannot be formed. Because Theorem 4.3 is the central optimality claim, the manuscript overclaims for the stated model class. Please add an explicit assumption that the conditional law of y given s is absolutely continuous with respect to Lebesgue measure, or reformulate Theorem 4.3 using conditional expectations/Radon-Nikodym derivatives so that discrete and mixed outputs are covered. In addition, the integrand in (4.9) implicitly requires E_s[p(y|s)] > 0 almost everywhere; this should also be stated explicitly.
  2. [Theorem 4.1 (statement, page 6)] The theorem says 'Let α1 ≥ ... ≥ α_j > τ (for some j ∈ [p]) be all the solutions' and then discusses the 'remaining p − j eigenvalues'. If (4.2) has no solutions in ]τ, ∞[, the statement as written does not cover the case j = 0, which is exactly the no-outlier regime discussed in the proof. Please either allow j = 0 explicitly or add a separate sentence stating that, when (4.2) has no solutions, all top-p eigenvalues converge to ζδ(λ̄δ). This is a presentation gap in a main theorem, but it is easy to repair.
minor comments (3)
  1. [Theorem 4.2 (page 7)] The eigenspace-invariance condition preceding (4.6) is an explicit additional assumption, and the authors note that it may be a proof artifact. Since the condition is not derived from the model assumptions, it would be helpful to state in the theorem or in a remark which natural classes of q are known to satisfy it (for example, permutation-invariant links via Proposition D.1) and whether there is a known example where it fails. This would clarify the scope of the overlap formula.
  2. [Remark 4.1 and Appendix G] The equivalence with the threshold of [TDD+24] is proved only under the conditions that the supremum is achieved by a rank-1 matrix and that the matrices E(y) are simultaneously diagonalizable. The main text is careful about this, but the appendix title 'Equivalence to [TDD+24]' could be read as unconditional. Consider a more guarded title or an explicit sentence in the appendix stating the exact hypotheses under which the two thresholds coincide.
  3. [Section 4, paragraph after Theorem 4.1] The phrase 'the remaining p − j eigenvalues' is slightly ambiguous when j = p; in that case there are no remaining eigenvalues and the statement is vacuous. Clarifying the notational convention for j = 0 and j = p would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the optimal threshold is derived by an explicit variational argument and Hölder saturation, not by fitting or by assuming the target.

full rationale

The derivation chain is self-contained. Theorem 4.3's optimal threshold is obtained by first using Theorem 4.2 to reduce weak recovery to the existence of a spectral outlier, then rewriting the outlier condition as max_u E[z(⟨s,u⟩²-1)/(1-z)] > 1/δ after the scaling λ̄δ=1, and finally bounding this maximum over preprocessing functions T using Cauchy–Schwarz. The optimal T* is constructed explicitly so that equality holds in the bound, which is the standard saturation argument; no parameter is fitted to data and no quantity entering the theorem is defined in terms of the theorem's conclusion. The comparison with the [TDD+24] threshold in Appendix G is a proof of equivalence, not an assumed premise: the paper derives its own expression (4.9) and then shows, under simultaneous diagonalizability, that it coincides with the [TDD+24] formula. Citations to [MM19], [LL20], and [BY12] supply technical tools (fixed-point equations, low-rank perturbation results, and eigenvalue interlacing arguments), but the multi-index spectral characterization is obtained from the paper's own Propositions A.7, B.2, B.3, and C.1 rather than imported as a black box. The only notable weakness is a hidden regularity assumption in Theorem 4.3: the statement uses the conditional density p(y|s) and Lebesgue integrals in (4.9)–(4.10), while assumptions (A1)–(A5) do not require y to be absolutely continuous, so discrete or mixed outputs are not covered by the optimal-threshold formula or by the proposed T*; this is a scope/correctness gap, not a circularity.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central derivations rely on standard random matrix theory and on the listed modeling assumptions. There are no fitted free parameters. The two ad-hoc conditions (density of y and eigenspace invariance) are the least justified inputs; the first is implicit, the second is flagged by the authors.

assumptions (6)
  • domain assumption Design matrix has i.i.d. standard Gaussian entries (Assumption A1).
    Required for the decomposition A = [S U] with independent blocks and for the [BY12] low-rank perturbation limits used in Propositions B.2 and C.1.
  • domain assumption Signal dimension p is fixed while n,d -> infinity with n/d -> delta (Assumption A3).
    The asymptotic equations for eigenvalues and overlaps are derived in the fixed-p regime; no bound is provided as p grows.
  • domain assumption Preprocessing T is bounded and P(z=0) < 1 (Assumption A5).
    Boundedness gives a finite right edge tau of the spectrum; P(z=0)<1 ensures rk(ZS)=p. The optimality of T* is proven only over this class.
  • ad hoc to paper The response y admits a conditional density p(y|s) with respect to Lebesgue measure.
    Theorem 4.3 and the optimal T* are written as integrals over dy against this density; the paper does not list this among (A1)-(A5).
  • ad hoc to paper Eigenspace invariance: the eigenspace E∞_k of R∞(alpha_k) is invariant in a neighbourhood of alpha_k (Theorem 4.2).
    Needed for the exact overlap formula (4.6); the authors state this appears to be a proof artifact and that no violating example is known.
  • domain assumption Signals w*_1,...,w*_p are linearly independent (Assumption A2).
    Used to guarantee a well-defined p-dimensional signal subspace and to set up the rotated representation (3.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery." pith.science (2026). https://pith.science/paper/TCZQENFJ

@misc{pith2026250201583,
  author       = {Pith},
  title        = {Pith review of: Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TCZQENFJ}},
  note         = {Machine review of arXiv:2502.01583}
}
abstract

Multi-index models provide a popular framework to investigate the learnability of functions with low-dimensional structure and, also due to their connections with neural networks, they have been object of recent intensive study. In this paper, we focus on recovering the subspace spanned by the signals via spectral estimators -- a family of methods routinely used in practice, often as a warm-start for iterative algorithms. Our main technical contribution is a precise asymptotic characterization of the performance of spectral methods, when sample size and input dimension grow proportionally and the dimension $p$ of the space to recover is fixed. Specifically, we locate the top-$p$ eigenvalues of the spectral matrix and establish the overlaps between the corresponding eigenvectors (which give the spectral estimators) and a basis of the signal subspace. Our analysis unveils a phase transition phenomenon in which, as the sample complexity grows, eigenvalues escape from the bulk of the spectrum and, when that happens, eigenvectors recover directions of the desired subspace. The precise characterization we put forward enables the optimization of the data preprocessing, thus allowing to identify the spectral estimator that requires the minimal sample size for weak recovery.

Figures

Figures reproduced from arXiv: 2502.01583 by the authors.

Figure 1
Figure 1. q(s1, s2, ε) = s1s2. Overlaps [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. q(ξ1, ξ2, ε) = |ξε|. Overlaps n [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spectral phase transitions in Gaussian multi-index models

    math.ST 2026-08 accept novelty 8.0 of 10

    In Gaussian multi-index models, the paper rigorously characterizes the bulk spectrum and BBP-type phase transition of matrix-valued spectral estimators, and proves the AMP-derived preprocessing map attains the minimal...

  2. Approximate Message Passing with Random Initialization for Phase Retrieval

    math.ST 2026-08 conditional novelty 7.0 of 10

    Randomly initialized Bayes-optimal AMP provably achieves the weak-recovery threshold δ=1/2 and arbitrarily accurate recovery for δ>1.13 in proportional-regime noiseless phase retrieval.

  3. The Generative Leap: Sharp Sample Complexity for Efficiently Learning Gaussian Multi-Index Models

    cs.LG 2025-06 conditional novelty 7.0 of 10

    For any Gaussian multi-index model, the generative leap exponent k⋆ sharply characterizes the sample complexity of efficient subspace recovery as Θ(d^(1∨k⋆/2)).

  4. Optimal Spectral Transitions in High-Dimensional Multi-Index Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Two linearized message-passing spectral estimators achieve the optimal weak-recovery threshold in Gaussian multi-index models, with a sharp BBP-like spectral phase transition at the critical sample complexity.

Reference graph

Works this paper leans on

56 extracted references · 47 canonical work pages · cited by 4 Pith papers

  1. [1]

    The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks

    Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz. The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks. In Proceedings of Thirty Fifth Conference on Learning Theory , volume 178 of Proceedings of Machine Learning Research , pages 4782--4887. PMLR, 02--05 Jul 2022

  2. [2]

    Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics

    Emmanuel Abbe, Enric Boix Adser \`a , and Theodor Misiakiewicz. Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics. In Proceedings of Thirty Sixth Conference on Learning Theory , volume 195 of Proceedings of Machine Learning Research , pages 2552--2623. PMLR, 12--15 Jul 2023

  3. [3]

    Community detection and stochastic block models: Recent developments

    Emmanuel Abbe. Community detection and stochastic block models: Recent developments. Journal of Machine Learning Research , 18(177):1--86, 2018

  4. [4]

    The committee machine: computational to statistical gaps in learning a two-layers neural network

    Benjamin Aubin, Antoine Maillard, Jean Barbier, Florent Krzakala, Nicolas Macris, and Lenka Zdeborov\' a . The committee machine: computational to statistical gaps in learning a two-layers neural network. J. Stat. Mech. Theory Exp. , (12):124023, 51, 2019

  5. [5]

    Learning sparse polynomial functions

    Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang. Learning sparse polynomial functions. In Proceedings of the T wenty- F ifth A nnual ACM - SIAM S ymposium on D iscrete A lgorithms , pages 500--510. ACM, New York, 2014

  6. [6]

    Online stochastic gradient descent on non-convex losses from high-dimensional inference

    Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath. Online stochastic gradient descent on non-convex losses from high-dimensional inference. Journal of Machine Learning Research , 22(106):1--51, 2021

  7. [7]

    Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices

    Jinho Baik, G\' e rard Ben Arous, and Sandrine P\' e ch\' e . Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices. Ann. Probab. , 33(5):1643--1697, 2005

  8. [8]

    On learning gaussian multi-index models with gradient flow

    Alberto Bietti, Joan Bruna, and Loucas Pillaud - Vivien. On learning gaussian multi-index models with gradient flow. CoRR , abs/2310.19793, 2023

Show all 56 references
  1. [9]

    Fundamental limits in structured principal component analysis and how to reach them

    Jean Barbier, Francesco Camilli, Marco Mondelli, and Manuel S\' a enz. Fundamental limits in structured principal component analysis and how to reach them. Proc. Natl. Acad. Sci. USA , 120(30):Paper No. e2302028120, 7, 2023

  2. [10]

    Optimal errors and phase transitions in high-dimensional generalized linear models

    Jean Barbier, Florent Krzakala, Nicolas Macris, L\' e o Miolane, and Lenka Zdeborov\' a . Optimal errors and phase transitions in high-dimensional generalized linear models. Proc. Natl. Acad. Sci. USA , 116(12):5451--5460, 2019

  3. [11]

    On sample eigenvalues in a generalized spiked population model

    Zhidong Bai and Jianfeng Yao. On sample eigenvalues in a generalized spiked population model. J. Multivariate Anal. , 106:167--177, 2012

  4. [12]

    Cand\`es

    Yuxin Chen and Emmanuel J. Cand\`es. Solving random quadratic systems of equations is nearly as easy as solving linear systems. Comm. Pure Appl. Math. , 70(5):822--883, 2017

  5. [13]

    Spectral methods for data science: A statistical perspective

    Yuxin Chen, Yuejie Chi, Jianqing Fan, and Cong Ma. Spectral methods for data science: A statistical perspective. Foundations and Trends® in Machine Learning , 14(5):566--806, 2021

  6. [14]

    Learning polynomials in few relevant dimensions

    Sitan Chen and Raghu Meka. Learning polynomials in few relevant dimensions. In Proceedings of Thirty Third Conference on Learning Theory , volume 125 of Proceedings of Machine Learning Research , pages 1161--1227. PMLR, 09--12 Jul 2020

  7. [15]

    The estimation error of general first order methods

    Michael Celentano, Andrea Montanari, and Yuchen Wu. The estimation error of general first order methods. In Proceedings of Thirty Third Conference on Learning Theory , volume 125 of Proceedings of Machine Learning Research , pages 1078--1141. PMLR, 09--12 Jul 2020

  8. [16]

    Hitting the high-dimensional notes: an ODE for SGD learning dynamics on GLM s and multi-index models

    Elizabeth Collins-Woodfin, Courtney Paquette, Elliot Paquette, and Inbar Seroussi. Hitting the high-dimensional notes: an ODE for SGD learning dynamics on GLM s and multi-index models. Inf. Inference , 13(4):Paper No. iaae028, 107, 2024

  9. [17]

    Analysis of spectral methods for phase retrieval with random orthogonal matrices

    Rishabh Dudeja, Milad Bakhshizadeh, Junjie Ma, and Arian Maleki. Analysis of spectral methods for phase retrieval with random orthogonal matrices. IEEE Trans. Inform. Theory , 66(8):5182--5203, 2020

  10. [18]

    Optimal spectral transitions in high-dimensional multi-index models

    Leonardo Defilippis, Yatin Dandi, Pierre Mergny, Florent Krzakala, and Bruno Loureiro. Optimal spectral transitions in high-dimensional multi-index models. arXiv preprint arXiv:2502.02545v1 , 2025

  11. [19]

    Nonlinear functional analysis

    Klaus Deimling. Nonlinear functional analysis . Springer-Verlag, Berlin, 1985

  12. [20]

    Dalalyan, Anatoly Juditsky, and Vladimir Spokoiny

    Arnak S. Dalalyan, Anatoly Juditsky, and Vladimir Spokoiny. A new algorithm for estimating the effective dimension-reduction subspace. Journal of Machine Learning Research , 9(53):1647--1678, 2008

  13. [21]

    Neural networks can learn representations with gradient descent

    Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi. Neural networks can learn representations with gradient descent. In Proceedings of Thirty Fifth Conference on Learning Theory , volume 178 of Proceedings of Machine Learning Research , pages 5413--5452. PMLR, 02--05 Jul 2022

  14. [22]

    Smoothing the landscape boosts the signal for sgd: Optimal sample complexity for learning single index models

    Alex Damian, Eshaan Nichani, Rong Ge, and Jason D Lee. Smoothing the landscape boosts the signal for sgd: Optimal sample complexity for learning single index models. In Advances in Neural Information Processing Systems , volume 36, pages 752--784. Curran Associates, Inc., 2023

  15. [23]

    Computational-statistical gaps in gaussian single-index models (extended abstract)

    Alex Damian, Loucas Pillaud-Vivien, Jason Lee, and Joan Bruna. Computational-statistical gaps in gaussian single-index models (extended abstract). In Proceedings of Thirty Seventh Conference on Learning Theory , volume 247 of Proceedings of Machine Learning Research , pages 12...

  16. [24]

    Order structure and topological methods in nonlinear partial differential equations

    Yihong Du. Order structure and topological methods in nonlinear partial differential equations. V ol. 1 , volume 2 of Series in Partial Differential Equations and Applications . World Scientific Publishing Co. Pte. Ltd., Hackensack, NJ, 2006. Maximum principles and applications

  17. [25]

    Learning functions of few arbitrary linear parameters in high dimensions

    Massimo Fornasier, Karin Schnass, and Jan Vybiral. Learning functions of few arbitrary linear parameters in high dimensions. Found. Comput. Math. , 12(2):229--262, 2012

  18. [26]

    Horn and Charles R

    Roger A. Horn and Charles R. Johnson. Matrix analysis . Cambridge University Press, Cambridge, second edition, 2013

  19. [27]

    Perturbation theory for linear operators

    Tosio Kato. Perturbation theory for linear operators . Classics in Mathematics. Springer-Verlag, Berlin, 1995. Reprint of the 1980 edition

  20. [28]

    M. G. Kre n and M. A. Rutman. Linear operators leaving invariant a cone in a B anach space. Uspehi Matem. Nauk (N.S.) , 3(1(23)):3--95, 1948

  21. [29]

    Wangyu Luo, Wael Alghamdi, and Yue M. Lu. Optimal spectral initialization for signal recovery with applications to phase retrieval. IEEE Trans. Signal Process. , 67(9):2347--2456, 2019

  22. [30]

    On eigenvalues of matrices dependent on a parameter

    Peter Lancaster. On eigenvalues of matrices dependent on a parameter. Numerische Mathematik , 6(1):377--387, 1964

  23. [31]

    Regression analysis under link violation

    Ker-Chau Li and Naihua Duan. Regression analysis under link violation. Ann. Statist. , 17(3):1009--1052, 1989

  24. [32]

    Sliced inverse regression for dimension reduction

    Ker-Chau Li. Sliced inverse regression for dimension reduction. J. Amer. Statist. Assoc. , 86(414):316--342, 1991. With discussion and a rejoinder by the author

  25. [33]

    On principal H essian directions for data visualization and dimension reduction: another application of S tein's lemma

    Ker-Chau Li. On principal H essian directions for data visualization and dimension reduction: another application of S tein's lemma. J. Amer. Statist. Assoc. , 87(420):1025--1039, 1992

  26. [34]

    Lu and Gen Li

    Yue M. Lu and Gen Li. Phase transitions of spectral initialization for high-dimensional non-convex estimation. Inf. Inference , 9(3):507--541, 2020

  27. [35]

    Generalized linear models

    Peter McCullagh. Generalized linear models. European J. Oper. Res. , 16(3):285--292, 1984

  28. [36]

    Lu, and Lenka Zdeborova

    Antoine Maillard, Florent Krzakala, Yue M. Lu, and Lenka Zdeborova. Construction of optimal spectral methods in phase retrieval. In Proceedings of the 2nd Mathematical and Scientific Machine Learning Conference , volume 145 of Proceedings of Machine Learning Research , pages 6...

  29. [37]

    Phase retrieval in high dimensions: Statistical and computational phase transitions

    Antoine Maillard, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborov\' a . Phase retrieval in high dimensions: Statistical and computational phase transitions. In Advances in Neural Information Processing Systems , volume 33, pages 11071--11082. Curran Associates, Inc., 2020

  30. [38]

    Fundamental limits of weak recovery with applications to phase retrieval

    Marco Mondelli and Andrea Montanari. Fundamental limits of weak recovery with applications to phase retrieval. Found. Comput. Math. , 19(3):703--773, 2019

  31. [39]

    Optimal combination of linear and spectral estimators for generalized linear models

    Marco Mondelli, Christos Thrampoulidis, and Ramji Venkataramanan. Optimal combination of linear and spectral estimators for generalized linear models. Found. Comput. Math. , 22(5):1513--1566, 2022

  32. [40]

    Estimation of low-rank matrices via approximate message passing

    Andrea Montanari and Ramji Venkataramanan. Estimation of low-rank matrices via approximate message passing. Ann. Statist. , 49(1):321--345, 2021

  33. [41]

    Approximate message passing with spectral initialization for generalized linear models

    Marco Mondelli and Ramji Venkataramanan. Approximate message passing with spectral initialization for generalized linear models. J. Stat. Mech. Theory Exp. , (11):Paper No. 114003, 41, 2022

  34. [42]

    Statistically optimal firstorder algorithms: a proof via orthogonalization

    Andrea Montanari and Yuchen Wu. Statistically optimal firstorder algorithms: a proof via orthogonalization. Inf. Inference , 13(4):Paper No. iaae027, 24, 2024

  35. [43]

    Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations

    Kazusato Oko, Yujin Song, Taiji Suzuki, and Denny Wu. Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations. In Proceedings of Thirty Seventh Conference on Learning Theory , volume 247 of Proceedings of Machine Le...

  36. [44]

    Yunwei Ren and Jason D. Lee. Learning orthogonal multi-index models: A fine-grained information exponent analysis. CoRR , abs/2410.09678, 2024

  37. [45]

    Fletcher

    Sundeep Rangan, Philip Schniter, and Alyson K. Fletcher. Vector approximate message passing. IEEE Trans. Inform. Theory , 65(10):6664--6684, 2019

  38. [46]

    Silverstein and Sang-Il Choi

    Jack W. Silverstein and Sang-Il Choi. Analysis of the limiting spectral distribution of large-dimensional random matrices. J. Multivariate Anal. , 54(2):295--309, 1995

  39. [47]

    A. Singer. Angular synchronization by eigenvectors and semidefinite programming. Appl. Comput. Harmon. Anal. , 30(1):20--36, 2011

  40. [48]

    High-dimensional single-index models: Link estimation and marginal inference

    Kazuma Sawaya, Yoshimasa Uematsu, and Masaaki Imaizumi. High-dimensional single-index models: Link estimation and marginal inference. arXiv preprint arXiv:2404.17812 , 2024

  41. [49]

    Fundamental limits of weak learnability in high-dimensional multi-index models

    Emanuele Troiani, Yatin Dandi, Leonardo Defilippis, Lenka Zdeborov \' a , Bruno Loureiro, and Florent Krzakala. Fundamental limits of weak learnability in high-dimensional multi-index models. CoRR , abs/2405.15480, 2024

  42. [50]

    Estimation in rotationally invariant generalized linear models via approximate message passing

    Ramji Venkataramanan, Kevin K \"o gler, and Marco Mondelli. Estimation in rotationally invariant generalized linear models via approximate message passing. In Proceedings of the 39th International Conference on Machine Learning , volume 162 of Proceedings of Machine Learning R...

  43. [51]

    Giannakis, and Yonina C

    Gang Wang, Georgios B. Giannakis, and Yonina C. Eldar. Solving systems of random quadratic equations via truncated amplitude flow. IEEE Trans. Inform. Theory , 64(2):773--794, 2018

  44. [52]

    Alternating minimization for mixed linear regression

    Xinyang Yi, Constantine Caramanis, and Sujay Sanghavi. Alternating minimization for mixed linear regression. In Proceedings of the 31st International Conference on Machine Learning , volume 32 of Proceedings of Machine Learning Research , pages 613--621, Bejing, China, 22--24 ...

  45. [53]

    On the identifiability of additive index models

    Ming Yuan. On the identifiability of additive index models. Statist. Sinica , 21(4):1901--1911, 2011

  46. [54]

    Spectral estimators for structured generalized linear models via approximate message passing (extended abstract)

    Yihan Zhang, Hong Chang Ji, Ramji Venkataramanan, and Marco Mondelli. Spectral estimators for structured generalized linear models via approximate message passing (extended abstract). In Proceedings of Thirty Seventh Conference on Learning Theory , volume 247 of Proceedings of...

  47. [55]

    Matrix denoising with doubly heteroscedastic noise: Fundamental limits and optimal spectral methods

    Yihan Zhang and Marco Mondelli. Matrix denoising with doubly heteroscedastic noise: Fundamental limits and optimal spectral methods. CoRR , abs/2405.13912, 2024

  48. [56]

    Precise asymptotics for spectral methods in mixed generalized linear models

    Yihan Zhang, Marco Mondelli, and Ramji Venkataramanan. Precise asymptotics for spectral methods in mixed generalized linear models. CoRR , abs/2211.11368, 2022

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.