Pith. sign in

REVIEW 2 major objections 4 minor 71 references

Spectral phase transitions in Gaussian multi-index models

T0 review · 2 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper proves that, in Gaussian multi-index models, the AMP-derived preprocessing map $T_*(y)=I_{p*}-C_*(y)^{-1}$ attains the optimal spectral weak-recovery threshold among all bounded matrix-valued preprocessing maps of any fixed…

desk verdict Serious, proof-heavy paper that substantially advances the spectral theory of Gaussian multi-index models; the optimality theorem is real but its scope is narrower than the abstract claims. read the letter →

arxiv 2608.12183 v1 pith:FQQW7YLV submitted 2026-08-12 math.ST math.PRstat.TH

classification math.STmath.PRstat.TH MSC 60B2062H1268Q87
keywords Gaussianmulti-indexmodelsmatrix-valuedspectralmethodsmatrix-weightedcovariancematricesBBPtransitionsweaksubspacerecoveryapproximatemessagepassingself-consistentequationsphase
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper studies spectral estimators for recovering a hidden $r$-dimensional subspace from nonlinear Gaussian observations, in the proportional regime where sample size and dimension grow at the same rate. It proves a complete spectral phase transition for the matrix-weighted covariance estimator $D_n=\frac{1}{n}\sum_i T(y_i)\otimes x_i x_i^\top$: below a deterministic threshold the largest eigenvalue stays at the bulk edge, above it an outlier separates and the associated eigenvector achieves weak recovery. Its central result is that the AMP-derived preprocessing map $T_*(y)=I_{p*}-C_*(y)^{-1}$ is optimal among all bounded symmetric matrix-valued preprocessing maps of any fixed dimension, with threshold $1/\alpha^*_c=\|\mathbb{E}_y[(C(y)-I_r)\otimes (C(y)-I_r)]\|_{\mathrm{op}}$ exactly matching the AMP weak-recovery threshold. If correct, this settles the general spectral conjecture that no fixed-dimensional bounded spectral preprocessing can beat the AMP threshold in these models.

What carries the argument

The engine is the matrix-weighted covariance matrix $D_n=\frac{1}{n}\sum_{i=1}^n T(y_i)\otimes x_i x_i^\top$ and the matrix Dyson equation $M_\alpha(z)^{-1}=-z I_p+\mathbb{E}_y\big[T(y)(I_p+\alpha^{-1}M_\alpha(z)T(y))^{-1}\big]$, which characterizes the deterministic bulk spectrum. The spectral transition is governed by the deterministic outlier function $h_\alpha(x)=\lambda_1\big(\mathbb{E}_y[T(y)(I_p+\alpha^{-1}M_\alpha(x)T(y))^{-1}\otimes C(y)]\big)-x$; the edge gap $\Delta_T(\alpha)=\lim_{x\downarrow \lambda_+(\alpha)}h_\alpha(x)$ decides whether an outlier separates. Optimality is carried by the Perron reduction of the operator $A[H]=\mathbb{E}_y[(C(y)-I_r)H(C(y)-I_r)]$, which produces the preprocessing map $T_*(y)=I_{p*}-C_*(y)^{-1}$ whose threshold is exactly $\alpha^*_c$.

What would settle it

Construct a two-index channel in which $C(y)$ is singular or arbitrarily close to singular on a set of positive probability, meaning that a bounded approximation of $T_*$ is the only admissible option; if the spectral transition of that approximation occurs strictly below $\alpha^*_c$, or if the formula $1/\alpha^*_c=\|\mathbb{E}[(C(y)-I_r)^{\otimes 2}]\|_{\mathrm{op}}$ fails to predict the empirical outlier location, the optimality theorem would be refuted. A direct check in a non-commuting two-index model would also settle whether the predicted threshold matches the AMP threshold numerically.

Watch

Extended reading notes

Core claim

The paper establishes a sharp, universal lower bound on the spectral weak-recovery threshold: for every fixed $p\ge 1$ and every bounded measurable preprocessing map $T$, one has $\alpha_{c,\min}(T)\ge \alpha^*_c$, where $1/\alpha^*_c=\|\mathbb{E}_y[(C(y)-I_r)\otimes(C(y)-I_r)]\|_{\mathrm{op}}$. Equality is attained by the preprocessing map $T_*(y)=I_{p*}-C_*(y)^{-1}$, constructed from the Perron eigenmatrix $H_*$ of the operator $A[H]=\mathbb{E}_y[(C(y)-I_r)H(C(y)-I_r)]$; for this map the lower and upper critical sampling ratios coincide, $\alpha_{c,\min}(T_*)=\alpha_{c,\max}(T_*)=\alpha^*_c$. Combined with the phase-transition theorem, this proves that the AMP weak-recovery threshold of Troiani et al. is attainable by a spectral method without side information, and that no bounded fixed-dimensional matrix-valued preprocessing can do better, thereby proving Conjecture 3.10 of Defilippis et al.

Load-bearing premise

The load-bearing premise is Assumption A.4: the conditional second moment $C(y)=\mathbb{E}[ss^\top|y]$ is bounded below by a positive constant times the identity almost surely and is not identically the identity; if some latent direction is determined essentially exactly, the map $T_*(y)=I_{p*}-C_*(y)^{-1}$ is not bounded and the optimality claim does not apply.

Editorial extensions

If this is right

  • Below the threshold $\alpha^*_c$, the largest eigenvalue of $D_n$ sticks to the bulk edge, so no spectral estimator of this class has a nonvanishing overlap with the hidden subspace.
  • Above $\alpha^*_c$, an outlier separates at a location given by the unique zero of $h_\alpha$, and every leading eigenvector achieves weak recovery of the latent subspace with overlap bounded away from zero.
  • The AMP weak-recovery threshold can be attained by a spectral method without side information or an informative initialization.
  • No choice of working dimension $p$ or bounded matrix-valued preprocessing rule can improve on $\alpha^*_c$, so the threshold depends only on the second-order fluctuations of the conditional second moment $C(y)$.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The optimality is proved inside the class of bounded, fixed-dimension preprocessing maps; it does not by itself rule out spectral estimators built from unbounded maps or from features whose dimension grows with $n$, so the 'no spectral method beats AMP' reading should be restricted to this class.
  • The threshold formula $1/\alpha^*_c=\|\mathbb{E}[(C(y)-I_r)^{\otimes 2}]\|_{\mathrm{op}}$ identifies the size of channel fluctuations as the sole resource for spectral recovery; a testable consequence is that any channel with $C(y)=I_r$ almost surely is spectrally unrecoverable at every finite $\alpha$.
  • The Perron-reduction construction suggests a general recipe: linearize the AMP update around its uninformative fixed point and form the resulting matrix-weighted covariance; the spectral transition will coincide with the AMP instability point whenever the conditional second-moment operator has a positive Perron eigenmatrix.
  • A natural extension not pursued in the paper is to replace Gaussian covariates by an orthogonally invariant design and check whether the same formula, with $C(y)$ and the design covariance, still predicts the transition.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper develops a random matrix theory for matrix-valued spectral estimators of the form D_n = (1/n) Σ_i T(y_i) ⊗ x_i x_i^T in Gaussian multi-index models, where T is a bounded symmetric matrix-valued preprocessing map of fixed dimension p. The main results are: (i) almost-sure convergence of the empirical spectral measure of D_n to a deterministic compactly supported law characterized by a matrix Dyson equation with a nonlinear self-energy; (ii) a BBP-type phase transition for the largest eigenvalue, with an explicit deterministic outlier equation and a sharp threshold in terms of the function h_α(x); (iii) weak recovery of the latent subspace from the leading eigenvector in the supercritical regime; and (iv) an optimality theorem stating that the AMP-derived preprocessing T_*(y) = I_{p*} - C_*(y)^{-1} is optimal among all bounded matrix-valued preprocessing maps of any fixed dimension, with threshold 1/α*_c = ||E_y[(C(y)-I_r)⊗(C(y)-I_r)]||_op, coinciding with the AMP weak-recovery threshold. The proofs are detailed and self-contained, with Sections 3–7 providing the technical arguments, and Appendices A–C giving heuristic motivation, a variational edge characterization, and auxiliary results.

Significance. If the main claims hold, this is a substantial contribution to high-dimensional statistics and random matrix theory. It resolves a natural conjecture about the reach of spectral methods in multi-index models, removes the simultaneous-diagonalizability assumption of earlier work, and establishes a rigorous optimality result for the AMP-derived preprocessing. The technical toolbox is impressive: a nonlinear matrix Dyson equation, a self-adjoint linearization for sign-indefinite covariance models, a spectral comparison theorem to a free Gaussian model, a Fock-space variational bound for the free edge, and a finite-dimensional reduction of the outlier problem. The proofs are structured, detailed, and appear internally consistent. The paper also honestly discloses the limitation that Proposition 1.7 does not characterize the intermediate regime [α_c,min, α_c,max] for general maps. The main weakness is the scope of the optimality theorem, which is conditional on Assumption A.4 and whose abstract formulation overstates the unqualified validity.

major comments (2)
  1. [Section 1.3, Theorem 1.10 and Assumption A.4] The optimality claim is conditional on Assumption A.4, which requires C(y) ⪰ c I_r almost surely. This condition fails for natural channels such as the noiseless single-index model y = s_1 with r=1, where C(y) = s_1^2 is positive but not uniformly bounded below. In such a channel the proposed optimal preprocessing T_*(y) = 1 - C(y)^{-1} is unbounded and therefore not admissible in the class of bounded preprocessing maps over which optimality is claimed. Consequently Theorem 1.10 does not establish the AMP-threshold coincidence or the general spectral conjecture for these channels. The paper does not discuss whether A.4 is technical or essential; in particular, it does not analyze whether a truncated or regularized version of T_* attains the same threshold α*_c. Since this is a restriction on the channel, not on the estimator, it is load-bearing for the headline claim. I ask the authors to qualify the statement of Theorem 1.10 and the abstract, and to add a discussion (or a rigorous approximation argument) addressing channels where C(y) is singular or only positive semidefinite with non-uniform lower bound.
  2. [Abstract and Introduction (claims about the general spectral conjecture)] The abstract states that the AMP-derived preprocessing is optimal among all bounded matrix-valued preprocessing maps and that its transition coincides with the AMP weak-recovery threshold, 'proving the general spectral conjecture of [Defilippis et al., 2025]'. As written, this suggests a fully general result. However, the proof in Section 7 requires Assumption A.4 for the definition and admissibility of T_*; outside that assumption the theorem does not apply. Moreover, the statement 'Theorem 1.10 proves Conjecture 3.10 of [26]' needs clarification: if the conjecture was originally stated for all Gaussian multi-index channels, the present paper proves a restricted version. The authors should explicitly state the precise class of channels for which the conjecture is proved and, if the original conjecture was broader, indicate the open case. This is more than a presentation issue because it affects the advertised scope of the main optimality result.
minor comments (4)
  1. [Abstract] The abstract should mention Assumption A.4 explicitly, or rephrase the optimality claim to 'under the nondegeneracy condition A.4', so that the conditional nature of the result is visible to the reader.
  2. [Section 6.3] The first paragraph contains a typo: 'two independent aprts' should be 'two independent parts'.
  3. [Section 7.2] There is a spurious period in the sentence 'We next show that a separating outlier. exists'; it should read 'a separating outlier exists'.
  4. [Section 1.1 and Section 5.2] The symbol ⊗ is used both for the Kronecker product and for the tensor product in the Fock-space construction. While the meaning is clear from context, a brief remark in the notation section would improve readability.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the optimal-threshold theorem and spectral phase transition are derived from the model, not fitted or imported.

full rationale

The paper's derivation chain is self-contained. The bulk law (Proposition 2.1) is proved from resolvent identities and leave-one-out estimates; the outlier equation (Proposition 2.5) is a deterministic consequence of the Schur complement and the limiting MDE; the phase transition (Theorem 1.4) follows from monotonicity of the deterministic function h_alpha; and the optimal threshold (Theorem 1.10) is proved by a universal lower bound (Lemma 7.1) plus an attainment argument for the specifically constructed map T_* (Lemmas 7.2, 7.6, 7.7). The threshold alpha*_c is defined by the channel's second-order fluctuation operator, not calibrated or fitted. The map T_* is constructed from the Perron eigenmatrix of A and is not assumed from the cited prior work; no load-bearing step cites [26] for a result used in the proof. The AMP threshold of [63] is an external benchmark, and the abstract's coincidence statement is an asserted comparison rather than a reduction used in the internal theorems. The only caveats are scope/qualification issues, not circularity: Assumption A.4 (C(y) ⪰ cI) is needed for T_* to be well-defined and bounded, so channels with conditionally singular C(y) fall outside the attainment claim; and the 'coincides with AMP weak-recovery threshold' statement is an external equivalence not proved inside the paper. These are correctness/qualification concerns, not reductions of the claimed results to their inputs.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The central results are parameter-free: no constants are fitted to data, and all quantities entering the thresholds (lambda_+(alpha), theta_alpha, alpha*_c) are derived from the model. The paper depends on four explicitly stated assumptions (A.1 through A.4) and on standard external theorems. No new physical or mathematical entities (particles, mediators, forces, dimensions) are introduced; the optimal preprocessing map is a derived object within the model.

assumptions (6)
  • domain assumption Proportional regime: n/d -> alpha in (0, infinity) with r, q, p fixed (Assumption A.1).
    All asymptotic statements (bulk law, phase transition, optimality) are proven only in this regime, stated at the start of Section 1.2.
  • domain assumption Preprocessing map T is measurable, bounded (||T(y)|| <= C_T a.s.), and nontrivial (Assumption A.2).
    Boundedness of T is used throughout Sections 4-6 for resolvent bounds, concentration estimates, and the spectral comparison theorem of [10].
  • domain assumption Population separation: lambda_1(E[T(y) tensor C(y)]) > lambda_1(E[T(y)]) (Assumption A.3).
    Used only in Lemma 6.6 and Proposition 1.7 to guarantee eventual supercriticality for large alpha; it is not needed for the main optimality theorem.
  • domain assumption Conditional second moment C(y) = E[ss^T|y] satisfies C(y) >= c I_r a.s. and P(C(y) != I_r) > 0 (Assumption A.4).
    Load-bearing for the optimality theorem: it makes the optimal preprocessing map T_* = I - C_*^{-1} in Definition 1.8 well-defined, bounded, and admissible.
  • standard math External theorems invoked as black boxes: spectral comparison theorem of Bandeira-Cipolloni-Schroeder-van Handel [10], matrix Nevanlinna-Herglotz representation [33], Earle-Hamilton fixed point theorem [34], Fock-space cumulant and norm bounds in the style of Lehner [38].
    Used for noise-spectrum confinement (Section 5), MDE existence and uniqueness (Section 3), and the variational upper edge bound (Appendix B).
  • domain assumption Multi-index channel model: x_i i.i.d. N(0, I_d) and y_i | x_i ~ P_*(.|W_*^T x_i) with W_* having orthonormal columns (Section 1.2).
    The entire problem and all definitions (D_n, weak recovery, C(y)) are built on this generative model; the Gaussianity is used throughout the random matrix arguments.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spectral phase transitions in Gaussian multi-index models." pith.science (2026). https://pith.science/paper/FQQW7YLV

@misc{pith2026260812183,
  author       = {Pith},
  title        = {Pith review of: Spectral phase transitions in Gaussian multi-index models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FQQW7YLV}},
  note         = {Machine review of arXiv:2608.12183}
}
abstract

Recovering a low-dimensional latent subspace from nonlinear observations of Gaussian covariates in high dimensions is a fundamental problem in feature learning. Here, we consider Gaussian multi-index models in which the covariates $\boldsymbol{x}_i \stackrel{\mathrm{i.i.d.}}{\sim} \mathcal{N}(0,\boldsymbol{I}_d)$ and the responses $\boldsymbol{y}_i$ depend on $\boldsymbol{x}_i$ only through its projection onto an unknown $r$-dimensional subspace. Earlier work based on approximate message passing (AMP) identified a sharp threshold for weak recovery [Troiani et al., 2025], raising the question of whether it can be attained, without side information, by a spectral method. We answer this affirmatively and develop a general random matrix theory for matrix-valued spectral estimators of the form \[\boldsymbol{D}_n=\frac{1}{n}\sum_{i=1}^n\boldsymbol{T}(\boldsymbol{y}_i)\otimes\boldsymbol{x}_i\boldsymbol{x}_i^\top,\] where $\boldsymbol{T}$ is an arbitrary bounded symmetric matrix-valued preprocessing map of fixed dimension. As $n,d \to \infty$ with $n/d\to\alpha$, we prove that the empirical spectral measure of $\boldsymbol{D}_n$ converges almost surely to a deterministic compactly supported distribution characterized by a matrix-valued self-consistent equation. We then establish a spectral phase transition for the largest eigenvalue: below threshold it sticks to the bulk edge, while above threshold an outlier emerges. We characterize the outlier location through a finite-dimensional deterministic equation and show that the associated spectral estimator achieves weak recovery of the latent subspace. Finally, we prove that the AMP-derived preprocessing of [Defilippis et al., 2025] is optimal among all bounded matrix-valued preprocessing maps of any fixed dimension. Its transition coincides with the AMP weak-recovery threshold, proving the general spectral conjecture of [Defilippis et al., 2025].

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

71 extracted references · 60 canonical work pages

  1. [26]

    Optimal spectral transitions in high-dimensional multi-index models

    L. Defilippis, Y. Dandi, P. Mergny, F. Krzakala, and B. Loureiro, “Optimal spectral transitions in high-dimensional multi-index models”, Advances in neural information processing systems, Vol. 38, edited by D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (2025), pp. 174966–175002

  2. [1]

    SGD learning on neural networks: leap complex- ity and saddle-to-saddle dynamics

    E. Abbe, E. B. Adserà, and T. Misiakiewicz, “SGD learning on neural networks: leap complex- ity and saddle-to-saddle dynamics”, Proceedings of thirty sixth conference on learning theory, Vol. 195, edited by G. Neu and L. Rosasco, Proceedings of Machine Learning Research (2023), pp. 2552–2623

  3. [2]

    Theory Related Fields173, 293–373 (2019)

    O.H.Ajanki,L.Erdős,andT.Krüger,Stability of the matrix Dyson equation and random matrices with correlations, Probab. Theory Related Fields173, 293–373 (2019)

  4. [3]

    J. Alt, L. Erdős, and T. Krüger,Local law for random Gram matrices, Electron. J. Probab.22, Paper No. 25, 41 (2017)

  5. [4]

    J. Alt, L. Erdős, and T. Krüger,The Dyson equation with linear self-energy: spectral bands, edges and cusps, Doc. Math.25, 1421–1539 (2020)

  6. [5]

    Overparametrization bends the landscape: BBP transitions at initialization in simple neural networks

    B. L. Annesi, D. Bocchi, and C. Cammarota, “Overparametrization bends the landscape: BBP transitions at initialization in simple neural networks”, The fourteenth international conference on learning representations (2026)

  7. [6]

    Arnaboldi, Y

    L. Arnaboldi, Y. Dandi, F. Krzakala, L. Pesce, and L. Stephan,Repetita iuvant: data repetition allows SGD to learn high-dimensional multi-index functions, arXiv preprint (2024),arXiv : 2405.15459

  8. [7]

    Aubin, A

    B. Aubin, A. Maillard, J. Barbier, F. Krzakala, N. Macris, and L. Zdeborová,The committee machine: computational to statistical gaps in learning a two-layers neural network, J. Stat. Mech. Theory Exp., 124023, 51 (2019)

Show all 71 references
  1. [8]

    Bai and J

    Z. Bai and J. W. Silverstein,Spectral analysis of large dimensional random matrices, Second, Springer Series in Statistics (Springer, New York, 2010), pp. xvi+551

  2. [9]

    J. Baik, G. Ben Arous, and S. Péché,Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices, Ann. Probab.33, 1643–1697 (2005)

  3. [10]

    Bandeira, G

    A. Bandeira, G. Cipolloni, D. Schröder, and R. van Handel,Matrix concentration inequalities and free probability II: Two-sided bounds and applications, Commun. Am. Math. Soc.6, 896–946 (2026)

  4. [11]

    Ben Arous, R

    G. Ben Arous, R. Gheissari, J. Huang, and A. Jagannath,Spectral alignment of stochastic gradient descent for high-dimensional classification tasks, Ann. Appl. Probab.35, 2767–2822 (2025)

  5. [12]

    Ben Arous, R

    G. Ben Arous, R. Gheissari, J. Huang, and A. Jagannath,Local geometry of high-dimensional mixture models: Effective spectral theory and dynamical transitions, arXiv preprint (2026),arXiv: 2502.15655

  6. [13]

    Ben Arous, R

    G. Ben Arous, R. Gheissari, and A. Jagannath,High-dimensional limit theorems for SGD: effective dynamics and critical scaling, Comm. Pure Appl. Math.77, 2030–2080 (2024)

  7. [14]

    Bietti, J

    A. Bietti, J. Bruna, and L. Pillaud-Vivien,On learning Gaussian multi-index models with gradient flow part I: general properties and two-timescale learning, Comm. Pure Appl. Math.78, 2354– 2435 (2025)

  8. [15]

    Bocchi, T

    D. Bocchi, T. Regimbeau, C. Lucibello, L. Saglietti, and C. Cammarota,Escape dynamics and implicit bias of one-pass SGD in overparameterized quadratic networks, arXiv preprint (2026), arXiv:2604.03068

  9. [16]

    Bolthausen,An iterative construction of solutions of the TAP equations for the Sherrington- Kirkpatrick model, Comm

    E. Bolthausen,An iterative construction of solutions of the TAP equations for the Sherrington- Kirkpatrick model, Comm. Math. Phys.325, 333–366 (2014)

  10. [17]

    Bonnaire, G

    T. Bonnaire, G. Biroli, and C. Cammarota,The role of the time-dependent Hessian in high- dimensional optimization, J. Stat. Mech. Theory Exp., Paper No. 083401, 32 (2025)

  11. [18]

    Bruna and D

    J. Bruna and D. Hsu,Survey on algorithms for multi-index models, Statist. Sci.40, 378–391 (2025)

  12. [19]

    Coeurdoux, G

    F. Coeurdoux, G. Ferré, and J.-P. Bouchaud,Random matrix theory of early-stopped gradient flow: a transient BBP scenario, arXiv preprint (2026),arXiv:2604.18450

  13. [20]

    Collins-Woodfin, C

    E. Collins-Woodfin, C. Paquette, E. Paquette, and I. Seroussi,Hitting the high-dimensional notes: an ODE for SGD learning dynamics on GLMs and multi-index models, Inf. Inference13, Paper No. iaae028, 107 (2024)

  14. [21]

    R. D. Cook,Save: a method for dimension reduction and graphics in regression, Communications in Statistics - Theory and Methods29, 2109–2121 (2000). Florent Krzakala, Pierre Mergny, and Vanessa Piccolo88

  15. [22]

    The generative leap: tight sample complexity for efficiently learninggaussian multi-indexmodels

    A. Damian, J. Lee, and J. Bruna, “The generative leap: tight sample complexity for efficiently learninggaussian multi-indexmodels”, Advancesin neuralinformationprocessing systems, Vol.38, edited by D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen...

  16. [23]

    Neural networks can learn representations with gradient descent

    A. Damian, J. Lee, and M. Soltanolkotabi, “Neural networks can learn representations with gradient descent”, Proceedings of thirty fifth conference on learning theory, Vol. 178, edited by P.-L. Loh and M. Raginsky, Proceedings of Machine Learning Research (2022), pp. 5413–5452

  17. [24]

    The benefits of reusing batches for gradient descent in two-layer networks: breaking the curse of informa- tion and leap exponents

    Y. Dandi, E. Troiani, L. Arnaboldi, L. Pesce, L. Zdeborová, and F. Krzakala, “The benefits of reusing batches for gradient descent in two-layer networks: breaking the curse of informa- tion and leap exponents”, Proceedings of the 41st international conference on machine learni...

  18. [25]

    Deep learning as neural low-degree filtering: a theory of hierarchical feature learning

    Y. Dandi, M. Vilucchio, L. Arnaboldi, H. Tabanelli, and F. Krzakala, “Deep learning as neural low-degree filtering: a theory of hierarchical feature learning”, High-dimensional learning dynamics 2026 (2026)

  19. [27]

    Defilippis, F

    L. Defilippis, F. Krzakala, B. Loureiro, and A. Maillard,Optimal scaling laws in learning hierar- chical multi-index models, arXiv preprint (2026),arXiv:2602.05846

  20. [28]

    H. Du, H. Hu, and S. Lepsveridze,Optimal spectral algorithms for correlated two-view models in high dimensions, arXiv preprint (2026),arXiv:2605.19364

  21. [29]

    The matrix Dyson equation and its applications for random matrices

    L. Erdős, “The matrix Dyson equation and its applications for random matrices”,Random matrices, Vol. 26, IAS/Park City Math. Ser. (Amer. Math. Soc., Providence, RI, 2019), pp. 75– 158

  22. [30]

    O. Y. Feng, R. Venkataramanan, C. Rush, and R. J. Samworth,A unifying tutorial on approximate message passing, Foundations and Trends in Machine Learning15, 335–536 (2022)

  23. [31]

    K. H. Fischer and J. A. Hertz,Spin glasses, Vol. 1, Cambridge Studies in Magnetism (Cambridge University Press, Cambridge, 1991), pp. x+408

  24. [32]

    Gerbelot and R

    C. Gerbelot and R. Berthier,Graph-based approximate message passing iterations, Inf. Inference 12, Paper No. iaad020, 67 (2023)

  25. [33]

    Gesztesy and E

    F. Gesztesy and E. Tsekanovskii,On matrix-valued Herglotz functions, Math. Nachr.218, 61–138 (2000)

  26. [34]

    J. W. Helton, R. Rashidi Far, and R. Speicher,Operator-valued semicircular elements: solving a quadratic matrix equation with positivity constraints, Int. Math. Res. Not., 1–15 (2007)

  27. [35]

    R. A. Horn and C. R. Johnson,Matrix analysis, Second (Cambridge University Press, Cambridge, 2013), pp. xviii+643

  28. [36]

    Javanmard and A

    A. Javanmard and A. Montanari,State evolution for general approximate message passing algo- rithms, with applications to spatial coupling, Inf. Inference2, 115–144 (2013)

  29. [37]

    Kovačević, Y

    F. Kovačević, Y. Zhang, and M. Mondelli,Spectral estimators for multi-index models: precise asymptotics and optimal weak recovery, arXiv preprint (2025),arXiv:2502.01583

  30. [38]

    Lehner,Computing norms of free operators with matrix coefficients, Amer

    F. Lehner,Computing norms of free operators with matrix coefficients, Amer. J. Math.121, 453–486 (1999)

  31. [39]

    Li,Sliced inverse regression for dimension reduction, J

    K.-C. Li,Sliced inverse regression for dimension reduction, J. Amer. Statist. Assoc.86, With discussion and a rejoinder by the author, 316–342 (1991)

  32. [40]

    Y. M. Lu and G. Li,Phase transitions of spectral initialization for high-dimensional non-convex estimation, Inf. Inference9, 507–541 (2020)

  33. [41]

    W. Luo, W. Alghamdi, and Y. M. Lu,Optimal spectral initialization for signal recovery with applications to phase retrieval, IEEE Trans. Signal Process.67, 2347–2456 (2019)

  34. [42]

    J. Ma, R. Dudeja, J. Xu, A. Maleki, and X. Wang,Spectral method for phase retrieval: an expectation propagation perspective, IEEE Trans. Inform. Theory67, 1332–1355 (2021). Spectral phase transitions in Gaussian multi-index models89

  35. [43]

    Construction of optimal spectral methods in phase retrieval

    A. Maillard, F. Krzakala, Y. M. Lu, and L. Zdeborová, “Construction of optimal spectral methods in phase retrieval”, Proceedings of the 2nd mathematical and scientific machine learning confer- ence, Vol. 145, edited by J. Bruna, J. Hesthaven, and L. Zdeborová, Proceedings of M...

  36. [44]

    Spectral phase transition and optimal PCA in block- structured spiked models

    P. Mergny, J. Ko, and F. Krzakala, “Spectral phase transition and optimal PCA in block- structured spiked models”, Proceedings of the 41st international conference on machine learning, Vol. 235, edited by R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlet...

  37. [45]

    J. A. Mingo and R. Speicher,Free probability and random matrices, Vol. 35, Fields Institute Monographs (Springer,NewYork;FieldsInstituteforResearchinMathematicalSciences,Toronto, ON, 2017), pp. xiv+336

  38. [46]

    Mondelli and A

    M. Mondelli and A. Montanari,Fundamental limits of weak recovery with applications to phase retrieval, Found. Comput. Math.19, 703–773 (2019)

  39. [47]

    Mondelli, C

    M. Mondelli, C. Thrampoulidis, and R. Venkataramanan,Optimal combination of linear and spectral estimators for generalized linear models, Found. Comput. Math.22, 1513–1566 (2022)

  40. [48]

    Mondelli and R

    M. Mondelli and R. Venkataramanan,Approximate message passing with spectral initialization for generalized linear models, J. Stat. Mech. Theory Exp., Paper No. 114003, 41 (2022)

  41. [49]

    Montanari and B

    A. Montanari and B. Saeed,Variational formulas for the spectrum of block Wishart matrices, arXiv preprint (2026),arXiv:2606.27774

  42. [50]

    Montanari and Z

    A. Montanari and Z. Wang,Phase transitions for feature learning in neural networks, arXiv preprint (2026),arXiv:2602.01434

  43. [51]

    Neural networks efficiently learn low-dimensional representations with SGD

    A. Mousavi-Hosseini, S. Park, M. Girotti, I. Mitliagkas, and M. A. Erdogdu, “Neural networks efficiently learn low-dimensional representations with SGD”, The eleventh international conference on learning representations (2023)

  44. [52]

    Nica and R

    A. Nica and R. Speicher,Lectures on the combinatorics of free probability, Vol. 335, Lon- don Mathematical Society Lecture Note Series (Cambridge University Press, Cambridge, 2006), pp. xvi+417

  45. [53]

    Parmaksiz and R

    E. Parmaksiz and R. van Handel,Computing extreme singular values of free operators, arXiv preprint (2025),arXiv:2510.23987

  46. [54]

    Spectral clustering of graphs with the bethe hessian

    A. Saade, F. Krzakala, and L. Zdeborová, “Spectral clustering of graphs with the bethe hessian”, Advances in neural information processing systems, Vol. 27, edited by Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Weinberger (2014)

  47. [55]

    Complex dynamics in simple neural networks: understanding gradient flow in phase retrieval

    S. Sarao Mannelli, G. Biroli, C. Cammarota, F. Krzakala, P. Urbani, and L. Zdeborová, “Complex dynamics in simple neural networks: understanding gradient flow in phase retrieval”, Advances in neural information processing systems, Vol. 33, edited by H. Larochelle, M. Ranzato, ...

  48. [56]

    Who is afraid of big bad minima? analysis of gradient-flow in spiked matrix-tensor models

    S. Sarao Mannelli, G. Biroli, C. Cammarota, F. Krzakala, and L. Zdeborová, “Who is afraid of big bad minima? analysis of gradient-flow in spiked matrix-tensor models”, Advances in neural information processing systems, Vol. 32, edited by H. Wallach, H. Larochelle, A. Beygelzim...

  49. [57]

    J. W. Silverstein and Z. D. Bai,On the empirical distribution of eigenvalues of a class of large- dimensional random matrices, J. Multivariate Anal.54, 175–192 (1995)

  50. [58]

    Learning gaussian multi-index models with gradient flow: time complexity and directional convergence

    B. Simsek, A. Bendjeddou, and D. Hsu, “Learning gaussian multi-index models with gradient flow: time complexity and directional convergence”, Proceedings of the 28th international conference on artificial intelligence and statistics, Vol. 258, edited by Y. Li, S. Mandt, S. Agr...

  51. [59]

    Subag,TAP approach for multispecies spherical spin glasses II: the free energy of the pure models, Ann

    E. Subag,TAP approach for multispecies spherical spin glasses II: the free energy of the pure models, Ann. Probab.51, 1004–1024 (2023)

  52. [60]

    Subag,TAP approach for multi-species spherical spin glasses I: General theory, Electron

    E. Subag,TAP approach for multi-species spherical spin glasses I: General theory, Electron. J. Probab.30, Paper No. 87, 32 (2025)

  53. [61]

    Tao,Topics in random matrix theory, Vol

    T. Tao,Topics in random matrix theory, Vol. 132, Graduate Studies in Mathematics (American Mathematical Society, Providence, RI, 2012), pp. x+282

  54. [62]

    D. J. Thouless, P. W. Anderson, and R. G. Palmer,Solution of ’solvable model of a spin glass’, The Philosophical Magazine: A Journal of Theoretical Experimental and Applied Physics35, 593–601 (1977). Florent Krzakala, Pierre Mergny, and Vanessa Piccolo90

  55. [63]

    Fundamental computational limits of weak learnability in high-dimensional multi-index models

    E. Troiani, Y. Dandi, L. Defilippis, L. Zdeborova, B. Loureiro, and F. Krzakala, “Fundamental computational limits of weak learnability in high-dimensional multi-index models”, Proceedings of the 28th international conference on artificial intelligence and statistics, Vol. 258...

  56. [64]

    Vershynin,High-dimensional probability, Vol

    R. Vershynin,High-dimensional probability, Vol. 47, Cambridge Series in Statistical and Proba- bilistic Mathematics, An introduction with applications in data science, With a foreword by Sara van de Geer (Cambridge University Press, Cambridge, 2018), pp. xiv+284

  57. [65]

    X. Yang, S. Sen, and Y. M. Lu,Sharp spectral thresholds for multi-view spiked Wigner models, arXiv preprint (2026),arXiv:2605.19894

  58. [66]

    Y. Q. Yin, Z. D. Bai, and P. R. Krishnaiah,On the limit of the largest eigenvalue of the large- dimensional sample covariance matrix, Probab. Theory Related Fields78, 509–521 (1988)

  59. [67]

    Zdeborová andF

    L. Zdeborová andF. Krzakala,Statistical physics of inference: thresholds and algorithms, Advances in Physics65, 453–552 (2016)

  60. [68]

    Zhang, Z

    B. Zhang, Z. Wang, H. Fu, and J. D. Lee,Neural networks learn generic multi-index models near information-theoretic limit, arXiv preprint (2025),arXiv:2511.15120

  61. [69]

    Zhang, H

    Y. Zhang, H. C. Ji, R. Venkataramanan, and M. Mondelli,Spectral estimators for structured generalized linear models via approximate message passing, Math. Stat. Learn.8, 193–304 (2025)

  62. [70]

    Zhang, H

    Y. Zhang, H. C. Ji, R. Venkataramanan, and M. Mondelli,Optimal estimation in orthogonally invariant generalized linear models: spectral initialization and approximate message passing, arXiv preprint (2026),arXiv:2602.09240

  63. [71]

    Zhang, M

    Y. Zhang, M. Mondelli, and R. Venkataramanan,Precise asymptotics for spectral methods in mixed generalized linear models, SIAM J. Math. Data Sci.8, 411–439 (2026)

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.