REVIEW 3 major objections 1 cited by
Online SGD on the nonlinear correlation objective recovers multiscale single-index features at Tensor-PCA sample complexity, with all layers trained on the same timescale.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 03:06 UTC pith:Z2WNFDZ2
load-bearing objection Solid same-timescale SGD analysis for a genuine multiscale hierarchical model, matching Tensor-PCA sample complexity; the softest link is mixed-incoherence propagation under the nonlinear field, not the depth-separation assumptions. the 3 major comments →
The Multiscale Single-Index Model: A Stylized Model for Hierarchical Feature Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For fixed depth K and local scale d, online spherical SGD on the nonlinear correlation loss, with all layers updated on the same timescale and started in a favorable basin of constant-in-d probability, achieves min_j m_j = 1-o_d(1) recovery of the planted multiscale directions with n = Õ(d^{K-1}) samples, the same complexity as the linear (Tensor-PCA) counterpart.
What carries the argument
The Edgeworth finite-rank hierarchy of the first Wiener chaos (and the companion staircase of singular-value plateaus in higher chaoses under non-cancellation): it refines the crude d^{-1/2} tensor perturbation into ordered spikes of size d^{-q/2}, so that product-scale overlap drifts dominate the nonlinear SGD dynamics throughout the search phase.
Load-bearing premise
Higher-order derivative moments of the link must not cancel among coalescence patterns of equal complexity, so the higher-chaos staircase plateaus stay visible; without that non-cancellation the depth-separation lower bounds collapse (and first-Hermite non-degeneracy is also required for the leading spike).
What would settle it
Construct a smooth activation with nonzero first Hermite coefficient whose higher Gaussian derivative moments cancel on every balanced coalescence chain of complexity ρ≥1; then the balanced flattening of a fixed higher chaos should lose its d^ρ-plateau, and a shallow network of width o(d^{(K-1)⌊R/2⌋}) should achieve constant L2 approximation error for large d, contradicting the claimed separation.
If this is right
- Depth is necessary for efficient approximation of MSIM targets whenever the Hermite energy decays slowly and higher-chaos non-cancellation holds.
- Same-timescale backpropagation on the correlation objective is sufficient; artificial layer-wise training is not required in this hierarchical model.
- Tensor unfolding of the empirical first Stein tensor recovers the planted directions at Õ(d^{⌈K/2⌉}) samples from the first chaos alone.
- For K>2 the correlation objective, not squared error, is the loss for which the analysis yields Tensor-PCA sample complexity; energy fluctuations of order d^{-1/2} swamp the MSE signal in mediocrity.
- Propagation of mixed incoherence keeps student directions delocalized, so nonlinear gradient errors remain lower-order relative to the product-scale drift.
Where Pith is reading between the lines
- If an analogous Edgeworth bookkeeping can be written for residual or attention architectures that lack strict scale separation, joint-timescale SGD may still succeed under comparable delocalization.
- The empirical failure of MSE for K>2 suggests a practical two-stage recipe: run correlation SGD to leave mediocrity, then switch to MSE for fine recovery.
- The staircase CP lower bounds supply a reusable template for proving depth separation on other multiscale Gaussian targets beyond the MSIM.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the Multiscale Single-Index Model (MSIM) of Oymak & Soltanolkotabi as a stylized hierarchical feature-learning model with physical scale separation. Under non-degeneracy (κ = E[φ'] ≠ 0) and delocalization of the planted directions, the first Wiener chaos is shown via Edgeworth expansions to be a finite-rank hierarchy of tensor spikes at scales d^{-q/2} (Theorem 3.1), making MSIM a nonlinear analogue of Tensor PCA. Higher chaoses are shown to have staircase singular-value plateaus under a non-cancellation assumption, yielding shallow-network approximation lower bounds (Theorem 4.18 / Corollary 4.20). The main algorithmic result is that online spherical SGD on the nonlinear correlation objective, with all layers on the same timescale and initialization in a favorable basin of constant-in-d probability, achieves min_j m_j = 1-o_d(1) recovery with n = Õ(d^{K-1}) samples (Theorems 6.21 and 6.24), matching the linear Tensor-PCA complexity. Spectral recovery via tensor unfolding is also recovered at the standard unfolding threshold.
Significance. If correct, the work supplies a clean, tractable hierarchical model in which both depth necessity (via chaos incompressibility) and sufficiency of same-timescale backpropagation (without layerwise training or artificial timescale separation) can be proved at the same sample complexity as the linear spiked-tensor baseline. The Edgeworth hierarchy of the first Stein tensor and the mixed-incoherence propagation argument are technically substantial contributions that go beyond the first-order perturbation used for spectral methods in prior work. The matching of Tensor-PCA sample complexity for nonlinear correlation SGD is a strong and falsifiable claim. The manuscript is self-contained, with detailed proofs for the hierarchy, staircase, population and online linear SGD, and the nonlinear comparison-plus-propagation pipeline; numerical experiments with normalized GeLU provide partial empirical support.
major comments (3)
- The central recovery claim (informal Thm 2.6; Thms 6.21 and 6.24) rests on propagation of the dynamic mixed-incoherence class J_{P,B} (Def 6.2) under the full nonlinear field F^{nl}, not only under the linear proxy (Thm 6.9 for population flow; Prop 6.20 and the stopped bootstrap in Thm 6.21 for online SGD). The relative-error control in Prop 6.7 / Cor 6.11 (and the product-drift lower bound ṁ_j ≳ ∏_{ℓ≠j} m_ℓ for K>2) fails if any mixed moment |S^{p,q}_j| or the coordinate envelope exits the class before weak recovery. The bootstrap is carefully written, but the paper never reports intermediate diagnostics of these mixed moments along the trajectories used in §7. Given that Table 1 already shows success fractions decaying with K (down to 3–6/10 at K=4), the manuscript should either (i) add mixed-moment diagnostics for the GeLU experiments confirming that J_{P,B} is preserved on successfu
- In §6.1 and Remark 6.26 the authors note that the MSE objective is not equivalent to correlation for K>1 and that energy fluctuations of order d^{-1/2} dominate the search-phase signal for K>2. Table 1 and Figure 2 confirm that MSE fails as K and d grow while correlation succeeds more often. The main theorem is therefore specific to the correlation objective. This is acceptable, but the abstract and informal Thm 2.6 should state explicitly that the Õ(d^{K-1}) guarantee is for the correlation loss, not for the more standard MSE loss used in practice; otherwise readers may over-read the claim as applying to vanilla teacher–student training.
- Depth separation (Thm 4.18 / Cor 4.20) depends on the higher-chaos non-cancellation Assumption 4.5/4.15 (staircase plateaus of multiplicity ~d^ρ at scale d^{-ρ/2}) and the slow Hermite-energy tail Assumption 4.19. Both are presented as natural, and Examples 4.6–4.7 give sufficient algebraic conditions, but the paper does not check them for the normalized GeLU used in §7, nor does it quantify how large a set of φ satisfies non-cancellation uniformly over the finite block {J_ε,…,L_ε}. Since depth separation is advertised as quantifying the benefit of depth, a short verification (or a generic-density argument) for at least one concrete activation would make the claim load-bearing rather than conditional on an unchecked algebraic hypothesis.
Circularity Check
No load-bearing circularity: main recovery and hierarchy claims are self-contained expansions under explicit non-degeneracy/incoherence assumptions; self-citations are non-essential background.
specific steps
-
self citation load bearing
[Related Works / Statement of Tool Use / Conclusions (companion [DB26])]
"In the companion paper [DB26] we analyze the MSE landscape in the only feasible setting K=2. ... Statement of Tool Use. We used GPT 5.5 Pro to proofread our results... The author assumes full responsibility..."
The companion [DB26] (same senior author) is cited for the MSE objective, which the present paper deliberately does not claim; the main nonlinear-correlation SGD theorems never invoke it. The GPT tool-use note is meta and does not enter any derivation. This is ordinary background self-citation, not a load-bearing uniqueness or ansatz import that forces the recovery claim.
full rationale
The paper's central claims (Edgeworth finite-rank hierarchy of the first Stein tensor in Thm 3.1; staircase singular values under higher-chaos non-cancellation in Thm 4.8; shallow CP lower bounds in Thm 4.18/Cor 4.20; online spherical SGD weak-to-strong recovery for nonlinear correlation in Thms 6.21/6.24) are derived from first-principles tools: Gaussian integration by parts/Stein identities, recursive Faà di Bruno + weighted Edgeworth expansions of the multiscale computation tree, singular-value analysis of copy-diagonal atoms, and stopped bootstrap + martingale/retraction controls for the spherical flow. These do not reduce the target statements (e.g., product-drift ṁ_j ≳ ∏_{ℓ≠j} m_ℓ or min m_j = 1-o_d(1) at n=Õ(d^{K-1})) to fitted quantities or to definitions of the same quantities. The MSIM model and the base non-degeneracy/incoherence assumptions are imported from the external reference [OS21] and restated explicitly (Assumptions 2.1–2.3); they are not defined in terms of the recovery claim. Higher-chaos non-cancellation (Ass. 4.5/4.15) is an explicit algebraic non-vanishing hypothesis on Gaussian derivative moments of ϕ, not derived from the depth-separation conclusion. Self-citations ([BH25] survey, companion [DB26] on MSE) appear only as related-work or future-work pointers and are not used to justify any step of the main proofs. Numerical experiments in §7 are pure illustration (success fractions under GeLU) and do not feed parameters back into theorems. Minor self-reference exists but is not load-bearing, yielding score 1 rather than 0.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption Non-degeneracy: E[φ'(G)]=κ≠0 (information exponent 1) together with Eφ=0, Eφ²=1 (Assumption 2.1).
- domain assumption Incoherence/delocalization of planted directions: ||θ_k||_∞ ≤ C√(log d / d) and higher ℓ_p bounds (Assumption 2.3).
- domain assumption Smoothness of φ: C^{L_K} with polynomial growth of derivatives (Assumption 2.2).
- ad hoc to paper Higher-chaos non-cancellation: leading coalescence patterns of equal crossing complexity do not cancel, so balanced flattenings retain staircase plateaus (Assumption 4.5/4.15).
- ad hoc to paper Slow Hermite-energy tail: sum_{R≥J} Λ̄_R^{(K)} ≳ J^{-p} (Assumption 4.19).
- domain assumption Favorable initialization basin of constant probability (positive overlaps of order d^{-1/2} and mixed-incoherence bounds).
invented entities (2)
-
Edgeworth finite-rank hierarchy of the first Stein tensor (G^{[q]}_θ)
independent evidence
-
Higher-chaos staircase singular-value plateaus of balanced flattenings
independent evidence
read the original abstract
We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}. Each layer extracts a shared single-index feature at one physical scale and passes it to the next, thus defining a tractable setting in which to study how deep architectures learn multiscale representations. Under non-degeneracy and delocalization assumptions on the link function and planted features respectively, for fixed depth $K$ and local scale $d$, the first Wiener chaos of the target behaves as a perturbed spiked tensor, where the perturbation of order $d^{-1/2}$ comes from the non-linearity -- revealing the MSIM as a natural non-linear analogue of the Tensor PCA model \cite{montanari2014statistical}. While this perturbative picture is sufficient to enable efficient spectral recovery based on Tensor unfolding (as already observed in \cite{oymak2021learning}), it is not precise enough for the analysis of backpropagation gradient-based methods. In this work, we address this limitation by performing a fine-grained analysis of the Wiener chaos using Edgeworth expansions. In the first chaos, this gives a finite-rank hierarchy at scales $d^{-q/2}$. In higher chaoses, balanced flattenings exhibit staircase singular-value plateaus of size $d^{-\rho/2}$ and multiplicity $d^{\rho}$ under a natural higher-chaos non-cancellation condition. Using this higher-chaos structure, and under an additional slow Hermite-energy tail condition, we first establish shallow-network approximation lower bounds, quantifying the benefit of depth in this model. Next, and most importantly, we prove that online SGD on the correlation objective, where all layers evolve in the same timescale, achieves $1 - o_d(1)$ recovery with $n = \widetilde{O}( d^{K-1})$ samples, recovering the same sample complexity as in the linear counterpart.
Figures
Forward citations
Cited by 1 Pith paper
-
Algorithmic Separation between Constant-Depth and Logarithmic-Depth Neural Networks
Logarithmic-depth networks trained by layerwise coordinate descent can learn hierarchical staircase Boolean functions that constant-depth networks with bounded spectral norms cannot approximate.
Reference graph
Works this paper leans on
-
[1]
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz. The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks. In Conference on Learning Theory , pages 4782--4887. PMLR, 2022
2022
-
[2]
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz. Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics. In The Thirty Sixth Annual Conference on Learning Theory , pages 2552--2623. PMLR, 2023
2023
-
[3]
Stochastic interpolants: A unifying framework for flows and diffusions
Michael S Albergo, Nicholas M Boffi, and Eric Vanden-Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint arXiv:2303.08797 , 2023
Pith/arXiv arXiv 2023
-
[4]
Algorithmic thresholds for tensor PCA
G \'e rard Ben Arous, Reza Gheissari, and Aukosh Jagannath. Algorithmic thresholds for tensor PCA . The Annals of Probability , 48(4):2052 -- 2087, 2020
2052
-
[5]
Online stochastic gradient descent on non-convex losses from high-dimensional inference, 2021
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath. Online stochastic gradient descent on non-convex losses from high-dimensional inference, 2021
2021
-
[6]
High-dimensional optimization for multi-spiked tensor pca
G \'e rard Ben Arous, C \'e dric Gerbelot, and Vanessa Piccolo. High-dimensional optimization for multi-spiked tensor pca. arXiv preprint arXiv:2408.06401 , 2024
Pith/arXiv arXiv 2024
-
[7]
Quality over quantity in attention layers: When adding more heads hurts
Noah Amsel, Gilad Yehudai, and Joan Bruna. Quality over quantity in attention layers: When adding more heads hurts. In Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu, editors, International Conference on Learning Representations , volume 2025, pages 62431--62483, 2025
2025
-
[8]
What Can ResNet Learn Efficiently, Going Beyond Kernels? 2019
Zeyuan Allen-Zhu and Yuanzhi Li. What Can ResNet Learn Efficiently, Going Beyond Kernels? 2019
2019
-
[9]
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath. Online stochastic gradient descent on non-convex losses from high-dimensional inference. Journal of Machine Learning Research (JMLR) , 22:106--1, 2021
2021
-
[10]
On learning gaussian multi-index models with gradient flow
Alberto Bietti, Joan Bruna, and Loucas Pillaud-Vivien. On learning gaussian multi-index models with gradient flow. arXiv preprint arXiv:2310.19793 , 2023
Pith/arXiv arXiv 2023
-
[11]
Survey on algorithms for multi-index models, 2025
Joan Bruna and Daniel Hsu. Survey on algorithms for multi-index models, 2025
2025
-
[12]
Optimal errors and phase transitions in high-dimensional generalized linear models
Jean Barbier, Florent Krzakala, Nicolas Macris, L \'e o Miolane, and Lenka Zdeborov \'a . Optimal errors and phase transitions in high-dimensional generalized linear models. Proceedings of the National Academy of Sciences , 116(12):5451--5460, 2019
2019
-
[13]
Invariant scattering convolution networks
Joan Bruna and St \'e phane Mallat. Invariant scattering convolution networks. IEEE Transactions on Pattern Analysis and Machine Intelligence , 35(8):1872--1886, 2013
2013
-
[14]
Resnets of all shapes and sizes: Convergence of training dynamics in the large-scale limit
Louis-Pierre Chaintron, L \'e na \" c Chizat, and Javier Maass. Resnets of all shapes and sizes: Convergence of training dynamics in the large-scale limit. arXiv preprint arXiv:2603.18168 , 2026
arXiv 2026
-
[15]
The hidden width of deep resnets: Tight error bounds and phase diagrams
L \'e na \" c Chizat. The hidden width of deep resnets: Tight error bounds and phase diagrams. arXiv preprint arXiv:2509.10167 , 2025
arXiv 2025
-
[16]
Asymptotics of feature learning in two-layer networks after one gradient-step
Hugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala, Yue M Lu, Lenka Zdeborov \'a , and Bruno Loureiro. Asymptotics of feature learning in two-layer networks after one gradient-step. arXiv preprint arXiv:2402.04980 , 2024
Pith/arXiv arXiv 2024
-
[17]
How deep neural networks learn compositional data: The random hierarchy model
Francesco Cagnetta, Leonardo Petrini, Umberto M Tomasini, Alessandro Favero, and Matthieu Wyart. How deep neural networks learn compositional data: The random hierarchy model. Physical Review X , 14(3):031001, 2024
2024
-
[18]
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations, 2019
2019
-
[19]
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua. On the expressive power of deep learning: A tensor analysis. In Conference on learning theory , pages 698--728. PMLR, 2016
2016
-
[20]
Mse analysis of online sgd for the multiscale single index model
Gordon Dai and Joan Bruna. Mse analysis of online sgd for the multiscale single index model. Preprint , 2026
2026
-
[21]
Donoho and Michael J
David L. Donoho and Michael J. Feldman. Sharp recovery thresholds of tensor pca spectral algorithms. In Advances in Neural Information Processing Systems , 2023
2023
-
[22]
Learning single-index models in gaussian space
Rishabh Dudeja and Daniel Hsu. Learning single-index models in gaussian space. In Conference On Learning Theory , pages 1887--1930. PMLR, 2018
1930
-
[23]
Statistical query lower bounds for tensor pca
Rishabh Dudeja and Daniel Hsu. Statistical query lower bounds for tensor pca. Journal of Machine Learning Research , 22(83):1--51, 2021
2021
-
[25]
Learning two-layer neural networks, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan. Learning two-layer neural networks, one (giant) step at a time. arXiv preprint arXiv:2305.18270 , 2023
Pith/arXiv arXiv 2023
-
[26]
The generative leap: Sharp sample complexity for efficiently learning gaussian multi-index models
Alex Damian, Jason D Lee, and Joan Bruna. The generative leap: Sharp sample complexity for efficiently learning gaussian multi-index models. arXiv preprint arXiv:2506.05500 , 2025
Pith/arXiv arXiv 2025
-
[27]
Computational-statistical gaps in gaussian single-index models
Alex Damian, Loucas Pillaud-Vivien, Jason D Lee, and Joan Bruna. Computational-statistical gaps in gaussian single-index models. arXiv preprint arXiv:2403.05529 , 2024
Pith/arXiv arXiv 2024
-
[28]
The computational advantage of depth: Learning high-dimensional hierarchical functions with gradient descent, 2025
Yatin Dandi, Luca Pesce, Lenka Zdeborová, and Florent Krzakala. The computational advantage of depth: Learning high-dimensional hierarchical functions with gradient descent, 2025
2025
-
[29]
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir. The power of depth for feedforward neural networks. In Conference on learning theory , pages 907--940. PMLR, 2016
2016
-
[30]
Sharp recovery thresholds of tensor pca spectral algorithms
Michael Feldman and David Donoho. Sharp recovery thresholds of tensor pca spectral algorithms. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems , volume 36, pages 56628--56640. Curran Associates, Inc., 2023
2023
-
[31]
Universality of high-dimensional scaling limits of stochastic gradient descent
Reza Gheissari and Aukosh Jagannath. Universality of high-dimensional scaling limits of stochastic gradient descent. arXiv preprint arXiv:2512.13634 , 2025
arXiv 2025
-
[32]
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma. Learning one-hidden-layer neural networks with landscape design. arXiv preprint arXiv:1711.00501 , 2017
Pith/arXiv arXiv 2017
-
[33]
A mathematical perspective on transformers
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet. A mathematical perspective on transformers. Bulletin of the American Mathematical Society , 62(3):427--479, 2025
2025
-
[34]
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015
2015
-
[35]
On the complexity of learning sparse functions with statistical and gradient queries
Nirmit Joshi, Theodor Misiakiewicz, and Nathan Srebro. On the complexity of learning sparse functions with statistical and gradient queries. arXiv preprint arXiv:2407.05622 , 2024
Pith/arXiv arXiv 2024
-
[36]
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar. Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods. arXiv preprint arXiv:1506.08473 , 2015
Pith/arXiv arXiv 2015
-
[37]
Bayesian inference with finitely wide neural networks
Chi-Ken Lu et al. Bayesian inference with finitely wide neural networks. Physical Review E , 108:014311, 2023
2023
-
[38]
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations , 2022
2022
-
[39]
Statistical-computational trade-offs in learning multi-index models via harmonic analysis
Hugo Latourelle-Vigeant and Theodor Misiakiewicz. Statistical-computational trade-offs in learning multi-index models via harmonic analysis. arXiv preprint arXiv:2602.09959 , 2026
arXiv 2026
-
[40]
Understanding deep convolutional networks
St \'e phane Mallat. Understanding deep convolutional networks. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences , 374(2065):20150203, 2016
2065
-
[41]
Fundamental limits of weak recovery with applications to phase retrieval
Marco Mondelli and Andrea Montanari. Fundamental limits of weak recovery with applications to phase retrieval. In Conference On Learning Theory , pages 1445--1450. PMLR, 2018
2018
-
[42]
A statistical model for tensor pca
Andrea Montanari and Emile Richard. A statistical model for tensor pca. Advances in neural information processing systems , 27, 2014
2014
-
[43]
Phase transitions for feature learning in neural networks, 2026
Andrea Montanari and Zihao Wang. Phase transitions for feature learning in neural networks, 2026
2026
-
[44]
The connection between approximation, depth separation and learnability in neural networks
Eran Malach, Gilad Yehudai, Shai Shalev-Schwartz, and Ohad Shamir. The connection between approximation, depth separation and learnability in neural networks. In Conference on Learning Theory , pages 3265--3295. PMLR, 2021
2021
-
[45]
Provable guarantees for nonlinear feature learning in three-layer neural networks
Eshaan Nichani, Alex Damian, and Jason D Lee. Provable guarantees for nonlinear feature learning in three-layer neural networks. Advances in Neural Information Processing Systems , 36, 2023
2023
-
[46]
Improving the gaussian approximation in neural networks: Para-gaussians and edgeworth expansions, 2024
Mihai Nica and Janosch Ortmann. Improving the gaussian approximation in neural networks: Para-gaussians and edgeworth expansions, 2024. NeurIPS Workshop on Mathematics of Modern Machine Learning
2024
-
[47]
Learning a deep convolutional neural network via tensor decomposition
Samet Oymak and Mahdi Soltanolkotabi. Learning a deep convolutional neural network via tensor decomposition. Information and Inference: A Journal of the IMA , 10(3):1031--1071, 2021
2021
-
[48]
Yunwei Ren, Yatin Dandi, Florent Krzakala, and Jason D. Lee. Provable learning of random hierarchy models and hierarchical shallow-to-deep chaining, 2026
2026
-
[49]
A statistical model for tensor pca
Emile Richard and Andrea Montanari. A statistical model for tensor pca. In Advances in Neural Information Processing Systems , 2014
2014
-
[50]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning , pages 2256--2265. pmlr, 2015
2015
-
[51]
Optimization-based separations for neural networks
Itay Safran and Jason Lee. Optimization-based separations for neural networks. In Po-Ling Loh and Maxim Raginsky, editors, Proceedings of Thirty Fifth Conference on Learning Theory , volume 178 of Proceedings of Machine Learning Research , pages 3--64. PMLR, 02--05 Jul 2022
2022
-
[52]
Serre, L
T. Serre, L. Wolf, and T. Poggio. Object recognition with features inspired by visual cortex. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05) , volume 2, pages 994--1000 vol. 2, 2005
2005
-
[53]
Deep learning of compositional targets with hierarchical spectral methods
Hugo Tabanelli, Yatin Dandi, Luca Pesce, and Florent Krzakala. Deep learning of compositional targets with hierarchical spectral methods. arXiv preprint arXiv:2602.10867 , 2026
arXiv 2026
-
[54]
Benefits of depth in neural networks
Matus Telgarsky. Benefits of depth in neural networks. In Conference on learning theory , pages 1517--1539. PMLR, 2016
2016
-
[55]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems , 30, 2017
2017
-
[56]
Zihao Wang, Eshaan Nichani, and Jason D. Lee. Learning hierarchical polynomials with three-layer neural networks, 2023
2023
-
[57]
Exponential separations in symmetric neural networks
Aaron Zweig and Joan Bruna. Exponential separations in symmetric neural networks. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems , volume 35, pages 33134--33145. Curran Associates, Inc., 2022
2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.