Pith. sign in

REVIEW 2 major objections 4 minor 67 references

Statistical Limits for Finite-Rank Tensor Estimation

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper proves an exact variational formula for the asymptotic free energy of a broad class of tensor estimation models with nonlinear observations, heteroskedastic noise, and covariates, and applies it to derive sharp statistical…

desk verdict A serious, mostly sound unified framework for q-wise tensor estimation; the applications are real, but the abstract's MMSE claim outruns the proofs and the key concentration condition is assumed rather than verified in the flagship examples. read the letter →

arxiv 2506.06749 v1 pith:POU6KKRE submitted 2025-06-07 cs.IT math.ITmath.STstat.TH

classification cs.ITmath.ITmath.STstat.TH MSC 62B1094A15
keywords tensorestimationfinite-ranktensorsfreeenergymutualinformationminimummean-squarederrorspikedmodelheteroskedasticnoiseq-adicassignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves an asymptotically exact formula for the information-theoretic limits of a broad class of finite-rank tensor estimation problems, where each observation depends on a known nonlinear function of q unknown parameters plus Gaussian noise, and where covariates, heteroskedasticity, and side information are allowed. The central result is that the free energy—which is equivalent to the mutual information—converges to the value of a max-min variational problem over finite-dimensional positive semidefinite matrices. Because differentiating the free energy yields the Bayes-optimal minimum mean-squared error, the formula gives sharp statistical thresholds rather than merely abstract limits. The framework is then applied to two settings that were previously open: spiked tensor estimation with non-identically distributed noise (including missing entries) and higher-order q-adic assignment problems, where an unknown permutation must be recovered from tensor-valued observations.

What carries the argument

The load-bearing object is the pair $(\Psi,\ \sup\text{-}\inf)$: the free-energy limit is expressed as the maximum over $Q$ and minimum over $S$ of the potential function $\Psi$. The proof mechanism is a two-step reduction: first, any sufficiently regular q-wise interaction model is approximated by a multilinear model whose free energy is close in quadratic Wasserstein distance (an optimal-transport metric on distributions), via Theorem 1 using the HWI inequality; second, the multilinear model is lifted, using one-hot encodings of covariates and tensor products, into a matrix tensor product (MTP) model whose free-energy limit is known. Orthogonal-projector identities map the lifted formula back to the compact max-min form, which strictly generalizes the single-maximum formulas that hold for positive semidefinite interactions.

What would settle it

Pick a q-wise interaction model satisfying the multilinear approximation condition but with parameters chosen so that the uniform log-density concentration fails (for example, covariates with heavy tails that make the augmented log-likelihood fluctuate wildly), simulate the normalized mutual information at a fixed SNR, and compare it to the max-min value of $\Psi$; a gap that persists as $n$ grows would refute Theorem 2. Equivalently, for a $q=3$ assignment problem with a kernel whose approximating series fails the concentration condition, check whether finite-$n$ mutual information approaches the predicted limit.

Watch

Extended reading notes

Core claim

The central claim is Theorem 2: for a multilinear q-wise interaction model satisfying a uniform concentration condition, the asymptotic free energy equals $\sup_{Q\in(S_+^d)^k}\inf_{S\in(S_+^d)^k}\Psi(Q,S)$, where $\Psi(Q,S)=H(S+\mathring{S})-\frac12\langle S,Q\rangle+\frac t2\sum_{\kappa\in[k]^q}\langle A_\kappa^\top A_\kappa,Q_{\kappa_1}\otimes\cdots\otimes Q_{\kappa_q}\rangle$; here $H$ is the relative entropy of a linear Gaussian channel, $\mathring{S}$ is the side-information matrix, and $S_+^d$ is the cone of positive semidefinite $d\times d$ matrices. This formula characterizes the mutual information between the parameters and the observations, and differentiation with respect to the SNR parameter $t$ yields the minimum mean-squared error. Theorems 3 and 4 show that the same max-min structure governs heteroskedastic spiked tensors, whose noise variances may vary arbitrarily and even be infinite (unobserved entries), and q-adic assignment problems, where the signal is a random permutation of known atoms.

Load-bearing premise

The load-bearing premise is that the log-density of the observation model, augmented with a linear side-information channel, concentrates uniformly on every compact set of parameters; without this concentration the max-min formula does not follow, and in the assignment application the premise is assumed for each approximating model rather than proved.

Editorial extensions

If this is right

  • The mutual information and the minimum mean-squared error for any q-wise interaction model satisfying the conditions can be computed by solving a max-min problem over positive semidefinite matrices, making the statistical limits explicit rather than qualitative.
  • For heteroskedastic spiked tensors, the formula yields exact detection and estimation thresholds for noise profiles that need not be positive definite and may include missing entries, extending the matrix results to every order q.
  • For q-adic assignment problems, the formula gives the asymptotic mutual information for recovering a random permutation from tensor-valued observations, extending the known quadratic case to higher-order matching.
  • Because the limiting free energy is convex and differentiable almost everywhere, the Bayes-optimal error curves follow by differentiation wherever the formula is differentiable; at boundary points the subdifferential gives lower bounds on the Bayes risk.
  • The known special cases—homogeneous spiked tensors, positive-semidefinite matrix models, and groupwise heteroskedastic matrices—are recovered as limits of the same potential function, so the framework acts as a common origin for those results.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same max-min mechanism should carry over to other multi-way inference problems expressible as q-wise interactions, such as hypergraph clustering and multidimensional assignment, whenever the corresponding approximation and concentration conditions can be verified.
  • Because the free energy is shown to be continuous in Wasserstein distance, the Gaussian-noise assumption is likely removable: standard universality arguments should extend the formulas to sub-Gaussian or other regular noise channels, though the paper only notes this direction.
  • The paper's discussion of boundary behavior suggests that in some models the Bayes risk may be discontinuous at zero side information, meaning the limiting MMSE might not be achieved by any algorithm operating at the boundary; this is a testable prediction for specific spiked tensor models.
  • The variational formula is finite-dimensional for multilinear models, so it could be turned into a numerical oracle: solving the max-min problem and comparing against message-passing, spectral, or sum-of-squares estimators would quantify statistical-to-computational gaps for these tensor problems.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes a unified framework for the analysis of high-dimensional tensor estimation problems with nonlinear observations, heteroskedastic noise, and covariates. The authors introduce a q-wise interaction model, prove a multilinear approximation result (Theorem 1), and then establish an asymptotic free energy formula via a max-min variational principle (Theorem 2) under an explicit concentration condition (Condition 2). The framework is applied to two new settings: heteroskedastic spiked tensor estimation (Theorem 3) and q-adic assignment problems (Theorem 4). The proofs rely on embedding into a matrix tensor product model and using the known MTP limit theorem of Chen-Mourrat-Xia.

Significance. If correct, the main results provide a genuine unification of several previously studied problems, such as spiked matrix/tensor models, heteroskedastic low-rank matrix estimation, and quadratic assignment problems, and they extend the known limit formulas to higher-order interactions, indefinite kernels, and arbitrary variance profiles. The paper gives precise theorems and detailed proofs, uses no fitted parameters, and explicitly cites the external MTP result that serves as its backbone. The applications to heteroskedastic tensors and q-adic assignment are new and interesting. The main caveat is that the central free energy theorem is conditional on a strong concentration hypothesis (Condition 2b) whose verification for the two applications is only asserted, not proved; this gap is load-bearing.

major comments (2)
  1. [Section 2.3, Condition 2; Appendix B.1; Appendix C.3] Condition 2b is a genuine, load-bearing hypothesis: the proof of Theorem 2 imports the MTP limit (Theorem 5) and therefore inherits its hypotheses, including the uniform L2 concentration of the log-density over compact subsets. For Theorem 3, the verification of Condition 2 is only stated in Appendix B.1 as 'verified via the Efron-Stein inequality (see e.g., [59, Appendix C])', but no argument is given that the supremum over the compact set C in Condition 2b is controlled. For Theorem 4, the proof in Appendix C.3 asserts that permutation invariance reduces the needed concentration to a decoupled i.i.d. model, but the reduction and the resulting bound are not supplied. Since the exact free-energy formulas for these applications rest on Condition 2b, the authors should provide complete proofs of the concentration property for the approximating models, or explicitly state the missing condition as an assumption in Theorems 3 and 4.
  2. [Abstract and Section 2.4] The abstract claims 'asymptotically exact formulas for the mutual information ... as well as the minimum mean-squared error', but Section 2.4 explicitly notes that the limiting overlap may be discontinuous at the boundary S = 0 and that only a subdifferential lower bound is available there. The main theorems are free energy (mutual information) results; the MMSE characterization holds almost everywhere in the side-information parameter and may fail at the boundary point relevant to the original model. This should be stated more carefully in the abstract and introduction to avoid overclaiming what is proven.
minor comments (4)
  1. [Section 4, model (20)] In the description of the approximating model for the assignment problem, the text says 'd-linear approximation model', but the model is q-wise multilinear in the q arguments; please replace 'd-linear' with 'q-linear' to avoid confusion with the feature dimension d.
  2. [Appendix A.2, statement of Theorem 5] In the sentence before Theorem 5, the notation 'FN (˚S, t)' appears instead of 'Fn(˚S, t)' in one place; this is a typo that should be corrected.
  3. [Appendix C.2, proof of Lemma 3] In the final displayed equation of the proof, the expression 'H(µ ∗ g ∥ g) − H(µ ∗ g ∥ g)' is a typo; the second term should be 'H(ν ∗ g ∥ g)'.
  4. [Equation (12) and surrounding text] The potential function Ψ(Q,S) in (12) is written with H(S + ˚S), but the relative entropy H in (11) is defined for the linear channel as depending on ˚S. The notation is clear from context, but it would help to explicitly state that H(·) is the pointwise limit of H_n and that the argument S + ˚S is an abuse of notation; a single sentence would remove ambiguity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the main free-energy theorem is derived from an external MTP limit, and the application theorems rest on explicit concentration assumptions rather than on a self-referential reduction.

full rationale

The central derivation is self-contained against external benchmarks. Theorem 2 (Appendix A.2) proves the variational formula by lifting the (d,k)-multilinear model into an MTP model and invoking Theorem 5, quoted from Chen–Mourrat–Xia [16], an external published result; the target free-energy formula is not assumed as an input. The potential function in Eq. (12) is written in terms of the linear-channel relative entropy H, which is a model ingredient, not the conclusion being derived. Theorem 3 obtains the heteroskedastic spiked tensor limit by constructing a stepping approximation under Condition 3 and then applying Theorem 2; the only self-citation, [59] for an Efron–Stein verification of Condition 2, is a standard concentration technique and is not load-bearing for the new variational structure. Theorem 4 for the q-adic assignment problem explicitly assumes Condition 2 for every approximating model (20), and Appendix C.3 only asserts that permutation invariance reduces the needed concentration to a decoupled model without supplying the proof; this is an unverified hypothesis and a missing proof, not a circular equivalence. No fitted parameter is renamed as a prediction, and no self-citation is used to forbid alternatives. The honest finding is no significant circularity, with the caveat that the application theorems are conditional on Condition 2b, which is genuinely assumed rather than proven for the general model.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The analysis introduces no fitted constants: all parameters (SNR t, variance profile sigma_alpha, basis functions) are inputs of the model. The cost is paid in regularity assumptions (Conditions 1-4) rather than in free parameters.

assumptions (6)
  • standard math HWI inequality of Otto-Villani
    Used in Proof of Theorem 1 to bound free energy difference via Wasserstein distance and Fisher information.
  • standard math MTP limit theorem of Chen-Mourrat-Xia (Theorem 5 / [16, Thm 1.1])
    Used as the starting point for the lifting argument in Proof of Theorem 2.
  • domain assumption Condition 1: existence of a multilinear approximation with o(n) squared Wasserstein error
    Ensures the free energy of the general model is close to that of a multilinear model; verified via L2 approximation (7) for the applications.
  • domain assumption Condition 2: pointwise convergence and differentiability of H_n, plus L2 concentration of log rho
    Technical regularity needed for the variational formula; verified only for special cases, assumed for the general model.
  • domain assumption Condition 3: inverse noise variance profile converges in L1 to a bounded limit
    Needed for the heteroskedastic tensor result and satisfied for finite-type or positive-definite profiles.
  • domain assumption Condition 4: the kernel f admits a finite-rank basis expansion with uniform error
    Needed for the q-adic assignment result; for q=2 PSD kernels it follows from Mercer's theorem.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Statistical Limits for Finite-Rank Tensor Estimation." pith.science (2026). https://pith.science/paper/POU6KKRE

@misc{pith2026250606749,
  author       = {Pith},
  title        = {Pith review of: Statistical Limits for Finite-Rank Tensor Estimation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/POU6KKRE}},
  note         = {Machine review of arXiv:2506.06749}
}
read the original abstract

This paper provides a unified framework for analyzing tensor estimation problems that allow for nonlinear observations, heteroskedastic noise, and covariate information. We study a general class of high-dimensional models where each observation depends on the interactions among a finite number of unknown parameters. Our main results provide asymptotically exact formulas for the mutual information (equivalently, the free energy) as well as the minimum mean-squared error in the Bayes-optimal setting. We then apply this framework to derive sharp characterizations of statistical thresholds for two novel scenarios: (1) tensor estimation in heteroskedastic noise that is independent but not identically distributed, and (2) higher-order assignment problems, where the goal is to recover an unknown permutation from tensor-valued observations.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

67 extracted references · 57 canonical work pages

  1. [1]

    Community detection and stochastic block models: Recent developments.Journal of Machine Learning Research, 18(177):1–86, 2018

    Emmanuel Abbe. Community detection and stochastic block models: Recent developments.Journal of Machine Learning Research, 18(177):1–86, 2018

  2. [2]

    Recovering communities in the general stochastic block model without knowing the parameters.Advances in neural information processing systems, 28, 2015

    Emmanuel Abbe and Colin Sandon. Recovering communities in the general stochastic block model without knowing the parameters.Advances in neural information processing systems, 28, 2015

  3. [3]

    Beyond pairwise clustering

    Sameer Agarwal, Jongwoo Lim, Lihi Zelnik-Manor, Pietro Perona, David Kriegman, and Serge Belongie. Beyond pairwise clustering. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), volume 2, pages 838–845. IEEE, 2005

  4. [4]

    Guaranteed non-orthogonal tensor decompo- sition via alternating rank-1 updates

    Animashree Anandkumar, Rong Ge, and Majid Janzamin. Guaranteed non-orthogonal tensor decompo- sition via alternating rank-1 updates. arXiv preprint arXiv:1402.5180, 2014

  5. [5]

    Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.The Annals of Probability, 33(5):1643–1697, 2005

    Jinho Baik, Gérard Ben Arous, and Sandrine Péché. Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.The Annals of Probability, 33(5):1643–1697, 2005

  6. [6]

    Approximation algorithms for multi- dimensional assignment problems with decomposable costs.Discrete Applied Mathematics, 49(1-3):25–50, 1994

    Hans-Jürgen Bandelt, Yves Crama, and Frits CR Spieksma. Approximation algorithms for multi- dimensional assignment problems with decomposable costs.Discrete Applied Mathematics, 49(1-3):25–50, 1994

  7. [7]

    Information-theoretic thresholds for community detection in sparse networks

    Jess Banks, Cristopher Moore, Joe Neeman, and Praneeth Netrapalli. Information-theoretic thresholds for community detection in sparse networks. InConference on Learning Theory, pages 383–416. PMLR, 2016

  8. [8]

    Behne and Galen Reeves

    Joshua K. Behne and Galen Reeves. Fundamental limits for rank-one matrix estimation with groupwise heteroskedasticity. InProceedings of The 25th International Conference on Artificial Intelligence and Statistics, pages 8650–8672. PMLR, May 2022

Show all 67 references
  1. [9]

    Algorithmic thresholds for tensor pca.The Annals of Probability, 48(4):2052–2087, 2020

    Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath. Algorithmic thresholds for tensor pca.The Annals of Probability, 48(4):2052–2087, 2020

  2. [10]

    The landscape of the spiked tensor model

    Gérard Ben Arous, Song Mei, Andrea Montanari, and Mihai Nica. The landscape of the spiked tensor model. Communications on Pure and Applied Mathematics, 72(11):2282–2330, 2019

  3. [11]

    The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices.Advances in Mathematics, 227(1):494–521, 2011

    Florent Benaych-Georges and Raj Rao Nadakuditi. The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices.Advances in Mathematics, 227(1):494–521, 2011

  4. [12]

    Giulio Biroli, Chiara Cammarota, and Federico Ricci-Tersenghi. How to iron out rough landscapes and get optimal performances: averaged gradient descent and its application to tensor pca.Journal of Physics A: Mathematical and Theoretical, 53(17):174003, 2020. 11

  5. [13]

    Recovering asymmetric communities in the stochastic block model.IEEE Transactions on Network Science and Engineering, 5(3):237–246, 2017

    Francesco Caltagirone, Marc Lelarge, and Léo Miolane. Recovering asymmetric communities in the stochastic block model.IEEE Transactions on Network Science and Engineering, 5(3):237–246, 2017

  6. [14]

    Differentiability and overlap concentration in optimal bayesian inference

    Hong-Bin Chen and Victor Issa. Differentiability and overlap concentration in optimal bayesian inference. arXiv preprint arXiv:2501.08786, 2025

  7. [15]

    Asymptotics of smoothed wasserstein distances.Potential Analysis, pages 1–25, 2021

    Hong-Bin Chen and Jonathan Niles-Weed. Asymptotics of smoothed wasserstein distances.Potential Analysis, pages 1–25, 2021

  8. [16]

    Statistical inference of finite-rank tensors

    Hongbin Chen, Jean-Christophe Mourrat, and Jiaming Xia. Statistical inference of finite-rank tensors. Annales Henri Lebesgue, 5:1161–1189, 2022

  9. [17]

    Asymptotic mutual information for the balanced binary stochastic block model.Information and Inference, 6(2):125–170, June 2017

    Yash Deshpande, Emmanuel Abbe, and Andrea Montanari. Asymptotic mutual information for the balanced binary stochastic block model.Information and Inference, 6(2):125–170, June 2017

  10. [18]

    Information-theoretically optimal sparse PCA

    Yash Deshpande and Andrea Montanari. Information-theoretically optimal sparse PCA. In2014 IEEE International Symposium on Information Theory, pages 2197–2201, Honolulu, HI, USA, June 2014. IEEE

  11. [19]

    Contextual Stochastic Block Models

    Yash Deshpande, Subhabrata Sen, Andrea Montanari, and Elchanan Mossel. Contextual Stochastic Block Models. InAdvances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018

  12. [20]

    Mutual information for symmetric rank-one matrix estimation: A proof of the replica formula.Advances in Neural Information Processing Systems, 29, 2016

    Mohamad Dia, Nicolas Macris, Florent Krzakala, Thibault Lesieur, Lenka Zdeborová, et al. Mutual information for symmetric rank-one matrix estimation: A proof of the replica formula.Advances in Neural Information Processing Systems, 29, 2016

  13. [21]

    On the rate of convergence in wasserstein distance of the empirical measure

    Nicolas Fournier and Arnaud Guillin. On the rate of convergence in wasserstein distance of the empirical measure. Probability theory and related fields, 162(3):707–738, 2015

  14. [22]

    An entropic interpolation proof of the hwi inequality.Stochastic Processes and their Applications, 130(2):907–923, 2020

    Ivan Gentil, Christian Léonard, Luigia Ripani, and Luca Tamanini. An entropic interpolation proof of the hwi inequality.Stochastic Processes and their Applications, 130(2):907–923, 2020

  15. [23]

    Estimating rank-one matrices with mismatched prior and noise: Universality and large deviations.Communications in Mathematical Physics, 406(1):9, 2024

    Alice Guionnet, Justin Ko, Florent Krzakala, and Lenka Zdeborová. Estimating rank-one matrices with mismatched prior and noise: Universality and large deviations.Communications in Mathematical Physics, 406(1):9, 2024

  16. [24]

    Low-rank matrix estimation with inhomogeneous noise

    Alice Guionnet, Justin Ko, Florent Krzakala, and Lenka Zdeborová. Low-rank matrix estimation with inhomogeneous noise. Information and Inference: A Journal of the IMA, 14(2):iaaf010, 2025

  17. [25]

    Estimation in Gaussian Noise: Properties of the Minimum Mean-Square Error.IEEE Transactions on Information Theory, 57(4):2371–2385, April 2011

    Dongning Guo, Yihong Wu, Shlomo Shamai, and Sergio Verdú. Estimation in Gaussian Noise: Properties of the Minimum Mean-Square Error.IEEE Transactions on Information Theory, 57(4):2371–2385, April 2011

  18. [26]

    Universality of regularized regression estimators in high dimensions.The Annals of Statistics, 51(4):1799–1823, 2023

    Qiyang Han and Yandi Shen. Universality of regularized regression estimators in high dimensions.The Annals of Statistics, 51(4):1799–1823, 2023

  19. [27]

    An optimal statistical and computational framework for generalized tensor estimation.The Annals of Statistics, 50(1):1–29, 2022

    Rungang Han, Rebecca Willett, and Anru R Zhang. An optimal statistical and computational framework for generalized tensor estimation.The Annals of Statistics, 50(1):1–29, 2022

  20. [28]

    Hopkins, Pravesh K

    Samuel B. Hopkins, Pravesh K. Kothari, Aaron Potechin, Prasad Raghavendra, Tselil Schramm, and David Steurer. The power of sum-of-squares for detecting hidden structures. In2017 IEEE 58th Annual Symposium on Foundations of Computer Science (FOCS), pages 720–731, 2017

  21. [29]

    Fast spectral algorithms from sum-of-squares proofs: tensor decomposition and planted sparse vectors

    Samuel B Hopkins, Tselil Schramm, Jonathan Shi, and David Steurer. Fast spectral algorithms from sum-of-squares proofs: tensor decomposition and planted sparse vectors. InProceedings of the forty-eighth annual ACM symposium on Theory of Computing, pages 178–191, 2016

  22. [30]

    Tensor principal component analysis via sum-of- square proofs

    Samuel B Hopkins, Jonathan Shi, and David Steurer. Tensor principal component analysis via sum-of- square proofs. InConference on Learning Theory, pages 956–1006. PMLR, 2015. 12

  23. [31]

    Hopkins and David Steurer

    Samuel B. Hopkins and David Steurer. Efficient bayesian estimation from few samples: Community detection and related problems. In2017 IEEE 58th Annual Symposium on Foundations of Computer Science (FOCS), pages 379–390, 2017

  24. [32]

    Sharp asymptotics of kernel ridge regression beyond the linear regime.arXiv preprint arXiv:2205.06798, 2022

    Hong Hu and Yue M Lu. Sharp asymptotics of kernel ridge regression beyond the linear regime.arXiv preprint arXiv:2205.06798, 2022

  25. [33]

    Power iteration for tensor pca.Journal of Machine Learning Research, 23(128):1–47, 2022

    Jiaoyang Huang, Daniel Z Huang, Qing Yang, and Guang Cheng. Power iteration for tensor pca.Journal of Machine Learning Research, 23(128):1–47, 2022

  26. [34]

    The nishimori line and bayesian statistics.Journal of Physics A: Mathematical and General, 32(21):3875, 1999

    Yukito Iba. The nishimori line and bayesian statistics.Journal of Physics A: Mathematical and General, 32(21):3875, 1999

  27. [35]

    Extending mercer’s expansion to indefinite and asymmetric kernels

    Sungwoo Jeong and Alex Townsend. Extending mercer’s expansion to indefinite and asymmetric kernels. arXiv preprint arXiv:2409.16453, 2024

  28. [36]

    On the distribution of the largest eigenvalue in principal components analysis.The Annals of statistics, 29(2):295–327, 2001

    Iain M Johnstone. On the distribution of the largest eigenvalue in principal components analysis.The Annals of statistics, 29(2):295–327, 2001

  29. [37]

    Statistical mechanics of low-rank tensor decomposition.Advances in Neural Information Processing Systems, 31, 2018

    Jonathan Kadmon and Surya Ganguli. Statistical mechanics of low-rank tensor decomposition.Advances in Neural Information Processing Systems, 31, 2018

  30. [38]

    Koopmans and Martin Beckmann

    Tjalling C. Koopmans and Martin Beckmann. Assignment problems and the location of economic activities. Econometrica, 25(1):53–76, 1957

  31. [39]

    Applications of the lindeberg principle in communications and statistical learning.IEEE transactions on information theory, 57(4):2440–2450, 2011

    Satish Babu Korada and Andrea Montanari. Applications of the lindeberg principle in communications and statistical learning.IEEE transactions on information theory, 57(4):2440–2450, 2011

  32. [40]

    Wein, and Afonso S

    Dmitriy Kunisky, Alexander S. Wein, and Afonso S. Bandeira. Notes on computational hardness of hypothesis testing: Predictions using the low-degree likelihood ratio. In Paula Cerejeiras and Michael Reissig, editors,Mathematical Analysis, its Applications and Computation, pages...

  33. [41]

    The quadratic assignment problem.Management science, 9(4):586–599, 1963

    Eugene L Lawler. The quadratic assignment problem.Management science, 9(4):586–599, 1963

  34. [42]

    Fundamental limits of symmetric low-rank matrix estimation.Probability Theory and Related Fields, 173(3):859–929, 2019

    Marc Lelarge and Léo Miolane. Fundamental limits of symmetric low-rank matrix estimation.Probability Theory and Related Fields, 173(3):859–929, 2019

  35. [43]

    Mmse of probabilistic low-rank matrix estimation: Universality with respect to the output channel

    Thibault Lesieur, Florent Krzakala, and Lenka Zdeborová. Mmse of probabilistic low-rank matrix estimation: Universality with respect to the output channel. In2015 53rd Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 680–687, 2015

  36. [44]

    Statistical and computational phase transitions in spiked tensor estimation

    Thibault Lesieur, Léo Miolane, Marc Lelarge, Florent Krzakala, and Lenka Zdeborová. Statistical and computational phase transitions in spiked tensor estimation. In2017 ieee international symposium on information theory (isit), pages 511–515. IEEE, 2017

  37. [45]

    Large networks and graph limits, volume 60

    László Lovász. Large networks and graph limits, volume 60. American Mathematical Soc., 2012

  38. [46]

    An equivalence principle for the spectrum of random inner-product kernel matrices with polynomial scalings.arXiv preprint arXiv:2205.06308, 2022

    Yue M Lu and Horng-Tzer Yau. An equivalence principle for the spectrum of random inner-product kernel matrices with polynomial scalings.arXiv preprint arXiv:2205.06308, 2022

  39. [47]

    Yuetian Luo, Garvesh Raskutti, Ming Yuan, and Anru R. Zhang. A sharp blockwise tensor perturbation bound for orthogonal iteration.Journal of Machine Learning Research, 22(179):1–48, 2021

  40. [48]

    Mutual Information in Community Detection with Covariate Information and Correlated Networks

    Vaishakhi Mayya and Galen Reeves. Mutual Information in Community Detection with Covariate Information and Correlated Networks. In2019 57th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 602–607, Monticello, IL, USA, September 2019. IEEE

  41. [49]

    Fundamental limits of non-linear low-rank matrix estimation.arXiv preprint arXiv:2403.04234, 2024

    Pierre Mergny, Justin Ko, Florent Krzakala, and Lenka Zdeborová. Fundamental limits of non-linear low-rank matrix estimation.arXiv preprint arXiv:2403.04234, 2024

  42. [50]

    Oxford University Press, 2009

    Marc Mezard and Andrea Montanari.Information, physics, and computation. Oxford University Press, 2009. 13

  43. [51]

    Fundamental limits of low-rank matrix estimation: the non-symmetric case.arXiv preprint arXiv:1702.00473, 2017

    Léo Miolane. Fundamental limits of low-rank matrix estimation: the non-symmetric case.arXiv preprint arXiv:1702.00473, 2017

  44. [52]

    Andrea Montanari, Daniel Reichman, and Ofer Zeitouni. On the limitation of spectral methods: From the gaussian hidden clique problem to rank-one perturbations of gaussian tensors.Advances in Neural Information Processing Systems, 28, 2015

  45. [53]

    A statistical model for tensor pca.Advances in neural information processing systems, 27, 2014

    Andrea Montanari and Emile Richard. A statistical model for tensor pca.Advances in neural information processing systems, 27, 2014

  46. [54]

    Statistical physics of spin glasses and information processing: an introduction

    Hidetoshi Nishimori. Statistical physics of spin glasses and information processing: an introduction. Clarendon Press, 2001

  47. [55]

    Generalization of an inequality by talagrand and links with the logarithmic sobolev inequality.Journal of Functional Analysis, 173(2):361–400, 2000

    Felix Otto and Cédric Villani. Generalization of an inequality by talagrand and links with the logarithmic sobolev inequality.Journal of Functional Analysis, 173(2):361–400, 2000

  48. [56]

    Wein, and Afonso S

    Amelia Perry, Alexander S. Wein, and Afonso S. Bandeira. Statistical limits of spiked tensor models. Annales de l’Institut Henri Poincaré, Probabilités et Statistiques, 56(1):230–264, 2 2020

  49. [57]

    Are gaussian data all you need? the extents and limits of universality in high-dimensional generalized linear estimation

    Luca Pesce, Florent Krzakala, Bruno Loureiro, and Ludovic Stephan. Are gaussian data all you need? the extents and limits of universality in high-dimensional generalized linear estimation. InInternational Conference on Machine Learning, pages 27680–27708. PMLR, 2023

  50. [58]

    The multidimensional assignment problem.Operations Research, 16(2):422–431, 1968

    William P Pierskalla. The multidimensional assignment problem.Operations Research, 16(2):422–431, 1968

  51. [59]

    Information-Theoretic Limits for the Matrix Tensor Product.IEEE Journal on Selected Areas in Information Theory, 1(3):777–798, November 2020

    Galen Reeves. Information-Theoretic Limits for the Matrix Tensor Product.IEEE Journal on Selected Areas in Information Theory, 1(3):777–798, November 2020

  52. [60]

    The Geometry of Community Detection via the MMSE Matrix

    Galen Reeves, Vaishakhi Mayya, and Alexander Volfovsky. The Geometry of Community Detection via the MMSE Matrix. In2019 IEEE International Symposium on Information Theory (ISIT), pages 400–404, Paris, France, July 2019. IEEE

  53. [61]

    Pfister, and Alex Dytso

    Galen Reeves, Henry D. Pfister, and Alex Dytso. Mutual Information as a Function of Matrix SNR for Linear Gaussian Channels. In2018 IEEE International Symposium on Information Theory (ISIT), pages 1754–1758, Vail, CO, USA, June 2018. IEEE

  54. [62]

    An empirical bayes approach to statistics

    Herbert E Robbins. An empirical bayes approach to statistics. InBreakthroughs in Statistics: Foundations and basic theory, pages 388–394. Springer, 1992

  55. [63]

    Optimal transport for applied mathematicians.Birkäuser, NY, 55(58-63):94, 2015

    Filippo Santambrogio. Optimal transport for applied mathematicians.Birkäuser, NY, 55(58-63):94, 2015

  56. [64]

    Solvable model of a spin-glass

    David Sherrington and Scott Kirkpatrick. Solvable model of a spin-glass. Physical review letters, 35(26):1792, 1975

  57. [65]

    Sharp analysis of power iteration for tensor pca.Journal of Machine Learning Research, 25(195):1–42, 2024

    Yuchen Wu and Kangjie Zhou. Sharp analysis of power iteration for tensor pca.Journal of Machine Learning Research, 25(195):1–42, 2024

  58. [66]

    Asymptotic mutual information in quadratic estimation problems over compact groups.arXiv preprint arXiv:2404.10169, 2024

    Kaylee Y Yang, Timothy LH Wee, and Zhou Fan. Asymptotic mutual information in quadratic estimation problems over compact groups.arXiv preprint arXiv:2404.10169, 2024

  59. [67]

    Learning with hypergraphs: Clustering, classification, and embedding.Advances in neural information processing systems, 19, 2006

    Dengyong Zhou, Jiayuan Huang, and Bernhard Schölkopf. Learning with hypergraphs: Clustering, classification, and embedding.Advances in neural information processing systems, 19, 2006. 14 A Proof of main results A.1 Proof of Theorem 1 Before presenting the proof, we recall the ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.