Pith. sign in

REVIEW 3 major objections 5 minor 56 references

Variational Bounds for Perceptron Learning from Structured Data

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read For perceptrons trained on Gaussian mixtures, the limiting free energy is trapped between two minimax variational bounds that differ only in the order of two scalar optimizations.

desk verdict Genuine progress: rigorous variational bounds for a broad perceptron class, but exactness is conditional on an unproven sup-inf exchange and a sum-rule typo should be fixed. read the letter →

arxiv 2608.04882 v1 pith:XJ7YOGSW submitted 2026-08-05 cs.LG cond-mat.dis-nnmath-phmath.MPstat.ML

classification cs.LGcond-mat.dis-nnmath-phmath.MPstat.ML MSC 82B4482C3268T05
keywords continuous-spinperceptronGaussianmixturedataquenchedpressurevariationalboundsadaptiveinterpolationlog-concavitygeneralizationerrorfixed-pointequations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish a variational characterization of the thermodynamic limit of a finite-temperature continuous-spin perceptron trained on a Gaussian mixture. The model covers a broad class of concave utility functions (losses) and log-concave separable priors on the weights. The authors prove lower and upper bounds on the quenched pressure; the bounds coincide whenever two scalar optimizations can be exchanged, in which case the pressure, ground-state energy, training loss, and generalization error all follow from one potential. A sympathetic reader should care because this gives a rigorous route to perceptron observables outside the Bayes-optimal setting, where the standard fixed-point identities are derived as stationarity conditions of a single variational function rather than assumed as a self-consistency system.

What carries the argument

The load-bearing object is the variational potential $\Phi_u(\rho,q,m,r,\delta,h)$, built from two scalar channel functions: a pattern channel $\Psi^\beta_u(\rho,q,m)$ and a spin channel $\Lambda_\phi(r,\delta,h)$, glued by quadratic and centroid-aligning terms. The central identity is the sum rule of Proposition 3, which expresses the interpolating pressure as an endpoint potential plus an integrated remainder. The argument is carried by adaptive interpolation: an interpolating Hamiltonian connects the true model at $t=0$ to a gas of independent scalar channels at $t=1$, and differentiating along the path yields the sum rule whose remainder vanishes by concentration estimates. Log-concavity, via Pr\'ekopa-Leindler and Brascamp-Lieb inequalities, supplies the convexity needed to apply Jensen and minimax arguments and, crucially, the sign inequality $\langle S_{12}\rangle_t\ge\langle S_{11}\rangle_t$ that keeps the lower-bound interpolation path inside the physical region.

What would settle it

Pick any admissible concave utility and log-concave prior, and compute the reduced potential $\Phi^\star_u(\rho,r)$ on a fine grid; if the surface has no saddle point in $(\rho,r)$, so that $\sup_\rho\inf_r$ differs from $\inf_r\sup_\rho$, then the two bounds cannot match and the exact variational formula in that regime would be false. A concrete candidate is logistic loss with smooth $L^1$ regularization at strong regularization strength, outside the parameter regimes plotted in the paper.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is Theorem 1: for any inverse temperature, sampling ratio, and signal strength satisfying hypotheses H1-H4, the liminf of the quenched pressure is at least $\sup_{\rho\ge 0}\inf_{r\ge 0}\Phi^\star_u(\rho,r)$, and the limsup is at most $\inf_{r\ge 0}\sup_{\rho\ge 0}\Phi^\star_u(\rho,r)$, where $\Phi^\star_u$ is the reduced variational potential obtained by optimizing the full potential over overlaps and conjugate fields. The two bounds differ only in the order of the maximum over the self-overlap and the minimum over the channel parameter; all remaining extrema are exchanged through convexity in $(\delta,h)$ and concavity in $(q,m)$. Whenever these two outer optimizations commute, the thermodynamic limit exists and equals the common variational value, and the fixed-point equations of the cavity computation are recovered as stationarity conditions of the same potential. The paper also derives the zero-temperature ground-state energy and, under local matching conditions, explicit formulas for training loss and generalization error, including the standard $Q(\lambda\tilde{m}/\sqrt{\tilde{\rho}})$ expression at zero temperature.

Load-bearing premise

The lower-bound construction needs the inequality $\langle S_{12}\rangle_t\ge\langle S_{11}\rangle_t$ to hold along the whole interpolation, a log-concavity consequence requiring the utility to be concave and the prior convex; and the exact formula additionally requires the supremum over $\rho$ and the infimum over $r$ to be exchangeable, which the paper does not prove in general.

Editorial extensions

If this is right

  • If the two outer optimizations commute, the thermodynamic limit of the quenched pressure exists and equals one scalar variational value, and this does not require uniqueness of the full stationary-point system.
  • The same potential yields the cavity fixed-point equations as stationarity conditions, so the usual obstruction to an exact formula moves from uniqueness of a fixed-point system to exchangeability of two explicitly identified optimizations.
  • At zero temperature the bounds reorganize into a ground-state variational principle, and the generalization error takes the form $Q(\lambda\tilde{m}/\sqrt{\tilde{\rho}})$, matching the known Gaussian-mixture classification result while allowing general log-concave priors.
  • For perturbed utilities, training loss and generalization error are obtained by differentiating the pressure, giving explicit scalar formulas in terms of the variational order parameters under local matching assumptions.
  • Numerical evaluation of the reduced potential for logistic loss with $L^2$ and smooth $L^1$ regularization shows a saddle structure in $(\rho,r)$ across several $\alpha,\beta,\kappa$, supporting commutativity in these regimes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If a general commutativity criterion were found, the same potential would unify the free energy, ground-state energy, training loss, and generalization error for the whole H1-H4 class, making the fixed-point equations a derived rather than assumed object.
  • A concrete computational probe suggested by the variational structure is to evaluate the gap $\sup_\rho\inf_r\Phi^\star_u(\rho,r)-\inf_r\sup_\rho\Phi^\star_u(\rho,r)$ for admissible losses outside the plotted set, such as quadratic or smoothed-hinge utilities; a nonzero gap would mark regimes where the exact formula fails while the bounds remain valid.
  • The reliance on log-concavity indicates the method likely extends to other log-concave prior families, but not to non-log-concave priors without a new sign mechanism to replace the inequality $\langle S_{12}\rangle_t\ge\langle S_{11}\rangle_t$.
  • The authors' multilayer discussion suggests low-rank multilayer perceptrons as a testbed where a finite set of overlaps might close the variational description, though this would require the convex structure to survive composition of layers.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies the finite-temperature quenched pressure of a continuous-spin perceptron trained on Gaussian-mixture data with labels, under a concave utility and a separable log-concave prior spin measure. The main result, Theorem 1, gives a lower bound liminf_N p_N >= sup_rho inf_r Phi*_u(rho,r) and an upper bound limsup_N p_N <= inf_r sup_rho Phi*_u(rho,r), where Phi*_u is a scalar variational potential obtained by optimizing four auxiliary parameters. The two bounds coincide whenever the outer sup/inf over (rho,r) commute; the paper then identifies the thermodynamic limit and, by differentiating the variational formula, derives the training loss (Corollary 1), generalization error (Corollary 2), and zero-temperature ground-state energy (Proposition 2). The proofs use adaptive interpolation with ODE-controlled paths, log-concavity (Prekopa-Leindler, Brascamp-Lieb), concentration estimates, and Sion's minimax theorem for the inner optimizations. Numerical saddle plots are provided for logistic and smooth-L1 losses.

Significance. If the bounds are correct, the paper supplies a unified finite-temperature variational formulation for a non-Bayes-optimal perceptron with general log-concave priors, going beyond the usual Gaussian-prior/cavity setting. The proof is substantial and largely self-contained: the interpolation sum rule, the convexity analysis, the ODE existence criterion, and the concentration appendices are detailed. The authors are also honest: the exact-solution claim is explicitly conditioned on the unproven exchange of two scalar optimizations, and the observable formulas carry further local-matching assumptions. The main value is therefore the unconditional two-sided bounds plus a plausible variational route to exactness; unconditional exactness itself remains open.

major comments (3)
  1. [Section 2, Theorem 1 and following paragraph] The exact-solution claim, and consequently Corollaries 1 and 2, the matching version of Proposition 2, and Remark 3, rest on the exchange sup_{rho>=0} inf_{r>=0} Phi*_u(rho,r) = inf_{r>=0} sup_{rho>=0} Phi*_u(rho,r). This equality is not proven, and it does not follow from the Sion-minimax steps in Section 3.3, which only exchange the inner (q,m) and (delta,h) optimizations. The paper explicitly acknowledges the gap, but because all formulas for the training loss, generalization error, and ground-state energy assume the exchange, this is the load-bearing point of the paper rather than a cosmetic limitation. The authors should either prove the exchange in a nontrivial regime (for example lambda=0 random labels, Gaussian prior, small alpha, or high temperature) or reformulate the abstract and Section 2 so that the 'solution of the model' is presented as a conjecture conditional on an explicit, unverified assumption.
  2. [Corollaries 1 and 2, Section 2] The matching assumption in Theorem 1 is used at a single parameter point for u, but the training-loss formula (2.21) requires the two bounds to match for u_gamma = gamma u on an interval gamma in (1-epsilon,1+epsilon), and the generalization-error formula (2.23) requires matching for all gamma in [0,eta) and t in [-eta,eta] of the perturbed utility gamma(u_y + t W_y). These are strictly stronger hypotheses than the matching checked for the unperturbed u, and the numerical section provides no check of them. Since (2.21) and (2.23) are presented as the paper's learning-theoretic output, the authors should either verify these local matching conditions in at least the numerical settings considered in Section 4, or state explicitly that these predictions are conjectural.
  3. [Section 4, Figures 1 and 2] The evidence for the required exchange is visual inspection of a saddle surface on a 50x50 grid. A saddle in the (rho,r) plane does not by itself imply sup_rho inf_r F = inf_r sup_rho F for a function with multiple stationary points, and no quantitative value of the difference between the two orders is reported. The sentence 'Consequently, the lower and upper bounds in Theorem 1 coincide in these regimes' is therefore stronger than the data support. Please report the numerical gap, and if it is nonzero at grid resolution, rephrase the conclusion as evidence for a conjecture.
minor comments (5)
  1. [Equation (3.15) versus (3.49)] The apparent factor-of-2 discrepancy between the sum rule (3.15) and the upper-bound expression (3.49) is resolved if (3.15) is read as beta/2 integral [delta_dot rho_dot - beta r_dot(rho_dot - q_dot)] dt; this reading is forced by the lower-bound chain (3.54). Please typeset the fraction unambiguously so that it is clear the 1/2 multiplies the whole bracket.
  2. [Section 2 and appendices] There are several typos and spelling inconsistencies: 'ceneterd' should be 'centered', 'nor impossible' should be 'not impossible', 'Prekopa-Leindler' should be 'Prekopa-Leindler', 'Gronwall' should be 'Gronwall', and 'Cauchy-Schwartz' should be 'Cauchy-Schwarz'.
  3. [Section 3.3, Eq. (3.50) and Eq. (3.57)] The set T_r is first introduced in (3.50) as 'a compact convex set to be chosen later', but the actual choice [0,alpha C] x [-sqrt(alpha r), sqrt(alpha r)] appears only after Eq. (3.57). Please define T_r explicitly before its first use.
  4. [Section 4, first paragraph] The claim that logistic regression with L2 regularization exhibits a learning transition at alpha=2 is stated without a citation at that point; please add a reference there.
  5. [Figures 1 and 2 captions] The color-scale ranges differ across panels, which makes visual comparison of the saddle structure difficult; please use a consistent scale or explain why the ranges differ.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the variational bounds are proved by adaptive interpolation from the model's own Hamiltonian, with the only gap being an explicitly stated, numerically checked exchange of optimizations rather than a concealed input.

full rationale

The paper's central result, Theorem 1, is a pair of minimax bounds obtained by an interpolation argument (Proposition 3 sum rule, upper/lower ODE choices, Jensen and Sion applications). The variational potential Phi is defined from the same utilities u_y and priors phi that appear in the Hamiltonian, but the bounds are derived, not assumed: the quenched pressure is expressed as an endpoint of an interpolating path plus a vanishing remainder, and the potential emerges from the t=1 endpoint and Jensen-type inequalities. No fitted parameter is renamed as a prediction; the stationarity equations are computed from Phi, and the zero-temperature generalization formula is checked against the independent Gordon-minimax result [16]. The only unproven ingredient is the exchange of sup_rho and inf_r (or its counterpart in the zero-temperature limit), which is explicitly flagged in the text ('Unfortunately, we were not able to prove the matching under sufficiently general criteria') and supported only by numerical saddle inspection. This is an honest limitation, not a circular step. Self-citations ([38,39] for adaptive interpolation outside Bayes-optimal setting, [49] for a standard derivative-convergence theorem in Corollary 1) are contextual or refer to established external results and do not carry the derivation of Theorem 1. Therefore the derivation chain is self-contained, and no claim reduces by construction to its own inputs.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The model has no fitted constants: α, β, λ, κ are inputs, and ρ, q, m, r, δ, h are optimization variables of a variational principle. The key non-standard premises are the log-concavity hypotheses and the unproven commutation of the outer optimizations.

assumptions (5)
  • domain assumption H1-H4: u_y concave, non-affine with growth bounds; φ convex, L-Lipschitz; bounded second derivatives; label-difference growth.
    Stated in Section 2; all convexity and concentration results rely on these hypotheses.
  • domain assumption P_θ satisfies a Poincaré inequality with constant c_θ and has unit second moment.
    Used for concentration of θ-dependent quantities in Appendix D.
  • domain assumption Proportional limit M/N → α with α > 0 fixed.
    The scaling regime of the model; all asymptotic statements are in this limit.
  • ad hoc to paper Commutation of sup_ρ and inf_r in Theorem 1 holds.
    Required for the thermodynamic limit to exist and for Corollaries 1 and 2; not proven, supported only numerically.
  • ad hoc to paper Local matching of the variational bounds for u_γ = γu and for perturbed utilities in Corollaries 1 and 2.
    The derivative identities for training and generalization errors require the variational formula to hold in a neighborhood of u.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Variational Bounds for Perceptron Learning from Structured Data." pith.science (2026). https://pith.science/paper/XJ7YOGSW

@misc{pith2026260804882,
  author       = {Pith},
  title        = {Pith review of: Variational Bounds for Perceptron Learning from Structured Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XJ7YOGSW}},
  note         = {Machine review of arXiv:2608.04882}
}
read the original abstract

We introduce a variational approach to a finite-temperature continuous-spin perceptron trained on a Gaussian mixture. The model allows for a broad class of concave utilities and log-concave separable prior measures on the spins. By combining the interpolation method with log-concavity and concentration estimates, we derive lower and upper minimax variational bounds for the limiting quenched pressure. Remarkably, the two bounds differ only in the order of optimization of two variational parameters, while all remaining extrema are controlled by the concave--convex structure of the variational potential. Whenever the two optimizations commute, the two bounds match and identify the solution of the model. The same potential yields the fixed-point equations as stationarity conditions and provides a unified route to the computation of the ground-state energy, training loss, and generalization error.

Figures

Figures reproduced from arXiv: 2608.04882 by the authors.

Figure 1
Figure 1. Numerical surfaces of the reduced variational potential Φ [PITH_FULL_IMAGE:figures/full_fig_p026_1.png] view at source ↗
Figure 2
Figure 2. Numerical surfaces of the reduced variational potential Φ [PITH_FULL_IMAGE:figures/full_fig_p027_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 34 canonical work pages

  1. [1]

    The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain

    Frank Rosenblatt. “The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain”. In:Psychological Review65.6 (1958), pp. 386–408. doi:10.1037/h0042519

  2. [2]

    Geometrical and Statistical Properties of Systems of Linear Inequalities with Applications in Pattern Recognition

    Thomas M. Cover. “Geometrical and Statistical Properties of Systems of Linear Inequalities with Applications in Pattern Recognition”. In:IEEE Transactions on Electronic ComputersEC-14.3 (1965), pp. 326–334.doi:10 . 1109 / PGEC . 1965 . 264137

  3. [3]

    The Space of Interactions in Neural Network Models

    Elizabeth Gardner. “The Space of Interactions in Neural Network Models”. In: Journal of Physics A: Mathematical and General21.1 (1988), pp. 257–270.doi: 10.1088/0305-4470/21/1/030

  4. [4]

    Optimal Storage Properties of Neural Network Models

    Elizabeth Gardner and Bernard Derrida. “Optimal Storage Properties of Neural Network Models”. In:Journal of Physics A: Mathematical and General21.1 (1988), pp. 271–284.doi:10.1088/0305-4470/21/1/031

  5. [5]

    Statistical Mechanics of Learning from Examples

    H. S. Seung, H. Sompolinsky, and N. Tishby. “Statistical Mechanics of Learning from Examples”. In:Physical Review A45.8 (1992), pp. 6056–6091.doi:10.1103/ PhysRevA.45.6056

  6. [6]

    The Statistical Mechanics of Learning a Rule

    Timothy L. H. Watkin, Albrecht Rau, and Michael Biehl. “The Statistical Mechanics of Learning a Rule”. In:Reviews of Modern Physics65.2 (1993), pp. 499–556.doi: 10.1103/RevModPhys.65.499

  7. [7]

    Cambridge University Press, 2001.doi:10.1017/CBO9781139164542

    Andreas Engel and Christian Van den Broeck.Statistical Mechanics of Learning. Cambridge University Press, 2001.doi:10.1017/CBO9781139164542

  8. [8]

    A Modern Maximum-Likelihood Theory for High-Dimensional Logistic Regression

    Pragya Sur and Emmanuel J. Cand` es. “A Modern Maximum-Likelihood Theory for High-Dimensional Logistic Regression”. In:Proceedings of the National Academy of Sciences116.29 (2019), pp. 14516–14525.doi:10.1073/pnas.1810420116

Show all 56 references
  1. [9]

    Precise Error Analy- sis of RegularizedM-Estimators in High Dimensions

    Christos Thrampoulidis, Ehsan Abbasi, and Babak Hassibi. “Precise Error Analy- sis of RegularizedM-Estimators in High Dimensions”. In:IEEE Transactions on Information Theory64.8 (2018), pp. 5592–5628.doi:10.1109/TIT.2018.2840720

  2. [10]

    Fundamental Lim- its of Ridge-Regularized Empirical Risk Minimization in High Dimensions

    Hossein Taheri, Ramtin Pedarsani, and Christos Thrampoulidis. “Fundamental Lim- its of Ridge-Regularized Empirical Risk Minimization in High Dimensions”. In:Pro- ceedings of the 24th International Conference on Artificial Intelligence and Statis- tics. Vol. 130. Proceedings of...

  3. [11]

    Generaliza- tion Error in High-Dimensional Perceptrons: Approaching Bayes Error with Convex Optimization

    Benjamin Aubin, Florent Krzakala, Yue M. Lu, and Lenka Zdeborov´ a. “Generaliza- tion Error in High-Dimensional Perceptrons: Approaching Bayes Error with Convex Optimization”. In:Advances in Neural Information Processing Systems. Vol. 33. Curran Associates, Inc., 2020.url:http...

  4. [12]

    Optimal errors and phase transitions in high-dimensional generalized linear mod- els

    Jean Barbier, Florent Krzakala, Nicolas Macris, L´ eo Miolane, and Lenka Zdeborov´ a. “Optimal errors and phase transitions in high-dimensional generalized linear mod- els”. In:Proceedings of the National Academy of Sciences116.12 (2019), pp. 5451– 5460.doi:10.1073/pnas.1802705116

  5. [13]

    Asymptotic Errors for Teacher– Student Convex Generalized Linear Models (or: How to Prove Kabashima’s Replica Formula)

    C´ edric Gerbelot, Alia Abbara, and Florent Krzakala. “Asymptotic Errors for Teacher– Student Convex Generalized Linear Models (or: How to Prove Kabashima’s Replica Formula)”. In:IEEE Transactions on Information Theory69.3 (2023), pp. 1824– 1852.doi:10.1109/TIT.2022.3222913

  6. [14]

    Xiaoyi Mai and Zhenyu Liao.High Dimensional Classification via Regularized and Unregularized Empirical Risk Minimization: Precise Error and Optimal Loss. 2020. arXiv:1905.13742 [stat.ML]

  7. [15]

    Sharp Guarantees and Optimal Performance for Inference in Binary and Gaussian-Mixture Models

    Hossein Taheri, Ramtin Pedarsani, and Christos Thrampoulidis. “Sharp Guarantees and Optimal Performance for Inference in Binary and Gaussian-Mixture Models”. In:Entropy23.2 (2021), p. 178.doi:10.3390/e23020178

  8. [16]

    The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture

    Francesca Mignacco, Florent Krzakala, Yue Lu, Pierfrancesco Urbani, and Lenka Zdeborova. “The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture”. In:Proceedings of the 37th International Conference on Ma- chine Learning. Ed. by Hal Daum´ e III...

  9. [17]

    Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-Dimensions

    Bruno Loureiro, Gabriele Sicuro, C´ edric Gerbelot, Alessandro Pacco, Florent Krza- kala, and Lenka Zdeborov´ a. “Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-Dimensions”. In:Advances in Neural Infor- mation Processing Systems. Ed. by ...

  10. [18]

    Gaussian Universality of Perceptrons with Random Labels

    Federica Gerace, Florent Krzakala, Bruno Loureiro, Ludovic Stephan, and Lenka Zdeborov´ a. “Gaussian Universality of Perceptrons with Random Labels”. In:Phys- ical Review E109.3 (2024), p. 034305.doi:10.1103/PhysRevE.109.034305

  11. [19]

    Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear Estimation

    Luca Pesce, Florent Krzakala, Bruno Loureiro, and Ludovic Stephan. “Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear Estimation”. In:Proceedings of the 40th International Confer- ence on Machine Learning. Vol. 202. P...

  12. [20]

    Universality laws for Gaussian mixtures in generalized linear models

    Yatin Dandi, Ludovic Stephan, Florent Krzakala, Bruno Loureiro, and Lenka Zde- borov´ a. “Universality laws for Gaussian mixtures in generalized linear models”. In: Journal of Statistical Mechanics: Theory and Experiment2024.10 (2024), p. 104015. doi:10.1088/1742-5468/ad65e7. ...

  13. [21]

    The Space of Interactions in Neural Networks: Gardner’s Computa- tion with the Cavity Method

    Marc M´ ezard. “The Space of Interactions in Neural Networks: Gardner’s Computa- tion with the Cavity Method”. In:Journal of Physics A: Mathematical and General 22.12 (1989), pp. 2181–2190.doi:10.1088/0305-4470/22/12/018

  14. [22]

    Intersecting Random Half-Spaces: Toward the Gardner–Derrida Formula

    Michel Talagrand. “Intersecting Random Half-Spaces: Toward the Gardner–Derrida Formula”. In:The Annals of Probability28.2 (2000), pp. 725–758.doi:10.1214/ aop/1019160259

  15. [23]

    On the Gaussian Perceptron at High Temperature

    Michel Talagrand. “On the Gaussian Perceptron at High Temperature”. In:Math- ematical Physics, Analysis and Geometry5.1 (2002), pp. 77–99.doi:10.1023/A: 1015840632110

  16. [24]

    Rigorous Solution of the Gardner Prob- lem

    Mariya Shcherbina and Brunello Tirozzi. “Rigorous Solution of the Gardner Prob- lem”. In:Communications in Mathematical Physics234 (2003), pp. 383–422.doi: 10.1007/s00220- 002- 0783-3.url:https://doi.org/10.1007/s00220-002- 0783-3

  17. [25]

    Gardner Formula for Ising Perceptron Models at Small Densities

    Erwin Bolthausen, Shuta Nakajima, Nike Sun, and Changji Xu. “Gardner Formula for Ising Perceptron Models at Small Densities”. In:Proceedings of the Thirty-Fifth Conference on Learning Theory. Vol. 178. Proceedings of Machine Learning Re- search. PMLR, 2022, pp. 1787–1911

  18. [26]

    Characterizing Finite-Dimensional Posterior Marginals in High-Dimensional GLMs via Leave-One-Out

    Manuel S´ aenz and Pragya Sur. “Characterizing Finite-Dimensional Posterior Marginals in High-Dimensional GLMs via Leave-One-Out”. In: (2025).doi:10.48550/arXiv. 2601.00091. arXiv:2601.00091 [math.ST]

  19. [27]

    Springer, 2010

    Michel Talagrand.Mean Field Models for Spin Glasses: Volume I: Basic Examples. Springer, 2010

  20. [28]

    The Thermodynamic Limit in Mean Field Spin Glass Models

    Francesco Guerra and Fabio Lucio Toninelli. “The Thermodynamic Limit in Mean Field Spin Glass Models”. In:Communications in Mathematical Physics230 (2002)

  21. [29]

    Broken Replica Symmetry Bounds in the Mean Field Spin Glass Model

    Francesco Guerra. “Broken Replica Symmetry Bounds in the Mean Field Spin Glass Model”. In:Communications in Mathematical Physics233 (2003)

  22. [30]

    The adaptive interpolation method: a simple scheme to prove replica formulas in Bayesian inference

    Jean Barbier and Nicolas Macris. “The adaptive interpolation method: a simple scheme to prove replica formulas in Bayesian inference”. In:Probability Theory and Related Fields174 (2019)

  23. [31]

    The adaptive interpolation method for proving replica formulas. Applications to the Curie–Weiss and Wigner spike models

    Jean Barbier and Nicolas Macris. “The adaptive interpolation method for proving replica formulas. Applications to the Curie–Weiss and Wigner spike models”. In: Journal of Physics A: Mathematical and Theoretical52.29 (2019), p. 294002.doi: 10.1088/1751-8121/ab2735

  24. [32]

    Oxford; New York: Oxford University Press, 2001

    Hidetoshi Nishimori.Statistical Physics of Spin Glasses and Information Processing: an Introduction. Oxford; New York: Oxford University Press, 2001

  25. [33]

    Griffiths Inequalities in the Nishimori Line

    Satoshi Morita, Hidetoshi Nishimori, and Pierluigi Contucci. “Griffiths Inequalities in the Nishimori Line”. In:Progress of Theoretical Physics Supplement157 (Jan. 2005), pp. 73–76.issn: 0375-9687.doi:10.1143/PTPS.157.73

  26. [34]

    Surface Terms on the Nishimori Line of the Gaussian Edwards–Anderson Model

    Pierluigi Contucci, Satoshi Morita, and Hidetoshi Nishimori. “Surface Terms on the Nishimori Line of the Gaussian Edwards–Anderson Model”. In:Journal of Statistical Physics122 (2006), pp. 303–312.doi:10.1007/s10955-005-8020-z

  27. [35]

    Strong Replica Symmetry in High-Dimensional Optimal Bayesian Inference

    Jean Barbier and Dmitry Panchenko. “Strong Replica Symmetry in High-Dimensional Optimal Bayesian Inference”. In:Communications in Mathematical Physics393 (2022), pp. 1199–1239.doi:10.1007/s00220-022-04387-w. 49

  28. [36]

    On Extensions of the Brunn–Minkowski and Pr´ ekopa–Leindler Theorems, Including Inequalities for Log Concave Functions, and with an Application to the Diffusion Equation

    Herm Jan Brascamp and Elliott H. Lieb. “On Extensions of the Brunn–Minkowski and Pr´ ekopa–Leindler Theorems, Including Inequalities for Log Concave Functions, and with an Application to the Diffusion Equation”. In:Journal of Functional Anal- ysis22.4 (1976), pp. 366–389.doi:1...

  29. [37]

    Strong replica symmetry for high-dimensional disordered log-concave Gibbs measures

    Jean Barbier, Dmitry Panchenko, and Manuel S´ aenz. “Strong replica symmetry for high-dimensional disordered log-concave Gibbs measures”. In:Information and Inference: A Journal of the IMA11.3 (2022), pp. 1079–1108.doi:10.1093/imaiai/ iaab027. arXiv:2009.12939

  30. [38]

    An inference prob- lem in a mismatched setting: a spin-glass model with Mattis interaction

    Francesco Camilli, Pierluigi Contucci, and Emanuele Mingione. “An inference prob- lem in a mismatched setting: a spin-glass model with Mattis interaction”. In:SciPost Phys.12 (4 2022), p. 125.doi:10.21468/SciPostPhys.12.4.125

  31. [39]

    The Onset of Parisi’s Complexity in a Mismatched Inference Problem

    Francesco Camilli, Pierluigi Contucci, and Emanuele Mingione. “The Onset of Parisi’s Complexity in a Mismatched Inference Problem”. In:Entropy26.1 (2024), p. 42. doi:10.3390/e26010042

  32. [40]

    The multi-species mean-field spin-glass on the Nishimori line

    Diego Alberici, Francesco Camilli, Pierluigi Contucci, and Emanuele Mingione. “The multi-species mean-field spin-glass on the Nishimori line”. In:Journal of Statistical Physics182.1 (2021), pp. 1–20

  33. [41]

    Information- Theoretic Reduction of Deep Neural Networks to Linear Models in the Overparametrized Proportional Regime

    Francesco Camilli, Daria Tieplova, Eleonora Bergamin, and Jean Barbier. “Information- Theoretic Reduction of Deep Neural Networks to Linear Models in the Overparametrized Proportional Regime”. In:Proceedings of the Thirty-Eighth Conference on Learning Theory. Vol. 291. Proceed...

  34. [42]

    Statistical Physics of Deep Learning: Optimal Learning of a Multilayer Perceptron near Interpolation

    Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, and Rudy Skerk. “Statistical Physics of Deep Learning: Optimal Learning of a Multilayer Perceptron near Interpolation”. In:Physical Review X16.3 (2026), p. 031014.doi: 10.1103/56sb-pdh6

  35. [43]

    The solution of the deep Boltzmann machine on the Nishimori line

    Diego Alberici, Francesco Camilli, Pierluigi Contucci, and Emanuele Mingione. “The solution of the deep Boltzmann machine on the Nishimori line”. In:Communications in Mathematical Physics (to appear)(July 2021).doi:10 . 1007 / s00220 - 021 - 04165-0

  36. [44]

    Statistical inference of finite-rank tensors

    Hongbin Chen, Jean-Christophe Mourrat, and Jiaming Xia. “Statistical inference of finite-rank tensors”. en. In:Annales Henri Lebesgue5 (2022), pp. 1161–1189.doi: 10.5802/ahl.146.url:https://www.numdam.org/articles/10.5802/ahl.146/

  37. [45]

    On General Minimax Theorems

    Maurice Sion. “On General Minimax Theorems”. In:Pacific Journal of Mathematics 8.1 (1958), pp. 171–176.doi:10.2140/pjm.1958.8.171

  38. [46]

    Performance of Bayesian linear regression in a model with mismatch

    Jean Barbier, Wei-Kuo Chen, Dmitry Panchenko, and Manuel S´ aenz. “Performance of Bayesian linear regression in a model with mismatch”. In:Information and In- ference: A Journal of the IMA14.3 (Sept. 2025), iaaf019.issn: 2049-8772.doi: 10.1093/imaiai/iaaf019

  39. [47]

    Some Inequalities for Gaussian Processes and Applications

    Yehoram Gordon. “Some Inequalities for Gaussian Processes and Applications”. In: Israel Journal of Mathematics50.4 (1985), pp. 265–289.doi:10.1007/BF02759761

  40. [48]

    The Gaussian Min- Max Theorem in the Presence of Convexity

    Christos Thrampoulidis, Samet Oymak, and Babak Hassibi. “The Gaussian Min- Max Theorem in the Presence of Convexity”. In:arXiv preprint arXiv:1408.4837 (2014). arXiv:1408.4837 [math.OC]. 50

  41. [49]

    On the Stability of the Quenched State in Mean Field Spin Glass Models

    Michael Aizenman and Pierluigi Contucci. “On the Stability of the Quenched State in Mean Field Spin Glass Models”. In:Journal of Statistical Physics92.5–6 (1998), pp. 765–783.doi:10.1023/A:1023080223894

  42. [50]

    Fundamental limits of over- parametrized shallow neural networks for supervised learning

    Francesco Camilli, Daria Tieplova, and Jean Barbier. “Fundamental limits of over- parametrized shallow neural networks for supervised learning”. In: (2023). arXiv: 2307.05635 [cs.LG].url:https://arxiv.org/abs/2307.05635

  43. [51]

    Statisti- cal mechanics of deep learning beyond the infinite-width limit

    S Ariosto, R Pacelli, M Pastore, F Ginelli, M Gherardi, and P Rotondo. “Statisti- cal mechanics of deep learning beyond the infinite-width limit”. In:arXiv preprint arXiv:2209.04882(2022)

  44. [52]

    Generalisation error in learning with random features and the hidden manifold model

    Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc M´ ezard, and Lenka Zde- borov´ a. “Generalisation error in learning with random features and the hidden manifold model”. In:International Conference on Machine Learning. PMLR. 2020, pp. 3452–3462

  45. [53]

    Bayes-optimal learning of deep random networks of extensive-width

    Hugo Cui, Florent Krzakala, and Lenka Zdeborov. “Bayes-optimal learning of deep random networks of extensive-width”. In:Proceedings of the 40th International Con- ference on Machine Learning. ICML’23. Honolulu, Hawaii, USA: JMLR.org, 2023

  46. [54]

    Spectral Dynamics of Learning in Restricted Boltzmann Machines

    Aur´ elien Decelle, Giancarlo Fissore, and Cyril Furtlehner. “Spectral Dynamics of Learning in Restricted Boltzmann Machines”. In:EPL (Europhysics Letters)119.6 (2017), p. 60001.doi:10.1209/0295-5075/119/60001

  47. [55]

    Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learn- ing

    Charles H. Martin and Michael W. Mahoney. “Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learn- ing”. In:Journal of Machine Learning Research22.165 (2021), pp. 1–73.doi:10. 48550/arXiv.1810.01075. arXiv:1810.01075 [cs.LG]

  48. [56]

    Cambridge University Press, 2018.doi:10.1017/9781108231596

    Roman Vershynin.High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press, 2018.doi:10.1017/9781108231596. 51

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.