REVIEW 2 major objections 6 minor 49 references
Allowing small training error keeps wide, searchable regions of binary-perceptron weight space alive past the zero-error hardness threshold, and those noisy regions still generalize.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Finite training error extends the overlap-gap threshold in binary perceptrons so wide, algorithmically reachable basins persist and still generalize where zero-error solutions are hard.
T0 review reviewed 2026-07-30 challenge →
load-bearing objection Solid finite-T extension of the binary-perceptron landscape program: frozen 1RSB survives for discontinuous losses, dense finite-energy regions push past zero-T OGP and still generalize, with the main caveat already flagged by the authors (RS cloned OGP is only an upper bound). the 2 major comments →
On the robustness of noisy solutions in non-convex neural networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Dense, algorithmically accessible regions of finite-energy configurations persist beyond the zero-temperature overlap-gap threshold, up to a larger threshold that grows with allowed training error; in the teacher-student setting those regions retain good generalization and are reachable by finite-temperature reinforced approximate message passing where zero-error solutions and teacher recovery are computationally hard.
What carries the argument
Finite-temperature m-clone free entropy (and its Legendre transform, the constrained entropy) whose vanishing and stationarity locate the finite-energy overlap-gap threshold α_OGP(ε); together with a boundary-layer criterion on the single-pattern Gibbs weight that decides whether freezing survives at positive temperature.
Load-bearing premise
The finite-temperature overlap-gap thresholds are computed under a replica-symmetric treatment of the cloned free entropy, which the paper itself notes produces an unphysical non-monotonic dependence on the number of clones and therefore only an upper bound.
What would settle it
A refined replica-symmetry-breaking calculation of the cloned entropy, or large-N runs of reinforced AMP, that push the algorithmic finite-error threshold past the reported RS α_OGP(ε) curves, or that find no accessible dense clusters once the true (non-RS) gap appears.
If this is right
- Tolerating a controlled training error expands the range of constraint densities at which wide basins remain algorithmically reachable.
- Smoothing or power-law vanishing of the single-pattern weight at the decision boundary can eliminate freezing of the equilibrium measure at positive temperature.
- In teacher-student learning, finite-energy wide regions can generalize well even when exact recovery of the teacher is information-theoretically or algorithmically impossible.
- Message-passing algorithms should be shaped to target wide finite-energy regions rather than isolated global minima once the zero-error landscape fractures.
Where Pith is reading between the lines
- The same finite-energy OGP construction should apply to other discrete non-convex CSPs (coloring, SAT) once an error-counting or soft loss is introduced.
- If wide finite-energy minima later absorb the teacher above the AMP threshold, that would explain why annealed or replicated searches succeed where pure zero-temperature search fails.
- Practical early-stopping and noise-injection schedules in deeper nets may be succeeding for the geometric reason identified here: they keep the optimizer inside a wide basin that has moved off the zero-loss manifold.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies binary perceptrons (storage and teacher–student) at finite temperature, where a positive training error is allowed and penalized. It shows that the frozen 1RSB equilibrium structure of the error-counting loss survives at every finite T, and derives a boundary-layer criterion BK(b,y) on the single-pattern Gibbs weight that predicts when smoothing removes freezing. It then extends the overlap-gap construction to finite energy, obtaining thresholds α_OGP(ε) that increase with allowed training error ε, and argues that dense finite-energy regions remain algorithmically accessible beyond the zero-T OGP line. In the teacher–student setting, constrained-clone calculations and a finite-T reinforced AMP (rAMP) algorithm are used to show that these wide finite-energy regions retain good generalization and can be reached in the hard regime where zero-error solutions and teacher recovery are computationally hard.
Significance. The work cleanly connects the zero-temperature OGP / robust-ensemble picture to the practically relevant finite-error regime. Strengths include: a careful frozen-1RSB analysis that recovers Horner’s dynamical picture and yields an explicit finite-N scaling Td ∼ N^{1/4}/√log N; a general, falsifiable smoothness criterion BK for freezing under different losses (jump, Horner soft penalty, logarithmic potential); recovery of known zero-T OGP values (≈0.784 storage, ≈1.154 teacher–student); and a concrete finite-T message-passing algorithm whose success locus tracks the predicted α_OGP(ε) trend. If the finite-energy accessibility and generalization claims hold under a sharper ansatz, the paper supplies a useful design principle—search wide finite-energy basins rather than isolated global minima—relevant beyond toy perceptrons.
major comments (2)
- [§IV, Eqs. (23)–(29), App. A2] §IV, Eqs. (23)–(29) and App. A2: the finite-T m-OGP thresholds are obtained under an RS ansatz for the cloned free entropy. As the paper notes (Fig. 2 right, Fig. 8), α_OGP(m) is non-monotonic in m, contradicting the combinatorial requirement that larger m can only harden the constrained problem. Consequently min_m α_OGP(m) is only an upper bound on the true OGP onset. This should be stated explicitly wherever α_OGP(ε) is compared to rAMP (Fig. 4 right, Fig. 6) and in the abstract/conclusions, so that the numerical agreement is not read as a sharp threshold test. A short 1RSB or annealed-bound check for small m (as done for SBP at T=0 in Table II) would substantially strengthen the claim.
- [§V, Fig. 6] §V and Fig. 6: the teacher–student phase diagram places finite-T rAMP points (fuchsia) deep into the hard region α_OGP < α < α_AMP with favorable generalization, but the manuscript does not report the corresponding success-probability curves, N-scaling, or hyperparameter optimization that are given for storage in Fig. 4. Without those data, the central claim that thermal noise “enables effective generalization” where zero-T solution finding and teacher recovery are hard rests mainly on the constrained-clone analytics (Fig. 5) plus schematic markers. Adding a storage-style success panel (or error bars / sample counts) for the teacher–student rAMP runs is load-bearing for that claim.
minor comments (6)
- [Abstract, §I] Abstract and §I: α_OGP(ε) is written both as α_OGP(ε) and α_OGP(T); pick one primary parameterization (training error vs temperature) and define the Legendre relation once.
- [Fig. 2] Fig. 2 caption: “α_OGP ≃ 1.950” for m=2, β=5 in the teacher–student setting looks inconsistent with the zero-T value ≈1.154 and with the right panel scale; please check the number and units.
- [§III, App. B1] Eq. (13) / (B14): the finite-N cutoff 1−q1 ≃ 1/N is heuristic; a one-sentence remark that the N^{1/4} scaling is not a controlled finite-size theory would help.
- [App. C, Fig. 4] App. C and Algorithm 1: state the stopping criterion, typical tmax, and the grid used to optimize (ρ, β) for the red points in Fig. 4 so the algorithmic thresholds are reproducible.
- [§II, throughout] Typos / wording: “having with N Ising weights”; “Thishastobedistinguished”; “onthecloned”; occasional missing spaces after commas in the arXiv text. A proofread pass is needed.
- [References] References [16,17] (Stojnic 2026) are cited for compatibility of OGP estimates; if still preprints, give arXiv IDs consistently with the rest of the bibliography.
Circularity Check
No significant circularity: finite-T freezing, OGP(ε), and generalization results are derived from standard replica saddles and numerics, not forced by fitted inputs or load-bearing self-citation.
full rationale
The load-bearing objects are constructed independently of the claimed outputs. The frozen-1RSB persistence and dynamical-temperature divergence follow from the m=1 1RSB saddle equations and the boundary-layer term BK(b,y) (App. B); the smoothness criterion is a direct asymptotic analysis of H(x,y) near decision boundaries for different K, not a rename of a prior fit. Finite-temperature m-OGP thresholds are defined by the Legendre entropy conditions sm(q1;β)=0 and ∂q1 sm=0 on the cloned free entropy (Eqs. 23–29, App. A2)—the standard OGP construction with β kept finite—and recover known zero-T values (αOGP≃1.154 teacher-student; ≃0.784 storage) as consistency checks rather than inputs. The RS non-monotonicity in m is an openly stated approximation artefact that makes min_m αOGP(m) an upper bound; that is a correctness limitation, not circularity. Teacher-student generalization of constrained clones is computed from the same cloned measure and compared to typical/Bayesian baselines. rAMP success curves use tuned (ρ,β) only to probe algorithmic reachability against those analytical thresholds; the theory is not fitted to the algorithm. Self-citations to Baldassi/Zecchina robust-ensemble and prior OGP work supply background geometry and zero-T benchmarks; the finite-T extensions and BK criterion are self-contained calculations. No step reduces a claimed prediction to its defining input by construction.
Axiom & Free-Parameter Ledger
free parameters (3)
- reinforcement rate ρ and inverse temperature β in rAMP =
optimized per ε_targ; schedule ρt=1−(1−ρ)^t
- clone number m and mutual overlap q1 in constrained measure =
m=2,3,4 primarily; q1≈0.9–0.98 in Fig. 5–6
- finite-N cutoff 1−q1 ≃ 1/N for Td scaling =
1−q1 ~ 1/N
axioms (6)
- domain assumption Replica method with RS and 1RSB ansätze computes the quenched free entropy and cloned entropy in the N→∞ limit.
- domain assumption m-OGP (forbidden intermediate overlaps among near-optimal m-tuples) obstructs stable algorithms; absence of OGP plus large local entropy indicates algorithmic accessibility.
- domain assumption Single-pattern Gibbs weights K_ABP / K_SBP are the error-counting factors e^{-βΘ(-s)} (and SBP analog); other K (Horner, s^γ) are alternative finite-T continuations.
- domain assumption Gaussian i.i.d. patterns and binary ±1 weights; storage labels Rademacher or teacher-planted signs.
- ad hoc to paper Boundary-layer term BK(b,y)→∞ as y→0 is necessary and sufficient for the frozen q1=1 solution in the m=1 1RSB dynamical analysis.
- standard math Standard Gaussian integrals, Parisi 1RSB block structure, and Legendre transform s=ϕ+βαε relating free entropy to entropy at finite β.
invented entities (2)
-
Boundary-layer freezing diagnostic BK(b,y)
independent evidence
-
Finite-temperature OGP threshold α_OGP(ε) [or α_OGP(β)]
independent evidence
Cite this review
Pith. "Pith review of On the robustness of noisy solutions in non-convex neural networks." pith.science (2026). https://pith.science/paper/JMILELZ4
@misc{pith2026260727000,
author = {Pith},
title = {Pith review of: On the robustness of noisy solutions in non-convex neural networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/JMILELZ4}},
note = {Machine review of arXiv:2607.27000}
}
abstract
Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare. At zero temperature this picture has been formalized in binary perceptrons through the overlap gap property (OGP), which limits algorithmic access to configurations with zero training error above a critical constraint density $\alpha_{\rm OGP}$. Here we extend this description to finite temperature, where a positive training error is allowed and statistically penalized. We first show that the frozen one-step replica-symmetry-breaking solution, dominating the zero temperature equilibrium measure, survives at any finite temperature. We furthermore derive a general criterion, based on the smoothness of the single-pattern Gibbs weight near the decision boundary, that determines when a finite-temperature relaxation of the loss removes freezing. We then extend the OGP construction to finite temperature and show that dense, algorithmically accessible regions of finite-energy configurations persist beyond $\alpha_{\rm OGP}$, up to a threshold $\alpha_{\rm OGP}(\epsilon)$ that grows with the allowed training error $\epsilon$. Finally, in the teacher-student setting, we show that these wide, finite-energy regions still retain good generalization. Using a finite energy message-passing algorithm, we demonstrate numerically that thermal noise enables effective generalization in the regime of constraint densities where both recovering the teacher and finding a zero temperature solution are computationally hard.
Figures
Reference graph
Works this paper leans on
-
[1]
Increasing the temperature generally worsens generalization, as expected, but the degradation is gradual
The same analysis can be repeated at finite temperature. Increasing the temperature generally worsens generalization, as expected, but the degradation is gradual. Wide regions can therefore be continuously followed from zero to positive temperature while retaining a generalization error substantially smaller than that of typical configurations. Searching ...
-
[2]
Baldassi, A
C. Baldassi, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina, Phys. Rev. Lett.115, 128101 (2015)
2015
-
[3]
C. Baldassi, C. Borgs, J. T. Chayes, A. Ingrosso, C. Lucibello, L. Saglietti, and R. Zecchina, Proceedings of the National Academy of Sciences113, E7655 (2016), https://www.pnas.org/doi/pdf/10.1073/pnas.1608103113
-
[4]
Chaudhari, A
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, Journal of Statistical Mechanics: Theory and Experiment2019, 124018 (2019)
2019
-
[5]
Gamarnik, Proceedings of the National Academy of Sciences (2021), 10.1073/pnas.2108492118
D. Gamarnik, Proceedings of the National Academy of Sciences (2021), 10.1073/pnas.2108492118
-
[6]
Gamarnik and A
D. Gamarnik and A. Jagannath, The Annals of Probability49, pp. 180 (2021)
2021
-
[7]
Gamarnik and M
D. Gamarnik and M. Sudan, The Annals of Probability45, 2353 (2017)
2017
-
[8]
Gamarnik, E
D. Gamarnik, E. C. Kizildag, W. Perkins, and C. Xu, in2022 IEEE 63rd Annual Symposium on Foundations of Computer Science (FOCS)(2022) pp. 576–587
2022
-
[9]
Baldassi, E
C. Baldassi, E. M. Malatesta, G. Perugini, and R. Zecchina, Phys. Rev. E108, 024310 (2023)
2023
-
[10]
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, ArXivabs/2001.08361(2020)
Pith/arXiv arXiv 2001
-
[11]
Hoffmann, S
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. W. Rae, and L. Sifre, inProceedings of the 36th International Conference on Neural Information Proc...
2022
-
[12]
W. Krauth and M. Mézard, Journal de Physique (1989), 10.1051/jphys:0198900500200305700
-
[13]
Aubin, W
B. Aubin, W. Perkins, and L. Zdeborová, Journal of Physics A: Mathematical and Theoretical52, 294003 (2019)
2019
-
[14]
H. Huang and Y. Kabashima, Physical Review E (2014), 10.1103/physreve.90.052813
-
[15]
Horner, Zeitschrift für Physik B Condensed Matter86, 291 (1992)
H. Horner, Zeitschrift für Physik B Condensed Matter86, 291 (1992)
1992
-
[16]
Mézard, G
M. Mézard, G. Parisi, and M. A. Virasoro,Spin glass theory and beyond(World Scientific, Singapore, 1987)
1987
-
[17]
Stojnic, arXiv preprint arXiv:2601.10628 (2026)
M. Stojnic, arXiv preprint arXiv:2601.10628 (2026)
arXiv 2026
-
[18]
Stojnic, arXiv preprint arXiv:2604.19712 (2026)
M. Stojnic, arXiv preprint arXiv:2604.19712 (2026)
Pith/arXiv arXiv 2026
-
[19]
Györgyi, Phys
G. Györgyi, Phys. Rev. A41, 7097(R) (1990)
1990
-
[20]
Gardner and B
E. Gardner and B. Derrida, Journal of Physics A: Mathematical and General22, 1983 (1989). 13
1983
-
[21]
E. Gardner and B. Derrida, Journal of Physics A: Mathematical and General (1988), 10.1088/0305-4470/21/1/031
-
[22]
Opper and D
M. Opper and D. Haussler, Phys. Rev. Lett.66, 2677 (1991)
1991
-
[23]
Baldassi, C
C. Baldassi, C. Lauditi, E. M. Malatesta, G. Perugini, and R. Zecchina, Phys. Rev. Lett.127, 278301 (2021)
2021
-
[24]
Perkins and C
W. Perkins and C. Xu, Random Structures & Algorithms64, 856 (2024)
2024
-
[25]
Barbier, A
D. Barbier, A. El Alaoui, F. Krzakala, and L. Zdeborová, Journal of Physics A: Mathematical and Theoretical57, 195202 (2024)
2024
-
[26]
Barbier, SciPost Phys.18, 115 (2025)
D. Barbier, SciPost Phys.18, 115 (2025)
2025
-
[27]
T. R. Kirkpatrick and D. Thirumalai, Phys. Rev. Lett.58, 2091 (1987)
2091
-
[28]
G. Catania, A. Decelle, and B. Seoane, Physical Review E (2024), 10.1103/physreve.109.065313
-
[29]
Straziota, E
D. Straziota, E. Demyanenko, C. Baldassi, and C. Lucibello, inAdvances in Neural Information Processing Systems, Vol. 38, edited by D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (Curran Associates, Inc., 2025) pp. 111927–111966
2025
-
[30]
A. Braunstein and R. Zecchina, Physical Review Letters96(2006), 10.1103/physrevlett.96.030201
-
[31]
C. Baldassi and A. Braunstein, Journal of Statistical Mechanics: Theory and Experiment2015(2015), 10.1088/1742- 5468/2015/08/p08008
doi:10.1088/1742- 2015
-
[32]
Barbier, arXiv preprint arXiv:2505.20954 (2025)
D. Barbier, arXiv preprint arXiv:2505.20954 (2025)
Pith/arXiv arXiv 2025
-
[33]
M. C. Angelini and F. Ricci-Tersenghi, Phys. Rev. X13, 021011 (2023)
2023
-
[34]
Budzynski, F
L. Budzynski, F. Ricci-Tersenghi, and G. Semerjian, Journal of Statistical Mechanics: Theory and Experiment2019, 023302 (2019)
2019
-
[35]
Budzynski and G
L. Budzynski and G. Semerjian, Journal of Statistical Mechanics: Theory and Experiment2020, 103406 (2020)
2020
-
[36]
M. C. Angelini, L. Budzynski, and F. Ricci-Tersenghi, Phys. Rev. E112, 064117 (2025)
2025
-
[37]
Benedetti, A
M. Benedetti, A. Bogdanov, E. M. Malatesta, M. Mézard, G. Perrupato, A. Rosen, N. I. Schwartzbach, and R. Zecchina, Journal of Statistical Mechanics: Theory and Experiment2025, 123303 (2025)
2025
-
[39]
C. Baldassi, F. Pittorino, and R. Zecchina, Proceedings of the National Academy of Sciences117, 161 (2020), https://www.pnas.org/doi/pdf/10.1073/pnas.1908636117
-
[40]
Benedetti, A
M. Benedetti, A. Bogdanov, E. M. Malatesta, M. Mézard, G. Perrupato, A. Rosen, N. I. Schwartzbach, and R. Zecchina, Phys. Rev. X16, 021051 (2026)
2026
-
[41]
E. M. Malatesta, arXiv preprint arXiv:2309.09240 (2023)
Pith/arXiv arXiv 2023
-
[42]
Parisi, Physics Letters A73, 203 (1979)
G. Parisi, Physics Letters A73, 203 (1979)
1979
-
[43]
Monasson, Phys
R. Monasson, Phys. Rev. Lett.75, 2847 (1995)
1995
-
[44]
Mézard and A
M. Mézard and A. Montanari,Information, Physics, and Computation(Oxford University Press, 2009). APPENDICES Appendices . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 A The free entropy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ...
2009
-
[45]
RS Ansatz We consider here the standard replica symmetric (RS) assumption on the student overlap matrix and its conjugate qab =δ ab + (1−δ ab)q ˆqab = (1−δ ab)ˆq; moreover we impose for the teacher-student overlaps ra =r ˆra = ˆr . A standard computation gives the free entropy (A4): SRS(q,ˆq, r,ˆr) =GRS S (q,ˆq, r,ˆr) +αGRS E (q, r),(A5a) GRS S (q,ˆq, r,ˆ...
-
[46]
1RSB Ansatz and entropy of them-cloned system A more refined parameterization of the student overlap matrixqab is in general needed to compute the equilibrium value of the free entropy [41]. We consider here a 1-step replica symmetry breaking (1RSB) Ansatz which consists in imposing qab =q 0 + (q1 −q 0)I (n,m) ab + (1−q 1)I (n,1) ab (A10a) ˆqab = ( ˆq0 + ...
-
[47]
The frozen 1RSB solution We specialize here the computation to the ABP for simplicity. Equation (B6) solved forˆq1 then reads: ˆq1 =α (1−e −β)2 1−q 1 Z Dz0 2H − rz0√ q0−r2 H √q0z0, √1−q 0 Z Dz1 G √q0z0+√q1−q0z1√1−q1 2 H √q0z0 + √q1 −q 0z1, √1−q 1 ,(B7) whereG(x)denotes the standard Gaussian density function. Substituting (B7) in (B5), we have∂ ˜ϕ ∂ˆq1 as ...
-
[48]
General criterion for freezing The previous computation for the ABP shows that the existence of the frozen solution is controlled by the behaviour of the energetic saddle-point equation (B6): ˆq1 =α Z Dz0 2H − rz0√ q0−r2 H √q0z0, √1−q 0 Z Dz1 ∂xH √q0z0 + √q1 −q 0z1, √1−q 1 2 H √q0z0 + √q1 −q 0z1, √1−q 1 =α Z Dz0 2H − rz0√ q0−r2 H √q0z0, √1−q 0 Z dx√q1 −q ...
-
[49]
Approximate Message Passing Consider the Gibbs measure induced by the partition function (5): pD (w) = QP µ=1 K(s µ(w)) ZD .(C1) The Approximate Message Passing (AMP) algorithm provides an iterative approximation to the local marginals of the measure (C1). AMP is obtained from the Belief Propagation (BP) equations using a Gaussian approximation which para...
-
[50]
Standard AMP estimates the local magnetizationsa i ≃ ⟨wi⟩of the Gibbs measure, but these magnetizations need not be close to±1
AMP + reinforcement (rAMP) The reinforcement AMP algorithm (rAMP) [29] is a modification of the AMP iteration designed to turn the soft marginal information computed by AMP into an actual binary configuration. Standard AMP estimates the local magnetizationsa i ≃ ⟨wi⟩of the Gibbs measure, but these magnetizations need not be close to±1. The role of reinfor...
This paper was first reviewed by grok-4.5 on July 30, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.