REVIEW 2 major objections 5 minor 54 references
Mixing labeled and unlabeled Hebbian kernels in a Hopfield network enlarges the retrieval region beyond either pure protocol, and thermodynamics cannot pick the mix—λ is a hyperparameter.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 15:39 UTC pith:J3LF7JUI
load-bearing objection Solid first RS theory of mixed supervised/unsupervised Hebbian Hopfield: interior λ enlarges retrieval, convexity makes λ a hyperparameter, and MC backs the SNR and phase claims. the 2 major comments →
Semi-supervised Hopfield model: Theoretical and Numerical results
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In the high-storage Hopfield model with synaptic matrix J^λ = λ J^(L) + (1−λ) J^(U), intermediate mixing strictly enlarges the replica-symmetric retrieval region relative to both pure protocols, while the quenched pressure is convex in λ, so any interior energy-equipartition point is a free-energy maximum—thermodynamics rejects the mixture and λ must be treated as an externally tuned hyperparameter optimized for retrieval capacity.
What carries the argument
Guerra interpolation of the quenched pressure after Gaussianizing the two channels (CLT for supervised noise, Carmona–Hu universality for unsupervised) and spectrally decomposing the correlated SW disorder into independent eigen-channels (two hybridized condensed modes, unsupervised bulk, supervised null space), which closes the RS self-consistency equations and the zero-temperature capacity map.
Load-bearing premise
The whole phase diagram and the claim that mixed learning wins rest on replica symmetry plus replacing the structured noise by Gaussians whose correlations are fully captured by that spectral split.
What would settle it
A Monte Carlo or RSB calculation at low temperature and intermediate load where pure supervision or pure unsupervised learning recovers a larger stable Mattis magnetization (or higher critical α) than the mixed rule at the paper’s λ⋆ would overturn the central ranking.
If this is right
- Semi-supervised Hebbian synapses can store more patterns than either pure channel at the same dataset quality and size.
- λ should be tuned to maximize retrieval capacity (closed form via SNR), not set to the labeled fraction, except for noiseless data.
- Below a quality threshold r_min = 1/√(M_L + M_U + 1) no mixing yields retrieval.
- The eigen-channel treatment of correlated multi-channel Hebbian noise can be reused for dense and hetero-associative networks.
Where Pith is reading between the lines
- If the network could slowly adapt λ during training, the capacity-optimal mix might be learned on-line rather than grid-searched—an extension the paper only flags as future work.
- The same convexity no-go likely applies to other convex combinations of labeled and unlabeled losses, suggesting free-energy criteria alone will not select semi-supervision in related mean-field models.
- Instability of RS at low T could shrink or reorder the mixed advantage; checking the de Almeida–Thouless line for this spectrum would test how much of the gain survives replica breaking.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper defines a semi-supervised Hopfield network whose couplings are the convex combination J^λ = λ J^(L) + (1−λ) J^(U) of supervised and unsupervised Hebbian kernels built from the same noisy archetypes, with mixing parameter λ ∈ [0,1]. A zero-temperature signal-to-noise analysis yields a closed-form one-step Mattis magnetization (3.15) and a learning threshold M⋆(α,r,λ). Guerra interpolation, after CLT/Carmona–Hu gaussianization and spectral decomposition of the correlated SW noise into independent eigen-channels (4.20), produces the RS quenched pressure (4.40), self-consistency equations, and zero-T phase diagrams. These show that intermediate λ strictly enlarges the retrieval region relative to both pure protocols. Convexity of the quenched pressure in λ (App. D) implies that any interior equipartition stationary point is a free-energy maximum, so thermodynamics rejects the mixture and λ is a hyperparameter; capacity maximization then supplies a closed-form λ⋆. All main predictions are checked against Monte Carlo simulations (Figs. 2–6).
Significance. The work fills a genuine gap: a solvable statistical-mechanical theory of semi-supervised Hebbian associative memory, as opposed to pure supervised/unsupervised protocols or to transductive label propagation on graphs. The eigen-channel treatment of correlated multi-channel disorder is a reusable technical contribution. The convexity argument cleanly separates thermodynamic stationarity from retrieval optimality and justifies treating λ as an externally tuned hyperparameter—an insight of practical as well as conceptual value. Strengths include explicit SNR formulae independent of the free-energy calculation, a model-independent convexity proof via Gibbs variance, closed-form λ⋆ and r_min, and extensive MC agreement on learning boundaries, phase diagrams, and m(α). Within the stated RS and large-dataset assumptions the central claims are well supported.
major comments (2)
- [Section 4, Figures 3–5; Section 6] Sec. 4 and Figs. 3–5: the capacity ranking “mixed outperforms pure” at low T rests on the RS self-consistency system (4.41)–(4.46) after gaussianization and the SW spectrum (4.20). The paper correctly flags RS as a limitation (Sec. 6), but the low-T insets (e.g. Fig. 3 left, Fig. 5) are precisely where AT-type instabilities are most likely. A short AT-line estimate or an explicit statement that the ranking is claimed only inside the RS-stable region would make the phase-diagram claim load-bearing-safe; the SNR one-step results (Sec. 3, Fig. 2) already support a mixed advantage independently of RS and should be cross-referenced more prominently when the RS diagrams are discussed.
- [Section 3, Equations (3.18)–(3.19), Figure 2] Sec. 3, Eq. (3.18)–(3.19) and the MC protocol in Fig. 2: the learning threshold uses the stability cut m^(1) > erf(θ) with θ = 1/√2, while the figure caption reports a retrieval threshold m^(1) = 0.9. These are not identical cuts. Please state explicitly which threshold is used for the solid curves in Fig. 2 and, if they differ, show that the qualitative λ-ordering of M⋆ is robust under reasonable θ (or under the m = 0.9 cut used numerically).
minor comments (5)
- [Section 4.3, Appendix B] Notation drift: ρ̂_L / ρ̂_U appear in (4.39) and App. B while ρ_L / ρ_U are used elsewhere; unify.
- [Figure 1, Remark 1] Fig. 1 is a schematic of N = 6; a brief caption note that self-couplings are omitted (Remark 1) would help non-specialist readers.
- [Section 4.2–4.3] In (4.12) and surrounding text, the signal block still carries finite-M factors while the large-dataset collapse ⟨n_L⟩ = ⟨n_U⟩ = ⟨m⟩ is used later; a one-sentence pointer to App. C.1 at first use would clarify the regime of each formula.
- [Sections 4–5] Typos / wording: “Intheunsupervisedchannel” (Sec. 4.1); “wehaveusedalso” and similar concatenated words in Sec. 4.3; “zero-temperaturemagnetizationcurves” in Sec. 5. A pass for spacing would help.
- [References] References [50] and the arXiv IDs in the 260x range look like placeholders or future citations; verify they are citable at submission.
Circularity Check
No significant circularity: SNR, Guerra RS pressure, convexity, and λ* are derived from the model definition and standard tools, not forced by fits or self-citation chains.
full rationale
The load-bearing claims are obtained from the Hamiltonian J^λ=λJ^(L)+(1-λ)J^(U) by independent calculations: one-step SNR magnetization (3.15) and threshold (3.18–3.19); Guerra interpolation with CLT/Carmona–Hu gaussianization and spectral decomposition of SW (4.20) yielding RS pressure (4.40) and zero-T equations (4.45–4.46); convexity of A in λ from the Gibbs variance of E_L-E_U for an affine Hamiltonian (App. D); and λ* by constrained maximization of α_SNR_c(λ) (5.9–5.17). Monte Carlo is used only for validation (Figs. 2–6), not to fit parameters that are then re-labeled as predictions. Citations to the same group’s pure supervised/unsupervised Hopfield results recover the λ→{0,1} limits as consistency checks and do not force the mixed-outperforms-pure ranking or the hyperparameter reading of λ. RS and universality are stated assumptions, not circular reductions. No step reduces a claimed prediction to its inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (3)
- λ (mixing hyperparameter) =
λ* closed form (5.17); interior ~0.5–0.86 depending on ρ
- θ (SNR stability threshold) =
1/√2 (and m>0.9 in Fig. 2 MC comparison)
- ρ_L, ρ_U (rescaled dataset noise / sizes)
axioms (6)
- domain assumption Replica-symmetric self-averaging of order parameters in the thermodynamic limit (4.30).
- domain assumption Carmona–Hu universality: unsupervised Rademacher noise may be replaced by Gaussians without changing quenched pressure (4.7).
- domain assumption Examples are multiplicative noise on Rademacher archetypes with quality r (2.1–2.4); labeled and unlabeled share the same archetypes.
- ad hoc to paper Synaptic rule is convex combination of supervised and unsupervised Hebb kernels with normalizations Γ_L, Γ_U (2.6–2.8).
- standard math Existence of thermodynamic quenched pressure limit (invoked via Guerra–Toninelli / Hopfield extensions).
- domain assumption One-step zero-T dynamics and SNR Gaussian approximation for m^(1) (3.5–3.8).
invented entities (2)
-
Semi-supervised mixed Hebbian coupling J^λ and mixing hyperparameter λ
independent evidence
-
Eigen-channels of the correlated SW noise matrix (μ±, bulk U, supervised null)
independent evidence
read the original abstract
In the daily practice of Machine Learning, fully labeled datasets are a luxury: labels demand expensive and time-consuming human annotation, whereas raw, unlabeled data can be harvested automatically and in bulk. Semi-supervised learning, where the network jointly exploits the few labeled and the many unlabeled examples at its disposal, is the standard answer to this asymmetry, yet a statistical mechanical theory of semi-supervised Hebbian learning is still lacking. In this paper we fill this gap for the Hopfield network: we prescribe a synaptic coupling given by the convex combination, weighted by a mixing parameter \lambda in [0,1], of the supervised and unsupervised Hebbian kernels built from the same archetypes, and we solve for the emergent computational capabilities of the resulting network. A signal-to-noise analysis yields the one-step Mattis magnetization and the learning threshold, i.e. the minimum dataset size for stable retrieval. Using Guerra's interpolation, we then derive the Replica Symmetric quenched pressure in the high-storage regime, treating the correlated disorder generated by the supervised and unsupervised channels through a particular eigen-channel decomposition. The resulting phase diagram shows that a mixed strategy outperforms both pure protocols. Finally, we prove that the quenched pressure is convex in \lambda, so thermodynamics cannot select an interior mixture: \lambda is therefore a learning hyperparameter. All the analytical findings are successfully checked against extensive Monte Carlo simulations.
Reference graph
Works this paper leans on
-
[1]
Agliari, L
E. Agliari, L. Albanese, F. Alemanno, A. Alessandrelli, A. Barra, F. Giannotti, D. Lotito, and D. Pedreschi. Dense hebbian neural networks: A replica symmetric picture of supervised learning. Physica A: Statistical Mechanics and its Applications, 626:129076, 2023
2023
-
[2]
Agliari, L
E. Agliari, L. Albanese, F. Alemanno, A. Alessandrelli, A. Barra, F. Giannotti, D. Lotito, and D. Pedreschi. Dense hebbian neural networks: A replica symmetric picture of unsupervised learning. Physica A: Statistical Mechanics and its Applications, 627:129143, 2023
2023
-
[3]
Agliari, F
E. Agliari, F. Alemanno, A. Barra, and A. Fachechi. Generalized guerra’s interpolation schemes for dense associative neural networks.Neural Networks, 128:254–267, 2020
2020
-
[4]
Agliari, F
E. Agliari, F. Alemanno, A. Barra, and G. D. Marzo. The emergence of a concept in shallow neural networks.Neural Networks, 148:232–253, 2022
2022
-
[5]
Agliari, A
E. Agliari, A. Alessandrelli, A. Barra, M. Centonze, and F. Ricci-Tersenghi. Networks of hebbian networks: more is different.Neural Networks, 194:108181, 2026
2026
-
[6]
Agliari, A
E. Agliari, A. Alessandrelli, A. Barra, M. S. Centonze, and F. Ricci-Tersenghi. Generalized hetero-associative neural networks.Journal of Statistical Mechanics: Theory and Experiment, 2025(1):013302, 2025
2025
-
[7]
E. Agliari, A. Barra, P. Bianco, A. Fachechi, and D. Pallara. The thermodynamic limit in mean field neural networks.arXiv preprint arXiv:2409.10145, 2024
Pith/arXiv arXiv 2024
-
[8]
Agliari, A
E. Agliari, A. Fachechi, and P. D. Mourao. The beneficial role of noises for disentanglement tasks in modular hebbian networks.Physica A: Statistical Mechanics and its Applications, 682:131134, 2026
2026
-
[9]
Agliari and G
E. Agliari and G. D. Marzo. Tolerance versus synaptic noise in dense associative memories.European Physical Journal Plus, 135, 11 2020
2020
-
[10]
Albanese, F
L. Albanese, F. Alemanno, A. Alessandrelli, and A. Barra. Replica symmetry breaking in dense hebbian neural networks.Journal of Statistical Physics, 189(2):1–41, 2022
2022
-
[11]
Albanese, A
L. Albanese, A. Alessandrelli, A. Annibale, and A. Barra. About the de Almeida–Thouless line in neural networks.Physica A: Statistical Mechanics and its Applications, 633:129372, 2024
2024
-
[12]
Albanese, A
L. Albanese, A. Barra, P. Bianco, F. Durante, and D. Pallara. Hebbian learning from first principles. Journal of Mathematical Physics, 65(11):113302, 2024. – 26 –
2024
-
[13]
Alemanno, M
F. Alemanno, M. Aquaro, I. Kanter, A. Barra, and E. Agliari. Supervised hebbian learning. Europhysics Letters, 141(1):11001, 2023
2023
-
[14]
A. Alessandrelli, A. Barra, A. Ladiana, A. Lepre, and F. Ricci-Tersenghi. Beyond disorder: Unveiling cooperativeness in multidirectional associative memories.arXiv preprint arXiv:2503.04454, 2025
Pith/arXiv arXiv 2025
-
[15]
Alessandrelli, A
A. Alessandrelli, A. Barra, A. Ladiana, A. Lepre, and F. Ricci-Tersenghi. Supervised and unsupervised protocols for hetero-associative neural networks.Physica A: Statistical Mechanics and its Applications, page 130871, 2025
2025
-
[16]
D. J. Amit.Modeling brain function: The world of attractor neural networks. Cambridge university press, 1989
1989
-
[17]
D. J. Amit, H. Gutfreund, and H. Sompolinsky. Storing infinite numbers of patterns in a spin-glass model of neural networks.Physical Review Letters, 55:1530–1533, 1985
1985
-
[18]
Baldi and S
P. Baldi and S. S. Venkatesh. Number of stable points for spin-glasses and neural networks of higher orders.Physical Review Letters, 58, 1987
1987
-
[19]
H. Bao, R. Zhang, and Y. Mao. The capacity of the dense associative memory networks. Neurocomputing, 469:198–208, 2022
2022
-
[20]
Barra, G
A. Barra, G. Genovese, and F. Guerra. The replica symmetric approximation of the analogical neural network.Journal of Statistical Physics, 140:784–796, 2010
2010
-
[21]
Bertoni, M
A. Bertoni, M. Frasca, and G. Valentini. COSNet: A cost sensitive neural network for semi-supervised learning in graphs. InMachine Learning and Knowledge Discovery in Databases (ECML PKDD), volume 6911 ofLecture Notes in Computer Science, pages 219–234. Springer, 2011
2011
-
[22]
Bovier and B
A. Bovier and B. Niederhauser. The spin-glass phase-transition in the hopfield model with p-spin interactions.Advances in Theoretical and Mathematical Physics, 5:1001–1046, 8 2001
2001
-
[23]
Carmona and Y
P. Carmona and Y. Hu. Universality in Sherrington-Kirkpatrick’s spin glass model.Annales de l’institut Henri Poincare (B) Probability and Statistics, 42, 2006
2006
-
[24]
Chapelle, B
O. Chapelle, B. Schölkopf, and A. Zien. Semi-supervised learning.IEEE Transactions on Neural Networks, 20(3):542–542, 2009
2009
-
[25]
A. C. C. Coolen, R. Kühn, and P. Sollich.Theory of neural information processing systems. OUP Oxford, 2005
2005
-
[26]
H. Cui, L. Saglietti, and L. Zdeborová. Large deviations of semisupervised learning in the stochastic block model.Physical Review E, 105(3):034108, 2022
2022
-
[27]
R. P. Feynman. Forces in molecules.Phys. Rev., 56:340–343, 1939
1939
-
[28]
Frasca, A
M. Frasca, A. Bertoni, M. Re, and G. Valentini. A neural network algorithm for semi-supervised node label learning from unbalanced data.Neural Networks, 43:84–98, 2013
2013
-
[29]
Fujii, H
T. Fujii, H. Ito, and S. Miyoshi. Statistical-mechanical analysis connecting supervised learning and semi-supervised learning.Journal of the Physical Society of Japan, 86(6):064003, 2017
2017
-
[30]
E. Gardner. Multiconnected neural network models.Journal of Physics A: General Physics, 20, 1987
1987
-
[31]
Genovese
G. Genovese. Universality in bipartite mean field spin glasses.Journal of Mathematical Physics, 53, 2012
2012
-
[32]
G. Getz, N. Shental, and E. Domany. Semi-supervised learning – a statistical physics approach. arXiv preprint cs/0604011, 2006. Presented at the 22nd ICML Workshop on Learning with Partially Classified Training Data, Bonn, 2005
Pith/arXiv arXiv 2006
-
[33]
Grandvalet and Y
Y. Grandvalet and Y. Bengio. Semi-supervised learning by entropy minimization. InAdvances in Neural Information Processing Systems, volume 17, pages 529–536. MIT Press, 2005
2005
-
[34]
F. Guerra. Broken replica symmetry bounds in the mean field spin glass model.Communications in Mathematical Physics, 233:1–12, 2003. – 27 –
2003
-
[35]
Guerra and F
F. Guerra and F. L. Toninelli. The thermodynamic limit in mean field spin glass models. Communications in Mathematical Physics, 230:71–79, 2002
2002
-
[36]
Güttinger
P. Güttinger. Das verhalten von atomen im magnetischen drehfeld.Z. Phys., 73:169–184, 1932
1932
-
[37]
Hellmann.Einführung in die Quantenchemie
H. Hellmann.Einführung in die Quantenchemie. Franz Deuticke, Leipzig, 1937
1937
-
[38]
Hertz, A
J. Hertz, A. Krogh, R. G. Palmer, and H. Horner.Introduction to the theory of neural computation. American Institute of Physics, 1991
1991
-
[39]
J. J. Hopfield. Neural networks and physical systems with emergent collective computational abilities.Proceedings of the National Academy of Sciences of the United States of America, 79:2554–2558, 1982
1982
-
[40]
D. P. Kingma, D. J. Rezende, S. Mohamed, and M. Welling. Semi-supervised learning with deep generative models. InAdvances in Neural Information Processing Systems, volume 27, pages 3581–3589. 2014
2014
-
[41]
Krotov and J
D. Krotov and J. Hopfield. Dense associative memory is robust to adversarial inputs.Neural Computation, 30:3151–3167, 2018
2018
-
[42]
Krotov and J
D. Krotov and J. J. Hopfield. Dense associative memory for pattern recognition.Advances in Neural Information Processing Systems, pages 1180–1188, 2016
2016
-
[43]
R. Latala. Some estimates of norms of random matrices.Proceedings of the American Mathematical Society, 133(5):1273–1282, 2005
2005
-
[44]
Mézard and A
M. Mézard and A. Montanari.Information, physics, and computation. Oxford University Press, 2009
2009
-
[45]
Mézard, G
M. Mézard, G. Parisi, and M. A. Virasoro.Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications, volume 9. World Scientific Publishing Company, 1987
1987
-
[46]
Nishimori.Statistical physics of spin glasses and information processing: an introduction
H. Nishimori.Statistical physics of spin glasses and information processing: an introduction. Number
-
[47]
Shalev-Shwartz and S
S. Shalev-Shwartz and S. Ben-David.Understanding machine learning: From theory to algorithms. Cambridge university press, 2014
2014
-
[48]
Sherrington and S
D. Sherrington and S. Kirkpatrick. Solvable model of a spin-glass.Physical review letters, 35(26):1792, 1975
1975
-
[49]
Talagrand
M. Talagrand. Rigorous results for mean field models for spin glasses.Theoretical computer science, 265(1-2):69–77, 2001
2001
-
[50]
V. Tutevych, R. Memmesheimer, L. Eichler, D. Pavlichenko, F. Schilke, R. Krudewig, and S. Behnke. Efficient image annotation via semi-supervised object segmentation with label propagation.arXiv preprint arXiv:2604.22992, 2026
Pith/arXiv arXiv 2026
-
[51]
J. E. Van Engelen and H. H. Hoos. A survey on semi-supervised learning.Machine learning, 109(2):373–440, 2020
2020
-
[52]
Ver Steeg, A
G. Ver Steeg, A. Galstyan, and A. E. Allahverdyan. Statistical mechanics of semi-supervised clustering in sparse graphs.Journal of Statistical Mechanics: Theory and Experiment, 2011(08):P08009, 2011
2011
-
[53]
Zhang, C
P. Zhang, C. Moore, and L. Zdeborová. Phase transitions in semisupervised clustering of sparse networks.Physical Review E, 90(5):052802, 2014. – 28 – A Evaluation of the momenta of the effective post synaptic potential In this section we compute in details the two main quantities of section 3:S:=E hi ξ1 i =E (h(L) i + h(U) i )ξ1 i andV:= Var hi ξ1 i , whe...
2014
-
[111]
Clarendon Press, 2001
2001
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.