Pith. sign in

REVIEW 2 major objections 5 minor 52 references

Exponential Capacity in Multilayer Hetero-Associative Neural Networks

T0 review · 2 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read A multilayer exponential Hopfield network stores exponentially many hetero-associations; each added layer multiplies capacity by a factor exponential in the layer size.

desk verdict A genuinely new multilayer hetero-associative exponential Hopfield model with a closed-form capacity rate and strong empirical support, but the rigorous case for the capacity claim has a real gap that the authors themselves concede. read the letter →

arxiv 2607.29554 v1 pith:XAZ3N4GZ submitted 2026-07-31 cond-mat.dis-nn stat.ML

classification cond-mat.dis-nnstat.ML
keywords associativememoryHopfieldnetworkshetero-associationexponentialstoragecapacitylargedeviationscavitymethodsurjectivemapgeneralisation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper extends exponential-capacity Hopfield networks from auto-association to hetero-association: distinct cue layers drive a distinct target layer through a symmetric energy. The central claim is that the perfectly aligned hetero-associative state is a fixed point of zero-temperature dynamics up to P_c ~ e^{N rho_L} stored patterns, with an explicit rate rho_L growing like L log 2, so binding more modalities multiplies capacity by an exponential amount. The analysis shows that only surjective functions of the cue can be stored at all, and that enlarging basins lowers the capacity rate without destroying its exponential character. On structured and real data the same closed forms describe capacity and basins, while generalisation to unseen cues stays real but bounded, controlled by the geometry of the encoding rather than by the storage rule. If correct, this yields a principled high-capacity content-addressable memory for many-to-one tasks and a sharp separation between memorisation and generalisation.

What carries the argument

The central object is the multilayer energy H = -N sum_mu exp[N sum_{a<b}(m^a_mu m^b_mu - 1)], where m^a_mu are per-layer Mattis overlaps. The exponential weight makes the field of each neuron a pattern sum dominated by the collectively retrieved pattern; the non-factorising per-pattern noise is evaluated by a large-deviation principle whose unique symmetric saddle m* = tanh(2(L-1)m*) sets the storage rate rho_L and prefactor K_L. Surjectivity of the stored rule follows from the field's single-valuedness: duplicate cues with distinct targets cancel in the target field and relax to their componentwise majority.

What would settle it

At a layer size where the transition is accessible (for example N around 10, L = 2), measure the one-step overlap versus P and compare with the Gaussian prediction (32), using the empirical per-pattern variance rather than the annealed K_L e^{-N rho_L}. If the transition sits at a P not exponentially close to e^{N rho_L}, or the recall curve deviates from erf beyond finite-size corrections, the Gaussian approximation fails below capacity. A second check: compute the exact third moment of the noise and see whether a Berry-Esseen bound can vanish in the regime P much smaller than e^{N rho_L}; if

Watch

Extended reading notes

Core claim

The paper introduces an energy H = -N sum_mu exp[N sum_{a<b} m^a_mu m^b_mu - N choose(L,2)] that is minimized exactly when every layer retrieves the pattern of the same index. A cavity signal-to-noise analysis, with the noise evaluated by a large-deviation saddle point on the symmetric ray of layer magnetisations, shows the aligned hetero-associative state is a fixed point up to P_c ~ e^{N rho_L} patterns, with rho_L = L[(L-1) - phi_L(x*)] ~ L log 2. The same field computation proves the stored rule must be a surjective function of the cue: a cue mapped to two targets returns their componentwise majority, and a target with no cue has an empty basin. Simulations on i.i.d., manifold, TCR/epito

Load-bearing premise

The load-bearing premise is that the per-pattern field is approximately Gaussian in the sub-critical regime, a step the paper's own Berry-Esseen bound does not certify below capacity; the further idealisation that layer datasets are independent, which the real data violate, is traded against empirical agreement.

Editorial extensions

If this is right

  • If correct, a hetero-associative memory can store an exponential number of surjective associations in N, with the capacity exponent growing linearly in the number of bound modalities.
  • Only single-valued, surjective many-to-one maps are storable; injectivity is neither required nor useful, and a cue with several targets is answered by their componentwise majority.
  • Widening the network increases capacity exponentially but shrinks the basins, so there is a quantitative trade-off between storage rate and robustness to corruption.
  • The same closed forms describe retrieval and basins for correlated, many-to-one real data, so the independence assumption is benign for memory performance.
  • Memorisation and generalisation are distinct capabilities: an exponential content-addressable repository can be near-perfect at recall while only modestly above chance at routing unseen cues, with the encoder's geometry setting the ceiling.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own cost analysis implies that the capacity theorem describes the shape of the energy landscape, not a device that could hold P_c patterns: at N = 64, L = 2 the predicted capacity already exceeds any physically realisable memory, so the practical content lives in the finite-N experiments and in the landscape's structural properties.
  • A natural follow-up is whether modifying the encoder or the exponent can convert part of the exponential memorisation budget into generalisation; the CLINC150 ablation shows target co-location in pattern space is what generalisation feeds on, so principled encoder design could improve it without collapsing to dense sampling.
  • The annealed-versus-typical gap computed here likely generalises to other product-of-overlap energies: whenever the signal or noise exponent is quadratic in the masks, the annealed average is dominated by rare favourable configurations, and the physically operative threshold is the annealed one.
  • The surjective-function constraint offers a design principle for federated or privacy-preserving associative memories: if each client contributes cues for a shared target, the stored rule stays well-posed exactly when the map is many-to-one, which is the regime real data naturally occupy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces a multilayer hetero-associative exponential Hopfield network with L layers of N binary neurons and energy H = −N Σ_μ exp[N Σ_{a<b}(m_a^μ m_b^μ − 1)]. The central claim is that the perfectly aligned state (σ^a = ξ^{1,a}) is a fixed point of the zero-temperature dynamics up to P_c ∼ e^{Nρ_L} stored patterns, with an explicit rate ρ_L = L[(L−1) − φ_L(x*)] that grows like L log 2. The paper also derives basin-of-attraction exponents under corrupted cues, identifies a surjective-function condition on storable rules, and tests the closed forms against i.i.d. Monte Carlo, a Hidden Manifold Model, VDJdb T-cell receptor triples, and CLINC150 intent data. It reports near-perfect memorisation with modest, geometry-limited generalisation.

Significance. If the main result is correct, it is a substantial extension of exponential-capacity associative memories to hetero-association, with a parameter-free closed-form rate obtained from a saddle-point evaluation of a large-deviation functional. The paper is unusually strong empirically: it ships reproducible code, makes falsifiable predictions with no fitted constants for the annealed curves, and includes careful null controls (label-permutation, out-of-scope, held-out-region) for the generalisation claims. The explicit annealed/typical distinction and the discussion of the exponential computational price of exponential capacity are also valuable. However, the theoretical status of the capacity rate is weaker than the abstract implies, and the synthetic data experiment contains a structural inconsistency with the paper's own storage condition; both points need attention before the claims are accepted at face value.

major comments (2)
  1. [§4–5, Eqs. (29)–(32); App. A, Eq. (A.16); Remark 4] The quantitative capacity rate ρ_L rests entirely on the Gaussian approximation of the stability variable X_i^a. The appended Berry–Esseen bound (A.16) has error C M_L/(σ_1√P), which at the transition P ∼ e^{Nρ_L} is C M_L/√K_L — a constant that does not vanish with N (for L=2, M_2≈3.19 and √K_2≈0.44). Remark 4 concedes this and states the bound is silent in the sub-critical regime where all experiments run. Since the stability condition is a tail event at a fixed signal-to-noise ratio, a constant Kolmogorov error does not control the failure probability; the exponential rate could in principle be affected. The large-deviation evaluation of the second moment (22)–(26) is exact at leading order, but the CLT step from that moment to Eqs. (29)–(32) is not. I request either a large-deviation/Chernoff bound on the sum of the bounded noise terms that recovers the rate, or an explicit downgrade
  2. [§7 and App. F, Eqs. (F.2)–(F.4); Fig. 3(c)] The Hidden Manifold Model construction does not enforce the function condition of Remark 1 on the stored cue set. The cue is ξ^{μ,a}=sign(F^a z_μ) while the target is ρ_{k(z_μ)} with k determined by the first n_bits signs of z_μ. Two different latents z, z′ can fall in the same cell of the hyperplane arrangement — hence have identical cue patterns — but have different first-n_bits sign patterns, yielding different targets. Once P exceeds Cover's count C(N,D), collisions are inevitable; Fig. 3(c) shows collision rates reaching ∼0.5. Such contradictory (same-cue, different-target) associations violate the single-valuedness condition, and by Eq. (17) the network would return a majority mixture rather than a stored association. The paper's collision measure counts duplicate cues only, not cue-target conflicts. Please filter the stored set to a function (as done for VDJdb and CLINC) or quanti
minor comments (5)
  1. [§5, Remark 2] The statement that 'the curves labelled typical below use the empirically calibrated per-pattern variance and are the ones that track the data' is confusing, because Fig. 2(a) shows the data sitting on the annealed curve. Please clarify which curves are parameter-free predictions and which are post-hoc fits, and avoid presenting fitted curves as theoretical predictions in the main text.
  2. [Fig. 1(a) caption] The caption text appears to contain a rendering artifact: 'Pc »e^{N½2}' should presumably read 'Pc ∼ e^{Nρ_2}'. Please check the figure source.
  3. [Table 1 / Eq. (33)] The exponential storage rate is defined as α := log P / N in Eq. (33), but Table 1 writes α = (1/N) log P with slightly different notation. Make the definitions uniform.
  4. [App. B, around Eq. (B.15)–(B.16)] The positivity argument for λ_∥ would be easier to follow with one extra sentence explaining why the crossing of tanh(x) with the line x/[2(L−1)] occurs at a point where the derivative of tanh is smaller than the line's slope.
  5. [App. C, Eq. (C.18)] The Gaussian integration leading to Eq. (C.18) is quite compressed; the condition (L−1)(1−r^2)<1 and the divergence otherwise deserve a few more lines of derivation for reproducibility.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the capacity rate is a parameter-free saddle-point result; the only empirical calibration is explicitly labeled 'typical' and is secondary.

full rationale

The central derivation chain is self-contained. The capacity exponent ρ_L is obtained in Section 4 and Appendix B by evaluating the single-pattern noise second moment (22) through Cramér's theorem and Varadhan's lemma; the saddle-point equation (23) and rate (24)/(31) contain no fitted constants and do not reference any dataset. The stability criterion (29)-(31) mathematically converts the computed noise variance into P_c ∼ e^{Nρ_L}, so the claimed capacity is a consequence of the model definition and the large-deviation computation, not a fit renamed as a prediction. The basin formula (40)-(41) likewise uses the tilted large-deviation signal (38) and the same parameter-free noise rate. The one disclosure that could look circular — 'the curves labelled “typical” below use the empirically calibrated per-pattern variance and are the ones that track the data' (Section 5, along with Remark 2) — is explicitly labeled empirical, is secondary to the annealed closed forms, and the paper states that real data track the annealed branch, not the calibrated typical branch. Self-citations to [19] and to the authors' earlier multidirectional work are contextual (direct ancestor, shared protocol, related architectures) rather than load-bearing, because the present paper re-derives the needed moments, saddle points, and prefactors in its appendices. The noted Gaussian-CLT/Berry-Esseen limitation (Remark 4, Appendix A.16) is a correctness-risk statement, not a circularity: the paper openly states that the bound is silent in the sub-critical regime and does not claim a rigorous sub-critical CLT control. Overall, the derivation does not reduce to its inputs or to a fitted parameter under another name.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central derivation assumes independent Rademacher layer datasets and a one-step Gaussian signal-to-noise criterion. No free parameters enter the rate ρ_L; the only fitted quantity in the paper is the 'typical' per-pattern variance used for some comparison curves. No new physical entities are postulated.

free parameters (1)
  • typical-branch per-pattern variance (empirically calibrated) = not reported numerically
    Remark 2 / Section 5: curves labelled 'typical' use an empirically calibrated per-pattern variance to track data; the central annealed closed forms do not use it.
assumptions (5)
  • domain assumption Layer datasets are mutually independent Rademacher patterns across layers, sites and indices.
    Model definition, Eq. (3) and footnote 1; drives the whole saddle-point analysis.
  • domain assumption Stored maps are restricted to surjective functions of the cue.
    Remark 1 shows non-function relations are not storable; all datasets are filtered to enforce this.
  • domain assumption Zero-temperature sequential-within-layer / parallel-across-layers single-flip dynamics.
    Algorithm 1; the capacity statement certifies one-step stability, not full multi-step relaxation.
  • domain assumption The noise sum is approximated as Gaussian by the CLT in the sub-critical regime.
    Section 4 and Appendix A; the paper's own Remark 4 says the Berry–Esseen bound is silent for P≪P_c.
  • standard math Cramér's theorem and Varadhan's lemma apply to the empirical magnetisations.
    Appendix B: standard large-deviation toolkit used to evaluate non-factorising noise average.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exponential Capacity in Multilayer Hetero-Associative Neural Networks." pith.science (2026). https://pith.science/paper/XAZ3N4GZ

@misc{pith2026260729554,
  author       = {Pith},
  title        = {Pith review of: Exponential Capacity in Multilayer Hetero-Associative Neural Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XAZ3N4GZ}},
  note         = {Machine review of arXiv:2607.29554}
}
abstract

Exponential Hopfield networks store a number of patterns that grows exponentially with the number of neurons, and in their classical formulation they are auto-associative: they complete a corrupted copy of a memory into the memory itself. Many of the tasks one wants such a network to perform are instead hetero-associative, mapping a cue to a different target. We introduce and analyse an exponential neural network of $L$ layers of $N$ binary neurons, each layer carrying its own dataset, whose energy is an exponential of the product of the per-layer Mattis overlaps, so that it is minimised precisely when every layer retrieves the pattern of the same index; the stored association must be a surjective function of the cue, and we show why nothing else can be stored at all. A cavity/signal-to-noise analysis, made exact at leading order by a large-deviation evaluation of the noise, shows that the aligned hetero-associative state is a fixed point of the zero-temperature dynamics up to a number of stored patterns $P_c\sim e^{N\rho_L}$, exponential in the layer size, with an explicit rate $\rho_L$ that grows like $L\log 2$; enlarging the basins of attraction lowers the rate but never destroys its exponential character. Comparing the theory with structured data we find that the exponential capacity and the predicted basins survive correlated, many-to-one patterns: the network is a near-perfect content-addressable memory. The same closed forms describe, without refitting, a synthetic manifold, real T-cell-receptor/epitope triples and natural-language intent data, so the mechanism is domain-universal. Generalisation to unseen cues, though significantly above chance, stays below memorisation, and it is the geometry of the encoding, rather than the data domain, that sets how far above chance it reaches. In this family, exponential storage and strong generalisation are distinct capabilities.

Figures

Figures reproduced from arXiv: 2607.29554 by the authors.

Figure 1
Figure 1. Exponential storage capacity. (a) One-step overlap m (1) 1 versus the number of stored patterns P for the i.i.d. ensemble at width L = 2 and layer sizes N = 8, 9, 10, 11, 12: solid lines are the closed form (32), markers Monte-Carlo (mean ± std over disorder realisations). The degradation transition tracks Pc ∼ eNρ2 and shifts right with N; because the capacity is exponential in N, only small N brings it into an acc… view at source ↗
Figure 2
Figure 2. Basins of attraction: larger basins cost rate. (a) Basin recovery at L = 3 (N = 30, P = 3.5 × 104 , i.i.d. ensemble): one-step overlap after the dynamics versus the cue overlap r, against the corrupted-cue prediction (40) built on the annealed signal µ1(r) (solid) and on the typical signal (dashed). The measured transition sits on the annealed curve, at r ≈0.47 next to r ann c , and nowhere near the typical threshol… view at source ↗
Figure 3
Figure 3. The data manifold. (a) Pattern overlap mµν against latent similarity ρz (density) with binned means (markers) lying on the arcsine law 2 π arcsin ρz (black); inset, a 2D PCA embedding of the cue patterns coloured by target region, showing the surjective clustering. (b) The cue-overlap standard deviation crosses over from the i.i.d. value 1/ √ N at αD = 1 to a strongly correlated regime as αD → 0, tracking the manifo… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Capacity, width and surjective compression. (a) One-step overlap versus load P (N = 10) for the i.i.d. ensemble and Hidden-Manifold ensembles at several αD: capacity falls as the manifold shrinks, and past capacity the manifold state relaxes onto its structural floor (…
Figure 5
Figure 5. Figure 5: Memorisation without generalisation. (a) Memorisation, generalisation and novel-region recall versus load (K = 8 regions, two held out), with chance (dotted) and majority (dashed). (b) Excess over chance versus coverage P/nseen: generalisation grows and stays well abov…
Figure 6
Figure 6. Figure 6: From database to binary patterns. (a) The cleaning funnel: 137,484 raw records reduce to 1,052 clean triples once the function filter enforces a single-valued receptor→epitope map. (b) Epitope cluster sizes (rank–frequency): strongly surjective, up to 146 receptors per…
Figure 7
Figure 7. Figure 7: Near-perfect memory, modest generalisation on real data. (a) Recall for the biological tasks: both chains name the epitope essentially without error, each single chain only partially, and the over-determined (β, epitope) → α perfectly. (b) Basins of attraction: the rea…
Figure 8
Figure 8. Figure 8: The same network on natural language (L = 2, N = 128, CLINC150). (a) PaCMAP embedding of the encoded utterances (grey: all 150 intents). Four intents are highlighted: the semantically affine pair credit score / improve credit score, whose clusters coincide and whose ta…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 4 linked inside Pith

  1. [1]

    J. J. Hopfield, Neural networks and physical systems with emergent collective computational abilities, Proceedings of the National Academy of Sciences 79 (8) (1982) 2554–2558

  2. [2]

    D. J. Amit, H. Gutfreund, H. Sompolinsky, Spin-glass models of neural networks, Physical Review A 32 (2) (1985) 1007–1018

  3. [3]

    D. J. Amit, H. Gutfreund, H. Sompolinsky, Storing infinite numbers of patterns in a spin-glass model of neural networks, Physical Review Letters 55 (14) (1985) 1530–1533

  4. [4]

    Baldi, S

    P. Baldi, S. S. Venkatesh, Number of stable points for spin-glasses and neural networks of higher orders, Physical Review Letters 58 (9) (1987) 913–916

  5. [5]

    Gardner, Multiconnected neural network models, Journal of Physics A: Mathematical and General 20 (11) (1987) 3453

    E. Gardner, Multiconnected neural network models, Journal of Physics A: Mathematical and General 20 (11) (1987) 3453. 43

  6. [6]

    Gardner, Spin glasses with p-spin interactions, Nuclear Physics B 257 (1985) 747–765

    E. Gardner, Spin glasses with p-spin interactions, Nuclear Physics B 257 (1985) 747–765

  7. [7]

    H. Bao, R. Zhang, Y. Mao, The capacity of the dense associative memory networks, Neuro- computing 469 (2022) 198–208

  8. [8]

    Krotov, J

    D. Krotov, J. J. Hopfield, Dense associative memory for pattern recognition, in: Advances in Neural Information Processing Systems, Vol. 29, 2016, pp. 1172–1180

Show all 52 references
  1. [9]

    Krotov, J

    D. Krotov, J. J. Hopfield, Dense associative memory is robust to adversarial inputs, Neural Computation 30 (12) (2018) 3151–3167

  2. [10]

    Krotov, A new frontier for Hopfield networks, Nature Reviews Physics 5 (7) (2023) 366–367

    D. Krotov, A new frontier for Hopfield networks, Nature Reviews Physics 5 (7) (2023) 366–367

  3. [11]

    Krotov, J

    D. Krotov, J. J. Hopfield, Large associative memory problem in neurobiology and machine learning, arXiv preprint arXiv:2008.06996 (2020)

  4. [12]

    Demircigil, J

    M. Demircigil, J. Heusel, M. Löwe, S. Upgang, F. Vermet, On a model of associative memory with huge storage capacity, Journal of Statistical Physics 168 (2) (2017) 288–299

  5. [13]

    Ramsauer, B

    H. Ramsauer, B. Schäfl, J. Lehner, P. Seidl, M. Widrich, T. Adler, L. Gruber, M. Holzleitner, M. Pavlović, G. K. Sandve, V. Greiff, D. Kreil, M. Kopp, G. Klambauer, J. Brandstetter, S. Hochreiter, Hopfield networks is all you need, in: International Conference on Learning Repr...

  6. [14]

    Hoover, Y

    B. Hoover, Y. Liang, B. Pham, R. Panda, H. Strobelt, D. H. Chau, M. Zaki, D. Krotov, Energy transformers, in: Advances in Neural Information Processing Systems, Vol. 36, 2024

  7. [15]

    Lucibello, M

    C. Lucibello, M. Mézard, Exponential capacity of dense associative memories, Physical Review Letters 132 (7) (2024) 077301

  8. [16]

    Derrida, Random-energy model: An exactly solvable model of disordered systems, Physical Review B 24 (5) (1981) 2613–2626

    B. Derrida, Random-energy model: An exactly solvable model of disordered systems, Physical Review B 24 (5) (1981) 2613–2626

  9. [17]

    Dohmatob, A different route to exponential storage capacity, in: Associative Memory and Hopfield Networks in 2023 (NeurIPS Workshop), 2023

    E. Dohmatob, A different route to exponential storage capacity, in: Associative Memory and Hopfield Networks in 2023 (NeurIPS Workshop), 2023

  10. [18]

    C. J. Hillar, N. M. Tran, Robust exponential memory in Hopfield networks, Journal of Mathematical Neuroscience 8 (1) (2018) 1–20

  11. [19]

    Albanese, A

    L. Albanese, A. Alessandrelli, A. Barra, P. Sollich, Yet another exponential Hopfield model, Neural Networks 186 (2026) 131223

  12. [20]

    Agliari, F

    E. Agliari, F. Alemanno, A. Barra, A. Fachechi, Generalized Guerra’s interpolation schemes for dense associative neural networks, Neural Networks 128 (2020) 254–267

  13. [21]

    Barra, M

    A. Barra, M. Beccaria, A. Fachechi, A new mechanical approach to handle generalized Hopfield neural networks, Neural Networks 106 (2018) 205–222

  14. [22]

    Albanese, A

    L. Albanese, A. Alessandrelli, A. Barra, et al., Hebbian learning from first principles, Journal of Mathematical Physics 65 (2024) 113302

  15. [23]

    M. S. Centonze, I. Kanter, A. Barra, Statistical mechanics of learning via reverberation in bidirectional associative memories, Physica A: Statistical Mechanics and its Applications 637 (2024) 129512

  16. [24]

    Agliari, A

    E. Agliari, A. Alessandrelli, A. Barra, M. S. Centonze, F. Ricci-Tersenghi, Generalized hetero- associative neural networks, Journal of Statistical Mechanics: Theory and Experiment 2025 (1) (2025) 013302

  17. [25]

    Alessandrelli, A

    A. Alessandrelli, A. Barra, A. Ladiana, A. Lepre, F. Ricci-Tersenghi, Supervised and unsuper- vised protocols for hetero-associative neural networks, Physica A: Statistical Mechanics and its Applications (2025) 130871ArXiv:2505.18796

  18. [26]

    Goldt, M

    S. Goldt, M. Mézard, F. Krzakala, L. Zdeborová, Modeling the influence of data structure on learning in neural networks: The hidden manifold model, Physical Review X 10 (4) (2020) 041044. 44

  19. [27]

    Gerace, B

    F. Gerace, B. Loureiro, F. Krzakala, M. Mézard, L. Zdeborová, Generalisation error in learning with random features and the hidden manifold model, in: International Conference on Machine Learning, 2020, pp. 3452–3462

  20. [28]

    Shugay, D

    M. Shugay, D. V. Bagaev, I. V. Zvyagin, R. M. A. Vroomans, J. C. Crawford, G. Dolton, E. A. Komech, A. L. Sycheva, A. E. Koneva, E. S. Egorov, et al., VDJdb: a curated database of T-cell receptor sequences of known antigen specificity, Nucleic Acids Research 46 (D1) (2018) D419–D427

  21. [29]

    D. V. Bagaev, R. M. A. Vroomans, J. Samir, U. Stervbo, C. Rius, G. Dolton, A. Greenshields- Watson, M. Attaf, E. S. Egorov, I. V. Zvyagin, et al., VDJdb in 2019: database extension, new analysis infrastructure and a T-cell receptor motif compendium, Nucleic Acids Research 48 (...

  22. [30]

    W. R. Atchley, J. Zhao, A. D. Fernandes, T. Drüke, Solving the protein sequence metric problem, Proceedings of the National Academy of Sciences 102 (18) (2005) 6395–6400

  23. [31]

    M. S. Charikar, Similarity estimation techniques from rounding algorithms, in: Proceedings of the 34th Annual ACM Symposium on Theory of Computing (STOC), 2002, pp. 380–388

  24. [32]

    Larson, A

    S. Larson, A. Mahendran, J. J. Peper, C. Clarke, A. Lee, P. Hill, J. K. Kummerfeld, K. Leach, M. A. Laurenzano, L. Tang, J. Mars, An evaluation dataset for intent classification and out-of-scope prediction, in: Proceedings of the 2019 Conference on Empirical Methods in Natural...

  25. [33]

    Negri, C

    M. Negri, C. Lauditi, G. Perugini, C. Lucibello, E. Malatesta, Storage and learning phase transitions in the random-features Hopfield model, Physical Review Letters 131 (25) (2023) 257301

  26. [34]

    Kalaj, C

    S. Kalaj, C. Lauditi, G. Perugini, C. Lucibello, E. M. Malatesta, M. Negri, Random fea- tures Hopfield networks generalize retrieval to previously unseen examples, arXiv preprint arXiv:2407.05658 (2024)

  27. [35]

    Onsager, Electric moments of molecules in liquids, Journal of the American Chemical Society 58 (8) (1936) 1486–1493

    L. Onsager, Electric moments of molecules in liquids, Journal of the American Chemical Society 58 (8) (1936) 1486–1493

  28. [36]

    Mézard, G

    M. Mézard, G. Parisi, M. A. Virasoro, SK model: the replica solution without replicas, Europhysics Letters 1 (2) (1986) 77–82

  29. [37]

    Guerra, The cavity method in the mean field spin glass model

    F. Guerra, The cavity method in the mean field spin glass model. Functional representations of thermodynamic variables, in: Advances in Dynamical Systems and Quantum Physics, World Scientific, Singapore, 1995, pp. 141–156

  30. [38]

    Barra, Irreducible free energy expansion and overlaps locking in mean field spin glasses, Journal of Statistical Physics 123 (3) (2006) 601–614

    A. Barra, Irreducible free energy expansion and overlaps locking in mean field spin glasses, Journal of Statistical Physics 123 (3) (2006) 601–614

  31. [39]

    Mézard, G

    M. Mézard, G. Parisi, M. A. Virasoro, Spin Glass Theory and Beyond, World Scientific, 1987

  32. [40]

    Pastur, M

    L. Pastur, M. Shcherbina, B. Tirozzi, On the replica symmetric equations for the Hopfield model, Journal of Mathematical Physics 40 (8) (1999) 3930–3947

  33. [41]

    Mézard, Mean-field message-passing equations in the Hopfield model and its generalizations, Physical Review E 95 (2) (2017) 022117

    M. Mézard, Mean-field message-passing equations in the Hopfield model and its generalizations, Physical Review E 95 (2) (2017) 022117

  34. [42]

    D. J. Amit, Modeling Brain Function: The World of Attractor Neural Networks, Cambridge University Press, Cambridge, 1989

  35. [43]

    A. C. C. Coolen, R. Kühn, P. Sollich, Theory of Neural Information Processing Systems, Oxford University Press, 2005

  36. [44]

    S. R. S. Varadhan, Asymptotic probabilities and differential equations, Communications on Pure and Applied Mathematics 19 (3) (1966) 261–286

  37. [45]

    Dembo, O

    A. Dembo, O. Zeitouni, Large Deviations Techniques and Applications, 2nd Edition, Springer, New York, 1998. 45

  38. [46]

    Alessandrelli, F

    A. Alessandrelli, F. Durante, A. Ladiana, A. Lepre, A federated many-to-one Hopfield model for associative neural networks, arXiv preprint arXiv:2603.19902 (2026)

  39. [47]

    Ladiana, Finite-size scaling of hetero-associative retrieval in continuous-signal-driven Ising spin systems, arXiv preprint arXiv:2605.14059 (2026)

    A. Ladiana, Finite-size scaling of hetero-associative retrieval in continuous-signal-driven Ising spin systems, arXiv preprint arXiv:2605.14059 (2026)

  40. [48]

    Agliari, A

    E. Agliari, A. Barra, A. Ladiana, A. Lepre, Thermodynamic binding: Freezing chimeric states in multi-modal associative memories, in: New Frontiers in Associative Memories – Workshop at ICLR, 2026

  41. [49]

    Fachechi, E

    A. Fachechi, E. Agliari, A. Barra, Dreaming neural networks: forgetting spurious memories and reinforcing pure ones, Neural Networks 112 (2019) 24–40

  42. [50]

    Agliari, F

    E. Agliari, F. Alemanno, A. Barra, A. Fachechi, Dreaming neural networks: rigorous results, Journal of Statistical Mechanics: Theory and Experiment 2019 (8) (2019) 083503

  43. [51]

    Barra, F

    A. Barra, F. Durante, A. Ladiana, M. M. Solazzo, Do Hopfield networks dream of stored patterns? A statistical-mechanical theory of dreaming in multidirectional associative memories, arXiv preprint arXiv:2605.13721 (2026)

  44. [52]

    Glanville, H

    J. Glanville, H. Huang, A. Nau, O. Hatton, L. E. Wagar, F. Rubelt, X. Ji, A. Han, S. M. Krams, C. Pettus, et al., Identifying specificity groups in the T-cell receptor repertoire, Nature 547 (7661) (2017) 94–98. 46

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.