Pith. sign in

REVIEW 3 major objections 3 minor 54 references

A Spin Glass Characterization of Neural Networks

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A single trained network can be characterized by the replica overlap statistics of an associated Hopfield-type spin glass.

desk verdict A per-instance replica-overlap spin-glass descriptor is a genuinely new idea, but the available text is only the abstract; send the full manuscript to a serious referee if it delivers the mapping and empirical checks. read the letter →

arxiv 2508.07397 v1 pith:37HGCW4N submitted 2025-08-10 cond-mat.dis-nn cs.AIcs.LG

classification cond-mat.dis-nncs.AIcs.LG
keywords spinglassreplicasymmetrybreakingHopfieldmodelfeedforwardneuralnetworkgeneralizationrobustnessinspection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes that a single trained feedforward neural network can be converted into a Hopfield-type spin glass, and that the overlap statistics of simulated replicas of this glass form a characteristic descriptor of that specific network. The descriptor is claimed to track the network's data fitting, capacity, generalization, and robustness, and to expose structural properties that loss and accuracy do not capture. If true, this gives a way to inspect individual networks rather than ensembles, with potential use in model auditing and safety verification.

What carries the argument

The central object is the Hopfield-type spin glass constructed from the trained FNN, together with the set of simulated replica samples used to compute overlaps. In this construction, the network's parameters are encoded as couplings of a spin system, and the overlap between two replica configurations measures how similar two equilibrium samples of that glass are. The replica-overlap distribution serves as the characteristic descriptor that is correlated with the network's properties.

What would settle it

Compare replica-overlap descriptors for two networks that have identical loss and accuracy on the training and test sets but differ in some hidden vulnerability (e.g., one is easily fooled by adversarial perturbations and the other is not). If their overlap statistics are indistinguishable, the descriptor fails to detect the structural difference it claims to expose. Alternatively, vary the construction's arbitrary choices (spin encoding, temperature, simulation length); if the correlation with robustness vanishes under mild variation, the effect is an artifact of the mapping rather than a pro

Watch

Extended reading notes

Core claim

The central claim is that the mapping from a trained FNN to a Hopfield-type spin glass yields a computable instance-level fingerprint: after constructing the spin glass from the network's parameters, one simulates multiple replica samples and computes their pairwise overlaps. These replica overlaps are claimed to be a characteristic descriptor of the FNN, in that they vary systematically with how well the network fits its data, its capacity, its generalization performance, and its robustness. The method is explicitly positioned against analytical studies of model ensembles; the novelty is that the descriptor applies to an individual trained network and reveals properties invisible to convent

Load-bearing premise

The load-bearing premise is that the way the network is mapped into a Hopfield-type spin glass keeps the computationally relevant structure of the network intact, and that the simulated replicas equilibrate so the overlap statistics are stable and meaningful rather than artifacts of the construction.

Editorial extensions

If this is right

  • The descriptor marks a shift from ensemble-level statistical mechanics to instance-level diagnostics: any single trained network can be characterized on its own.
  • Because the overlap statistics are claimed to expose structure not captured by loss or accuracy, they can distinguish networks that conventional metrics treat as equal.
  • The paper reports empirical correlations with data fitting, capacity, generalization, and robustness, making the descriptor a candidate early-warning signal for network performance.
  • The stated practical direction is that the method can support model inspection, safety verification, and detection of hidden vulnerabilities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the overlap descriptor is stable under reasonable choices of temperature and spin encoding, it could serve as a coarse-grained 'state variable' for a network, enabling phase-diagram-like maps of training dynamics.
  • The method might be extended to recurrent or convolutional architectures by specifying analogous coupling constructions, which would broaden its applicability beyond feedforward networks.
  • A natural stress test: compare overlap statistics of networks trained on the same data with different random seeds; if the descriptor varies widely across seeds, it may be capturing idiosyncratic initialization rather than generalizable structure.
  • One could attempt to use the descriptor as a training objective or regularizer, e.g., steering a network toward a target overlap profile, though the paper does not propose this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper claims to construct a Hopfield-type spin glass model from a given feedforward neural network and to use overlaps between simulated replica samples as a characteristic descriptor that tracks data fitting, capacity, generalization, and robustness. The abstract states that this connection is empirically demonstrated and suggests applications to model inspection, safety verification, and vulnerability detection. The submitted manuscript, however, contains only the abstract: no equations, no construction of the spin-glass Hamiltonian, no description of the replica simulation protocol, no datasets, no experimental results, and no comparison against conventional metrics. The central claim is therefore asserted rather than supported.

Significance. The proposed idea is potentially interesting: an instance-level spin-glass fingerprint would go beyond existing ensemble-level replica analyses and could give a computable structural descriptor for a single trained network. If the mapping from network to spin glass is non-arbitrary and the replica overlaps are equilibrated, the claimed correlations with generalization and robustness would be a useful addition to the statistical-mechanics-of-neural-networks literature. However, as submitted, the manuscript provides no evidence that the descriptor is well-defined, stable, or more informative than loss and accuracy. The conceptual direction is worth exploring, but the present text is not a verifiable research contribution.

major comments (3)
  1. [Abstract] The central construction is unspecified. The abstract states that 'A Hopfield-type spin glass model is constructed from a given FNN' but gives no Hamiltonian, no spin encoding (binary, continuous, or otherwise), no rule for converting network weights into couplings, and no specification of the replica temperature or equilibration criterion. Without these details, 'overlaps between simulated replica samples' is undefined. This is load-bearing: the claimed descriptor cannot be evaluated, and any correlation with network properties could depend on arbitrary choices in the mapping or on out-of-equilibrium sampling.
  2. [Abstract] The claim that the connection to fitting, capacity, generalization, and robustness is 'empirically demonstrated' is not supported by any reported evidence. The manuscript contains no datasets, architectures, training protocols, baselines, error bars, or quantitative comparisons with loss and accuracy. The phrase 'Preliminary results suggests' is a hedge, not a demonstration. A journal submission must include a reproducibility-ready experimental section, and its absence is a decisive shortcoming for the central empirical claim.
  3. [Abstract] There is a risk of circularity: the descriptor is computed from a trained network and then correlated with properties of that same network. The manuscript must show that the overlaps do not simply re-encode loss or accuracy. At minimum, this requires demonstrating invariance under function-preserving transformations (e.g., hidden-unit permutations) and reporting convergence of overlap estimates across independent replica simulations. None of these checks are described.
minor comments (3)
  1. [Abstract] Typo: 'Preliminary results suggests' should be 'Preliminary results suggest'.
  2. [Abstract] The phrase 'statistical mechanics characterization' is vague; the manuscript should define what is meant by a characterization in terms of spin-glass observables.
  3. [Abstract] The relation to 'prior analytical studies that focus on model ensembles' should be supported with specific references, since the claimed novelty is instance-level computability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identifiable from the available text; the abstract proposes an empirical descriptor without equations that could exhibit a reduction.

full rationale

The manuscript as provided contains only the abstract; the full text has no equations, no Hamiltonian, no spin encoding, no temperature schedule, no fitting procedure, and no explicit prediction identities. Under the governing rules, circularity may be claimed only when the paper itself exhibits a specific reduction, e.g. a fitted parameter later renamed as a prediction or a definition that imports the target result. The abstract states that a Hopfield-type spin glass model is constructed from a given FNN and that replica overlaps serve as a characteristic descriptor, and it reports an empirical connection to generalization and robustness. This is a proposed computational characterization, not a derivation whose output equals its input by construction. The absence of a specified Hamiltonian or equilibration checks is an evidential incompleteness that could undermine the empirical claims, but it is not itself circularity. No self-citation or imported uniqueness theorem appears in the provided text. Therefore the honest finding is no significant circularity, score 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 1 invented entities

Because only the abstract was available, every ledger entry is inferred from the abstract's wording. The central pattern is that the paper rests on an asserted weight-to-spin-glass mapping and an asserted link from replica overlaps to generalization, with the empirical generality of that link assumed from 'preliminary results'.

free parameters (3)
  • replica simulation temperature (thermal noise level)
    Spin glass overlaps are computed from a Gibbs distribution that requires a temperature. The abstract never states how this temperature is chosen or whether the reported correlations are stable across temperatures.
  • replica count and sampling protocol
    Overlap statistics depend on the number of simulated replicas and the Monte Carlo schedule. These hyperparameters shape the descriptor and are not reported in the abstract.
  • scaling/discretization in the FNN-to-spin-glass mapping
    Mapping continuous network weights to spin couplings involves a discretization or normalization choice. Any such hand-set choice is a free parameter of the descriptor construction, invisible from the abstract.
assumptions (3)
  • domain assumption A trained FNN can be represented as a Hopfield-type spin glass whose equilibrium replica overlap statistics are meaningful descriptors of the network's computational properties.
    This is the paper's central modeling step as stated in the abstract ('A Hopfield-type spin glass model is constructed from a given feedforward neural network'). It is asserted, not derived, and no construction details are visible.
  • domain assumption Simulated replica samples reach equilibrium and their overlaps are statistically stable across runs.
    Replica overlap is only a valid descriptor if sampling is equilibrated and reproducible; the abstract provides no Monte Carlo protocol, temperature, or convergence analysis.
  • domain assumption The empirical connections between the descriptor and data fitting, capacity, generalization, and robustness generalize beyond the particular networks and datasets tested.
    The abstract reports 'preliminary results' of empirical demonstration; without visible dataset breadth or error analysis, the generality of the reported connections is assumed.
invented entities (1)
  • replica overlap descriptor
    purpose: A per-instance statistical fingerprint of a trained FNN, computed as overlaps between simulated replicas of the network-derived spin glass, intended to summarize fitting, capacity, generalization, and robustness.
    The descriptor is manufactured by the paper's own construction and carries no falsifiable handle outside the paper; its value is validated only by the in-paper correlations it reports.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Spin Glass Characterization of Neural Networks." pith.science (2026). https://pith.science/paper/37HGCW4N

@misc{pith2026250807397,
  author       = {Pith},
  title        = {Pith review of: A Spin Glass Characterization of Neural Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/37HGCW4N}},
  note         = {Machine review of arXiv:2508.07397}
}
read the original abstract

This work presents a statistical mechanics characterization of neural networks, motivated by the replica symmetry breaking (RSB) phenomenon in spin glasses. A Hopfield-type spin glass model is constructed from a given feedforward neural network (FNN). Overlaps between simulated replica samples serve as a characteristic descriptor of the FNN. The connection between the spin-glass description and commonly studied properties of the FNN -- such as data fitting, capacity, generalization, and robustness -- has been investigated and empirically demonstrated. Unlike prior analytical studies that focus on model ensembles, this method provides a computable descriptor for individual network instances, which reveals nontrivial structural properties that are not captured by conventional metrics such as loss or accuracy. Preliminary results suggests its potential for practical applications such as model inspection, safety verification, and detection of hidden vulnerabilities.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

54 extracted references · 48 canonical work pages

  1. [1]

    Amit, Hanoch Gutfreund, and Haim Sompolinsky

    Daniel J. Amit, Hanoch Gutfreund, and Haim Sompolinsky. Spin-glass models of neural networks. Physical Review A , 32(2):1007--1018, 1985

  2. [2]

    Symmetry & Critical Points

    Yossi Arjevani. Symmetry and critical points. ArXiv , abs/2408.14445, 2024

  3. [3]

    Complexity of random smooth functions on the high-dimensional sphere

    Antonio Auffinger and Gérard Ben Arous. Complexity of random smooth functions on the high-dimensional sphere. The Annals of Probability , 41(6):4214--4247, 2013

  4. [4]

    Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli

    Yasaman Bahri, Jonathan Kadmon, Jeffrey Pennington, Sam S. Schoenholz, Jascha Sohl-Dickstein, and Surya Ganguli. Statistical mechanics of deep learning. Annual Review of Condensed Matter Physics , 11:501--528, 2020

  5. [5]

    Curtis, and Jorge Nocedal

    Léon Bottou, Frank E. Curtis, and Jorge Nocedal. Optimization methods for large-scale machine learning. SIAM Review , 60(2):223--311, 2018

  6. [6]

    Bray and David S

    Alan J. Bray and David S. Dean. Statistics of critical points of gaussian fields on large-dimensional spaces. Physical Review Letters , 98(15):150201, 2007

  7. [7]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...

  8. [8]

    Quantum langevin dynamics for optimization

    Zherui Chen, Yuchen Lu, Hao Wang, Yizhou Liu, and Tongyang Li. Quantum langevin dynamics for optimization. Communications in Mathematical Physics , 406(3):52, 2025

Show all 54 references
  1. [9]

    Landscape analysis for shallow neural networks: Complete classification of critical points for affine target functions

    Patrick Cheridito, Arnulf Jentzen, and Florian Rossmannek. Landscape analysis for shallow neural networks: Complete classification of critical points for affine target functions. Journal of Nonlinear Science , 32(5):64, 2022

  2. [10]

    The loss surfaces of multilayer networks

    Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun. The loss surfaces of multilayer networks. In Proceedings of the 18th International Conference on Artificial Intelligence and Statistics (AISTATS) , volume 38, pages 192--204. PMLR, 2015

  3. [11]

    Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio

    Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio. Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In Advances in Neural Information Processing Systems 28 (NeurIPS 2014) , pages 293...

  4. [12]

    Unifying grokking and double descent

    Xander Davies, Lauro Langosco, and David Krueger. Unifying grokking and double descent. arXiv preprint arXiv:2303.06173 , 2023

  5. [13]

    Towards a mathematical understanding of neural network-based machine learning: What we know and what we don't

    Weinan E, Chao Ma, Lei Wu, and Stephan Wojtowytsch. Towards a mathematical understanding of neural network-based machine learning: What we know and what we don't. CSIAM Transactions on Applied Mathematics , 1(4):561--615, 2020

  6. [14]

    Engel and C

    A. Engel and C. Van den Broeck. Statistical Mechanics of Learning . Cambridge University Press, 2001

  7. [15]

    Entropy and mutual information in models of deep neural networks

    Marylou Gabri \`e , Andre Manoel, Cl \'e ment Luneau, Jean Barbier, Nicolas Macris, Florent Krzakala, and Lenka Zdeborov \'a . Entropy and mutual information in models of deep neural networks. In Advances in Neural Information Processing Systems 31 (NeurIPS 2018) , pages 1821-...

  8. [16]

    The space of interactions in neural network models

    Elizabeth Gardner. The space of interactions in neural network models. Journal of Physics A: Mathematical and General , 21(1):257--270, 1988

  9. [17]

    Roy J. Glauber. Time-dependent statistics of the ising model. Journal of Mathematical Physics , 4(2):294--307, 1963

  10. [18]

    Flat minima

    Sepp Hochreiter and J \"u rgen Schmidhuber. Flat minima. Neural Computation , 9(1):1--42, 1997

  11. [19]

    Hopfield

    John J. Hopfield. Neural networks and physical systems with emergent collective computational abilities. Proceedings of the National Academy of Sciences , 79(8):2554--2558, 1982

  12. [20]

    I. T. Jolliffe. Principal Component Analysis . Springer Series in Statistics. Springer, New York, 2nd edition, 2002

  13. [21]

    Schmidt, and Michael Riis Andersen

    Mikkel Jordahn, Jonas Vestergaard Jensen, Mikkel N. Schmidt, and Michael Riis Andersen. On local posterior structure in deep ensembles, 2025

  14. [22]

    Inference from correlated patterns: a unified theory for perceptron learning and linear vector channels

    Yoshiyuki Kabashima. Inference from correlated patterns: a unified theory for perceptron learning and linear vector channels. Journal of Physics: Conference Series , 95, 2008

  15. [23]

    mingpt: A minimal pytorch re-implementation of gpt

    Andrej Karpathy. mingpt: A minimal pytorch re-implementation of gpt. https://github.com/karpathy/minGPT, 2020. GitHub repository

  16. [24]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 , 2014

  17. [25]

    Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll \'a r, and Ross Girshick. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vis...

  18. [26]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009

  19. [27]

    Deep learning

    Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. Nature , 521(7553):436--444, 2015

  20. [28]

    Gradient-based learning applied to document recognition

    Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE , 86(11):2278--2324, 1998

  21. [29]

    Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein

    Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein. Deep neural networks as gaussian processes. In 6th International Conference on Learning Representations (ICLR) , 2018

  22. [30]

    Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations

    Qianxiao Li, Cheng Tai, and Weinan E. Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations. Journal of Machine Learning Research , 20(40):1--47, 2019

  23. [31]

    Hoffman, and David M

    Stephan Mandt, Matthew D. Hoffman, and David M. Blei. Stochastic gradient descent as approximate bayesian inference. Journal of Machine Learning Research , 18(134):1--35, 2017

  24. [32]

    Information, Physics, and Computation

    Marc M \'e zard and Andrea Montanari. Information, Physics, and Computation . Oxford Graduate Texts. Oxford University Press, 2009

  25. [33]

    Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications , volume 9 of World Scientific Lecture Notes in Physics

    Marc M\' e zard, Giorgio Parisi, and Miguel Angel Virasoro. Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications , volume 9 of World Scientific Lecture Notes in Physics . World Scientific, 1987

  26. [34]

    Bridging lottery ticket and grokking: Understanding grokking from inner structure of networks

    Gouki Minegishi, Yusuke Iwasawa, and Yutaka Matsuo. Bridging lottery ticket and grokking: Understanding grokking from inner structure of networks. Transactions on Machine Learning Research , 2025

  27. [35]

    Radford M. Neal. Bayesian Learning for Neural Networks , volume 118 of Lecture Notes in Statistics . Springer, 1996

  28. [36]

    Pytorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K \"o pf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner,...

  29. [37]

    Exponential expressivity in deep neural networks through transient chaos

    Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli. Exponential expressivity in deep neural networks through transient chaos. In Advances in Neural Information Processing Systems 29 (NeurIPS 2016) , pages 3368--3376, 2016

  30. [38]

    On the expressive power of deep neural networks

    Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein. On the expressive power of deep neural networks. In Doina Precup and Yee Whye Teh, editors, Proceedings of the 34th International Conference on Machine Learning , volume 70 of Proceedings of Mac...

  31. [39]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. arXiv preprint arXiv:2112.10752 , 2022

  32. [40]

    Singularity of the H essian in deep learning

    Levent Sagun, L \' e on Bottou, and Yann LeCun. Singularity of the H essian in deep learning. In Proceedings of the 5th International Conference on Learning Representations (ICLR) , 2017

  33. [41]

    Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein

    Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein. Deep information propagation. In Proceedings of the 5th International Conference on Learning Representations (ICLR) , 2017

  34. [42]

    What Is Life? The Physical Aspect of the Living Cell

    Erwin Schr \"o dinger. What Is Life? The Physical Aspect of the Living Cell . Cambridge University Press, Cambridge, 1944. Based on lectures delivered at Trinity College Dublin in February 1943

  35. [43]

    Solvable model of a spin-glass

    David Sherrington and Scott Kirkpatrick. Solvable model of a spin-glass. Physical Review Letters , 35(26):1792--1796, 1975

  36. [44]

    Su, and Michael I

    Bin Shi, Weijie J. Su, and Michael I. Jordan. On learning rates and schrödinger operators. Journal of Machine Learning Research , 24(379):1--53, 2023

  37. [45]

    Mean Field Models for Spin Glasses: Volume I: Basic Examples , volume 54 of Ergebnisse der Mathematik und ihrer Grenzgebiete

    Michel Talagrand. Mean Field Models for Spin Glasses: Volume I: Basic Examples , volume 54 of Ergebnisse der Mathematik und ihrer Grenzgebiete. 3. Folge / A Series of Modern Surveys in Mathematics . Springer, 2011

  38. [46]

    Opening the black box: Predicting the trainability of deep neural networks with reconstruction entropy

    Yanick Thurn, Ro Jefferson, and Johanna Erdmenger. Opening the black box: Predicting the trainability of deep neural networks with reconstruction entropy. arXiv preprint arXiv:2406.12916 , 2024

  39. [47]

    Pereira, and William Bialek

    Naftali Tishby, Fernando C. Pereira, and William Bialek. The information bottleneck method. In Proceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing , pages 368--377, 1999

  40. [48]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems , volume 30, pages 5998--6008, 2017

  41. [49]

    Tent: Fully test-time adaptation by entropy minimization

    Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. Tent: Fully test-time adaptation by entropy minimization. In International Conference on Learning Representations (ICLR) , 2021

  42. [50]

    Schoenholz, and Jeffrey Pennington

    Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S. Schoenholz, and Jeffrey Pennington. Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks. arXiv preprint arXiv:1806.05393 , 2018. Published in ICML 2018

  43. [51]

    Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions

    Ning Yang, Chao Tang, and Yuhai Tu. Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions. Physical Review Letters , 130(23):237101, 2023

  44. [52]

    Statistical physics of inference: Thresholds and algorithms

    Lenka Zdeborov \'a and Florent Krzakala. Statistical physics of inference: Thresholds and algorithms. Advances in Physics , 65(5):453--552, 2016

  45. [53]

    Understanding deep learning requires rethinking generalization

    Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. In Proceedings of the International Conference on Learning Representations (ICLR) , 2017

  46. [54]

    Edge of chaos as a guiding principle for modern neural network training

    Lin Zhang, Ling Feng, Kan Chen, and Choy Heng Lai. Edge of chaos as a guiding principle for modern neural network training. arXiv preprint arXiv:2107.09437 , 2021

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.