Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:46.143996Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2505.23489.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:46.143996Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
79 of 79 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2c5405bf-6824-4880-a578-313cde9c8185 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training TherML: Thermodynamics of Machine Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3243b1bb-8435-4704-81a1-041a34cd44dd · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training SGD with Large Step Sizes Learns Sparse Features
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 517aaa9f-c710-4dc6-b5a0-076d38697fad · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdbb3af-20d8-42ed-897c-9cb699d81e6e · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Implicit gradient regularization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b82d6ee-f178-472c-ba2f-663d27710d13 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Science, 116(32):15849–15854, 2019
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 288f3599-7164-4196-a634-048bf1de99ce · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Practical recommendations for gradient-based training of deep architectures
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fefda19d-bf62-476f-9342-a81d024e532b · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Language Models are Few-Shot Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3849048-9ae6-4931-9f0a-257e1a7a22f1 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic gradient descent performs variational infer- ence, converges to limit cycles for deep networks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa6adc37-e9a1-4333-81b1-036b7b5ec6e9 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Entropy-SGD: Biasing gradient descent into wide valleys
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3b06684-69c2-45c8-8162-1cc2a6d9f69b · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Convergence diagnostics for stochastic gradient descent with constant learning rate
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 961ad5b2-8ec7-4673-bfc3-d1623dc07ab7 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc138851-290c-4b68-8622-d741abb879fc · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic collapse: How gra- dient noise attracts SGD dynamics towards simpler subnetworks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7422fef2-f593-4b54-942f-20077021b893 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Symbolic discovery of optimization algorithms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16d05eca-ae36-4b1d-8361-dbe6d71bc8d8 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Gradient descent on neural networks typically occurs at the edge of stability
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b1a750f-4c4b-471b-ad38-a82be506ce70 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training URL https://constructor.tech/products/ research-platform
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 966fca58-2736-4112-81f1-87aba74e8ed3 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Determining intrinsic dimension and entropy of high-dimensional shape spaces.Modeling and Simulation in Science, Engineering and Technology, pages 231–252,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a6b142d-03bf-4d07-b6f5-f8c36e64625b · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Why do we need weight decay in modern deep learning? InAdvances in Neural Information Processing Systems, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91df84ab-a8df-416c-aae8-7be012c5531e · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training BERT: Pre-training of deep bidirectional transformers for language understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0918aa1a-5731-47e4-8aa1-ed4a0a3bc1fa · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Essentially no barriers in neural network energy landscape
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a19d1065-d8b1-46db-8f53-43c663845821 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76675590-5794-4a6a-8824-09adfc9074bd · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training A free-energy principle for representation learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8a40732-1810-4aba-8157-d1c6b296f298 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Fixed-time stable gradient flows: Applications to continuous- time optimization.IEEE Transactions on Automatic Control, 66(5):2002–2015, 2021
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b483a68-5559-4e66-8690-821feaafcd8d · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 576d4845-946c-43f6-86f6-1da01bcffbb9 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic training is not necessary for generalization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e52ca576-43c7-40b2-81f5-4ab1eecbe448 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Abrupt learning in transformers: A case study on matrix completion
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93369438-8a66-4d8b-96bc-4e3ecda53feb · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep Residual Learning for Image Recognition
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73431ffa-d71a-4c42-8e26-5b6a9fba3ed9 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Three Factors Influencing Minima in SGD
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c204db-ca40-4720-a820-5da35deaa3fd · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Scaling Laws for Neural Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2015c824-c544-4b57-b3cf-11d281e975b5 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Kingma and Jimmy Ba
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 231877f2-6a5b-4758-aa3a-d070439aed62 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Training scale-invariant neural networks on the sphere can happen in three regimes
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d1962e9-77e7-4320-8ddd-65a345c5c553 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Big Transfer (BiT): General Visual Representation Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee4deaa-30ca-43ae-a0b2-0d42a487b4d4 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training CIFAR-10 (canadian institute for advanced research)
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32b601c5-e9d5-4cd7-9a81-e58de55d9de6 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training CIFAR-100 (canadian institute for advanced research)
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19d3eb95-9c45-4b8a-be47-56a9d7b1d7f9 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8891303-7c11-46e5-a72a-ea91e4b653c2 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca1d1278-3582-4bfe-94dc-f0c1cdf38cfe · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Few-shot adaptation of multi-modal foundation models: A survey.Artificial Intelli- gence Review, 57(10):268, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f575c99c-1242-4d96-af71-e05728df534e · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Understanding why neural networks generalize well through GSNR of parameters
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8241ca34-06a4-40c1-914d-6141f0766561 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c6e4e48-6a9e-4715-b430-90b2723e4a71 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Towards understanding grokking: An effective theory of representation learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 827a797e-dcfe-4010-befe-06364c2ae4b7 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the periodic behavior of neural network training with batch normalization and weight decay
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebebddeb-4929-498a-92c9-45e220a39e4c · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Decoupled weight decay regularization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 828f5cd0-057f-4141-a457-5f9059fdb4be · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d1edcc-38ab-401e-b0ef-e3fa9d5b9094 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic Gradient Descent as Approximate Bayesian Inference
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbff6cb6-7569-4135-971f-733035b16bc0 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Phase transitions in the mini-batch size for sparse and dense two-layer neural networks.Machine Learning: Science and Technology, 5(1): 015015, 2024
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3d4a475-5726-4fab-a6a2-aa39e708545e · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic Gradient Descent on Separable Data: Exact Convergence with a Fixed Learning Rate
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c781f22f-abda-448a-8f2d-7ad84fdf099d · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Bayesian Free Energy of Deep ReLU Neural Network in Overparametrized Cases
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56031d6d-5de4-4f21-af93-d6af4d27f522 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep double descent: Where bigger models and more data hurt
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ffc897f-72ea-47a6-bea6-e32e919e7dbe · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training LR0.FM: Low-resolution zero-shot classification benchmark for foundation models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1de94719-1f6a-4d71-bf51-888bc13cb8d8 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Loss landscape: SGD has a better view
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d099425f-38a9-4378-810d-088fb2301e2c · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f347c5-f46a-48a5-9ea0-de6afe4c3d50 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Accelerating Large Batch Training via Gradient Signal to Noise Ratio (GSNR)
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4970ed-0469-497f-8a72-61ba9cb8f5d3 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Learning transferable visual models from natural language supervision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c01f840f-b890-48a8-9bd6-8d601ae36896 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Where do large learning rates lead us? InAdvances in Neural Information Processing Systems, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e38b7c4-cdf5-48af-8f49-a64a26b5ae39 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the different regimes of stochastic gradient descent
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bff367d-c3eb-4f2e-9e81-ef15edefd85c · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d61f117-34a8-45cb-ae16-2c4881bc476d · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the Generalization Benefit of Noise in Stochastic Gradient Descent
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6fec504-538a-4709-8a10-96d443fc7f7d · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training On the origin of implicit regular- ization in stochastic gradient descent
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53d40a8b-67e4-44b8-a71e-b3fa1970713e · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.Transactions on Machine Learning Research, 2023
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f553db86-34a5-4eb8-bc03-012018789d0a · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unleashing the power of gradient signal-to-noise ratio for zero-shot NAS
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328c658d-3707-4117-8e0a-39fdf8de6db0 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep learning and the information bottleneck principle,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6382461-7c03-4953-8b2a-bf1fc8f2f61e · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training The discovery of superconductivity.Physics Today, 63(9):38–43,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c81ca0d2-a013-4afd-9237-66511ad32f92 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training A survey of basic thermodynamics, 2004
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a41a7fe5-488a-4fe5-99fa-cfca3b4be752 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e7e1725-9700-4565-b8c1-8edda39592b0 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Bayesian learning via stochastic gradient langevin dynamics
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d38b50ff-745b-4cc9-a359-5708d3fc7698 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Towards few- shot adaptation of foundation models via multitask finetuning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc4b5a34-80a4-473f-af50-29d8dd8d6472 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Fluctuation-dissipation relations for stochastic gradient descent
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation faf4f39e-dd6a-4079-a3d0-c9e1dd180e17 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training How Does Learning Rate Decay Help Modern Neural Networks?
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3fab21-76e1-4d7e-ba52-3c684d880759 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Saxe, Madhu S
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a29911-f563-4fdb-99d5-79ca2eeb6d7f · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Strength of minibatch noise in SGD
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23174278-5a2e-4102-a24b-ad149a96c37a · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Stochastic gradient descent opti- mizes over-parameterized deep ReLU networks, 2018
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81c054d8-a7b3-4c13-be06-505ae14b05ce · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training g., unit sphere in our case), can be interpreted as fixed volume
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93744545-3594-4ef4-80f8-fe9de8266bcb · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72fb72a9-53b7-49f6-8bd8-c24002a51a42 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training An additional justification for using the Helmholtz free energy arises from the stationary distributions of SGD
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e36e8c0-bfb9-47c7-831e-035b7aa6f484 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e721c2-df99-4bf3-a9f3-95e2df384f1c · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work
Reference 2010
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b7de796-ebd5-4559-89b7-2c70f9b31ada · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Deep Learning and the Information Bottleneck Principle
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc848c5b-e058-4feb-9c30-8229cef5be8e · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Essentially No Barriers in Neural Network Energy Landscape
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 541da72f-9f5e-4f2a-b0ea-81eb4a9d7495 · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed43d29b-cfc1-4ed2-a31a-9f7e8047b64f · outbound
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.