REVIEW 4 major objections 5 minor 1 cited by
Anomalous diffusion dynamics of learning in deep neural networks
T0 review · 4 major / 5 minor · reviewed 2026-08-27 · deepseek-v4-flash
Pith's one-line read Stochastic gradient descent starts with superdiffusive, jumpy motion and later becomes subdiffusive as it reaches flat regions of the loss landscape.
desk verdict The paper's headline superdiffusion claim is likely an artifact of unsubtracted deterministic drift in the time-averaged MSD; the fractal-landscape mechanism is not well supported, but the empirical effort and heavy-tailed gradient analysis deserve a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing tools are two scaling laws. The first is the time-averaged mean squared displacement $\Delta r^2(t_w,\tau)$ of the weights (Eq. 3), whose log-log slope $\alpha$ classifies the walk as superdiffusive ($1<\alpha<2$), normal ($\alpha=1$), or subdiffusive ($0<\alpha<1$). The second is the mean squared loss $\mathrm{MSL}(\Delta r)$ of pairs of points on the trajectory with end-to-end distance $\Delta r$ (Eq. 9): if $\mathrm{MSL} \propto \Delta r^{2H}$, the landscape is fractal-like in the sense of a random function with scaling exponent $H$. Together they link trajectory statistics (how far SGD moves) to landscape geometry (how loss varies with displacement), and the paper uses their joint evolution to argue the causal chain: fractal terrain produces heavy-tailed gradients, which produce intermittent big jumps, which produce superdiffusion and basin escape.
What would settle it
Compute the mean squared loss for pairs of weights sampled uniformly at random from a small ball around a point on the early-training trajectory, rather than pairs selected along the SGD path. If the resulting MSL no longer follows a power law (or has a different exponent from $2H \approx 1.8$) while the trajectory-based MSL still does, the fractal inference is an artifact of the sampling and the claimed mechanism for superdiffusion fails.
Extended reading notes
Core claim
The central claim is that stochastic gradient descent performs a time-inhomogeneous random walk on the loss landscape: for early training intervals the time-averaged mean squared displacement scales as $\Delta r^2(\tau) \propto \tau^{\alpha}$ with $\alpha > 1$ at long lag times (superdiffusion), while $\alpha < 1$ at short lag times; after a crossover time $t_0$, the process becomes purely subdiffusive with $\alpha \approx 0.78$. The paper attributes this to the loss function being fractal-like in the explored region: the mean squared loss between pairs of weights separated by distance $\Delta r$ follows $\mathrm{MSL} \propto \Delta r^{2H}$ with exponent $1.8$ early in training, flattening to a constant later, indicating the optimizer moves from rough fractal terrain to a flat basin. It supports the mechanism with a two-dimensional model driven by Gaussian noise alone, where a fractal landscape reproduces the same superdiffusion-to-subdiffusion transition and enables escapes from local minima.
Load-bearing premise
The analysis assumes that the points an SGD run passes through fairly represent the surrounding loss landscape, so that scaling laws computed along one trajectory reveal the landscape's own fractal structure rather than the optimizer's idiosyncratic route.
Editorial extensions
If this is right
- Superdiffusion gives a concrete, measurable account of how SGD avoids poor local minima: the loss landscape itself intermittently launches the optimizer on long jumps, so escape does not require non-Gaussian gradient noise.
- The crossover from superdiffusion to subdiffusion marks when the optimizer reaches a flat region; monitoring the diffusion exponent $\alpha(t_w)$ can serve as a training-progress diagnostic.
- Deeper networks shorten the superdiffusive scale range $\tau_0$, which the paper interprets as a reason they are harder to train; shortcut connections lengthen it, offering a dynamical explanation for ResNet's trainability.
- Batch size and learning rate do not change the qualitative anomalous-diffusion behavior, so the phenomenon is a property of the landscape–optimizer interaction rather than of particular hyperparameter settings.
Reading between the lines
- If the fractal picture transfers, then architectures or data augmentations that increase the fractal dimension of the explored loss region should prolong the superdiffusive phase and, with it, the ability to escape sharp minima; this is a design rule the paper does not state.
- A testable extension is to compute the MSL scaling on pairs of weights sampled uniformly in a small ball rather than along the SGD trajectory; matching exponents would confirm the landscape, not the optimizer's path, is fractal.
- The superdiffusion-to-subdiffusion transition may also explain learning-rate decay heuristics: annealing would simply shift $t_0$ earlier, accelerating entry into the stabilizing subdiffusive regime.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies the dynamics of vanilla SGD in deep neural networks through the lens of anomalous diffusion. Using the time-averaged mean squared displacement (MSD) of the raw weight trajectory, the authors report a crossover from subdiffusion at short lags to superdiffusion at longer lags during the early training phase, evolving to pure subdiffusion as training progresses. They relate this behavior to heavy-tailed minibatch gradients, infer a fractal-like structure of the loss landscape from the scaling of mean squared loss differences along the same trajectory, and propose a 2D toy model with fractal landscapes to argue that the fractal structure causes the superdiffusive learning dynamics and enables escape from sharp minima.
Significance. If the reported superdiffusion is genuine, the work would contribute a new physical characterization of deep learning dynamics and a mechanism linking landscape geometry to optimization. The authors provide open source code and report stability analyses across architectures and hyperparameters. However, the validity of the central empirical claim is jeopardized by the lack of de-trending in the MSD analysis, and the causal link to fractal landscapes relies on an explicit ergodicity assumption and a toy model whose selection already builds in the desired outcome. These issues must be resolved before the conclusions can be accepted.
major comments (4)
- [III.A, Eq. (3)] The time-averaged MSD is computed on the raw weight trajectory without subtracting the mean displacement or any linear trend within the interval [tw, tw+T]. Because the update rule (Eq. 2) contains the deterministic drift term -eta nabla L(w_t), the MSD includes a contribution proportional to tau^2 that dominates at large lags. The observed crossover from subdiffusion at short tau to superdiffusion at long tau, and the alpha = 2 values reported for small learning rates in Fig. 4(c), are exactly the signature of ballistic drift. The authors should repeat the analysis after centering the trajectory (e.g., subtracting the mean displacement or a fitted velocity) and report whether the superdiffusive exponent alpha_2 > 1 persists. Without this control, the claim that SGD exhibits anomalous superdiffusion is not established.
- [III.D, Eq. (9)] The MSL-based fractal characterization uses pairs of weights sampled along the same SGD trajectory that is used to measure the diffusion exponents. The paper explicitly takes ergodicity for granted, but this assumption is load-bearing: if the trajectory is not representative of the landscape, the inferred fractal scaling and the heavy-tailed gradient distribution in Section III.B may be properties of the optimizer's sampling rather than intrinsic properties of the loss function. The causal chain 'fractal landscape -> heavy-tailed gradients -> superdiffusion' therefore rests on circular evidence. The authors should validate the ergodicity assumption or provide an independent measurement of the landscape structure (e.g., random sampling of a neighborhood around the trajectory).
- [III.E, Fig. 8] The toy model selects generated fractal landscapes that 'contain a wide minimum' before running the dynamics, which builds in the conclusion that the optimizer reaches a wide flat basin. The control experiments also do not isolate the fractal property: the convex paraboloid yields an MSD exponent close to 2 (Fig. 8(f)), which is again the ballistic signature discussed in the first comment, and the smoothed random landscape is a different perturbation that may destroy both the fractal character and the wide-minimum geometry. Since the model is not derived from DNNs and the authors concede its generalization is limited, the claim that fractal-like landscapes are 'essential' for the observed dynamics is stronger than the evidence supports.
- [III.A and Fig. 4] The statement that the anomalous diffusion dynamics are insensitive to batch size and learning rate is overstated. For eta = 0.001, Fig. 4(c) shows a single-segment MSD with alpha = 2 over the observed 500 epochs, and no transition from superdiffusion to subdiffusion occurs within that window. This case is particularly important because alpha = 2 is the drift limit; the authors should either qualify the insensitivity claim or explicitly reconcile this result with the drift confound from the first comment.
minor comments (5)
- [II.B, Eq. (3)] The notation is inconsistent: Eq. (3) writes Delta r^2(t_w, tau) while the text and figures often use Delta r^2(tau); clarify the dependence on t_w throughout. Also, 'T is the length of the time interval' but is used as a number of steps; define the units explicitly.
- [III.A, Fig. 2] The caption of Fig. 2 says 'The color scheme is the same as in (b)' but then refers to (a) and (b); specify the actual colormap and the mapping from t_w to colors to avoid confusion.
- [III.A, Eq. (5)] The threshold t_0 is defined as the first curve whose RMSE falls below 0.03, and tau_0 similarly depends on the RMSE < 0.03 criterion. Report the sensitivity of the fitted exponents and the regime boundaries to this threshold, since the distinction between regimes is central to the paper.
- [III.C, Fig. 6(b)] The reported fractal dimension range D_f in [1.32, 2.67] is very wide; please provide the raw scaling data, the fitting range, and the fit quality (e.g., R^2 or confidence intervals) so the reader can judge whether the power law is well determined.
- [III.D, Fig. 7 and Hausdorff dimension] The box-counting calculation of the Hausdorff dimension of the projected 2D loss landscape is described in one sentence; provide details on the box sizes, the number of samples, and the uncertainty of the estimated dimension (approximately 1.8).
Circularity Check
No significant circularity: the MSD, gradient-tail, and MSL measurements are independent observables; the fractal-landscape model is a forward demonstration rather than a fitted prediction.
full rationale
The paper's central empirical content is observational: the time-averaged MSD (Eq. 3) of the SGD weight trajectory is measured as a function of lag time tau, the minibatch gradient distribution is fitted independently to a Levy stable law (Section III B), and the MSL (Eq. 9) is a separate statistic binned by end-to-end distance Delta-r. None of these quantities is defined in terms of the others: MSD depends on tau, MSL depends on Delta-r, and the stable fit is to raw gradient values. The fractal-like landscape claim is tested by the MSL power law (Eq. 8), and the toy model (Eq. 10) uses independently generated fractal landscapes as inputs, showing that heavy-tailed gradients and superdiffusion can arise from such landscapes without inserting the target behavior into the inputs. That is a constructive simulation, not a fitted parameter renamed as a prediction. The paper explicitly disclaims the toy model's direct generalizability to DNNs ('as our 2D model is not directly derived from DNN models, the generalization of this mechanism to DNNs is limited'). The main identified weakness is the ergodicity assumption in Section III D, stated as 'we take it for granted,' which means trajectory-based MSL estimates are assumed to represent the landscape's intrinsic scaling. This is a validity and identifiability limitation, not a circular reduction: the empirical observations would remain meaningful even if the causal interpretation were wrong, and no equation in the paper reduces to another by construction. The only self-citation (Ref. [46]) is a speculative future direction and is not load-bearing. Potential concerns about unsubtracted drift in the MSD are correctness risks, not circularity, and were not grounds for a circularity score.
Assumptions & free parameters
free parameters (7)
- diffusion exponent alpha_2 (superdiffusive regime) =
reported as > 1, exact value not given
- diffusion exponent alpha_3 (subdiffusive regime) =
0.78
- threshold tau0 =
not specified numerically
- threshold t0 =
21000 for ResNet-14 (batch 1024, lr 0.1)
- Levy stability parameter alpha_dist =
1.46528 (95% CI [1.46509, 1.46546]) for first regime; 1.09 at tw=1; 1.58 at tw=23001
- MSL power-law exponent =
1.8
- fractal dimension of toy landscape =
in [1,2], example value not stated
assumptions (6)
- domain assumption The loss landscape can be modeled as a random function whose increments have variance scaling as a power law of distance (fractal definition).
- domain assumption Ergodicity: sampling pairs of points along a single SGD trajectory in a window is equivalent to sampling random pairs from the landscape.
- domain assumption Time-averaged MSD over a finite window correctly characterizes the diffusion type even when the process is non-stationary (training is not in equilibrium).
- domain assumption The heavy-tailed stable fit is reliable with the available sample sizes.
- ad hoc to paper The 2D fractal toy landscape with a selected wide minimum captures the geometry relevant to high-dimensional DNN loss landscapes.
- domain assumption Flat (wide) minima generalize better than sharp minima, so subdiffusion consolidating residence in flat regions is beneficial.
Cite this review
Pith. "Pith review of Anomalous diffusion dynamics of learning in deep neural networks." pith.science (2026). https://pith.science/paper/4ZJWPXNS
@misc{pith2026200910588,
author = {Pith},
title = {Pith review of: Anomalous diffusion dynamics of learning in deep neural networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/4ZJWPXNS}},
note = {Machine review of arXiv:2009.10588}
}
read the original abstract
Learning in deep neural networks (DNNs) is implemented through minimizing a highly non-convex loss function, typically by a stochastic gradient descent (SGD) method. This learning process can effectively find good wide minima without being trapped in poor local ones. We present a novel account of how such effective deep learning emerges through the interactions of the SGD and the geometrical structure of the loss landscape. Rather than being a normal diffusion process (i.e. Brownian motion) as often assumed, we find that the SGD exhibits rich, complex dynamics when navigating through the loss landscape; initially, the SGD exhibits anomalous superdiffusion, which attenuates gradually and changes to subdiffusion at long times when the solution is reached. Such learning dynamics happen ubiquitously in different DNNs such as ResNet and VGG-like networks and are insensitive to batch size and learning rate. The anomalous superdiffusion process during the initial learning phase indicates that the motion of SGD along the loss landscape possesses intermittent, big jumps; this non-equilibrium property enables the SGD to escape from sharp local minima. By adapting the methods developed for studying energy landscapes in complex physical systems, we find that such superdiffusive learning dynamics are due to the interactions of the SGD and the fractal-like structure of the loss landscape. We further develop a simple model to demonstrate the mechanistic role of the fractal loss landscape in enabling the SGD to effectively find global minima. Our results thus reveal the effectiveness of deep learning from a novel perspective and have implications for designing efficient deep neural networks.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
The paper models SGD dynamics as a percolation process with discrete scale invariance, showing that architectural symmetries cause subnetworks to merge in blocks leading to predictable variance cascades.
Reference graph
Works this paper leans on
-
[1]
LeCun, Y
Y. LeCun, Y. Bengio, and G. Hinton, Nature 521, 436 (2015)
2015
-
[2]
When αdist = 2, the distribution is Gaus- sian
The probability density function (PDF) decays with a power-law tail|x|−αdist−1 which is slower compared to Gaussian distributions; thus the distribution is heavy- tailed [30]. When αdist = 2, the distribution is Gaus- sian. β is the skewness parameter. In particular, for a symmetric L´ evyα-stable (SαS) random variable X, i.e. X∼S αS, the skewness paramet...
- [3]
-
[4]
Z. C. Lipton, in The International Conference on Learn- ing Representations (ICLR) workshop (2016)
work page 2016
-
[5]
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun, in Artificial Intelligence and Statistics (2015) pp. 192–204
work page 2015
-
[6]
T. J. Sejnowski, Proceedings of the National Academy of Sciences , 201907373 (2020)
work page 2020
-
[7]
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein, in Advances in Neural Information Processing Systems (2018) pp. 6389–6399
work page 2018
- [8]
Show all 52 references
-
[9]
Welling and Y
M. Welling and Y. W. Teh, in Proceedings of the 28th In- ternational Conference on Machine Learning (2011) pp. 681–688
2011
-
[10]
Jastrz¸ ebski, Z
S. Jastrz¸ ebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey, arXiv:1711.04623 (2017)
2017 arXiv
-
[11]
Baity-Jesi, L
M. Baity-Jesi, L. Sagun, M. Geiger, S. Spigler, G. B. Arous, C. Cammarota, Y. LeCun, M. Wyart, and G. Biroli, in International Conference on Machine Learn- ing (PMLR, 2018) pp. 314–323
2018
-
[12]
Chaudhari and S
P. Chaudhari and S. Soatto, in Information Theory and Applications (ITA) Workshop (IEEE, 2018) pp. 1–10
2018
-
[13]
Simsekli, L
U. Simsekli, L. Sagun, and M. Gurbuzbalaban, in Inter- national Conference on Machine Learning (PMLR, 2019) pp. 5827–5837
2019
-
[14]
Geiger, S
M. Geiger, S. Spigler, S. d’Ascoli, L. Sagun, M. Baity- Jesi, G. Biroli, and M. Wyart, Physical Review E 100, 12115 (2019)
2019
-
[15]
K. He, X. Zhang, S. Ren, and J. Sun, in Proceedings of the IEEE conference on computer vision and pattern recognition (2016) pp. 770–778
2016
-
[16]
Feng and Y
Y. Feng and Y. Tu, Proceedings of the National Academy of Sciences 118 (2021), 10.1073/pnas.2015617118
2021 doi
-
[17]
H. J. Hwang, R. A. Riggleman, and J. C. Crocker, Na- ture Materials 15, 1031 (2016)
2016
-
[18]
Charbonneau, J
P. Charbonneau, J. Kurchan, G. Parisi, P. Urbani, and F. Zamponi, Nature Communications 5, 3725 (2014)
2014
-
[19]
P. Cao, M. P. Short, and S. Yip, Proceedings of the National Academy of Sciences , 201907317 (2019)
2019
-
[20]
Jin and H
Y. Jin and H. Yoshino, Nature Communications 8, 14935 (2017)
2017
-
[21]
Bronstein, Y
I. Bronstein, Y. Israel, E. Kepten, S. Mai, Y. Shav-Tal, E. Barkai, and Y. Garini, Physical Review Letters 103, 018102 (2009)
2009
-
[22]
Golding and E
I. Golding and E. C. Cox, Physical Review Letters 96, 098102 (2006)
2006
-
[23]
Metzler and J
R. Metzler and J. Klafter, Physics Reports 339, 1 (2000)
2000
-
[24]
Zaburdaev, S
V. Zaburdaev, S. Denisov, and J. Klafter, Reviews of Modern Physics 87, 483 (2015)
2015
-
[25]
G. M. Viswanathan, S. V. Buldyrev, S. Havlin, M. G. E. Da Luz, E. P. Raposo, and H. E. Stanley, Nature 401, 911 (1999)
1999
-
[26]
T. H. Solomon, E. R. Weeks, and H. L. Swinney, Physical Review Letters 71, 3975 (1993)
1993
-
[27]
L. G. Alves, D. B. Scariot, R. R. Guimar˜ aes, C. V. Naka- mura, R. S. Mendes, and H. V. Ribeiro, PLoS One 11, e0152092 (2016)
2016
-
[28]
Dieterich, R
P. Dieterich, R. Klages, R. Preuss, and A. Schwab, Pro- ceedings of the National Academy of Sciences 105, 459 13 (2008)
2008
-
[29]
J. P. Nolan, Univariate Stable Distributions: Models for Heavy Tailed Data (Springer International Publishing,
-
[30]
K. A. Sankararaman, S. De, Z. Xu, W. R. Huang, and T. Goldstein, in Proceedings of the 37th International Conference on Machine Learning (2020)
2020
-
[31]
M. F. Barnsley, R. L. Devaney, B. B. Mandelbrot, H.- O. Peitgen, D. Saupe, R. F. Voss, Y. Fisher, and M. McGuire, The Science of Fractal Images (Springer, 1988)
1988
-
[32]
Klafter and I
J. Klafter and I. M. Sokolov, First Steps in Random Walks: from Tools to Applications (Oxford University Press, 2011)
2011
-
[33]
Hochreiter and J
S. Hochreiter and J. Schmidhuber, Neural Computation 9, 1 (1997)
1997
-
[34]
G. B. Sorkin, Algorithmica 6, 367 (1991)
1991
-
[35]
Baldassi, F
C. Baldassi, F. Pittorino, and R. Zecchina, Proceedings of the National Academy of Sciences 117, 161 (2020)
2020
-
[36]
Sagun, U
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bot- tou, in The International Conference on Learning Repre- sentations (ICLR) workshop (2018)
2018
-
[37]
Chaudhari, A
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, Journal of Statistical Mechanics: Theory and Experiment 2019, 124018 (2019)
2019
-
[38]
Ghorbani, S
B. Ghorbani, S. Krishnan, and Y. Xiao, in Proceedings of the 36th International Conference on Machine Learn- ing, Proceedings of Machine Learning Research, Vol. 97, edited by K. Chaudhuri and R. Salakhutdinov (PMLR,
-
[39]
Douillet, MATLAB Central File Exchange (2020)
N. Douillet, MATLAB Central File Exchange (2020)
2020
-
[40]
J. Li, Q. Du, and C. Sun, Pattern Recognition 42, 2460 (2009)
2009
-
[41]
Zhang, S
J. Zhang, S. P. Karimireddy, A. Veit, S. Kim, S. J. Reddi, S. Kumar, and S. Sra, arXiv:1912.03194 (2019)
2019 arXiv
-
[42]
C. H. Martin and M. W. Mahoney, arXiv:1810.01075 (2019)
2019 arXiv
-
[43]
Metzler, J.-H
R. Metzler, J.-H. Jeon, A. G. Cherstvy, and E. Barkai, Physical Chemistry Chemical Physics 16, 24128 (2014)
2014
-
[44]
Panigrahi, R
A. Panigrahi, R. Somani, N. Goyal, and P. Netrapalli, arXiv:1910.09626 (2019)
2019 arXiv
-
[45]
Hausdorff Dimension, Stochastic Differen- tial Equations, and Generalization in Neural Networks,
U. S ¸im¸ sekli, O. Sener, G. Deligiannidis, and M. A. Erdogdu, “Hausdorff Dimension, Stochastic Differen- tial Equations, and Generalization in Neural Networks,” (2020)
2020
-
[46]
D. S. Grebenkov, Physical Review E 99, 032133 (2019)
2019
-
[47]
J. A. Costa and A. O. Hero, in European Signal Process- ing Conference (2004) pp. 369–372
2004
-
[48]
Wardak and P
A. Wardak and P. Gong, Physical Review Research 3, 013083 (2021)
2021
- [49]
-
[50]
Levina and P
E. Levina and P. J. Bickel, in Neural Information Pro- cessing Systems (2004) pp. 777–784
2004
-
[52]
Santurkar, D
S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry, in Ad- vances in Neural Information Processing Systems (2018) pp. 2483–2493
2018
-
[128]
(b) Similar to (a) but for α in the second regime (when tw = t0)
(a) The diffusion exponents α on larger lag times ( τ > τ0) when tw = 1 as a function of minibatch size. (b) Similar to (a) but for α in the second regime (when tw = t0). (c) The crossover τ0 as a function of minibatch size. Here, τ0 is the lag time when the MSD curve transitio...
Reviewed August 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.