Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Anomalous diffusion dynamics of learning in deep neural networks

T0 review · 4 major / 5 minor · reviewed 2026-08-27 · deepseek-v4-flash

Pith's one-line read Stochastic gradient descent starts with superdiffusive, jumpy motion and later becomes subdiffusive as it reaches flat regions of the loss landscape.

desk verdict The paper's headline superdiffusion claim is likely an artifact of unsubtracted deterministic drift in the time-averaged MSD; the fractal-landscape mechanism is not well supported, but the empirical effort and heavy-tailed gradient analysis deserve a serious referee. read the letter →

arxiv 2009.10588 v2 pith:4ZJWPXNS submitted 2020-09-22 cs.LG stat.ML

classification cs.LGstat.ML MSC 68T0782C31
keywords anomalousdiffusionstochasticgradientdescentlosslandscapefractalsuperdiffusionheavy-tailedgradientsmeansquareddisplacementdeepneuralnetworks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the learning dynamics of deep networks are not Brownian but anomalously diffusive: early in training the SGD optimizer moves by intermittent long jumps (superdiffusion), and late in training it slows to subdiffusion as it enters flat regions. The authors claim this two-phase behavior appears across ResNets and VGG-like networks, is insensitive to batch size and learning rate, and is caused by the fractal-like structure of the loss landscape. If correct, this reframes escape from sharp minima: the rugged geometry itself supplies heavy-tailed gradients that occasionally throw the optimizer out of a basin, so no special noise model is needed.

What carries the argument

The load-bearing tools are two scaling laws. The first is the time-averaged mean squared displacement $\Delta r^2(t_w,\tau)$ of the weights (Eq. 3), whose log-log slope $\alpha$ classifies the walk as superdiffusive ($1<\alpha<2$), normal ($\alpha=1$), or subdiffusive ($0<\alpha<1$). The second is the mean squared loss $\mathrm{MSL}(\Delta r)$ of pairs of points on the trajectory with end-to-end distance $\Delta r$ (Eq. 9): if $\mathrm{MSL} \propto \Delta r^{2H}$, the landscape is fractal-like in the sense of a random function with scaling exponent $H$. Together they link trajectory statistics (how far SGD moves) to landscape geometry (how loss varies with displacement), and the paper uses their joint evolution to argue the causal chain: fractal terrain produces heavy-tailed gradients, which produce intermittent big jumps, which produce superdiffusion and basin escape.

What would settle it

Compute the mean squared loss for pairs of weights sampled uniformly at random from a small ball around a point on the early-training trajectory, rather than pairs selected along the SGD path. If the resulting MSL no longer follows a power law (or has a different exponent from $2H \approx 1.8$) while the trajectory-based MSL still does, the fractal inference is an artifact of the sampling and the claimed mechanism for superdiffusion fails.

Watch

Extended reading notes

Core claim

The central claim is that stochastic gradient descent performs a time-inhomogeneous random walk on the loss landscape: for early training intervals the time-averaged mean squared displacement scales as $\Delta r^2(\tau) \propto \tau^{\alpha}$ with $\alpha > 1$ at long lag times (superdiffusion), while $\alpha < 1$ at short lag times; after a crossover time $t_0$, the process becomes purely subdiffusive with $\alpha \approx 0.78$. The paper attributes this to the loss function being fractal-like in the explored region: the mean squared loss between pairs of weights separated by distance $\Delta r$ follows $\mathrm{MSL} \propto \Delta r^{2H}$ with exponent $1.8$ early in training, flattening to a constant later, indicating the optimizer moves from rough fractal terrain to a flat basin. It supports the mechanism with a two-dimensional model driven by Gaussian noise alone, where a fractal landscape reproduces the same superdiffusion-to-subdiffusion transition and enables escapes from local minima.

Load-bearing premise

The analysis assumes that the points an SGD run passes through fairly represent the surrounding loss landscape, so that scaling laws computed along one trajectory reveal the landscape's own fractal structure rather than the optimizer's idiosyncratic route.

Editorial extensions

If this is right

  • Superdiffusion gives a concrete, measurable account of how SGD avoids poor local minima: the loss landscape itself intermittently launches the optimizer on long jumps, so escape does not require non-Gaussian gradient noise.
  • The crossover from superdiffusion to subdiffusion marks when the optimizer reaches a flat region; monitoring the diffusion exponent $\alpha(t_w)$ can serve as a training-progress diagnostic.
  • Deeper networks shorten the superdiffusive scale range $\tau_0$, which the paper interprets as a reason they are harder to train; shortcut connections lengthen it, offering a dynamical explanation for ResNet's trainability.
  • Batch size and learning rate do not change the qualitative anomalous-diffusion behavior, so the phenomenon is a property of the landscape–optimizer interaction rather than of particular hyperparameter settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the fractal picture transfers, then architectures or data augmentations that increase the fractal dimension of the explored loss region should prolong the superdiffusive phase and, with it, the ability to escape sharp minima; this is a design rule the paper does not state.
  • A testable extension is to compute the MSL scaling on pairs of weights sampled uniformly in a small ball rather than along the SGD trajectory; matching exponents would confirm the landscape, not the optimizer's path, is fractal.
  • The superdiffusion-to-subdiffusion transition may also explain learning-rate decay heuristics: annealing would simply shift $t_0$ earlier, accelerating entry into the stabilizing subdiffusive regime.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Request a human review

A listed scientist reviews the paper for a fee and the review publishes here regardless of verdict. See the reviewers or get listed.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript studies the dynamics of vanilla SGD in deep neural networks through the lens of anomalous diffusion. Using the time-averaged mean squared displacement (MSD) of the raw weight trajectory, the authors report a crossover from subdiffusion at short lags to superdiffusion at longer lags during the early training phase, evolving to pure subdiffusion as training progresses. They relate this behavior to heavy-tailed minibatch gradients, infer a fractal-like structure of the loss landscape from the scaling of mean squared loss differences along the same trajectory, and propose a 2D toy model with fractal landscapes to argue that the fractal structure causes the superdiffusive learning dynamics and enables escape from sharp minima.

Significance. If the reported superdiffusion is genuine, the work would contribute a new physical characterization of deep learning dynamics and a mechanism linking landscape geometry to optimization. The authors provide open source code and report stability analyses across architectures and hyperparameters. However, the validity of the central empirical claim is jeopardized by the lack of de-trending in the MSD analysis, and the causal link to fractal landscapes relies on an explicit ergodicity assumption and a toy model whose selection already builds in the desired outcome. These issues must be resolved before the conclusions can be accepted.

major comments (4)
  1. [III.A, Eq. (3)] The time-averaged MSD is computed on the raw weight trajectory without subtracting the mean displacement or any linear trend within the interval [tw, tw+T]. Because the update rule (Eq. 2) contains the deterministic drift term -eta nabla L(w_t), the MSD includes a contribution proportional to tau^2 that dominates at large lags. The observed crossover from subdiffusion at short tau to superdiffusion at long tau, and the alpha = 2 values reported for small learning rates in Fig. 4(c), are exactly the signature of ballistic drift. The authors should repeat the analysis after centering the trajectory (e.g., subtracting the mean displacement or a fitted velocity) and report whether the superdiffusive exponent alpha_2 > 1 persists. Without this control, the claim that SGD exhibits anomalous superdiffusion is not established.
  2. [III.D, Eq. (9)] The MSL-based fractal characterization uses pairs of weights sampled along the same SGD trajectory that is used to measure the diffusion exponents. The paper explicitly takes ergodicity for granted, but this assumption is load-bearing: if the trajectory is not representative of the landscape, the inferred fractal scaling and the heavy-tailed gradient distribution in Section III.B may be properties of the optimizer's sampling rather than intrinsic properties of the loss function. The causal chain 'fractal landscape -> heavy-tailed gradients -> superdiffusion' therefore rests on circular evidence. The authors should validate the ergodicity assumption or provide an independent measurement of the landscape structure (e.g., random sampling of a neighborhood around the trajectory).
  3. [III.E, Fig. 8] The toy model selects generated fractal landscapes that 'contain a wide minimum' before running the dynamics, which builds in the conclusion that the optimizer reaches a wide flat basin. The control experiments also do not isolate the fractal property: the convex paraboloid yields an MSD exponent close to 2 (Fig. 8(f)), which is again the ballistic signature discussed in the first comment, and the smoothed random landscape is a different perturbation that may destroy both the fractal character and the wide-minimum geometry. Since the model is not derived from DNNs and the authors concede its generalization is limited, the claim that fractal-like landscapes are 'essential' for the observed dynamics is stronger than the evidence supports.
  4. [III.A and Fig. 4] The statement that the anomalous diffusion dynamics are insensitive to batch size and learning rate is overstated. For eta = 0.001, Fig. 4(c) shows a single-segment MSD with alpha = 2 over the observed 500 epochs, and no transition from superdiffusion to subdiffusion occurs within that window. This case is particularly important because alpha = 2 is the drift limit; the authors should either qualify the insensitivity claim or explicitly reconcile this result with the drift confound from the first comment.
minor comments (5)
  1. [II.B, Eq. (3)] The notation is inconsistent: Eq. (3) writes Delta r^2(t_w, tau) while the text and figures often use Delta r^2(tau); clarify the dependence on t_w throughout. Also, 'T is the length of the time interval' but is used as a number of steps; define the units explicitly.
  2. [III.A, Fig. 2] The caption of Fig. 2 says 'The color scheme is the same as in (b)' but then refers to (a) and (b); specify the actual colormap and the mapping from t_w to colors to avoid confusion.
  3. [III.A, Eq. (5)] The threshold t_0 is defined as the first curve whose RMSE falls below 0.03, and tau_0 similarly depends on the RMSE < 0.03 criterion. Report the sensitivity of the fitted exponents and the regime boundaries to this threshold, since the distinction between regimes is central to the paper.
  4. [III.C, Fig. 6(b)] The reported fractal dimension range D_f in [1.32, 2.67] is very wide; please provide the raw scaling data, the fitting range, and the fit quality (e.g., R^2 or confidence intervals) so the reader can judge whether the power law is well determined.
  5. [III.D, Fig. 7 and Hausdorff dimension] The box-counting calculation of the Hausdorff dimension of the projected 2D loss landscape is described in one sentence; provide details on the box sizes, the number of samples, and the uncertainty of the estimated dimension (approximately 1.8).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the MSD, gradient-tail, and MSL measurements are independent observables; the fractal-landscape model is a forward demonstration rather than a fitted prediction.

full rationale

The paper's central empirical content is observational: the time-averaged MSD (Eq. 3) of the SGD weight trajectory is measured as a function of lag time tau, the minibatch gradient distribution is fitted independently to a Levy stable law (Section III B), and the MSL (Eq. 9) is a separate statistic binned by end-to-end distance Delta-r. None of these quantities is defined in terms of the others: MSD depends on tau, MSL depends on Delta-r, and the stable fit is to raw gradient values. The fractal-like landscape claim is tested by the MSL power law (Eq. 8), and the toy model (Eq. 10) uses independently generated fractal landscapes as inputs, showing that heavy-tailed gradients and superdiffusion can arise from such landscapes without inserting the target behavior into the inputs. That is a constructive simulation, not a fitted parameter renamed as a prediction. The paper explicitly disclaims the toy model's direct generalizability to DNNs ('as our 2D model is not directly derived from DNN models, the generalization of this mechanism to DNNs is limited'). The main identified weakness is the ergodicity assumption in Section III D, stated as 'we take it for granted,' which means trajectory-based MSL estimates are assumed to represent the landscape's intrinsic scaling. This is a validity and identifiability limitation, not a circular reduction: the empirical observations would remain meaningful even if the causal interpretation were wrong, and no equation in the paper reduces to another by construction. The only self-citation (Ref. [46]) is a speculative future direction and is not load-bearing. Potential concerns about unsubtracted drift in the MSD are correctness risks, not circularity, and were not grounds for a circularity score.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

The central empirical claim rests on fitted diffusion and stability exponents and on the ergodicity assumption for the fractal landscape measurement. The toy model adds modeling choices, including the selection of fractal landscapes that contain a wide minimum.

free parameters (7)
  • diffusion exponent alpha_2 (superdiffusive regime) = reported as > 1, exact value not given
    Fitted to time-averaged MSD curves for lag time tau > tau0 in the early training phase; the claim that the regime is superdiffusive depends on this exponent.
  • diffusion exponent alpha_3 (subdiffusive regime) = 0.78
    Fitted to MSD curves in the late phase (tw > t0). Reported alpha=0.78 and used to claim subdiffusion.
  • threshold tau0 = not specified numerically
    Chosen as the smallest lag time such that the power-law fit RMSE < 0.03; this threshold separates the subdiffusive and superdiffusive segments of the MSD and is data-dependent.
  • threshold t0 = 21000 for ResNet-14 (batch 1024, lr 0.1)
    Identified as the first interval whose MSD fit has RMSE < 0.03; defines the transition from the superdiffusion-dominated first regime to the subdiffusive second regime.
  • Levy stability parameter alpha_dist = 1.46528 (95% CI [1.46509, 1.46546]) for first regime; 1.09 at tw=1; 1.58 at tw=23001
    Maximum likelihood fit of minibatch gradients to a symmetric Levy alpha-stable distribution; used to support the heavy-tailed gradient claim.
  • MSL power-law exponent = 1.8
    Fitted to the mean squared loss versus end-to-end distance in the fractal-like landscape claim.
  • fractal dimension of toy landscape = in [1,2], example value not stated
    Chosen when generating 2D fractal landscapes, and only landscapes containing a wide minimum are selected.
assumptions (6)
  • domain assumption The loss landscape can be modeled as a random function whose increments have variance scaling as a power law of distance (fractal definition).
    Used in Section III.D to justify MSL power-law fit, following the fractal definition from Barnsley et al. and the simulated annealing work by Sorkin.
  • domain assumption Ergodicity: sampling pairs of points along a single SGD trajectory in a window is equivalent to sampling random pairs from the landscape.
    Explicitly stated in Section III.D: 'The equivalence of these two procedures is referred to as ergodicity and we take it for granted.' This is load-bearing for the fractal landscape measurement.
  • domain assumption Time-averaged MSD over a finite window correctly characterizes the diffusion type even when the process is non-stationary (training is not in equilibrium).
    The paper uses time-averaged MSD because ensemble averaging is impossible; this assumes interpretability of time averages for SGD, which is not proven.
  • domain assumption The heavy-tailed stable fit is reliable with the available sample sizes.
    Maximum likelihood fitting of Levy stable distributions is known to be difficult; the paper reports extremely narrow confidence intervals, suggesting large sample sizes but no robustness checks.
  • ad hoc to paper The 2D fractal toy landscape with a selected wide minimum captures the geometry relevant to high-dimensional DNN loss landscapes.
    The authors acknowledge in Section IV that the 2D model is not directly derived from DNN models, limiting generalization.
  • domain assumption Flat (wide) minima generalize better than sharp minima, so subdiffusion consolidating residence in flat regions is beneficial.
    This is taken from prior work (Hochreiter and Schmidhuber; Chaudhari et al.) and used in the interpretation of the late subdiffusive phase.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Anomalous diffusion dynamics of learning in deep neural networks." pith.science (2026). https://pith.science/paper/4ZJWPXNS

@misc{pith2026200910588,
  author       = {Pith},
  title        = {Pith review of: Anomalous diffusion dynamics of learning in deep neural networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4ZJWPXNS}},
  note         = {Machine review of arXiv:2009.10588}
}
read the original abstract

Learning in deep neural networks (DNNs) is implemented through minimizing a highly non-convex loss function, typically by a stochastic gradient descent (SGD) method. This learning process can effectively find good wide minima without being trapped in poor local ones. We present a novel account of how such effective deep learning emerges through the interactions of the SGD and the geometrical structure of the loss landscape. Rather than being a normal diffusion process (i.e. Brownian motion) as often assumed, we find that the SGD exhibits rich, complex dynamics when navigating through the loss landscape; initially, the SGD exhibits anomalous superdiffusion, which attenuates gradually and changes to subdiffusion at long times when the solution is reached. Such learning dynamics happen ubiquitously in different DNNs such as ResNet and VGG-like networks and are insensitive to batch size and learning rate. The anomalous superdiffusion process during the initial learning phase indicates that the motion of SGD along the loss landscape possesses intermittent, big jumps; this non-equilibrium property enables the SGD to escape from sharp local minima. By adapting the methods developed for studying energy landscapes in complex physical systems, we find that such superdiffusive learning dynamics are due to the interactions of the SGD and the fractal-like structure of the loss landscape. We further develop a simple model to demonstrate the mechanistic role of the fractal loss landscape in enabling the SGD to effectively find global minima. Our results thus reveal the effectiveness of deep learning from a novel perspective and have implications for designing efficient deep neural networks.

Figures

Figures reproduced from arXiv: 2009.10588 by the authors.

Figure 1
Figure 1. FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6 [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7 [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: (d). Some long-range steps make the SGD opti￾mizer escape the minimum; however, some of them jump to a lower altitude and then leave the minimum. This example illustrates how the fractal landscape assists the SGD optimizer escape local minima. As we only use Gaussian n…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

    cs.LG 2026-09 conditional novelty 7.0 of 10

    The paper models SGD dynamics as a percolation process with discrete scale invariance, showing that architectural symmetries cause subnetworks to merge in blocks leading to predictable variance cascades.

Reference graph

Works this paper leans on

52 extracted references · 47 canonical work pages · cited by 1 Pith paper

  1. [1]

    LeCun, Y

    Y. LeCun, Y. Bengio, and G. Hinton, Nature 521, 436 (2015)

  2. [2]

    When αdist = 2, the distribution is Gaus- sian

    The probability density function (PDF) decays with a power-law tail|x|−αdist−1 which is slower compared to Gaussian distributions; thus the distribution is heavy- tailed [30]. When αdist = 2, the distribution is Gaus- sian. β is the skewness parameter. In particular, for a symmetric L´ evyα-stable (SαS) random variable X, i.e. X∼S αS, the skewness paramet...

  3. [3]

    Hoffer, I

    E. Hoffer, I. Hubara, and D. Soudry, in Advances in Neural Information Processing Systems (2017) pp. 1731– 1741

  4. [4]

    Z. C. Lipton, in The International Conference on Learn- ing Representations (ICLR) workshop (2016)

  5. [5]

    Choromanska, M

    A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun, in Artificial Intelligence and Statistics (2015) pp. 192–204

  6. [6]

    T. J. Sejnowski, Proceedings of the National Academy of Sciences , 201907373 (2020)

  7. [7]

    H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein, in Advances in Neural Information Processing Systems (2018) pp. 6389–6399

  8. [8]

    Becker, Y

    S. Becker, Y. Zhang, and A. A. Lee, Physical Review Letters 124 (2020)

Show all 52 references
  1. [9]

    Welling and Y

    M. Welling and Y. W. Teh, in Proceedings of the 28th In- ternational Conference on Machine Learning (2011) pp. 681–688

  2. [10]

    Jastrz¸ ebski, Z

    S. Jastrz¸ ebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey, arXiv:1711.04623 (2017)

  3. [11]

    Baity-Jesi, L

    M. Baity-Jesi, L. Sagun, M. Geiger, S. Spigler, G. B. Arous, C. Cammarota, Y. LeCun, M. Wyart, and G. Biroli, in International Conference on Machine Learn- ing (PMLR, 2018) pp. 314–323

  4. [12]

    Chaudhari and S

    P. Chaudhari and S. Soatto, in Information Theory and Applications (ITA) Workshop (IEEE, 2018) pp. 1–10

  5. [13]

    Simsekli, L

    U. Simsekli, L. Sagun, and M. Gurbuzbalaban, in Inter- national Conference on Machine Learning (PMLR, 2019) pp. 5827–5837

  6. [14]

    Geiger, S

    M. Geiger, S. Spigler, S. d’Ascoli, L. Sagun, M. Baity- Jesi, G. Biroli, and M. Wyart, Physical Review E 100, 12115 (2019)

  7. [15]

    K. He, X. Zhang, S. Ren, and J. Sun, in Proceedings of the IEEE conference on computer vision and pattern recognition (2016) pp. 770–778

  8. [16]

    Feng and Y

    Y. Feng and Y. Tu, Proceedings of the National Academy of Sciences 118 (2021), 10.1073/pnas.2015617118

  9. [17]

    H. J. Hwang, R. A. Riggleman, and J. C. Crocker, Na- ture Materials 15, 1031 (2016)

  10. [18]

    Charbonneau, J

    P. Charbonneau, J. Kurchan, G. Parisi, P. Urbani, and F. Zamponi, Nature Communications 5, 3725 (2014)

  11. [19]

    P. Cao, M. P. Short, and S. Yip, Proceedings of the National Academy of Sciences , 201907317 (2019)

  12. [20]

    Jin and H

    Y. Jin and H. Yoshino, Nature Communications 8, 14935 (2017)

  13. [21]

    Bronstein, Y

    I. Bronstein, Y. Israel, E. Kepten, S. Mai, Y. Shav-Tal, E. Barkai, and Y. Garini, Physical Review Letters 103, 018102 (2009)

  14. [22]

    Golding and E

    I. Golding and E. C. Cox, Physical Review Letters 96, 098102 (2006)

  15. [23]

    Metzler and J

    R. Metzler and J. Klafter, Physics Reports 339, 1 (2000)

  16. [24]

    Zaburdaev, S

    V. Zaburdaev, S. Denisov, and J. Klafter, Reviews of Modern Physics 87, 483 (2015)

  17. [25]

    G. M. Viswanathan, S. V. Buldyrev, S. Havlin, M. G. E. Da Luz, E. P. Raposo, and H. E. Stanley, Nature 401, 911 (1999)

  18. [26]

    T. H. Solomon, E. R. Weeks, and H. L. Swinney, Physical Review Letters 71, 3975 (1993)

  19. [27]

    L. G. Alves, D. B. Scariot, R. R. Guimar˜ aes, C. V. Naka- mura, R. S. Mendes, and H. V. Ribeiro, PLoS One 11, e0152092 (2016)

  20. [28]

    Dieterich, R

    P. Dieterich, R. Klages, R. Preuss, and A. Schwab, Pro- ceedings of the National Academy of Sciences 105, 459 13 (2008)

  21. [29]

    J. P. Nolan, Univariate Stable Distributions: Models for Heavy Tailed Data (Springer International Publishing,

  22. [30]

    K. A. Sankararaman, S. De, Z. Xu, W. R. Huang, and T. Goldstein, in Proceedings of the 37th International Conference on Machine Learning (2020)

  23. [31]

    M. F. Barnsley, R. L. Devaney, B. B. Mandelbrot, H.- O. Peitgen, D. Saupe, R. F. Voss, Y. Fisher, and M. McGuire, The Science of Fractal Images (Springer, 1988)

  24. [32]

    Klafter and I

    J. Klafter and I. M. Sokolov, First Steps in Random Walks: from Tools to Applications (Oxford University Press, 2011)

  25. [33]

    Hochreiter and J

    S. Hochreiter and J. Schmidhuber, Neural Computation 9, 1 (1997)

  26. [34]

    G. B. Sorkin, Algorithmica 6, 367 (1991)

  27. [35]

    Baldassi, F

    C. Baldassi, F. Pittorino, and R. Zecchina, Proceedings of the National Academy of Sciences 117, 161 (2020)

  28. [36]

    Sagun, U

    L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bot- tou, in The International Conference on Learning Repre- sentations (ICLR) workshop (2018)

  29. [37]

    Chaudhari, A

    P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, Journal of Statistical Mechanics: Theory and Experiment 2019, 124018 (2019)

  30. [38]

    Ghorbani, S

    B. Ghorbani, S. Krishnan, and Y. Xiao, in Proceedings of the 36th International Conference on Machine Learn- ing, Proceedings of Machine Learning Research, Vol. 97, edited by K. Chaudhuri and R. Salakhutdinov (PMLR,

  31. [39]

    Douillet, MATLAB Central File Exchange (2020)

    N. Douillet, MATLAB Central File Exchange (2020)

  32. [40]

    J. Li, Q. Du, and C. Sun, Pattern Recognition 42, 2460 (2009)

  33. [41]

    Zhang, S

    J. Zhang, S. P. Karimireddy, A. Veit, S. Kim, S. J. Reddi, S. Kumar, and S. Sra, arXiv:1912.03194 (2019)

  34. [42]

    C. H. Martin and M. W. Mahoney, arXiv:1810.01075 (2019)

  35. [43]

    Metzler, J.-H

    R. Metzler, J.-H. Jeon, A. G. Cherstvy, and E. Barkai, Physical Chemistry Chemical Physics 16, 24128 (2014)

  36. [44]

    Panigrahi, R

    A. Panigrahi, R. Somani, N. Goyal, and P. Netrapalli, arXiv:1910.09626 (2019)

  37. [45]

    Hausdorff Dimension, Stochastic Differen- tial Equations, and Generalization in Neural Networks,

    U. S ¸im¸ sekli, O. Sener, G. Deligiannidis, and M. A. Erdogdu, “Hausdorff Dimension, Stochastic Differen- tial Equations, and Generalization in Neural Networks,” (2020)

  38. [46]

    D. S. Grebenkov, Physical Review E 99, 032133 (2019)

  39. [47]

    J. A. Costa and A. O. Hero, in European Signal Process- ing Conference (2004) pp. 369–372

  40. [48]

    Wardak and P

    A. Wardak and P. Gong, Physical Review Research 3, 013083 (2021)

  41. [49]

    Goldt, M

    S. Goldt, M. Mezard, F. Krzakala, and L. Zdeborova, arXiv:1909.11500v3 (2019)

  42. [50]

    Levina and P

    E. Levina and P. J. Bickel, in Neural Information Pro- cessing Systems (2004) pp. 777–784

  43. [52]

    Santurkar, D

    S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry, in Ad- vances in Neural Information Processing Systems (2018) pp. 2483–2493

  44. [128]

    (b) Similar to (a) but for α in the second regime (when tw = t0)

    (a) The diffusion exponents α on larger lag times ( τ > τ0) when tw = 1 as a function of minibatch size. (b) Similar to (a) but for α in the second regime (when tw = t0). (c) The crossover τ0 as a function of minibatch size. Here, τ0 is the lag time when the MSD curve transitio...

Pith tools

Reviewed August 27, 2026 · model on record in the stance chip above.