Pith. sign in

REVIEW 115 references

Adversarial Training from Mean Field Perspective

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A mean field framework for random ReLU networks yields adversarial-loss bounds and predicts that adversarial training shrinks weights, hurts vanilla depth, and is rescued by residual connections and width.

arxiv 2505.14021 v1 pith:L6LHFZOW submitted 2025-05-20 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords adversarialtrainingmeanboundsexamplesfieldframeworknetwork
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Imagine a randomly initialized deep network with many neurons per layer. This paper analyzes what happens if you train it to resist small adversarial perturbations. The key move is to write the network output near an input as a linear function: an input-dependent Jacobian matrix times the input, plus an offset. The authors claim that for ReLU-like networks, this Jacobian and offset behave like random Gaussian objects whose probability law does not depend on the specific input. That simplification lets them compute upper bounds for how much the output can change under ℓp-bounded perturbations, for several choices of p and q, and the bounds match simulations when the network is wide.

From there, the paper studies training dynamics. It replaces the adversarial loss by its upper bound, epsilon times a constant times the L-th power of the weight variance parameter. Under gradient flow, this makes weight variance decrease linearly in time. A decreasing weight variance is bad for plain deep networks, because they need a specific variance to avoid vanishing gradients. The paper therefore concludes that vanilla networks can become untrainable under adversarial training, while residual networks are safe because their trainability condition has no lower bound. It also derives that the Fisher-Rao capacity drops linearly in time, with degradation accelerated by depth and slowed by width.

The caveats are serious: the dynamic theorems rely on a surrogate adversarial loss rather than the true max over perturbations, and one key independence lemma in the appendix is not rigorously proved. The empirical checks cover early training on MNIST and Fashion-MNIST, not the full generality claimed in the title.

Extended reading notes

Core claim

Theorem 4.1 states that for any fixed input xin, the input-output Jacobian J(xin) and offset a(xin) of a random ReLU network are independent, with i.i.d. Gaussian entries whose variances do not depend on xin (J entries have variance omega^L/d). If true, the whole network reduces to two Gaussians, and the paper uses this to upper bound adversarial loss for various norm pairs and to derive trainability and capacity theorems for adversarial training.

Load-bearing premise

Assumption 5.3 replaces the true adversarial loss by its upper bound: Ladv := epsilon * beta_{p,q} * omega(t)^{L/2}. All dynamic results (weight variance decay, vanilla untrainability, Fisher-Rao capacity loss) are then consequences of training on this surrogate loss. If the surrogate does not track the actual maximum over perturbations during training, then Theorems 5.4, 5.7, 5.9, and their residual counterparts do not describe real adversarial training.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical or architectural entities. The central claims rest on standard mean field Gaussian assumptions, the gradient independence assumption, and especially the surrogate adversarial loss of Assumption 5.3. The unproved asymptotic independence in Lemma E.8 is a further load-bearing premise.

assumptions (5)
  • domain assumption Width N is sufficiently large so that central limit theorem applies and entries of J and a are Gaussian.
    Stated in Sec. 3.1 and used throughout Thm 4.1; standard mean field assumption.
  • domain assumption Assumption 5.2: for 0 <= t <= T << N, model parameters remain independent and Gaussian with variances sigma_w^2(t)/N and sigma_b^2(t).
    Introduced in Sec. 5.2; needed to keep the mean field Gaussian structure valid during training.
  • domain assumption Gradient independence Assumption B.1 is applied to the standard loss Lstd and to the network output in Sec. 5.4.
    Borrowed from [105] and used in Lemmas G.8, G.12, and G.13; the paper states it is not applied to the adversarial loss.
  • ad hoc to paper Assumption 5.3: Ladv := epsilon * beta_{p,q} * omega(t)^{L/2} is used as the adversarial loss in training dynamics.
    This surrogate loss is the core premise of Thms 5.4, 5.7, and 5.9; it is not derived from the actual max over perturbations.
  • ad hoc to paper Lemma E.8: phi'(w^T x) w and x are independent for sufficiently large m.
    Used to establish Thm 4.1; the proof is informal and the statement is false at finite m.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Adversarial Training from Mean Field Perspective." pith.science (2026). https://pith.science/paper/L6LHFZOW

@misc{pith2026250514021,
  author       = {Pith},
  title        = {Pith review of: Adversarial Training from Mean Field Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/L6LHFZOW}},
  note         = {Machine review of arXiv:2505.14021}
}
abstract

Although adversarial training is known to be effective against adversarial examples, training dynamics are not well understood. In this study, we present the first theoretical analysis of adversarial training in random deep neural networks without any assumptions on data distributions. We introduce a new theoretical framework based on mean field theory, which addresses the limitations of existing mean field-based approaches. Based on this framework, we derive (empirically tight) upper bounds of $\ell_q$ norm-based adversarial loss with $\ell_p$ norm-based adversarial examples for various values of $p$ and $q$. Moreover, we prove that networks without shortcuts are generally not adversarially trainable and that adversarial training reduces network capacity. We also show that network width alleviates these issues. Furthermore, we present the various impacts of the input and output dimensions on the upper bounds and time evolution of the weight variance.

Figures

Figures reproduced from arXiv: 2505.14021 by the authors.

Figure 1
Figure 1. Distribution of J(x in)1,1 in the vanilla ReLU network with d = 1, 000, K = 1, N = 5, 000, L = 10, σ 2 w = 2, and σ 2 b = 0.01. The blue histogram represents the experimental re￾sults (10,000-time samplings), and the orange curve is predicted by Thm 4.1 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Adversarial loss (Eq. (3)) in vanilla networks with N = 40, 000, K = 100, L = 3, and ϵ = 0.1. We generated 100 adversarial examples for each input dimension. The blue curves and bands represent the mean and standard deviation of the adversarial loss, respectively, whereas the orange curves (upper bounds) are predicted based on Thm 5.1. Some samples slightly exceed the upper bounds because we used the finite network … view at source ↗
Figure 4
Figure 4. Heat map of the training accuracy of vanilla networks with p = ∞, q = ∞, and ϵ = 0.3. The dashed lines represent the condition of T in Thm 5.7 with m = 0.0001. In standard training, high accuracy is obtained across all the depths and widths (cf. Fig. A19). 6 Experimental results We validate Thms 5.1, 5.4 and 5.7 via numerical experiments. The vanilla ReLU networks were initialized with σ 2 w = 2 and σ 2 b = 0.01 to … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

115 extracted references · 67 canonical work pages

  1. [1]

    Alayrac, J

    J.-B. Alayrac, J. Uesato, P.-S. Huang, A. Fawzi, R. Stanforth, and P. Kohli. Are labels required for improving adversarial robustness? In NeurIPS, volume 32, 2019

  2. [2]

    Amsaleg, J

    L. Amsaleg, J. Bailey, D. Barbe, S. Erfani, M. E. Houle, V . Nguyen, and M. Radovanovi´c. The vulnerability of learning to adversarial perturbation increases with intrinsic dimensionality. In WIFS, pages 1–6, 2017

  3. [3]

    C. Anil, J. Lucas, and R. Grosse. Sorting out lipschitz function approximation. In ICML, pages 291–301, 2019

  4. [4]

    Arora, S

    S. Arora, S. S. Du, W. Hu, Z. Li, R. R. Salakhutdinov, and R. Wang. On exact computation with an infinitely wide neural net. In NeurIPS, volume 32, 2019

  5. [5]

    Athalye, N

    A. Athalye, N. Carlini, and D. Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In ICML, pages 274–283, 2018

  6. [6]

    Awasthi, N

    P. Awasthi, N. Frank, and M. Mohri. Adversarial learning guarantees for linear hypotheses and neural networks. In ICML, pages 431–441, 2020

  7. [7]

    Bartlett, S

    P. Bartlett, S. Bubeck, and Y . Cherapanamjeri. Adversarial examples in multi-layer random relu networks. In NeurIPS, volume 34, pages 9241–9252, 2021

  8. [8]

    P. L. Bartlett, D. J. Foster, and M. J. Telgarsky. Spectrally-normalized margin bounds for neural networks. In NeurIPS, volume 30, 2017. 10

Show all 115 references
  1. [9]

    P. L. Bartlett and S. Mendelson. Rademacher and gaussian complexities: Risk bounds and structural results. JMLR, 3(Nov):463–482, 2002

  2. [10]

    Blumenfeld, D

    Y . Blumenfeld, D. Gilboa, and D. Soudry. A mean field theory of quantized deep networks: The quantization-depth trade-off. In NeurIPS, volume 32, 2019

  3. [11]

    Bubeck, Y

    S. Bubeck, Y . Cherapanamjeri, G. Gidel, and R. Tachet des Combes. A single gradient step finds adversarial examples on random two-layers neural networks. In NeurIPS, volume 34, pages 10081–10091, 2021

  4. [12]

    Carlini and D

    N. Carlini and D. Wagner. Adversarial examples are not easily detected: Bypassing ten detection methods. In ACM WS, pages 3–14, 2017

  5. [13]

    Carlini and D

    N. Carlini and D. Wagner. Towards evaluating the robustness of neural networks. In SSP, pages 39–57, 2017

  6. [14]

    Carmon, A

    Y . Carmon, A. Raghunathan, L. Schmidt, P. Liang, and J. C. Duchi. Unlabeled data improves adversarial robustness. In NeurIPS, 2019

  7. [15]

    M. Chen, J. Pennington, and S. Schoenholz. Dynamical isometry and a mean field theory of RNNs: Gating enables signal propagation in recurrent neural networks. In ICML, pages 873–882, 2018

  8. [16]

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In ICML, pages 1597–1607, 2020

  9. [17]

    Cho and L

    Y . Cho and L. Saul. Kernel methods for deep learning. In NeurIPS, volume 22, 2009

  10. [18]

    Cisse, P

    M. Cisse, P. Bojanowski, E. Grave, Y . Dauphin, and N. Usunier. Parseval networks: Improving robustness to adversarial examples. In ICML, pages 854–863, 2017

  11. [19]

    Cohen, E

    J. Cohen, E. Rosenfeld, and Z. Kolter. Certified adversarial robustness via randomized smoothing. In ICML, pages 1310–1320, 2019

  12. [20]

    Croce and M

    F. Croce and M. Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In ICML, pages 2206–2216, 2020

  13. [21]

    Damianou and N

    A. Damianou and N. D. Lawrence. Deep gaussian processes. In AISTATS, pages 207–215, 2013

  14. [22]

    A. Daniely. SGD learns the conjugate kernel class of the network. In NeurIPS, volume 30, 2017

  15. [23]

    Daniely, R

    A. Daniely, R. Frostig, and Y . Singer. Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity. In NeurIPS, volume 29, 2016

  16. [24]

    Daniely and H

    A. Daniely and H. Schacham. Most relu networks suffer from ell ˆ2 adversarial perturbations. In NeurIPS, volume 33, pages 6629–6636, 2020

  17. [25]

    De Palma, B

    G. De Palma, B. Kiani, and S. Lloyd. Adversarial robustness guarantees for random deep neural networks. In ICML, pages 2522–2534, 2021

  18. [26]

    L. Deng. The MNIST database of handwritten digit images for machine learning research. Signal Process. Mag., 29(6):141–142, 2012

  19. [27]

    Z. Deng, L. Zhang, K. V odrahalli, K. Kawaguchi, and J. Y . Zou. Adversarial training helps transfer learning via better representations. In NeurIPS, volume 34, pages 25179–25191, 2021

  20. [28]

    G. W. Ding, Y . Sharma, K. Y . C. Lui, and R. Huang. MMA training: Direct input space margin maximization through adversarial training. In ICLR, 2020

  21. [29]

    Dobriban, H

    E. Dobriban, H. Hassani, D. Hong, and A. Robey. Provable tradeoffs in adversarially robust classification. arXiv:2006.05161, 2020. 11

  22. [30]

    Fawzi, H

    A. Fawzi, H. Fawzi, and O. Fawzi. Adversarial vulnerability for any classifier. In NeurIPS, volume 31, 2018

  23. [31]

    Fawzi, O

    A. Fawzi, O. Fawzi, and P. Frossard. Analysis of classifiers’ robustness to adversarial pertur- bations. ML, 107(3):481–508, 2018

  24. [32]

    Fawzi, S.-M

    A. Fawzi, S.-M. Moosavi-Dezfooli, and P. Frossard. Robustness of classifiers: from adversarial to random noise. In NeurIPS, volume 29, 2016

  25. [33]

    Fukushima

    K. Fukushima. Cognitron: A self-organizing multilayered neural network. Biol. Cybern., 20(3):121–136, 1975

  26. [34]

    Galloway, A

    A. Galloway, A. Golubeva, T. Tanay, M. Moussa, and G. W. Taylor. Batch normalization is a cause of adversarial vulnerability. In ICML WS, 2019

  27. [35]

    R. Gao, T. Cai, H. Li, C.-J. Hsieh, L. Wang, and J. D. Lee. Convergence of adversarial training in overparametrized neural networks. In NeurIPS, volume 32, 2019

  28. [36]

    Gilboa, B

    D. Gilboa, B. Chang, M. Chen, G. Yang, S. S. Schoenholz, E. H. Chi, and J. Pennington. Dynamical isometry and a mean field theory of LSTMs and GRUs. arXiv:1901.08987, 2019

  29. [37]

    Gilmer, L

    J. Gilmer, L. Metz, F. Faghri, S. S. Schoenholz, M. Raghu, M. Wattenberg, and I. Goodfellow. Adversarial spheres. In ICLR WS, 2018

  30. [38]

    I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. In ICLR, 2015

  31. [39]

    Gowal, K

    S. Gowal, K. D. Dvijotham, R. Stanforth, R. Bunel, C. Qin, J. Uesato, R. Arandjelovic, T. Mann, and P. Kohli. Scalable verified training for provably robust image classification. In ICCV, pages 4842–4851, 2019

  32. [40]

    Gowal, S.-A

    S. Gowal, S.-A. Rebuffi, O. Wiles, F. Stimberg, D. A. Calian, and T. A. Mann. Improving robustness using generated data. In NeurIPS, 2021

  33. [41]

    Hayou, A

    S. Hayou, A. Doucet, and J. Rousseau. On the selection of initialization and activation function for deep neural networks. arXiv:1805.08266, 2018

  34. [42]

    K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016

  35. [43]

    Hein and M

    M. Hein and M. Andriushchenko. Formal guarantees on the robustness of a classifier against adversarial manipulation. In NeurIPS, volume 30, 2017

  36. [44]

    J. Hron, Y . Bahri, J. Sohl-Dickstein, and R. Novak. Infinite attention: NNGP and NTK for deep attention networks. In ICML, pages 4376–4386, 2020

  37. [45]

    Huang, Z

    G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger. Densely connected convolutional networks. In CVPR, pages 4700–4708, 2017

  38. [46]

    Jacot, F

    A. Jacot, F. Gabriel, and C. Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In NeurIPS, volume 31, 2018

  39. [47]

    Javanmard, M

    A. Javanmard, M. Soltanolkotabi, and H. Hassani. Precise tradeoffs in adversarial training for linear regression. In COLT, pages 2034–2078, 2020

  40. [48]

    Kannan, A

    H. Kannan, A. Kurakin, and I. Goodfellow. Adversarial logit pairing. arXiv:1803.06373, 2018

  41. [49]

    Karakida, S

    R. Karakida, S. Akaho, and S.-i. Amari. Universal statistics of Fisher information in deep neural networks: Mean field approach. In AISTATS, pages 1032–1041, 2019

  42. [50]

    Khim and P.-L

    J. Khim and P.-L. Loh. Adversarial risk bounds via function transformation.arXiv:1810.09519, 2018

  43. [51]

    D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In CVPR, 2015. 12

  44. [52]

    Krizhevsky

    A. Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009

  45. [53]

    Krizhevsky, I

    A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, pages 1097–1105, 2012

  46. [54]

    J. Lee, Y . Bahri, R. Novak, S. S. Schoenholz, J. Pennington, and J. Sohl-Dickstein. Deep neural networks as gaussian processes. In ICLR, 2018

  47. [55]

    J. Lee, L. Xiao, S. Schoenholz, Y . Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. In NeurIPS, volume 32, 2019

  48. [56]

    Liang, T

    T. Liang, T. Poggio, A. Rakhlin, and J. Stokes. Fisher-rao metric, geometry, and complexity of neural networks. In AISTAT, pages 888–896, 2019

  49. [57]

    A. L. Maas, A. Y . Hannun, A. Y . Ng, et al. Rectifier nonlinearities improve neural network acoustic models. In ICML, volume 30, page 3, 2013

  50. [58]

    Madry, A

    A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks. In ICLR, 2018

  51. [59]

    A. G. d. G. Matthews, M. Rowland, J. Hron, R. E. Turner, and Z. Ghahramani. Gaussian process behaviour in wide deep neural networks. In ICLR, 2018

  52. [60]

    Miyato, T

    T. Miyato, T. Kataoka, M. Koyama, and Y . Yoshida. Spectral normalization for generative adversarial networks. In ICLR, 2018

  53. [61]

    Montanari and Y

    A. Montanari and Y . Wu. Adversarial examples in random neural networks with general activations. arXiv:2203.17209, 2022

  54. [62]

    Najafi, S.-i

    A. Najafi, S.-i. Maeda, M. Koyama, and T. Miyato. Robustness to adversarial perturbations in learning from incomplete data. In NeurIPS, volume 32, 2019

  55. [63]

    Nakkiran

    P. Nakkiran. Adversarial robustness may be at odds with simplicity. arXiv:1901.00532, 2019

  56. [64]

    R. M. Neal. Priors for infinite networks. In Bayesian Learning for Neural Networks, pages 29–53. Springer, 1996

  57. [65]

    Neyshabur, S

    B. Neyshabur, S. Bhojanapalli, D. McAllester, and N. Srebro. Exploring generalization in deep learning. In NeurIPS, volume 30, 2017

  58. [66]

    Neyshabur, R

    B. Neyshabur, R. R. Salakhutdinov, and N. Srebro. Path-SGD: Path-normalized optimization in deep neural networks. In NeurIPS, volume 28, 2015

  59. [67]

    Neyshabur, R

    B. Neyshabur, R. Tomioka, and N. Srebro. Norm-based capacity control in neural networks. In COLT, pages 1376–1401, 2015

  60. [68]

    Novak, L

    R. Novak, L. Xiao, J. Lee, Y . Bahri, G. Yang, J. Hron, D. A. Abolafia, J. Pennington, and J. Sohl-Dickstein. Bayesian deep convolutional networks with many channels are gaussian processes. In ICLR, 2019

  61. [69]

    Pennington, S

    J. Pennington, S. Schoenholz, and S. Ganguli. Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice. In NeurIPS, volume 30, 2017

  62. [70]

    Pennington, S

    J. Pennington, S. Schoenholz, and S. Ganguli. The emergence of spectral universality in deep networks. In AISTATS, pages 1924–1932, 2018

  63. [71]

    Poole, S

    B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli. Exponential expressivity in deep neural networks through transient chaos. In NeurIPS, volume 29, 2016

  64. [72]

    Rade and S.-M

    R. Rade and S.-M. Moosavi-Dezfooli. Helper-based adversarial training: Reducing excessive margin to achieve a better accuracy vs. robustness trade-off. In ICML, 2021

  65. [73]

    Raghu, B

    M. Raghu, B. Poole, J. Kleinberg, S. Ganguli, and J. Sohl-Dickstein. On the expressive power of deep neural networks. In ICML, pages 2847–2854, 2017. 13

  66. [74]

    Raghunathan, J

    A. Raghunathan, J. Steinhardt, and P. Liang. Certified defenses against adversarial examples. In ICLR, 2018

  67. [75]

    Raghunathan, S

    A. Raghunathan, S. M. Xie, F. Yang, J. Duchi, and P. Liang. Understanding and mitigating the tradeoff between robustness and accuracy. In ICML, 2020

  68. [76]

    Raghunathan, S

    A. Raghunathan, S. M. Xie, F. Yang, J. C. Duchi, and P. Liang. Adversarial training can hurt generalization. In ICML WS, 2019

  69. [77]

    Rebuffi, S

    S.-A. Rebuffi, S. Gowal, D. A. Calian, F. Stimberg, O. Wiles, and T. Mann. Fixing data augmentation to improve adversarial robustness. arXiv:2103.01946, 2021

  70. [78]

    K. Roth, Y . Kilcher, and T. Hofmann. Adversarial training is a form of data-dependent operator norm regularization. In NeurIPS, volume 33, pages 14973–14985, 2020

  71. [79]

    A. M. Saxe, J. L. McClelland, and S. Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. In ICLR, 2014

  72. [80]

    Schmidt, S

    L. Schmidt, S. Santurkar, D. Tsipras, K. Talwar, and A. Madry. Adversarially robust general- ization requires more data. In NeurIPS, 2018

  73. [81]

    S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein. Deep information propagation. In ICLR, 2017

  74. [82]

    Shafahi, W

    A. Shafahi, W. R. Huang, C. Studer, S. Feizi, and T. Goldstein. Are adversarial examples inevitable? In ICLR, 2019

  75. [83]

    Shafahi, M

    A. Shafahi, M. Najibi, A. Ghiasi, Z. Xu, J. Dickerson, C. Studer, L. S. Davis, G. Taylor, and T. Goldstein. Adversarial training for free! In NeurIPS, 2019

  76. [84]

    Simon-Gabriel, Y

    C.-J. Simon-Gabriel, Y . Ollivier, L. Bottou, B. Schölkopf, and D. Lopez-Paz. First-order adversarial vulnerability of neural networks and input dimension. In ICML, pages 5809–5817, 2019

  77. [85]

    Simonyan and A

    K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2014

  78. [86]

    Sinha, H

    A. Sinha, H. Namkoong, R. V olpi, and J. Duchi. Certifying some distributional robustness with principled adversarial training. In ICLR, 2018

  79. [87]

    Sompolinsky, A

    H. Sompolinsky, A. Crisanti, and H.-J. Sommers. Chaos in random neural networks. Phys. Rev. Lett., 61(3):259, 1988

  80. [88]

    D. Su, H. Zhang, H. Chen, J. Yi, P.-Y . Chen, and Y . Gao. Is robustness the cost of accuracy?–a comprehensive study on the robustness of 18 deep image classification models. In ECCV, pages 631–648, 2018

  81. [89]

    Szegedy, W

    C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. In ICLR, 2014

  82. [90]

    Tramer, N

    F. Tramer, N. Carlini, W. Brendel, and A. Madry. On adaptive attacks to adversarial example defenses. In NeurIPS, volume 33, pages 1633–1645, 2020

  83. [91]

    J. A. Tropp. Topics in sparse approximation. The University of Texas at Austin, 2004

  84. [92]

    Tsipras, S

    D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry. Robustness may be at odds with accuracy. In ICLR, 2019

  85. [93]

    Tsuzuku, I

    Y . Tsuzuku, I. Sato, and M. Sugiyama. Lipschitz-margin training: Scalable certification of perturbation invariance for deep neural networks. In NeurIPS, volume 31, 2018

  86. [94]

    Y . Wang, D. Zou, J. Yi, J. Bailey, X. Ma, and Q. Gu. Improving adversarial robustness requires revisiting misclassified examples. In ICLR, 2020

  87. [95]

    Williams

    C. Williams. Computing with infinite networks. In NeurIPS, volume 9, 1996. 14

  88. [96]

    Wong and Z

    E. Wong and Z. Kolter. Provable defenses against adversarial examples via the convex outer adversarial polytope. In ICML, pages 5286–5295, 2018

  89. [97]

    E. Wong, L. Rice, and J. Z. Kolter. Fast is better than free: Revisiting adversarial training. In ICLR, 2020

  90. [98]

    B. Wu, J. Chen, D. Cai, X. He, and Q. Gu. Do wider neural networks really help adversarial robustness? In NeurIPS, volume 34, pages 7054–7067, 2021

  91. [99]

    H. Xiao, K. Rasul, and R. V ollgraf. Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. arXiv:1708.07747, 2017

  92. [100]

    L. Xiao, Y . Bahri, J. Sohl-Dickstein, S. Schoenholz, and J. Pennington. Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks. In ICML, pages 5393–5402, 2018

  93. [101]

    L. Xiao, J. Pennington, and S. Schoenholz. Disentangling trainability and generalization in deep neural networks. In ICML, pages 10462–10472, 2020

  94. [102]

    Y . Xing, Q. Song, and G. Cheng. On the generalization properties of adversarial training. In AISTATS, pages 505–513, 2021

  95. [103]

    G. Yang. Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation. arXiv:1902.04760, 2019

  96. [104]

    G. Yang, J. Pennington, V . Rao, J. Sohl-Dickstein, and S. S. Schoenholz. A mean field theory of batch normalization. In ICLR, 2019

  97. [105]

    Yang and S

    G. Yang and S. Schoenholz. Mean field residual networks: On the edge of chaos. In NeurIPS, volume 30, 2017

  98. [106]

    Yang and S

    G. Yang and S. S. Schoenholz. Deep mean field theory: Layerwise variance and width variation as methods to control gradient explosion. OpenReview, 2018

  99. [107]

    D. Yin, R. Kannan, and P. Bartlett. Rademacher complexity for adversarially robust general- ization. In ICML, pages 7085–7094, 2019

  100. [108]

    Yoshida and T

    Y . Yoshida and T. Miyato. Spectral norm regularization for improving the generalizability of deep learning. arXiv:1705.10941, 2017

  101. [109]

    Zagoruyko and N

    S. Zagoruyko and N. Komodakis. Wide residual networks. In BMVC, pages 1–12, 2016

  102. [110]

    R. Zhai, T. Cai, D. He, C. Dan, K. He, J. Hopcroft, and L. Wang. Adversarially robust generalization just requires more unlabeled data. arXiv:1906.00555, 2019

  103. [111]

    Zhang, T

    D. Zhang, T. Zhang, Y . Lu, Z. Zhu, and B. Dong. You only propagate once: Accelerating adversarial training via maximal principle. In NeurIPS, 2019

  104. [112]

    Zhang, D

    H. Zhang, D. Yu, Y . Lu, and D. He. Adversarial noises are linearly separable for (nearly) random neural networks. arXiv:2206.04316, 2022

  105. [113]

    Zhang, Y

    H. Zhang, Y . Yu, J. Jiao, E. Xing, L. El Ghaoui, and M. Jordan. Theoretically principled trade-off between robustness and accuracy. In ICML, pages 7472–7482, 2019

  106. [114]

    Zhang, X

    J. Zhang, X. Xu, B. Han, G. Niu, L. Cui, M. Sugiyama, and M. Kankanhalli. Attacks which do not kill training make adversarial learning stronger. In ICML, pages 11278–11287, 2020

  107. [115]

    W (t)2 − W (t) nX i=1 ∂Li(xin) ∂W (t) dt + O(dt2) # (A136) = σ2 w(t) − N EW ∈W

    Y . Zhang, O. Plevrakis, S. S. Du, X. Li, Z. Song, and S. Arora. Over-parameterized adversarial training: An analysis overcoming the curse of dimensionality. In NeurIPS, volume 33, pages 679–688, 2020. 15 Table A2: Notation. While h(l) is a function that takes xin as input, we...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.