REVIEW 5 major objections 4 minor 233 references
Principles of Lipschitz continuity in neural networks
T0 review · 5 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Training noise is not a nuisance: it irreversibly inflates a neural network's worst-case input sensitivity, a phenomenon this thesis derives from first principles.
desk verdict A mathematically serious thesis whose Chapter 4 SDE framework is original but rests on an unvalidated diffusion approximation and in-sample validation; Chapters 2 and 3 are solid enough to justify referee time. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Jordan–Wielandt embedding turns a rectangular weight matrix into a self-adjoint block operator, allowing refined Kato analytic perturbation theory to yield closed-form higher-order Fréchet derivatives of singular values. The new closed-form singular-value Hessian (expressed in Kronecker-product form) supplies the second-order term needed for Itô calculus on the spectral norm. These two pieces produce the layer-wise SDE and its network-level aggregation.
What would settle it
Train a small MLP with a large learning rate under heavy-tailed gradient noise and compare the measured layer-wise Lipschitz trajectories' variance to the SDE prediction; systematic mismatch — or finding the noise–curvature term to be negative — would falsify the central claim. A simpler check: train with increasing label noise and verify the predicted monotone decrease of the final spectral-norm product; any non-monotone ordering falsifies the drift decomposition.
Extended reading notes
Core claim
The central result is a system of SDEs for the layer-wise spectral-norm Lipschitz bound. For each layer ℓ, dK(ℓ)/K(ℓ) = (μ(ℓ) + κ(ℓ)) dt + λ(ℓ)ᵀ dB(ℓ), where κ(ℓ) = η/(2σ₁)⟨H_op, Σ⟩ ≥ 0 is an entropy-production term coupling gradient noise covariance to the curvature of the operator norm, and λ(ℓ) is a diffusion intensity. Writing Z = Σ_ℓ log K(ℓ), the network bound K = e^Z, the thesis derives network-level drift, diffusion, and statistics. It claims that supervision noise shrinks the drift, that mini-batch trajectory does not affect variance for large batch size, that relative fluctuations of the bound grow unboundedly near convergence, and that batch size controls variance.
Load-bearing premise
Mini-batch SGD is replaced by a Gaussian diffusion whose covariance is measured from the very runs being predicted, and no two nonzero singular values ever become equal during training.
Editorial extensions
If this is right
- The Lipschitz bound has an irreducible upward drift produced by the interaction of SGD noise with the curvature of the spectral norm, even when the gradient flow itself would shrink it.
- Label noise during training systematically lowers the final Lipschitz bound, because supervision noise shrinks the optimization-induced drift.
- For sufficiently large batches, the variance of the bound becomes independent of the particular mini-batch trajectory.
- Batch size directly scales the diffusion intensity, providing a practical control knob for the fluctuation of robustness during training.
- Relative fluctuations of the bound grow without bound in the near-convergence regime, predicting a distinctive late-training signature.
Reading between the lines
- A cheap, theory-agnostic check: train the same architecture with increasing label noise and measure the final spectral-norm product; the framework predicts a monotone decrease, reverse ordering would undercut the drift decomposition.
- The same singular-value Hessian machinery could generate SDEs for other spectral functionals (stable rank, effective dimension, von Neumann entropy of the Gram matrix), extending the framework beyond Lipschitz bounds.
- The near-convergence divergence prediction suggests that late-training spectral-norm spikes, sometimes blamed on optimization failure, may be an intrinsic diffusion-dominated phenomenon — a testable hypothesis for training diagnostics.
- The Gaussian diffusion approximation is the fragile point; measuring heavy-tailedness of mini-batch gradient noise during large-learning-rate training would delimit the regime where the SDE predictions hold.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This thesis, compiled from four papers and centered on Lipschitz continuity in neural networks, investigates two directions: an internal one (the temporal evolution of a spectral-norm Lipschitz bound during training) and an external one (how Lipschitz continuity modulates frequency signal propagation and robustness). Chapter 2 provides a survey with corrected activation-function Lipschitz constants, a sum-over-paths bound for DAG networks, and ℓ2/conjugate certified robustness radii. Chapter 3 develops an operator-theoretic perturbation framework for singular-value derivatives via the Jordan–Wielandt embedding and Kato expansions. Chapter 4 models mini-batch SGD as a Gaussian diffusion SDE and derives layer-wise and network-level dynamics for the log Lipschitz bound, with drift μ, noise–curvature production κ ≥ 0, and diffusion λ, yielding predictions on label noise, batch size, mini-batch trajectory, and near-convergence behavior. Chapters 5 and 6 connect Lipschitz bounds to Fourier frequency analysis and a Shapley-based spectral robustness score.
Significance. If fully established, Chapter 4's SDE framework would be a novel unifying account of Lipschitz dynamics during training, with concrete falsifiable predictions. The thesis has several verifiable strengths: reproducible code links, correct spot-checked constants (softmax 1/2, sigmoid 1/4), a DAG sum-over-paths bound, and a closed-form singular-value Hessian not previously in the literature. However, the central dynamics result rests on an unvalidated diffusion approximation and in-sample validation; Chapter 3's arbitrary-order perturbation theorem also appears incorrect for n ≥ 3. These issues do not necessarily invalidate the n = 1, 2 formulas used in Chapter 4, but they place the manuscript in major-revision territory.
major comments (5)
- [Thm 3.3.3 (Eq. 3.55)] Theorem 3.3.3 is not the correct eigenvalue coefficient for n ≥ 3. For T(x) = T0 + xV, Eq. (3.55) reduces to λ^(3) = <VSVSV>; the standard Rayleigh–Schrödinger expansion gives λ^(3) = <VSVSV> − <V><VS²V> (checkable in the 2×2 case T0 = diag(0,Δ), V = [[a,b],[c,d]]). The proof drops higher-order-pole contributions in the residue step (Eqs. 3.110–3.115), which do not vanish. Consequently the claimed arbitrary-order singular-value derivatives are not established. The n = 1, 2 cases used in Chapter 4 appear unaffected, but the theorem and all 'arbitrary-order' claims must be corrected or restricted to n ≤ 2.
- [Def. 4.3.2; §4.7, Fig. 4.2] The central SDE replaces discrete mini-batch SGD with dvecθ = −∇L dt + √η Σ^{1/2} dB without any error analysis: no bound in η, batch size, or gradient-noise tails, and no justification for Gaussianity. The validation estimates Σ_t from the same runs whose K(t) is then 'predicted'; agreement in Fig. 4.2 is therefore a consistency check of Itô calculus, not an out-of-sample test. To support the Chapter 4 conclusions, provide a formal approximation theorem (weak/strong error) and validate on independent runs — e.g., covariance estimated from one window and the trajectory predicted on a disjoint window — or on synthetic SGD with known noise.
- [Assumption 3.2.2; Lemmas 4.6.2–4.6.3] The dynamics coefficients require simplicity of all non-zero singular values; the operator-norm Jacobian/Hessian require the top singular value to stay simple along the whole training trajectory. The thesis acknowledges this (Sec. 4.9), but no experiment reports the spectral gap. At initialization and under parameter symmetries spectral collisions occur, so the SDE coefficients are undefined at those times. Add empirical spectral-gap diagnostics or a nonsmooth extension; otherwise the derived drift/diffusion do not describe the actual trajectory.
- [§4.8.3, Prop. 4.8.1] The 'unbounded growth near convergence' prediction assumes Σ_t stays non-degenerate as ∇L → 0. For clean, overparameterized networks at interpolation, every mini-batch gradient is zero at the minimizer, so Σ_t → 0 and both λ and κ vanish; relative fluctuations of K need not diverge. The claim should explicitly state the non-degeneracy condition on Σ_t (e.g., label-noise-driven dynamics) or be restricted to that regime.
- [§2.2.9; §4.8] The statement that the Lipschitz bound 'irreversibly increases' because κ_Z ≥ 0 is not implied: κ_Z is only one additive component of the d log K drift, and μ_Z can be negative and dominate. Reformulate as a claim about the nonnegative noise–curvature contribution rather than monotonicity of K(t).
minor comments (4)
- [Table 2.1] The table lists the Swish and GELU constants as ≈1.1; since exact expressions are derived in Appendix 2.B.4–2.B.5, reporting the closed forms would be clearer.
- [Prop. 2.2.17] The SVD statement assumes distinct positive singular values (σ1 > σ2 > ... > σr), but the spectral norm is defined without a simplicity requirement; use ≥.
- [Eq. (3.141)] The expression D^n σ_k[dA,...,dA] = n! lim_{x→0} x^n σ_k^{(n)} is dimensionally awkward; if σ_k^{(n)} is the Taylor coefficient, the derivative is simply n! σ_k^{(n)}. Please clarify the notation.
- [Appendix 2.B.4] Typo: 'LiprSiwshpxqs' should read 'LiprSwishpxqs'. Also, the proof of the DAG bound in Theorem 2.2.20 could state explicitly that the modules h_v are assumed 1-Lipschitz in their inputs for the constant C_{u→v} = Lip[h_v] to be well-defined per edge.
Circularity Check
No significant circularity: the core derivation applies Itô's lemma to measured, not fitted, quantities; self-citations are to independently published work.
full rationale
The central derivation chain (Ch. 4, summarized in §2.2.9, Eqs. 2.116–2.119) takes the vectorized SDE for continuous-time SGD (Def. 4.3.2) as a stated modeling assumption and applies Itô's lemma to the layer-wise spectral-norm bound K^(ℓ)(t)=‖θ^(ℓ)(t)‖_op. The resulting drift μ^(ℓ), noise–curvature term κ^(ℓ), and diffusion λ^(ℓ) are explicit functions of the measured gradient, the measured batch-gradient covariance Σ_t, and the singular vectors/operator-norm Hessian of the current weight matrix. No parameter is fitted to the Lipschitz trajectory K(t), so the derived SDE is not equivalent to its output by construction. The operator-norm Hessian used in Ch. 4 comes from Ch. 3, which is an independently peer-reviewed publication (Róisín Luo et al., 2025c, JMAA) with stated simplicity assumptions; citing it is legitimate external support, not a load-bearing self-citation chain. The acknowledged simplicity assumption (Assumption 3.2.2) is a limitation that can invalidate the theory at spectral collisions, but an unvalidated or restrictive assumption is a correctness risk, not circularity. Similarly, measuring Σ from the same training runs when validating the SDE is an in-sample consistency check rather than an out-of-sample test, and this weakens the evidence without making the prediction identical to its input: K(t) is not used to determine Σ or any free parameter. The thesis's self-citations are to its own compiled articles and do not function as an external-authority uniqueness argument. I therefore find no circular step that meets the required standard of exhibiting a specific reduction of a claimed prediction to its own fitted or definitional input.
Assumptions & free parameters
free parameters (3)
- learning rate η =
small positive constant (standard SGD setting)
- layer-wise gradient-noise covariance Σ_t^(ℓ) (and its square root) =
estimated per layer from mini-batch gradients (Props. 4.4.1–4.4.3)
- spectral-band partition and coalition design (Ch. 6) =
ℓ∞-ball over ℓ2-ball banding; sample counts (Appendix 6.C.1, Figs. 6.15–6.16)
assumptions (5)
- domain assumption Discrete mini-batch SGD is modeled as the diffusion dvecθ = −vec∇L dt + √η Σ^{1/2} dB (Def. 4.3.2)
- domain assumption All non-zero singular values of every weight matrix are simple throughout training (Assumption 3.2.2)
- domain assumption The network's true Lipschitz constant is tracked via the upper bound K(t) = Π_ℓ ‖θ^(ℓ)(t)‖ (Prop. 4.5.1, Def. 4.5.3)
- domain assumption Input domain is convex so the tight constant equals sup_x ‖∇f(x)‖ (Remark 2.2.6)
- standard math Standard analytic toolbox: Kato perturbation theory, Jordan–Wielandt embedding spectra, Itô's lemma, Popoviciu's inequality, Rademacher complexity bounds
invented entities (2)
-
Noise–curvature entropy production κ^(ℓ)(t) (and aggregate κ_Z)
independent evidence
-
Spectral Robustness Score (SRS)
independent evidence
Cite this review
Pith. "Pith review of Principles of Lipschitz continuity in neural networks." pith.science (2026). https://pith.science/paper/A3RAY2ZJ
@misc{pith2026260204078,
author = {Pith},
title = {Pith review of: Principles of Lipschitz continuity in neural networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/A3RAY2ZJ}},
note = {Machine review of arXiv:2602.04078}
}
read the original abstract
Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable in artificial intelligence. Yet, despite these advances, critical challenges remain -- most notably, ensuring robustness to small input perturbations and generalization to out-of-distribution data. These critical challenges underscore the need to understand the underlying fundamental principles that govern robustness and generalization. Among the theoretical tools available, Lipschitz continuity plays a pivotal role in governing the fundamental properties of neural networks related to robustness and generalization. It quantifies the worst-case sensitivity of network's outputs to small input perturbations. While its importance is widely acknowledged, prior research has predominantly focused on empirical regularization approaches based on Lipschitz constraints, leaving the underlying principles less explored. This thesis seeks to advance a principled understanding of the principles of Lipschitz continuity in neural networks within the paradigm of machine learning, examined from two complementary perspectives: an internal perspective -- focusing on the temporal evolution of Lipschitz continuity in neural networks during training (i.e., training dynamics); and an external perspective -- investigating how Lipschitz continuity modulates the behavior of neural networks with respect to features in the input data, particularly its role in governing frequency signal propagation (i.e., modulation of frequency signal propagation).
Figures
Figures from the paper (37 more)
Reference graph
Works this paper leans on
-
[1]
K. Aas, M. Jullum, and A. L land. Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artificial Intelligence, 298 0 (C), Sept. 2021. ISSN 0004-3702. doi:10.1016/j.artint.2021.103502. URL https://doi.org/10.1016/j.artint.2021.103502
arXiv 2021
-
[2]
Absil, R
P.-A. Absil, R. Mahony, and R. Sepulchre. Optimization Algorithms on Matrix Manifolds. Princeton University Press, USA, 2007. ISBN 0691132984
2007
-
[3]
Amerehi and P
F. Amerehi and P. Healy. Label augmentation for neural networks robustness. In V. Lomonaco, S. Melacci, T. Tuytelaars, S. Chandar, and R. Pascanu, editors, Proceedings of The 3rd Conference on Lifelong Learning Agents, volume 274 of Proceedings of Machine Learning Research, pages 620--640. PMLR, 29 Jul--01 Aug 2025. URL https://proceedings.mlr.press/v274/...
2025
-
[4]
C. Anil, J. Lucas, and R. Grosse. Sorting out L ipschitz function approximation. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 291--301. PMLR, 09--15 Jun 2019. URL https://proceedings.mlr.press/v97/anil19a.html
2019
-
[5]
R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Silver, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. Barham, T. Hennigan, B. Lee, F. Viola, M. Reynolds, Y. Xu, R. Doherty, E. Collins, C. Meyer, E. Ruthe...
arXiv 2025
-
[6]
Applebaum
D. Applebaum. L\'evy Processes and Stochastic Calculus. Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2nd edition, 2009
2009
-
[7]
Arjovsky, A
M. Arjovsky, A. Shah, and Y. Bengio. Unitary evolution recurrent neural networks. In M. F. Balcan and K. Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1120--1128, New York, New York, USA, 20--22 Jun 2016. PMLR. URL https://proceedings.mlr.press/v48...
2016
-
[8]
Arjovsky, S
M. Arjovsky, S. Chintala, and L. Bottou. W asserstein generative adversarial networks. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 214--223. PMLR, 06--11 Aug 2017. URL https://proceedings.mlr.press/v70/arjovsky17a.html
2017
Show all 233 references
-
[9]
R. J. Aumann and M. Maschler. Game theoretic analysis of a bankruptcy problem from the Talmud . Journal of economic theory, 36 0 (2): 0 195--213, 1985
1985
-
[10]
Aumann and Y
Y. Aumann and Y. Dombb. The efficiency of fair division with connected pieces. ACM Transactions on Economics and Computation (TEAC), 3 0 (4): 0 1--16, 2015
2015
-
[11]
L. J. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization. 2016. URL http://arxiv.org/abs/1607.06450
2016 arXiv
-
[12]
T. Bai, J. Luo, J. Zhao, B. Wen, and Q. Wang. Recent advances in adversarial training for adversarial robustness. In Z.-H. Zhou, editor, Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21 , pages 4312--4321. International Joint Con...
2021 doi
-
[13]
Bansal, X
N. Bansal, X. Chen, and Z. Wang. Can we gain more from orthogonality regularizations in training deep CNN s? In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS'18, page 4266–4276, Red Hook, NY, USA, 2018. Curran Associates Inc
2018
-
[14]
P. L. Bartlett, D. J. Foster, and M. Telgarsky. Spectrally-normalized margin bounds for neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, page 6241–6250, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN ...
2017
-
[15]
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba. Network dissection: Quantifying interpretability of deep visual representations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6541--6549, 2017
2017
-
[16]
Behrmann, W
J. Behrmann, W. Grathwohl, R. T. Q. Chen, D. Duvenaud, and J.-H. Jacobsen. Invertible residual networks. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, ...
2019
-
[17]
Bereska and S
L. Bereska and S. Gavves. Mechanistic interpretability for AI safety --- A review. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/forum?id=ePUVetPKu6. Survey Certification, Expert Certification
2024
-
[18]
Binder, G
A. Binder, G. Montavon, S. Lapuschkin, K.-R. M \"u ller, and W. Samek. Layer-wise relevance propagation for neural networks with local renormalization layers. In Artificial Neural Networks and Machine Learning--ICANN 2016: 25th International Conference on Artificial Neural Net...
2016
-
[19]
Bj\" o rck and C. Bowie. An iterative algorithm for computing the best estimate of an orthogonal matrix. SIAM Journal on Numerical Analysis, 8 0 (2): 0 358--364, 1971. doi:10.1137/0708036. URL https://doi.org/10.1137/0708036
1971 doi
-
[20]
S. Boyd, J. Duchi, M. Pilanci, and L. Vandenberghe. Notes for EE364b : Subgradients. Stanford University, 2022. URL https://stanford.edu/class/ee364b/lectures/subgradients_notes.pdf. Lecture notes for Spring 2021--22; Accessed on January 2nd 2026
2022
-
[21]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020
1901
-
[22]
N. J. Calkin, E. Y. S. Chan, R. M. Corless, D. J. Jeffrey, and P. W. Lawrence. A fractal eigenvector. The American Mathematical Monthly, 129 0 (6): 0 503--523, 2022
2022
-
[23]
Carlini and D
N. Carlini and D. Wagner. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy (SP), pages 39--57. IEEE, 2017
2017
-
[24]
Castin, P
V. Castin, P. Ablin, and G. Peyr\' e . How smooth is attention? In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR, 2024
2024
-
[25]
A. Cayley. Sur quelques propri \'e t \'e s des d \'e terminants gauches. Journal f\"ur die reine und angewandte Mathematik, 1846
-
[26]
Chaudhari, A
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina. Entropy- SGD : Biasing gradient descent into wide valleys. Journal of Statistical Mechanics: Theory and Experiment, 2019 0 (12): 0 124018, 2019
2019
-
[27]
Chen and R
L. Chen and R. Ng. On the marriage of Lp -norms and edit distance. In Proceedings of the Thirtieth International Conference on Very Large Data Bases - Volume 30, VLDB '04, page 792–803. VLDB Endowment, 2004. ISBN 0120884690
2004
-
[28]
R. T. Q. Chen, J. Behrmann, D. K. Duvenaud, and J.-H. Jacobsen. Residual flows for invertible generative modeling. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. C...
2019
-
[29]
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning, pages 1597--1607. PMLR, 2020
2020
-
[30]
Chernodub and D
A. Chernodub and D. Nowicki. Norm-preserving orthogonal permutation linear unit activation functions ( OPLU ), 2017. URL https://arxiv.org/abs/1604.02313
2017 arXiv
-
[31]
Chowdhery, S
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Au...
2023
-
[32]
Cisse, P
M. Cisse, P. Bojanowski, E. Grave, Y. Dauphin, and N. Usunier. Parseval networks: improving robustness to adversarial examples. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML'17, page 854–863. PMLR, 2017
2017
-
[33]
F. H. Clarke. Generalized gradients and applications. Transactions of the American Mathematical Society, 205: 0 247--262, 1975
1975
-
[34]
F. H. Clarke. Optimization and Nonsmooth Analysis, volume 5 of Classics in Applied Mathematics. SIAM, Philadelphia, PA, second edition, 1990
1990
-
[35]
Clevert, T
D. Clevert, T. Unterthiner, and S. Hochreiter. Fast and accurate deep network learning by exponential linear units ( ELU s). In Y. Bengio and Y. LeCun, editors, 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conferenc...
2016 arXiv
-
[36]
Coates, A
A. Coates, A. Ng, and H. Lee. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 215--223. JMLR Workshop and Conference Proceedings, 2011
2011
-
[37]
I. C. Covert, S. Lundberg, and S.-I. Lee. Understanding global feature contributions with additive importance measures. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS '20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN ...
2020
-
[38]
G. Cybenko. Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals and Systems, 2 0 (4): 0 303--314, 1989. doi:10.1007/BF02551274. URL https://doi.org/10.1007/BF02551274
1989 doi
-
[39]
Damian, E
A. Damian, E. Nichani, and J. D. Lee. Self-stabilization: The implicit bias of gradient descent at the edge of stability. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=nhKHA59gXz
2023
-
[40]
DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li...
2025 arXiv
-
[41]
Devaguptapu, D
C. Devaguptapu, D. Agarwal, G. Mittal, P. Gopalani, and V. N. Balasubramanian. On adversarial robustness: A neural architecture search perspective. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 152--161, 2021
2021
-
[42]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), pages 4171--4186, 2019
2019
-
[43]
L. Dinh, J. Sohl-Dickstein, and S. Bengio. Density estimation using real NVP . In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=HkpbnH9lx
2017
-
[44]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning R...
2021 arXiv
-
[45]
Dugas, Y
C. Dugas, Y. Bengio, F. B\' e lisle, C. Nadeau, and R. Garcia. Incorporating second-order functional knowledge for better option pricing. In T. Leen, T. Dietterich, and V. Tresp, editors, Advances in Neural Information Processing Systems, volume 13. MIT Press, 2000
2000
-
[46]
Dunford and J
N. Dunford and J. T. Schwartz. Linear operators, part 1: general theory. John Wiley & Sons, 1988
1988
-
[47]
Edelman and N
A. Edelman and N. R. Rao. Random matrix theory. Acta numerica, 14: 0 233--297, 2005
2005
-
[48]
Elsken, J
T. Elsken, J. H. Metzen, and F. Hutter. Neural architecture search: A survey. Journal of Machine Learning Research, 20 0 (55): 0 1--21, 2019. URL http://jmlr.org/papers/v20/18-598.html
2019
-
[49]
Ethics guidelines for trustworthy AI , 2019
European Commission . Ethics guidelines for trustworthy AI , 2019
2019
-
[50]
Regulation (EU) 2024/1689 of the european parliament and of the council on harmonised rules on artificial intelligence ( AI Act ), 2024
European Parliament and Council . Regulation (EU) 2024/1689 of the european parliament and of the council on harmonised rules on artificial intelligence ( AI Act ), 2024
2024
-
[51]
U. Fano. Description of states in quantum mechanics by density matrix and operator techniques. Reviews of modern physics, 29 0 (1): 0 74, 1957
1957
-
[52]
Fazlyab, A
M. Fazlyab, A. Robey, H. Hassani, M. Morari, and G. Pappas. Efficient and accurate estimation of L ipschitz constants for deep neural networks. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Pro...
2019
-
[53]
Fazlyab, T
M. Fazlyab, T. Entesari, A. Roy, and R. Chellappa. Certified robustness via dynamic margin maximization and improved Lipschitz regularization. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=LDhhi8HBO3
2023
-
[54]
Flatow and D
D. Flatow and D. Penner. On the robustness of convnets to training on noisy labels. Technical report, Stanford University, 2017
2017
-
[55]
J. N. Franklin. Matrix theory. Courier Corporation, 2000
2000
-
[56]
Frenay and M
B. Frenay and M. Verleysen. Classification in the presence of label noise: A survey. IEEE Transactions on Neural Networks and Learning Systems, 25 0 (5): 0 845--869, 2014. doi:10.1109/TNNLS.2013.2292894
2014
-
[57]
Gamba, H
M. Gamba, H. Azizpour, and M. Bjorkman. On the Lipschitz constant of deep networks and double descent. In 34th British Machine Vision Conference 2023, BMVC 2023, Aberdeen, UK, November 20-24, 2023 . BMVA, 2023. URL https://papers.bmvc2023.org/0871.pdf
2023
-
[58]
Ghorbani, J
A. Ghorbani, J. Wexler, J. Y. Zou, and B. Kim. Towards automatic concept-based explanations. In H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch \' e - Buc, E. B. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Ne...
2019
-
[59]
Ghosh, H
A. Ghosh, H. Kumar, and P. S. Sastry. Robust loss functions under label noise for deep neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 31, 2017
2017
-
[60]
Golowich, A
N. Golowich, A. Rakhlin, and O. Shamir. Size-independent sample complexity of neural networks. In S. Bubeck, V. Perchet, and P. Rigollet, editors, Proceedings of the 31st Conference On Learning Theory, volume 75 of Proceedings of Machine Learning Research, pages 297--299. PMLR...
2018
-
[61]
G. H. Golub and C. F. Van Loan. Matrix computations. JHU press, 2013
2013
-
[62]
Gonon, N
A. Gonon, N. Brisebarre, E. Riccietti, and R. Gribonval. A rescaling-invariant Lipschitz bound based on path-metrics for modern ReLU network parameterizations. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=T8VLY1KuOz
2025
-
[63]
Goodfellow, J
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems (NeurIPS), pages 2672--2680, 2014
2014
-
[64]
Goodfellow, Y
I. Goodfellow, Y. Bengio, and A. Courville. Deep learning. MIT Press, 2016
2016
-
[65]
I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. In Y. Bengio and Y. LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , 2015. URL htt...
2015 arXiv
-
[66]
J. Gou, B. Yu, S. J. Maybank, and D. Tao. Knowledge distillation: A survey. International Journal of Computer Vision, 129 0 (6): 0 1789--1819, 2021
2021
-
[67]
H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree. Regularisation of neural networks by enforcing L ipschitz continuity. Machine Learning, 110: 0 393--416, 2021
2021
-
[68]
Grosse and J
R. Grosse and J. Martens. A Kronecker --factored approximate Fisher matrix for convolution layers. In International Conference on Machine Learning, pages 573--582. PMLR, 2016
2016
-
[69]
Gulrajani, F
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville. Improved training of Wasserstein GANs . In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. ...
2017
-
[70]
K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026--1034, 2015
2015
-
[71]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770--778, 2016
2016
-
[72]
Hein and M
M. Hein and M. Andriushchenko. Formal guarantees on the robustness of a classifier against adversarial manipulation. Advances in neural information processing systems, 30, 2017
2017
-
[73]
Hendrycks and T
D. Hendrycks and T. Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HJz6tiCqYm
2019
-
[74]
Hendrycks and K
D. Hendrycks and K. Gimpel. Gaussian error linear units ( GELUs ). 2016. URL https://arxiv.org/abs/1606.08415
2016 arXiv
-
[75]
Hendrycks, M
D. Hendrycks, M. Mazeika, D. Wilson, and K. Gimpel. Using trusted data to train deep networks on labels corrupted by severe noise. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, ...
2018
-
[76]
Hendrycks, M
D. Hendrycks, M. Mazeika, and T. Dietterich. Deep anomaly detection with outlier exposure. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyxCxhRcY7
2019
-
[77]
Hendrycks, S
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF international conference on computer vision, ...
2021
-
[78]
Hinton, N
G. Hinton, N. Srivastava, and K. Swersky. Neural networks for machine learning lecture (6a): Overview of mini-batch gradient descent. URL https://www.cs.toronto.edu/ tijmen/csc321/slides/lecture_slides_lec6.pdf
-
[79]
G. E. Hinton and R. R. Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313 0 (5786): 0 504--507, 2006
2006
-
[80]
J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 6840--6851. Curran Associates, Inc., 2020. URL https://proceed...
2020
-
[81]
Hochreiter and J
S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural Computation, 9 0 (8): 0 1735--1780, 1997. doi:10.1162/neco.1997.9.8.1735
1997 doi
-
[82]
R. A. Horn. The Hadamard product. In Proc. Symp. Appl. Math, volume 40, pages 87--169, 1990
1990
-
[83]
R. A. Horn and C. R. Johnson. Matrix analysis. Cambridge University Press, 2012
2012
-
[84]
X. Hu, R. Zheng, J. Wang, C. H. Leung, Q. Wu, and X. Xie. SpecFormer : Guarding vision transformer robustness via maximum singular value penalization. In European Conference on Computer Vision, pages 345--362. Springer, 2024
2024
-
[85]
Huang, X
L. Huang, X. Liu, B. Lang, A. W. Yu, Y. Wang, and B. Li. Orthogonal weight normalization: solution to optimization over multiple dependent Stiefel manifolds in deep neural networks. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth In...
2018
-
[86]
Ilyas, S
A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry. Adversarial examples are not bugs, they are features. Advances in neural information processing systems, 32, 2019
2019
-
[87]
K. It \^o . On stochastic differential equations. Number 4. American Mathematical Soc., 1951
1951
-
[88]
A. K. Jain. Fundamentals of digital image processing. Englewood Cliffs, NJ: Prentice Hall, 1989
1989
-
[89]
Jastrzębski, Z
S. Jastrzębski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey. On the relation between the sharpest directions of DNN loss and the SGD step length. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=SkgEaj05t7
2019
-
[90]
Jordan and A
M. Jordan and A. G. Dimakis. Exactly computing the local L ipschitz constant of ReLU networks. In International Conference on Machine Learning (ICML), pages 4985--4994. PMLR, 2020. URL https://arxiv.org/abs/2002.11572
2020 arXiv
-
[91]
Karatzas and S
I. Karatzas and S. Shreve. Brownian motion and stochastic calculus, volume 113. Springer Science & Business Media, 2012
2012
-
[92]
T. Kato. Perturbation theory for linear operators. Springer, 1995
1995
-
[93]
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang. On large-batch training for deep learning: Generalization gap and sharp minima. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=H1oyRlYgg
2017
-
[94]
Khromov and S
G. Khromov and S. P. Singh. Some fundamental aspects about L ipschitz continuity of neural networks. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=5jWsW08zUh
2024
-
[95]
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, and R. sayres. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors ( TCAV ). In J. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machi...
2018
-
[96]
H. Kim, G. Papamakarios, and A. Mnih. The L ipschitz constant of self-attention. In M. Meila and T. Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5562--5571. PMLR, 18--24 Jul ...
2021
-
[97]
D. P. Kingma and J. Ba. Adam: A method for stochastic optimization, 2017. URL https://arxiv.org/abs/1412.6980
2017 arXiv
-
[98]
T. N. Kipf and M. Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=SJU4ayYgl
2017
-
[99]
C. Klamler. Fair division. Handbook of group decision and negotiation, pages 183--202, 2010
2010
-
[100]
P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang. Concept bottleneck models. In International Conference on Machine Learning, pages 5338--5348. PMLR, 2020
2020
-
[101]
T. G. Kolda and B. W. Bader. Tensor decompositions and applications. SIAM review, 51 0 (3): 0 455--500, 2009
2009
-
[102]
Kolek, D
S. Kolek, D. A. Nguyen, R. Levie, J. Bruna, and G. Kutyniok. Cartoon explanations of image classifiers. In European Conference on Computer Vision, pages 443--458. Springer, 2022
2022
-
[103]
T. W. K \"o rner. Fourier analysis. Cambridge University Press, 2014. ISBN 9781107049949
2014
-
[104]
Krizhevsky, I
A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012
2012
-
[105]
Lakkaraju, E
H. Lakkaraju, E. Kamar, R. Caruana, and J. Leskovec. Faithful and customizable explanations of black box models. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pages 131--138, 2019
2019
-
[106]
V. I. Levenshtein. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10: 0 707--710, 1966
1966
-
[107]
A. S. Lewis and H. S. Sendov. Nonsmooth analysis of singular values. Part I : Theory. Set-Valued Analysis, 13 0 (3): 0 213--241, 2005
2005
-
[108]
Lezcano-Casado and D
M. Lezcano-Casado and D. Mart\' nez-Rubio. Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume...
2019
-
[109]
Li and R.-C
C.-K. Li and R.-C. Li. A note on eigenvalues of perturbed Hermitian matrices. Linear algebra and its applications, 395: 0 183--190, 2005
2005
-
[110]
J. Li, D. Li, C. Xiong, and S. Hoi. BLIP : Bootstrapping language-image pre-training for unified vision-language understanding and generation. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, editors, Proceedings of the 39th International Conference ...
2022
-
[111]
Q. Li, S. Haque, C. Anil, J. Lucas, R. B. Grosse, and J.-H. Jacobsen. Preventing gradient attenuation in Lipschitz constrained convolutional networks. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Informat...
2019
-
[112]
Q. Li, C. Tai, and E. Weinan. Stochastic modified equations and dynamics of stochastic gradient algorithms I : Mathematical foundations. Journal of Machine Learning Research, 20 0 (40): 0 1--47, 2019 b
2019
-
[113]
Li and C
Y. Li and C. Xu. Trade-off between robustness and accuracy of vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7558--7568, 2023
2023
-
[114]
Li and Y
Y. Li and Y. Yuan. Convergence analysis of two-layer neural networks with ReLU activation. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, I...
2017
-
[115]
Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou. A survey of convolutional neural networks: analysis, applications, and prospects. IEEE transactions on neural networks and learning systems, 33 0 (12): 0 6999--7019, 2021
2021
-
[116]
Z. Li, T. Wang, and S. Arora. What happens after SGD reaches zero loss? --- A mathematical framework. In International Conference on Learning Representations, 2022 b . URL https://openreview.net/forum?id=siCt4xZn5Ve
2022
-
[117]
Lipman, R
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=PqvMRDCJT9t
2023
-
[118]
Lukasik, S
M. Lukasik, S. Bhojanapalli, A. Menon, and S. Kumar. Does label smoothing mitigate label noise? In H. D. III and A. Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 6448--6458. P...
2020
-
[119]
S. M. Lundberg and S.-I. Lee. A unified approach to interpreting model predictions. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017
2017
-
[120]
C. Luo, Q. Lin, W. Xie, B. Wu, J. Xie, and L. Shen. Frequency-driven imperceptible adversarial attack on semantic similarity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15315--15324, 2022
2022
-
[121]
A. L. Maas, A. Y. Hannun, and A. Y. Ng. Rectifier nonlinearities improve neural network acoustic models. In ICML Workshop on Deep Learning for Audio, Speech and Language Processing, 2013
2013
-
[122]
Macdonald, S
J. Macdonald, S. W \"a ldchen, S. Hauch, and G. Kutyniok. A rate-distortion framework for explaining neural network decisions. arXiv preprint arXiv:1905.11092, 2019
1905 arXiv
-
[123]
Madry, A
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks. International Conference on Learning Representations (ICLR), 2018
2018
-
[124]
J. R. Magnus and H. Neudecker. Matrix differential calculus with applications in statistics and econometrics. John Wiley & Sons, 2019
2019
-
[125]
Malladi, K
S. Malladi, K. Lyu, A. Panigrahi, and S. Arora. On the SDEs and scaling rules for adaptive gradient algorithms. Advances in Neural Information Processing Systems, 35: 0 7697--7711, 2022
2022
-
[126]
Mandt, M
S. Mandt, M. D. Hoffman, D. M. Blei, et al. Continuous-time limit of stochastic gradient descent revisited. NIPS-2015, 2015
2015
-
[127]
Mandt, M
S. Mandt, M. D. Hoffman, and D. M. Blei. Stochastic gradient descent as approximate Bayesian inference. Journal of Machine Learning Research, 18 0 (134): 0 1--35, 2017
2017
-
[128]
V. A. Mar c enko and L. A. Pastur. Distribution of eigenvalues for some sets of random matrices. Mathematics of the USSR-Sbornik, 1 0 (4): 0 457, 1967
1967
-
[129]
J. Martens. New insights and perspectives on the natural gradient method. Journal of Machine Learning Research, 21 0 (146): 0 1--76, 2020
2020
-
[130]
A. Maurer. A vector-contraction inequality for Rademacher complexities. In Algorithmic Learning Theory: 27th International Conference, ALT 2016, Bari, Italy, October 19-21, 2016, Proceedings, page 3–17, Berlin, Heidelberg, 2016. Springer-Verlag. ISBN 978-3-319-46378-0. doi:10....
2016 doi
-
[131]
o sung. ZAMM-Journal of Applied Mathematics and Mechanics/Zeitschrift f \
R. Mises and H. Pollaczek-Geiringer. Praktische verfahren der gleichungsaufl \"o sung. ZAMM-Journal of Applied Mathematics and Mechanics/Zeitschrift f \"u r Angewandte Mathematik und Mechanik , 9 0 (1): 0 58--77, 1929
1929
-
[132]
Miyato, T
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=B1QRgziT-
2018
-
[133]
Modas, S.-M
A. Modas, S.-M. Moosavi-Dezfooli, and P. Frossard. Sparsefool: A few pixels make a big difference. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9079--9088, 2019. doi:10.1109/CVPR.2019.00930
2019
-
[134]
Mohri, A
M. Mohri, A. Rostamizadeh, and A. Talwalkar. Foundations of machine learning. MIT Press, 2018
2018
-
[135]
P. Nair. Softmax is 1/2 - Lipschitz : A tight bound across all _p norms. arXiv preprint arXiv:2510.23012, 2025
2025
-
[136]
Nair and G
V. Nair and G. E. Hinton. Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on International Conference on Machine Learning, ICML'10, page 807–814, Madison, WI, USA, 2010. Omnipress. ISBN 9781605589077
2010
-
[137]
Natarajan, I
N. Natarajan, I. S. Dhillon, P. K. Ravikumar, and A. Tewari. Learning with noisy labels. In C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger, editors, Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013
2013
-
[138]
F. Navarro. Necessary players, Myerson fairness and the equal treatment of equals. Annals of Operations Research, 280: 0 111--119, 2019
2019
-
[139]
Neelakantan, L
A. Neelakantan, L. Vilnis, Q. V. Le, I. Sutskever, L. Kaiser, K. Kurach, and J. Martens. Adding gradient noise improves learning for very deep networks. 2015. URL https://arxiv.org/abs/1511.06807
2015 arXiv
-
[140]
Neyshabur, R
B. Neyshabur, R. Tomioka, and N. Srebro. Norm-based capacity control in neural networks. In P. Grünwald, E. Hazan, and S. Kale, editors, Proceedings of The 28th Conference on Learning Theory, volume 40 of Proceedings of Machine Learning Research, pages 1376--1401, Paris, Franc...
2015
-
[141]
Neyshabur, S
B. Neyshabur, S. Bhojanapalli, D. Mcallester, and N. Srebro. Exploring generalization in deep learning. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran ...
2017
-
[142]
Nguyen, A
A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune. Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems...
2016
-
[143]
Nguyen, J
A. Nguyen, J. Yosinski, and J. Clune. Understanding neural networks via feature visualization: A survey. Explainable AI: interpreting, explaining and visualizing deep learning, pages 55--76, 2019
2019
-
[144]
M. A. Nielsen and I. L. Chuang. Quantum computation and quantum information. Cambridge University Press, 2010
2010
-
[145]
Artificial intelligence risk management framework, 2023
NIST. Artificial intelligence risk management framework, 2023. URL https://www.nist.gov/itl/ai-risk-management-framework
2023
-
[146]
Oksendal
B. Oksendal. Stochastic differential equations: an introduction with applications. Springer Science & Business Media, 2013
2013
-
[147]
Achiam, S
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L....
2024 arXiv
-
[148]
A. V. Oppenheim. Applications of digital signal processing. Prentice-Hall, 1978. ISBN 0130391158, 9780130391155
1978
-
[149]
T. Pang, X. Yang, Y. Dong, H. Su, and J. Zhu. Bag of tricks for adversarial training. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Xb8xvrtB8Ce
2021
-
[150]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K\" o pf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. PyTorch : an imperative style, hig...
2019
-
[151]
Patrini, A
G. Patrini, A. Rozza, A. K. Menon, R. Nock, and L. Qu. Making deep neural networks robust to label noise: A loss correction approach. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2233--2241, 2017. doi:10.1109/CVPR.2017.240
2017 doi
-
[152]
Paul and P.-Y
S. Paul and P.-Y. Chen. Vision transformers are robust learners. In Proceedings of the AAAI conference on Artificial Intelligence, volume 36, pages 2071--2081, 2022
-
[153]
Perozzi, R
B. Perozzi, R. Al-Rfou, and S. Skiena. DeepWalk : online learning of social representations. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '14, page 701–710, New York, NY, USA, 2014. Association for Computing Machine...
2014
-
[154]
Perugachi-Diaz, J
Y. Perugachi-Diaz, J. Tomczak, and S. Bhulai. Invertible DenseNets with concatenated LipSwish . In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 17246--17257. Curran Associates,...
2021
-
[155]
Pomponi, S
J. Pomponi, S. Scardapane, and A. Uncini. Pixle: a fast and effective black-box attack based on rearranging pixels. In 2022 International Joint Conference on Neural Networks (IJCNN), pages 1--7. IEEE, 2022
2022
-
[156]
X. Qi, J. Wang, Y. Chen, Y. Shi, and L. Zhang. LipsFormer : Introducing Lipschitz continuity to vision transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=cHf1DcCwcH3
2023
-
[157]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners, 2019
2019
-
[158]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. In M. Meila and T. Zhang, editors, Proceedings of the 38th Interna...
2021
-
[159]
Rahaman, A
N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y. Bengio, and A. Courville. On the spectral bias of neural networks. In International Conference on Machine Learning, pages 5301--5310. PMLR, 2019
2019
-
[160]
Ramachandran, B
P. Ramachandran, B. Zoph, and Q. V. Le. Searching for activation functions, 2018. URL https://openreview.net/forum?id=SkBYYyZRZ
2018
-
[161]
J. W. S. Rayleigh. The theory of sound, Volume One. Courier Corporation, 2013
2013
-
[162]
Recht, R
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar. Do ImageNet classifiers generalize to ImageNet ? In International Conference on Machine Learning (ICML), pages 5389--5400. PMLR, 2019
2019
-
[163]
F. Rellich. Perturbation theory of eigenvalue problems. CRC Press, 1969. ISBN 0677006802, 9780677006802
1969
-
[164]
Rezende and S
D. Rezende and S. Mohamed. Variational inference with normalizing flows. In F. Bach and D. Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1530--1538, Lille, France, 07--09 Jul 20...
2015
-
[165]
M. T. Ribeiro, S. Singh, and C. Guestrin. `` Why Should I Trust You ?'': Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '16, page 1135–1144, New York, NY, USA, 2016. Assoc...
2016
-
[166]
Robbins and S
H. Robbins and S. Monro. A stochastic approximation method. The Annals of Mathematical Statistics, 22 0 (3): 0 400--407, 1951. ISSN 00034851. URL http://www.jstor.org/stable/2236626
1951
-
[167]
E. A. Rocamora, G. Chrysos, and V. Cevher. Certified robustness under bounded Levenshtein distance. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=cd79pbXi4N
2025
-
[168]
A. E. Roth. The Shapley value: essays in honor of Lloyd S. Shapley. Cambridge University Press, 1988
1988
-
[169]
W. Rudin. Principles of Mathematical Analysis. McGraw-Hill, 3 edition, 1976
1976
-
[170]
W. Rudin. Real and Complex Analysis. McGraw-Hill, New York, 3rd edition, 1987. ISBN 9780070542341
1987
-
[171]
J. J. Sakurai and J. Napolitano. Modern quantum mechanics. Cambridge University Press, 2020
2020
-
[172]
Schr \"o dinger
E. Schr \"o dinger. Quantisierung als eigenwertproblem. Annalen der physik, 385 0 (13): 0 437--490, 1926
1926
-
[173]
Sedghi, V
H. Sedghi, V. Gupta, and P. M. Long. The singular values of convolutional layers. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJevYoA9Fm
2019
-
[174]
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra. Grad-CAM : Visual explanations from deep networks via gradient-based localization. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 618--626, 2017. doi:10.1109/ICCV.2017.74
2017 doi
-
[175]
Shalev-Shwartz and S
S. Shalev-Shwartz and S. Ben-David. Understanding machine learning: From theory to algorithms. Cambridge University Press, 2014
2014
-
[176]
O. M. Shalit. Dilation theory: a guided tour. In Operator theory, functional analysis and applications, pages 551--623. Springer, 2021
2021
-
[177]
R. Shao, Z. Shi, J. Yi, P.-Y. Chen, and C.-J. Hsieh. On the adversarial robustness of vision transformers. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=lE7K4n1Esk
2022
-
[178]
Shrikumar, P
A. Shrikumar, P. Greenside, and A. Kundaje. Learning important features through propagating activation differences. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research,...
2017
-
[179]
Simsekli, L
U. Simsekli, L. Sagun, and M. Gurbuzbalaban. A tail-index analysis of stochastic gradient noise in deep neural networks. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Lea...
2019
-
[180]
Simsekli, L
U. Simsekli, L. Zhu, Y. W. Teh, and M. Gurbuzbalaban. Fractional underdamped L angevin dynamics: Retargeting SGD with momentum under heavy-tailed gradient noise. In H. D. III and A. Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 11...
2020
-
[181]
Smilkov, N
D. Smilkov, N. Thorat, B. Kim, F. Vi \'e gas, and M. Wattenberg. Smoothgrad: removing noise by adding noise. 2017. URL https://arxiv.org/abs/1706.03825
2017 arXiv
-
[182]
Sokolić, R
J. Sokolić, R. Giryes, G. Sapiro, and M. R. D. Rodrigues. Robust large margin deep neural networks. IEEE Transactions on Signal Processing, 65 0 (16): 0 4265--4280, 2017. doi:10.1109/TSP.2017.2708039
2017
-
[183]
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=PxTIG12RRHS
2021
-
[184]
M. Spivak. Calculus on manifolds: a modern approach to classical theorems of advanced calculus. CRC press, 2018
2018
-
[185]
E. M. Stein and R. Shakarchi. Fourier analysis: an introduction, volume 1. Princeton University Press, 2011
2011
-
[186]
G. W. Stewart and J.-g. Sun. Matrix perturbation theory. Academic Press, 1990
1990
-
[187]
Stoica, R
P. Stoica, R. L. Moses, et al. Spectral analysis of signals, volume 452. Pearson Prentice Hall Upper Saddle River, NJ, 2005
2005
-
[188]
W. Su, S. Boyd, and E. J. Cand \`e s. A differential equation for modeling nesterov's accelerated gradient method: Theory and insights. Journal of Machine Learning Research, 17 0 (153): 0 1--43, 2016. URL http://jmlr.org/papers/v17/15-084.html
2016
-
[189]
Sucholutsky, R
I. Sucholutsky, R. M. Battleday, K. M. Collins, R. Marjieh, J. Peterson, P. Singh, U. Bhatt, N. Jacoby, A. Weller, and T. L. Griffiths. On the informativeness of supervision signals. In R. J. Evans and I. Shpitser, editors, Proceedings of the Thirty-Ninth Conference on Uncerta...
-
[190]
Sundararajan, A
M. Sundararajan, A. Taly, and Q. Yan. Axiomatic attribution for deep networks. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 3319--3328. PMLR, 06--11 Aug 2...
2017
-
[191]
R. S. Sutton, A. G. Barto, et al. Reinforcement learning: An introduction, volume 1. MIT Press, 1998
1998
-
[192]
Szegedy, W
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. J. Goodfellow, and R. Fergus. Intriguing properties of neural networks. In Y. Bengio and Y. LeCun, editors, 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, C...
2014 arXiv
-
[193]
Tan and Q
M. Tan and Q. Le. E fficient N et: Rethinking model scaling for convolutional neural networks. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 6105...
2019
-
[194]
T. Tao. Topics in random matrix theory, volume 132. American Mathematical Soc., 2012
2012
-
[195]
Taori, A
R. Taori, A. Dave, V. Shankar, N. Carlini, B. Recht, and L. Schmidt. Measuring robustness to natural distribution shifts in image classification. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume...
2020
-
[196]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[197]
C. A. Tracy and H. Widom. Level-spacing distributions and the Airy kernel. Communications in Mathematical Physics, 159: 0 151--174, 1994
1994
-
[198]
Trockman and J
A. Trockman and J. Z. Kolter. Orthogonalizing convolutional layers with the Cayley transform. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Pbj8H_jEHYv
2021
-
[199]
Tsipras, S
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry. Robustness may be at odds with accuracy. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=SyxAb30cY7
2019
-
[200]
Tsuzuku and I
Y. Tsuzuku and I. Sato. On the structural sensitivity of deep convolutional networks to the directions of Fourier basis functions. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 51--60, 2019. doi:10.1109/CVPR.2019.00014
2019
-
[201]
Tsuzuku, I
Y. Tsuzuku, I. Sato, and M. Sugiyama. Lipschitz-margin training: scalable certification of perturbation invariance for deep neural networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS'18, page 6542–6551, Red Hook, NY, USA...
2018
-
[202]
Drimbarean, J
R\'ois\'in Luo , A. Drimbarean, J. McDermott, and C. O'Riordan. Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization. In The 35th British Machine Vision Conference, BMVC 2024, Glasgow, UK, November 25-28, 2024 . BMVA, 2024 a . URL https://arxiv.org/abs/2408.00923
2024 arXiv
-
[203]
McDermott, and C
R\'ois\'in Luo , J. McDermott, and C. O'Riordan. Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition. Transactions on Machine Learning Research (TMLR), 2024 b . ISSN 2835-8856. URL https://arxiv.org/abs/2408.01139. Pres...
2024 arXiv
-
[204]
McDermott, C
R\'ois\'in Luo , J. McDermott, C. Gagn\'e, Q. Sun, and C. O'Riordan. Optimization-Induced Dynamics of L ipschitz Continuity in Neural Networks . Manuscript is under review at Journal of Machine Learning Research (JMLR), 2025 a . URL https://arxiv.org/abs/2506.18588
2025
-
[205]
McDermott, and C
R\'ois\'in Luo , J. McDermott, and C. O'Riordan. Lipschitz Continuity in Deep Learning: A Systematic Review of Theoretical Foundations, Estimation Methods, Regularization Approaches and Certifiable Robustness. Manuscript is under review at Transactions on Machine Learning Rese...
2025
-
[206]
O'Riordan, and J
R\'ois\'in Luo , C. O'Riordan, and J. McDermott. Higher-Order Singular-Value Derivatives of Real Rectangular Matrices. Journal of Mathematical Analysis and Applications (JMAA), page 130236, 2025 c . ISSN 0022-247X. doi:10.1016/j.jmaa.2025.130236. URL https://doi.org/10.1016/j....
2025
-
[207]
Van Den Berg, L
R. Van Den Berg, L. Hasenclever, J. M. Tomczak, and M. Welling. Sylvester normalizing flows for variational inference. UAI, 2018
2018
-
[208]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing...
2017
-
[209]
Veličković, G
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio. Graph attention networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJXMpikCZ
2018
-
[210]
Villani et al
C. Villani et al. Optimal transport: old and new, volume 338. Springer, 2008
2008
-
[211]
Virmaux and K
A. Virmaux and K. Scaman. Lipschitz regularity of deep neural networks: analysis and efficient estimation. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31. Curran Associ...
2018
-
[212]
Vorontsov, C
E. Vorontsov, C. Trabelsi, S. Kadoury, and C. Pal. On orthogonality and learning recurrent networks with long term dependencies. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learn...
2017
-
[213]
Vuckovic, A
J. Vuckovic, A. Baratin, and R. T. d. Combes. A mathematical theory of attention. 2020. URL https://arxiv.org/abs/2007.02876
2020 arXiv
-
[214]
H. Wang, X. Wu, Z. Huang, and E. P. Xing. High-frequency component helps explain the generalization of convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8684--8694, 2020
2020
-
[215]
Welling and Y
M. Welling and Y. W. Teh. Bayesian learning via stochastic gradient Langevin dynamics. In Proceedings of the 28th International Conference on International Conference on Machine Learning, ICML'11, page 681–688, Madison, WI, USA, 2011. Omnipress. ISBN 9781450306195
2011
-
[216]
L. Weng, H. Zhang, H. Chen, Z. Song, C.-J. Hsieh, L. Daniel, D. Boning, and I. Dhillon. Towards fast computation of certified robustness for R e LU networks. In International Conference on Machine Learning, pages 5276--5285. PMLR, 2018 a
2018
-
[217]
T.-W. Weng, H. Zhang, P.-Y. Chen, J. Yi, D. Su, Y. Gao, C.-J. Hsieh, and L. Daniel. Evaluating the robustness of neural networks: An extreme value theory approach. In International Conference on Learning Representations, 2018 b . URL https://openreview.net/forum?id=BkUHlMZ0b
2018
-
[218]
Wisdom, T
S. Wisdom, T. Powers, J. Hershey, J. Le Roux, and L. Atlas. Full-capacity unitary recurrent neural networks. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URL ...
2016
-
[219]
T. Xiao, X. Wang, A. A. Efros, and T. Darrell. What should not be contrastive in contrastive learning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=CZ8Y3NzuVzO
2021
-
[220]
K. Xu, W. Hu, J. Leskovec, and S. Jegelka. How powerful are graph neural networks? In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019 a . URL https://openreview.net/forum?id=ryGs6iA5Km
2019
-
[221]
Z.-Q. J. Xu, Y. Zhang, and Y. Xiao. Training behavior of deep neural network in frequency domain. In Neural Information Processing: 26th International Conference, ICONIP 2019, Sydney, NSW, Australia, December 12–15, 2019, Proceedings, Part I, page 264–274, Berlin, Heidelberg, ...
2019 doi
-
[222]
Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, and Z. Ma. Frequency principle: Fourier analysis sheds light on deep neural networks. Communications in Computational Physics, 28 0 (5): 0 1746–1767, Nov. 2020. doi:10.4208/cicp.OA-2020-0085. URL https://global-sci.com/index.php/cicp/art...
2020 doi
-
[223]
D. Yin, R. Kannan, and P. Bartlett. Rademacher complexity for adversarially robust generalization. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages ...
2019
-
[224]
K. Yosida. Functional analysis. Springer Science & Business Media, 2012
2012
-
[225]
Yudin, A
N. Yudin, A. Gaponov, S. Kudriashov, and M. Rakhuba. Pay attention to attention distribution: A new local Lipschitz bound for transformers. 2025. URL https://arxiv.org/abs/2507.07814
2025 arXiv
-
[226]
M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus. Deconvolutional networks. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 2528--2535, 2010. doi:10.1109/CVPR.2010.5539957
2010
-
[227]
Zhang, D
B. Zhang, D. Jiang, D. He, and L. Wang. Rethinking Lipschitz neural networks and certified robustness: a Boolean function perspective. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red Hook, NY, USA, 2022. Curran Associ...
2022
-
[228]
Zhang, S
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals. Understanding deep learning (still) requires rethinking generalization. Commun. ACM, 64 0 (3): 0 107–115, Feb. 2021. ISSN 0001-0782. doi:10.1145/3446776. URL https://doi.org/10.1145/3446776
2021 doi
-
[229]
Zhang, X
Z. Zhang, X. Shu, B. Yu, T. Liu, J. Zhao, Q. Li, and L. Guo. Distilling knowledge from well-informed soft labels for neural relation extraction. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9620--9627, 2020
2020
-
[230]
Zheng, Y
S. Zheng, Y. Song, T. Leung, and I. Goodfellow. Improving the robustness of deep neural networks via stability training. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4480--4488, 2016. doi:10.1109/CVPR.2016.485
2016 doi
-
[231]
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba. Learning deep features for discriminative localization. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2921--2929, 2016. doi:10.1109/CVPR.2016.319
2016 doi
-
[232]
D. Zhou, Z. Yu, E. Xie, C. Xiao, A. Anandkumar, J. Feng, and J. M. Alvarez. Understanding the robustness in vision transformers. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, editors, Proceedings of the 39th International Conference on Machine Lea...
2022
-
[233]
Z. Zhu, J. Wu, B. Yu, L. Wu, and J. Ma. The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learn...
2019
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.