Pith. sign in

REVIEW 5 major objections 4 minor 233 references

Principles of Lipschitz continuity in neural networks

T0 review · 5 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Training noise is not a nuisance: it irreversibly inflates a neural network's worst-case input sensitivity, a phenomenon this thesis derives from first principles.

desk verdict A mathematically serious thesis whose Chapter 4 SDE framework is original but rests on an unvalidated diffusion approximation and in-sample validation; Chapters 2 and 3 are solid enough to justify referee time. read the letter →

arxiv 2602.04078 v2 pith:A3RAY2ZJ submitted 2026-02-03 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML MSC 15A1847A5560H1068T07
keywords LipschitzcontinuitytrainingdynamicsstochasticdifferentialequationsspectralnormsingularvalueHessianperturbationtheoryrobustnessgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This thesis attempts to establish the first theoretical framework for how the Lipschitz constant of a neural network evolves during training. It claims that each layer's spectral-norm Lipschitz bound obeys a stochastic differential equation, with a deterministic gradient-alignment drift, a non-negative 'noise–curvature' drift that grows the bound irreversibly, and a diffusion driven by mini-batch sampling noise. Aggregating the layer equations yields predictions: supervision noise lowers the bound, batch size controls its variance, and relative fluctuations blow up near convergence. If correct, this turns ad hoc observations about robustness into quantitative, testable laws.

What carries the argument

The Jordan–Wielandt embedding turns a rectangular weight matrix into a self-adjoint block operator, allowing refined Kato analytic perturbation theory to yield closed-form higher-order Fréchet derivatives of singular values. The new closed-form singular-value Hessian (expressed in Kronecker-product form) supplies the second-order term needed for Itô calculus on the spectral norm. These two pieces produce the layer-wise SDE and its network-level aggregation.

What would settle it

Train a small MLP with a large learning rate under heavy-tailed gradient noise and compare the measured layer-wise Lipschitz trajectories' variance to the SDE prediction; systematic mismatch — or finding the noise–curvature term to be negative — would falsify the central claim. A simpler check: train with increasing label noise and verify the predicted monotone decrease of the final spectral-norm product; any non-monotone ordering falsifies the drift decomposition.

Watch

Extended reading notes

Core claim

The central result is a system of SDEs for the layer-wise spectral-norm Lipschitz bound. For each layer ℓ, dK(ℓ)/K(ℓ) = (μ(ℓ) + κ(ℓ)) dt + λ(ℓ)ᵀ dB(ℓ), where κ(ℓ) = η/(2σ₁)⟨H_op, Σ⟩ ≥ 0 is an entropy-production term coupling gradient noise covariance to the curvature of the operator norm, and λ(ℓ) is a diffusion intensity. Writing Z = Σ_ℓ log K(ℓ), the network bound K = e^Z, the thesis derives network-level drift, diffusion, and statistics. It claims that supervision noise shrinks the drift, that mini-batch trajectory does not affect variance for large batch size, that relative fluctuations of the bound grow unboundedly near convergence, and that batch size controls variance.

Load-bearing premise

Mini-batch SGD is replaced by a Gaussian diffusion whose covariance is measured from the very runs being predicted, and no two nonzero singular values ever become equal during training.

Editorial extensions

If this is right

  • The Lipschitz bound has an irreducible upward drift produced by the interaction of SGD noise with the curvature of the spectral norm, even when the gradient flow itself would shrink it.
  • Label noise during training systematically lowers the final Lipschitz bound, because supervision noise shrinks the optimization-induced drift.
  • For sufficiently large batches, the variance of the bound becomes independent of the particular mini-batch trajectory.
  • Batch size directly scales the diffusion intensity, providing a practical control knob for the fluctuation of robustness during training.
  • Relative fluctuations of the bound grow without bound in the near-convergence regime, predicting a distinctive late-training signature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A cheap, theory-agnostic check: train the same architecture with increasing label noise and measure the final spectral-norm product; the framework predicts a monotone decrease, reverse ordering would undercut the drift decomposition.
  • The same singular-value Hessian machinery could generate SDEs for other spectral functionals (stable rank, effective dimension, von Neumann entropy of the Gram matrix), extending the framework beyond Lipschitz bounds.
  • The near-convergence divergence prediction suggests that late-training spectral-norm spikes, sometimes blamed on optimization failure, may be an intrinsic diffusion-dominated phenomenon — a testable hypothesis for training diagnostics.
  • The Gaussian diffusion approximation is the fragile point; measuring heavy-tailedness of mini-batch gradient noise during large-learning-rate training would delimit the regime where the SDE predictions hold.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. This thesis, compiled from four papers and centered on Lipschitz continuity in neural networks, investigates two directions: an internal one (the temporal evolution of a spectral-norm Lipschitz bound during training) and an external one (how Lipschitz continuity modulates frequency signal propagation and robustness). Chapter 2 provides a survey with corrected activation-function Lipschitz constants, a sum-over-paths bound for DAG networks, and ℓ2/conjugate certified robustness radii. Chapter 3 develops an operator-theoretic perturbation framework for singular-value derivatives via the Jordan–Wielandt embedding and Kato expansions. Chapter 4 models mini-batch SGD as a Gaussian diffusion SDE and derives layer-wise and network-level dynamics for the log Lipschitz bound, with drift μ, noise–curvature production κ ≥ 0, and diffusion λ, yielding predictions on label noise, batch size, mini-batch trajectory, and near-convergence behavior. Chapters 5 and 6 connect Lipschitz bounds to Fourier frequency analysis and a Shapley-based spectral robustness score.

Significance. If fully established, Chapter 4's SDE framework would be a novel unifying account of Lipschitz dynamics during training, with concrete falsifiable predictions. The thesis has several verifiable strengths: reproducible code links, correct spot-checked constants (softmax 1/2, sigmoid 1/4), a DAG sum-over-paths bound, and a closed-form singular-value Hessian not previously in the literature. However, the central dynamics result rests on an unvalidated diffusion approximation and in-sample validation; Chapter 3's arbitrary-order perturbation theorem also appears incorrect for n ≥ 3. These issues do not necessarily invalidate the n = 1, 2 formulas used in Chapter 4, but they place the manuscript in major-revision territory.

major comments (5)
  1. [Thm 3.3.3 (Eq. 3.55)] Theorem 3.3.3 is not the correct eigenvalue coefficient for n ≥ 3. For T(x) = T0 + xV, Eq. (3.55) reduces to λ^(3) = <VSVSV>; the standard Rayleigh–Schrödinger expansion gives λ^(3) = <VSVSV> − <V><VS²V> (checkable in the 2×2 case T0 = diag(0,Δ), V = [[a,b],[c,d]]). The proof drops higher-order-pole contributions in the residue step (Eqs. 3.110–3.115), which do not vanish. Consequently the claimed arbitrary-order singular-value derivatives are not established. The n = 1, 2 cases used in Chapter 4 appear unaffected, but the theorem and all 'arbitrary-order' claims must be corrected or restricted to n ≤ 2.
  2. [Def. 4.3.2; §4.7, Fig. 4.2] The central SDE replaces discrete mini-batch SGD with dvecθ = −∇L dt + √η Σ^{1/2} dB without any error analysis: no bound in η, batch size, or gradient-noise tails, and no justification for Gaussianity. The validation estimates Σ_t from the same runs whose K(t) is then 'predicted'; agreement in Fig. 4.2 is therefore a consistency check of Itô calculus, not an out-of-sample test. To support the Chapter 4 conclusions, provide a formal approximation theorem (weak/strong error) and validate on independent runs — e.g., covariance estimated from one window and the trajectory predicted on a disjoint window — or on synthetic SGD with known noise.
  3. [Assumption 3.2.2; Lemmas 4.6.2–4.6.3] The dynamics coefficients require simplicity of all non-zero singular values; the operator-norm Jacobian/Hessian require the top singular value to stay simple along the whole training trajectory. The thesis acknowledges this (Sec. 4.9), but no experiment reports the spectral gap. At initialization and under parameter symmetries spectral collisions occur, so the SDE coefficients are undefined at those times. Add empirical spectral-gap diagnostics or a nonsmooth extension; otherwise the derived drift/diffusion do not describe the actual trajectory.
  4. [§4.8.3, Prop. 4.8.1] The 'unbounded growth near convergence' prediction assumes Σ_t stays non-degenerate as ∇L → 0. For clean, overparameterized networks at interpolation, every mini-batch gradient is zero at the minimizer, so Σ_t → 0 and both λ and κ vanish; relative fluctuations of K need not diverge. The claim should explicitly state the non-degeneracy condition on Σ_t (e.g., label-noise-driven dynamics) or be restricted to that regime.
  5. [§2.2.9; §4.8] The statement that the Lipschitz bound 'irreversibly increases' because κ_Z ≥ 0 is not implied: κ_Z is only one additive component of the d log K drift, and μ_Z can be negative and dominate. Reformulate as a claim about the nonnegative noise–curvature contribution rather than monotonicity of K(t).
minor comments (4)
  1. [Table 2.1] The table lists the Swish and GELU constants as ≈1.1; since exact expressions are derived in Appendix 2.B.4–2.B.5, reporting the closed forms would be clearer.
  2. [Prop. 2.2.17] The SVD statement assumes distinct positive singular values (σ1 > σ2 > ... > σr), but the spectral norm is defined without a simplicity requirement; use ≥.
  3. [Eq. (3.141)] The expression D^n σ_k[dA,...,dA] = n! lim_{x→0} x^n σ_k^{(n)} is dimensionally awkward; if σ_k^{(n)} is the Taylor coefficient, the derivative is simply n! σ_k^{(n)}. Please clarify the notation.
  4. [Appendix 2.B.4] Typo: 'LiprSiwshpxqs' should read 'LiprSwishpxqs'. Also, the proof of the DAG bound in Theorem 2.2.20 could state explicitly that the modules h_v are assumed 1-Lipschitz in their inputs for the constant C_{u→v} = Lip[h_v] to be well-defined per edge.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the core derivation applies Itô's lemma to measured, not fitted, quantities; self-citations are to independently published work.

full rationale

The central derivation chain (Ch. 4, summarized in §2.2.9, Eqs. 2.116–2.119) takes the vectorized SDE for continuous-time SGD (Def. 4.3.2) as a stated modeling assumption and applies Itô's lemma to the layer-wise spectral-norm bound K^(ℓ)(t)=‖θ^(ℓ)(t)‖_op. The resulting drift μ^(ℓ), noise–curvature term κ^(ℓ), and diffusion λ^(ℓ) are explicit functions of the measured gradient, the measured batch-gradient covariance Σ_t, and the singular vectors/operator-norm Hessian of the current weight matrix. No parameter is fitted to the Lipschitz trajectory K(t), so the derived SDE is not equivalent to its output by construction. The operator-norm Hessian used in Ch. 4 comes from Ch. 3, which is an independently peer-reviewed publication (Róisín Luo et al., 2025c, JMAA) with stated simplicity assumptions; citing it is legitimate external support, not a load-bearing self-citation chain. The acknowledged simplicity assumption (Assumption 3.2.2) is a limitation that can invalidate the theory at spectral collisions, but an unvalidated or restrictive assumption is a correctness risk, not circularity. Similarly, measuring Σ from the same training runs when validating the SDE is an in-sample consistency check rather than an out-of-sample test, and this weakens the evidence without making the prediction identical to its input: K(t) is not used to determine Σ or any free parameter. The thesis's self-citations are to its own compiled articles and do not function as an external-authority uniqueness argument. I therefore find no circular step that meets the required standard of exhibiting a specific reduction of a claimed prediction to its own fitted or definitional input.

Assumptions & free parameters 3 free parameters · 5 assumptions · 2 invented entities

The ledger reflects that the load-bearing contributions are derivations relative to a standard toolbox (Kato perturbation theory, Jordan–Wielandt embedding, Itô calculus, Rademacher complexity). The genuinely fragile premises are empirical/modeling choices: the SDE lift of SGD, the simplicity of singular values throughout training, and the identification of the network's Lipschitz constant with the spectral-product upper bound. No free parameters are fitted to the Lipschitz trajectory itself; covariances and gradients feeding the SDE are measured from training runs, which keeps the circularity burden moderate. The SRS band-partition and characteristic-function choices are design parameters whose ex-ante status is not documented.

free parameters (3)
  • learning rate η = small positive constant (standard SGD setting)
    Input to the SDE system (Def. 4.3.2); taken from the training setup, not fitted to the Lipschitz trajectory.
  • layer-wise gradient-noise covariance Σ_t^(ℓ) (and its square root) = estimated per layer from mini-batch gradients (Props. 4.4.1–4.4.3)
    Measured from the same training runs whose Lipschitz evolution is then 'predicted'; not fitted to the target quantity, but data-derived inputs.
  • spectral-band partition and coalition design (Ch. 6) = ℓ∞-ball over ℓ2-ball banding; sample counts (Appendix 6.C.1, Figs. 6.15–6.16)
    Design choices for the spectral game; the thesis does not show these were fixed before observing the correlation results.
assumptions (5)
  • domain assumption Discrete mini-batch SGD is modeled as the diffusion dvecθ = −vec∇L dt + √η Σ^{1/2} dB (Def. 4.3.2)
    Load-bearing for all of Ch. 4: the Wiener process with empirical covariance replaces minibatch noise; no error bounds or validity regime (learning rate, gradient-noise tails) are given.
  • domain assumption All non-zero singular values of every weight matrix are simple throughout training (Assumption 3.2.2)
    Required for C^∞ differentiability of singular values and for the closed-form Jacobian/Hessian (Lemmas 3.5.1, 3.6.1); singular-value collisions occur in real training.
  • domain assumption The network's true Lipschitz constant is tracked via the upper bound K(t) = Π_ℓ ‖θ^(ℓ)(t)‖ (Prop. 4.5.1, Def. 4.5.3)
    The derived dynamics concern the spectral-product bound, which can be exponentially loose in depth; the thesis's framing ('evolution of Lipschitz continuity') presupposes the bound is representative.
  • domain assumption Input domain is convex so the tight constant equals sup_x ‖∇f(x)‖ (Remark 2.2.6)
    Used in Ch. 2 to equate the global Lipschitz constant with the supremum of gradient norms for the numerical validation of activation constants.
  • standard math Standard analytic toolbox: Kato perturbation theory, Jordan–Wielandt embedding spectra, Itô's lemma, Popoviciu's inequality, Rademacher complexity bounds
    Invoked at Thms 3.3.2/3.3.3, Thm 3.1.2, §4.6, Appendix 2.B.6, §2.2.8; accepted without proof.
invented entities (2)
  • Noise–curvature entropy production κ^(ℓ)(t) (and aggregate κ_Z) independent evidence
    purpose: Decomposes the SDE drift for the log Lipschitz bound into a deterministic optimization part and a non-negative noise-curvature part; underpins the claims of irreversible growth and near-convergence unboundedness (Eq. 2.118, §4.8.3–4.8.4).
    Not postulated blindly: under the SDE assumptions it is η/(2σ_1)<H_op, Σ> with H_op the convex spectral-norm Hessian and Σ PSD, so κ≥0 is a theorem; its magnitude is computable from checkpoints, making the growth and blow-up predictions falsifiable in principle.
  • Spectral Robustness Score (SRS) independent evidence
    purpose: Assigns Shapley-value importances to frequency-band coalitions in a spectral game, yielding a scalar robustness score claimed to correlate with corruption and adversarial errors (Ch. 6, Def. 6.3.4, §6.4.1).
    Defined from models and data alone; the correlation claims against mean corruption error and adversarial prediction error are falsifiable measurements reported across architectures in §6.4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Principles of Lipschitz continuity in neural networks." pith.science (2026). https://pith.science/paper/A3RAY2ZJ

@misc{pith2026260204078,
  author       = {Pith},
  title        = {Pith review of: Principles of Lipschitz continuity in neural networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A3RAY2ZJ}},
  note         = {Machine review of arXiv:2602.04078}
}
read the original abstract

Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable in artificial intelligence. Yet, despite these advances, critical challenges remain -- most notably, ensuring robustness to small input perturbations and generalization to out-of-distribution data. These critical challenges underscore the need to understand the underlying fundamental principles that govern robustness and generalization. Among the theoretical tools available, Lipschitz continuity plays a pivotal role in governing the fundamental properties of neural networks related to robustness and generalization. It quantifies the worst-case sensitivity of network's outputs to small input perturbations. While its importance is widely acknowledged, prior research has predominantly focused on empirical regularization approaches based on Lipschitz constraints, leaving the underlying principles less explored. This thesis seeks to advance a principled understanding of the principles of Lipschitz continuity in neural networks within the paradigm of machine learning, examined from two complementary perspectives: an internal perspective -- focusing on the temporal evolution of Lipschitz continuity in neural networks during training (i.e., training dynamics); and an external perspective -- investigating how Lipschitz continuity modulates the behavior of neural networks with respect to features in the input data, particularly its role in governing frequency signal propagation (i.e., modulation of frequency signal propagation).

Figures

Figures reproduced from arXiv: 2602.04078 by the authors.

Figure 1
Figure 1. Thesis Thematic Structure . . . . . . . . . . . . . . . . . . . . . [PITH_FULL_IMAGE:figures/full_fig_p019_1.png] view at source ↗
Figure 5
Figure 5. Frequency Signals in Training Data Shape Loss Landscape Flatness [PITH_FULL_IMAGE:figures/full_fig_p019_5.png] view at source ↗
Figure 6
Figure 6. Examples of Spectral Coalitions . . . . . . . . . . . . . . . . . . [PITH_FULL_IMAGE:figures/full_fig_p020_6.png] view at source ↗
Figures from the paper (37 more)
Figure 1.1
Figure 1.1. Figure 1.1: Principles of Lipschitz Continuity Manifest in Paradigm of Machine Learning. In the paradigm of machine learning, we consider data x P X, supervision signal y P Y, a learned function f : X ˆ Θ Ñ Y˜ — where Y˜ is output space — parameter￾ized by θ pℓq P Rdℓ and Θ :“ R…
Figure 1.2
Figure 1.2. Figure 1.2: Thesis Thematic Structure. A high-level conceptual diagram outlining the thematic connections between chapters. This thesis investigates the principles of Lipschitz continuity in neural networks through a structured approach guided by three primary research questions…
Figure 2.1
Figure 2.1. Figure 2.1: Example of Feedforward DAG Network. There are four computational paths: s Ñ u1 Ñ v Ñ t, s Ñ u2 Ñ v Ñ t, s Ñ u3 Ñ v Ñ t, and s Ñ t. A node is a module in the DAG neural network f. activation functions, combining Proposition 2.2.16 and Proposition 2.2.18 immediately yi…
Figure 2.2
Figure 2.2. Figure 2.2: Topology of non-biconnected DAG network. If a DAG is separated into two sub￾DAGs by the removal of a vertex ai, then ai is referred to as a cut vertex (or artic￾ulation point), and the DAG is said to be non-biconnected. This diagram shows the topology for a DAG that …
Figure 2.3
Figure 2.3. Figure 2.3: A residual module mpxq “ x ` ϕpxq consists of a non-identity unit ϕ : x ÞÑ ϕpxq and an identity skip connection unit s : x ÞÑ x. 2.2.8 Complexity-Theoretic Generalization Bound Notations. Let H be a hypothesis family and H Q h : X Ñ Y be a hypothesis. Let ℓ : Y ˆ Y Ñ…
Figure 3.1
Figure 3.1. Figure 3.1: Theoretical Framework for Infinitesimal Spectral Variations. We employ Kato’s analytic perturbation theory for self-adjoint operators (Kato, 1995). For a rectangular matrix A, we construct its Jordan–Wielandt embedding (Theorem 3.1.2), a block self￾adjoint operator t…
Figure 3.2
Figure 3.2. Figure 3.2: Numerical Experiments for Singular-Value Jacobian. This experiment compares the singular-value Jacobian derived from our framework with that obtained via Py￾Torch’s auto–differentiation. The error ϵ is measured as the ℓ2-norm between the theoretical and ground-truth …
Figure 3.3
Figure 3.3. Figure 3.3: Numerical Experiments for Singular-Value Hessian. This experiment compares the singular-value Hessian derived from our framework with that obtained via Py￾Torch’s auto–differentiation. The error ϵ is measured as the ℓ2-norm between the theoretical and ground-truth re…
Figure 3.4
Figure 3.4. Figure 3.4: Errors for Singular-Value Hessian. Random matrix entries are sampled i.i.d. from Np0, 1q and Ur0, 1s, respectively. For each singular-value index k “ 1, 2, . . . , r, the error ϵ is computed over 500 trials and visualized using an unnormalized histogram density. All …
Figure 4.1
Figure 4.1. Figure 4.1: Optimization-induced dynamics. During the training, the network parameters, starting from θ0, moves towards a solution θA or θB as shown in the loss land￾scape (a), driven by optimization process. Accordingly, this dynamics, driven by the optimization, induces the ev…
Figure 4.2
Figure 4.2. Figure 4.2: Numerical validation of our mathematical framework. The theoretical Lipschitz constants computed using our framework closely agree with empirical observations. To validate our framework, we train a five-layer ConvNet on CIFAR-10 and CIFAR-100 across multiple config￾u…
Figure 4.3
Figure 4.3. Figure 4.3: Dynamics near convergence. We profile both layer-specific and network-specific dynamics over 344, 370 steps (1766 epochs) on CIFAR-10. At the end of training, the final training loss and test loss are 9.75 ˆ 10´3 and 2.22, respectively; the final training accuracy an…
Figure 4.5
Figure 4.5. Figure 4.5: Predicted effect of batch size on the variance of Lipschitz continu￾ity. 1.0 0.8 0.6 0.4 0.2 Gradient magnitude 4.5 5.0 5.5 6.0 6.5 Lipschitz constant perturbation normal uniform zero [PITH_FULL_IMAGE:figures/full_fig_p186_4_5.png]
Figure 4.6
Figure 4.6. Figure 4.6: Predicted effect of gradient mag￾nitude and perturbation. 0.0 0.2 0.4 0.6 0.8 1.0 Label noise level 2 4 6 8 Lipschitz constant [PITH_FULL_IMAGE:figures/full_fig_p186_4_6.png]
Figure 4.8
Figure 4.8. Figure 4.8 [PITH_FULL_IMAGE:figures/full_fig_p197_4_8.png]
Figure 5.1
Figure 5.1. Figure 5.1: Directional frequency analysis of a three-layer MLP ( [PITH_FULL_IMAGE:figures/full_fig_p206_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Spectral bound of Lipschitz continuity with synthesized functions. (a)–(c): Three functions fpxq generated using different random seeds, each as a sum of sinusoidal components with randomly sampled frequencies and amplitudes. (d)-(f): Fourier transforms ˆfpxq of fpxq…
Figure 5.3
Figure 5.3. Figure 5.3: Lemma 5.3.4: Spectral-Band Perturbation Bound states that the output change due to a perturbation at a frequency point ζ, within the ball Bδpζq of energy ε in the frequency domain of the input space, is proportionally bounded above by the supremum Mδ of the spectral …
Figure 5.4
Figure 5.4. Figure 5.4: We train a ConvNet (see [PITH_FULL_IMAGE:figures/full_fig_p214_5_4.png]
Figure 5.5
Figure 5.5. Figure 5.5: Empirical and theoretical Lipschitz constants with respect to frequency signals in training data. The results show that models trained on datasets with stronger high￾frequency components associate with larger Lipschitz constants. Lipschitz constant Kf . The frequency…
Figure 5.6
Figure 5.6. Figure 5.6: Frequency Signals in Training Data Shape Loss Landscape Flatness of a ConvNet on CIFAR-10 ( [PITH_FULL_IMAGE:figures/full_fig_p219_5_6.png]
Figure 6.1
Figure 6.1. Figure 6.1: Power-law-like energy spectral density (ESD) distribution of natural images over the frequency. The signal spectrum is divided into M bands (from I0 to IM´1). Each spectral band is a robustness band. adversarial attacks often take place in inferences. This research p…
Figure 6.2
Figure 6.2. Figure 6.2: Spectral SNR characterization with respect to multiple corruptions and adversarial attacks. The corruptions include: white noise, Poisson noise, Salt-and-pepper noise, and Gaussian blur. The adversarial attacks include: FGSM (Goodfellow et al., 2015), PGD (Madry et a…
Figure 6.3
Figure 6.3. Figure 6.3: Understanding the role of spectral signals. We train a resnet18 on three datasets de￾rived from STL10 (Coates et al., 2011): r0, 1s contains full-frequency signals; r0, 0.3s only contains low-frequency signals with a cut-off frequency by 0.3; and r0.3, 1s only contai…
Figure 6.4
Figure 6.4. Figure 6.4: Framework of applying Shapley value theory. Spectral coalition filtering creates spectral coalitions over X. Each coalition contains a unique combination of spectral signals, in which some spectral bands are present and others are absent. The coali￾tions are fed into…
Figure 6.5
Figure 6.5. Figure 6.5: Spectral coalition filtering. In this example, the mask map TprIq (i.e. transfer func￾tion) only allows to pass the signals present in the spectral coalition tI0, I2u. The M is 4 and the absences are assigned to zeros. The images after spectral coalition filtering (x…
Figure 6.6
Figure 6.6. Figure 6.6: An example of a complete 2M spectral coalitions. This example shows 16 spec￾tral coalitions with M “ 4. Each coalition provides various information relevant to decisions. Each image is a coalition. Each coalition contains a unique spectral signal combination. We use …
Figure 6.7
Figure 6.7. Figure 6.7: Spectral importance distributions (SIDs) of trained models and un-trained models. The experimental models are pre-trained on ImageNet. We also include the models with random weights as a control marked by the blue box. We have noticed that: (1) The spectral importanc…
Figure 6.8
Figure 6.8. Figure 6.8: The spectral robustness scores (SRS), measured with I-ASIDE, correlate to the mean corruption errors (mCE) in the literature (Hendrycks and Dietterich, 2019). 6.4 Experiments We design experiments to show the dual functionality of I-ASIDE, which can not only measure …
Figure 6.9
Figure 6.9. Figure 6.9: The spectral robustness scores (SRS), measured with I-ASIDE, correlate with the mean prediction errors (mPE) in adversarial attacks. The circle sizes in (b) are pro￾portional to the SRS. (mCE). Let x be some clean image and x ˚ be the perturbed image. For a classifie…
Figure 6.10
Figure 6.10. Figure 6.10: The spectral robustness scores (SRS), measured with I-ASIDE, correlate to the mean prediction errors (mPE) in corruptions. The circle sizes in (b) are propor￾tional to the SRS. from the literature (Hendrycks and Dietterich, 2019). The mCE scores are measured on a co…
Figure 6.11
Figure 6.11. Figure 6.11: How do architectural elements affect robustness? The left figure is to answer: “Does model parameter size play a role on robustness?”. The right figure, a t-SNE projection of SIDs, is to answer: “Are vision transformers more robust than convo￾lutional neural network…
Figure 6.12
Figure 6.12. Figure 6.12: How do models respond to label noise? Our results show that models trained with higher label noise levels tend to use spectral signals uniformly, i.e. without a prefer￾ence for robust (low-frequency) features. label noise into CIFAR-10 and study its impact on model …
Figure 6.13
Figure 6.13. Figure 6.13: Three absence assignment strategies: (1) Assigning the spectral absences with con￾stant zeros (Zeroing), (2) assigning the spevtral absences with Gaussian noise (Com￾plex Gaussian) and (3) randomly sampling spectral components from the same im￾age datasets (Replacem…
Figure 6.14
Figure 6.14. Figure 6.14: Information quantity relationship. This shows the theoretical information quantity relationship between what the characteristic function v measures and the mutual information IpX ’ rI, Yq. For a given coalition rI, a dataset xX, Yy and a classifier Q, the v measures…
Figure 6.15
Figure 6.15. Figure 6.15: Two spectral band partitioning schemes. This shows the motivation we choose ℓ8 ball over ℓ2 ball in partitioning the frequency domain into the M bands (i.e., M ‘spectral players’) over 2D Fourier spectrum. The frequency data density of the spectral players with ℓ8 r…
Figure 6.16
Figure 6.16. Figure 6.16: Convergence of relative estimation errors converge with respect to the numbers of samples K. The errors are measured by: 1 M||Ψ pi`1q pvq ´ Ψ piq pvq||1 where Ψ piq pvq denotes the i-th measured spectral importance distribution with respect to charac￾teristic functi…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

233 extracted references · 1 canonical work pages

  1. [1]

    K. Aas, M. Jullum, and A. L land. Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artificial Intelligence, 298 0 (C), Sept. 2021. ISSN 0004-3702. doi:10.1016/j.artint.2021.103502. URL https://doi.org/10.1016/j.artint.2021.103502

  2. [2]

    Absil, R

    P.-A. Absil, R. Mahony, and R. Sepulchre. Optimization Algorithms on Matrix Manifolds. Princeton University Press, USA, 2007. ISBN 0691132984

  3. [3]

    Amerehi and P

    F. Amerehi and P. Healy. Label augmentation for neural networks robustness. In V. Lomonaco, S. Melacci, T. Tuytelaars, S. Chandar, and R. Pascanu, editors, Proceedings of The 3rd Conference on Lifelong Learning Agents, volume 274 of Proceedings of Machine Learning Research, pages 620--640. PMLR, 29 Jul--01 Aug 2025. URL https://proceedings.mlr.press/v274/...

  4. [4]

    C. Anil, J. Lucas, and R. Grosse. Sorting out L ipschitz function approximation. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 291--301. PMLR, 09--15 Jun 2019. URL https://proceedings.mlr.press/v97/anil19a.html

  5. [5]

    R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Silver, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. Barham, T. Hennigan, B. Lee, F. Viola, M. Reynolds, Y. Xu, R. Doherty, E. Collins, C. Meyer, E. Ruthe...

  6. [6]

    Applebaum

    D. Applebaum. L\'evy Processes and Stochastic Calculus. Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2nd edition, 2009

  7. [7]

    Arjovsky, A

    M. Arjovsky, A. Shah, and Y. Bengio. Unitary evolution recurrent neural networks. In M. F. Balcan and K. Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1120--1128, New York, New York, USA, 20--22 Jun 2016. PMLR. URL https://proceedings.mlr.press/v48...

  8. [8]

    Arjovsky, S

    M. Arjovsky, S. Chintala, and L. Bottou. W asserstein generative adversarial networks. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 214--223. PMLR, 06--11 Aug 2017. URL https://proceedings.mlr.press/v70/arjovsky17a.html

Show all 233 references
  1. [9]

    R. J. Aumann and M. Maschler. Game theoretic analysis of a bankruptcy problem from the Talmud . Journal of economic theory, 36 0 (2): 0 195--213, 1985

  2. [10]

    Aumann and Y

    Y. Aumann and Y. Dombb. The efficiency of fair division with connected pieces. ACM Transactions on Economics and Computation (TEAC), 3 0 (4): 0 1--16, 2015

  3. [11]

    L. J. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization. 2016. URL http://arxiv.org/abs/1607.06450

  4. [12]

    T. Bai, J. Luo, J. Zhao, B. Wen, and Q. Wang. Recent advances in adversarial training for adversarial robustness. In Z.-H. Zhou, editor, Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21 , pages 4312--4321. International Joint Con...

  5. [13]

    Bansal, X

    N. Bansal, X. Chen, and Z. Wang. Can we gain more from orthogonality regularizations in training deep CNN s? In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS'18, page 4266–4276, Red Hook, NY, USA, 2018. Curran Associates Inc

  6. [14]

    P. L. Bartlett, D. J. Foster, and M. Telgarsky. Spectrally-normalized margin bounds for neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, page 6241–6250, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN ...

  7. [15]

    D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba. Network dissection: Quantifying interpretability of deep visual representations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6541--6549, 2017

  8. [16]

    Behrmann, W

    J. Behrmann, W. Grathwohl, R. T. Q. Chen, D. Duvenaud, and J.-H. Jacobsen. Invertible residual networks. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, ...

  9. [17]

    Bereska and S

    L. Bereska and S. Gavves. Mechanistic interpretability for AI safety --- A review. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/forum?id=ePUVetPKu6. Survey Certification, Expert Certification

  10. [18]

    Binder, G

    A. Binder, G. Montavon, S. Lapuschkin, K.-R. M \"u ller, and W. Samek. Layer-wise relevance propagation for neural networks with local renormalization layers. In Artificial Neural Networks and Machine Learning--ICANN 2016: 25th International Conference on Artificial Neural Net...

  11. [19]

    Bj\" o rck and C. Bowie. An iterative algorithm for computing the best estimate of an orthogonal matrix. SIAM Journal on Numerical Analysis, 8 0 (2): 0 358--364, 1971. doi:10.1137/0708036. URL https://doi.org/10.1137/0708036

  12. [20]

    S. Boyd, J. Duchi, M. Pilanci, and L. Vandenberghe. Notes for EE364b : Subgradients. Stanford University, 2022. URL https://stanford.edu/class/ee364b/lectures/subgradients_notes.pdf. Lecture notes for Spring 2021--22; Accessed on January 2nd 2026

  13. [21]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020

  14. [22]

    N. J. Calkin, E. Y. S. Chan, R. M. Corless, D. J. Jeffrey, and P. W. Lawrence. A fractal eigenvector. The American Mathematical Monthly, 129 0 (6): 0 503--523, 2022

  15. [23]

    Carlini and D

    N. Carlini and D. Wagner. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy (SP), pages 39--57. IEEE, 2017

  16. [24]

    Castin, P

    V. Castin, P. Ablin, and G. Peyr\' e . How smooth is attention? In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR, 2024

  17. [25]

    A. Cayley. Sur quelques propri \'e t \'e s des d \'e terminants gauches. Journal f\"ur die reine und angewandte Mathematik, 1846

  18. [26]

    Chaudhari, A

    P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina. Entropy- SGD : Biasing gradient descent into wide valleys. Journal of Statistical Mechanics: Theory and Experiment, 2019 0 (12): 0 124018, 2019

  19. [27]

    Chen and R

    L. Chen and R. Ng. On the marriage of Lp -norms and edit distance. In Proceedings of the Thirtieth International Conference on Very Large Data Bases - Volume 30, VLDB '04, page 792–803. VLDB Endowment, 2004. ISBN 0120884690

  20. [28]

    R. T. Q. Chen, J. Behrmann, D. K. Duvenaud, and J.-H. Jacobsen. Residual flows for invertible generative modeling. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. C...

  21. [29]

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning, pages 1597--1607. PMLR, 2020

  22. [30]

    Chernodub and D

    A. Chernodub and D. Nowicki. Norm-preserving orthogonal permutation linear unit activation functions ( OPLU ), 2017. URL https://arxiv.org/abs/1604.02313

  23. [31]

    Chowdhery, S

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Au...

  24. [32]

    Cisse, P

    M. Cisse, P. Bojanowski, E. Grave, Y. Dauphin, and N. Usunier. Parseval networks: improving robustness to adversarial examples. In Proceedings of the 34th International Conference on Machine Learning - Volume 70, ICML'17, page 854–863. PMLR, 2017

  25. [33]

    F. H. Clarke. Generalized gradients and applications. Transactions of the American Mathematical Society, 205: 0 247--262, 1975

  26. [34]

    F. H. Clarke. Optimization and Nonsmooth Analysis, volume 5 of Classics in Applied Mathematics. SIAM, Philadelphia, PA, second edition, 1990

  27. [35]

    Clevert, T

    D. Clevert, T. Unterthiner, and S. Hochreiter. Fast and accurate deep network learning by exponential linear units ( ELU s). In Y. Bengio and Y. LeCun, editors, 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conferenc...

  28. [36]

    Coates, A

    A. Coates, A. Ng, and H. Lee. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 215--223. JMLR Workshop and Conference Proceedings, 2011

  29. [37]

    I. C. Covert, S. Lundberg, and S.-I. Lee. Understanding global feature contributions with additive importance measures. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS '20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN ...

  30. [38]

    G. Cybenko. Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals and Systems, 2 0 (4): 0 303--314, 1989. doi:10.1007/BF02551274. URL https://doi.org/10.1007/BF02551274

  31. [39]

    Damian, E

    A. Damian, E. Nichani, and J. D. Lee. Self-stabilization: The implicit bias of gradient descent at the edge of stability. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=nhKHA59gXz

  32. [40]

    DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li...

  33. [41]

    Devaguptapu, D

    C. Devaguptapu, D. Agarwal, G. Mittal, P. Gopalani, and V. N. Balasubramanian. On adversarial robustness: A neural architecture search perspective. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 152--161, 2021

  34. [42]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), pages 4171--4186, 2019

  35. [43]

    L. Dinh, J. Sohl-Dickstein, and S. Bengio. Density estimation using real NVP . In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=HkpbnH9lx

  36. [44]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning R...

  37. [45]

    Dugas, Y

    C. Dugas, Y. Bengio, F. B\' e lisle, C. Nadeau, and R. Garcia. Incorporating second-order functional knowledge for better option pricing. In T. Leen, T. Dietterich, and V. Tresp, editors, Advances in Neural Information Processing Systems, volume 13. MIT Press, 2000

  38. [46]

    Dunford and J

    N. Dunford and J. T. Schwartz. Linear operators, part 1: general theory. John Wiley & Sons, 1988

  39. [47]

    Edelman and N

    A. Edelman and N. R. Rao. Random matrix theory. Acta numerica, 14: 0 233--297, 2005

  40. [48]

    Elsken, J

    T. Elsken, J. H. Metzen, and F. Hutter. Neural architecture search: A survey. Journal of Machine Learning Research, 20 0 (55): 0 1--21, 2019. URL http://jmlr.org/papers/v20/18-598.html

  41. [49]

    Ethics guidelines for trustworthy AI , 2019

    European Commission . Ethics guidelines for trustworthy AI , 2019

  42. [50]

    Regulation (EU) 2024/1689 of the european parliament and of the council on harmonised rules on artificial intelligence ( AI Act ), 2024

    European Parliament and Council . Regulation (EU) 2024/1689 of the european parliament and of the council on harmonised rules on artificial intelligence ( AI Act ), 2024

  43. [51]

    U. Fano. Description of states in quantum mechanics by density matrix and operator techniques. Reviews of modern physics, 29 0 (1): 0 74, 1957

  44. [52]

    Fazlyab, A

    M. Fazlyab, A. Robey, H. Hassani, M. Morari, and G. Pappas. Efficient and accurate estimation of L ipschitz constants for deep neural networks. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Pro...

  45. [53]

    Fazlyab, T

    M. Fazlyab, T. Entesari, A. Roy, and R. Chellappa. Certified robustness via dynamic margin maximization and improved Lipschitz regularization. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=LDhhi8HBO3

  46. [54]

    Flatow and D

    D. Flatow and D. Penner. On the robustness of convnets to training on noisy labels. Technical report, Stanford University, 2017

  47. [55]

    J. N. Franklin. Matrix theory. Courier Corporation, 2000

  48. [56]

    Frenay and M

    B. Frenay and M. Verleysen. Classification in the presence of label noise: A survey. IEEE Transactions on Neural Networks and Learning Systems, 25 0 (5): 0 845--869, 2014. doi:10.1109/TNNLS.2013.2292894

  49. [57]

    Gamba, H

    M. Gamba, H. Azizpour, and M. Bjorkman. On the Lipschitz constant of deep networks and double descent. In 34th British Machine Vision Conference 2023, BMVC 2023, Aberdeen, UK, November 20-24, 2023 . BMVA, 2023. URL https://papers.bmvc2023.org/0871.pdf

  50. [58]

    Ghorbani, J

    A. Ghorbani, J. Wexler, J. Y. Zou, and B. Kim. Towards automatic concept-based explanations. In H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch \' e - Buc, E. B. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Ne...

  51. [59]

    Ghosh, H

    A. Ghosh, H. Kumar, and P. S. Sastry. Robust loss functions under label noise for deep neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 31, 2017

  52. [60]

    Golowich, A

    N. Golowich, A. Rakhlin, and O. Shamir. Size-independent sample complexity of neural networks. In S. Bubeck, V. Perchet, and P. Rigollet, editors, Proceedings of the 31st Conference On Learning Theory, volume 75 of Proceedings of Machine Learning Research, pages 297--299. PMLR...

  53. [61]

    G. H. Golub and C. F. Van Loan. Matrix computations. JHU press, 2013

  54. [62]

    Gonon, N

    A. Gonon, N. Brisebarre, E. Riccietti, and R. Gribonval. A rescaling-invariant Lipschitz bound based on path-metrics for modern ReLU network parameterizations. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=T8VLY1KuOz

  55. [63]

    Goodfellow, J

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems (NeurIPS), pages 2672--2680, 2014

  56. [64]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, and A. Courville. Deep learning. MIT Press, 2016

  57. [65]

    I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. In Y. Bengio and Y. LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , 2015. URL htt...

  58. [66]

    J. Gou, B. Yu, S. J. Maybank, and D. Tao. Knowledge distillation: A survey. International Journal of Computer Vision, 129 0 (6): 0 1789--1819, 2021

  59. [67]

    H. Gouk, E. Frank, B. Pfahringer, and M. J. Cree. Regularisation of neural networks by enforcing L ipschitz continuity. Machine Learning, 110: 0 393--416, 2021

  60. [68]

    Grosse and J

    R. Grosse and J. Martens. A Kronecker --factored approximate Fisher matrix for convolution layers. In International Conference on Machine Learning, pages 573--582. PMLR, 2016

  61. [69]

    Gulrajani, F

    I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville. Improved training of Wasserstein GANs . In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. ...

  62. [70]

    K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE international conference on computer vision, pages 1026--1034, 2015

  63. [71]

    K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770--778, 2016

  64. [72]

    Hein and M

    M. Hein and M. Andriushchenko. Formal guarantees on the robustness of a classifier against adversarial manipulation. Advances in neural information processing systems, 30, 2017

  65. [73]

    Hendrycks and T

    D. Hendrycks and T. Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HJz6tiCqYm

  66. [74]

    Hendrycks and K

    D. Hendrycks and K. Gimpel. Gaussian error linear units ( GELUs ). 2016. URL https://arxiv.org/abs/1606.08415

  67. [75]

    Hendrycks, M

    D. Hendrycks, M. Mazeika, D. Wilson, and K. Gimpel. Using trusted data to train deep networks on labels corrupted by severe noise. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, ...

  68. [76]

    Hendrycks, M

    D. Hendrycks, M. Mazeika, and T. Dietterich. Deep anomaly detection with outlier exposure. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyxCxhRcY7

  69. [77]

    Hendrycks, S

    D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF international conference on computer vision, ...

  70. [78]

    Hinton, N

    G. Hinton, N. Srivastava, and K. Swersky. Neural networks for machine learning lecture (6a): Overview of mini-batch gradient descent. URL https://www.cs.toronto.edu/ tijmen/csc321/slides/lecture_slides_lec6.pdf

  71. [79]

    G. E. Hinton and R. R. Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313 0 (5786): 0 504--507, 2006

  72. [80]

    J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 6840--6851. Curran Associates, Inc., 2020. URL https://proceed...

  73. [81]

    Hochreiter and J

    S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural Computation, 9 0 (8): 0 1735--1780, 1997. doi:10.1162/neco.1997.9.8.1735

  74. [82]

    R. A. Horn. The Hadamard product. In Proc. Symp. Appl. Math, volume 40, pages 87--169, 1990

  75. [83]

    R. A. Horn and C. R. Johnson. Matrix analysis. Cambridge University Press, 2012

  76. [84]

    X. Hu, R. Zheng, J. Wang, C. H. Leung, Q. Wu, and X. Xie. SpecFormer : Guarding vision transformer robustness via maximum singular value penalization. In European Conference on Computer Vision, pages 345--362. Springer, 2024

  77. [85]

    Huang, X

    L. Huang, X. Liu, B. Lang, A. W. Yu, Y. Wang, and B. Li. Orthogonal weight normalization: solution to optimization over multiple dependent Stiefel manifolds in deep neural networks. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth In...

  78. [86]

    Ilyas, S

    A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry. Adversarial examples are not bugs, they are features. Advances in neural information processing systems, 32, 2019

  79. [87]

    K. It \^o . On stochastic differential equations. Number 4. American Mathematical Soc., 1951

  80. [88]

    A. K. Jain. Fundamentals of digital image processing. Englewood Cliffs, NJ: Prentice Hall, 1989

  81. [89]

    Jastrzębski, Z

    S. Jastrzębski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey. On the relation between the sharpest directions of DNN loss and the SGD step length. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=SkgEaj05t7

  82. [90]

    Jordan and A

    M. Jordan and A. G. Dimakis. Exactly computing the local L ipschitz constant of ReLU networks. In International Conference on Machine Learning (ICML), pages 4985--4994. PMLR, 2020. URL https://arxiv.org/abs/2002.11572

  83. [91]

    Karatzas and S

    I. Karatzas and S. Shreve. Brownian motion and stochastic calculus, volume 113. Springer Science & Business Media, 2012

  84. [92]

    T. Kato. Perturbation theory for linear operators. Springer, 1995

  85. [93]

    N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang. On large-batch training for deep learning: Generalization gap and sharp minima. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=H1oyRlYgg

  86. [94]

    Khromov and S

    G. Khromov and S. P. Singh. Some fundamental aspects about L ipschitz continuity of neural networks. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=5jWsW08zUh

  87. [95]

    B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, and R. sayres. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors ( TCAV ). In J. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machi...

  88. [96]

    H. Kim, G. Papamakarios, and A. Mnih. The L ipschitz constant of self-attention. In M. Meila and T. Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 5562--5571. PMLR, 18--24 Jul ...

  89. [97]

    D. P. Kingma and J. Ba. Adam: A method for stochastic optimization, 2017. URL https://arxiv.org/abs/1412.6980

  90. [98]

    T. N. Kipf and M. Welling. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=SJU4ayYgl

  91. [99]

    C. Klamler. Fair division. Handbook of group decision and negotiation, pages 183--202, 2010

  92. [100]

    P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang. Concept bottleneck models. In International Conference on Machine Learning, pages 5338--5348. PMLR, 2020

  93. [101]

    T. G. Kolda and B. W. Bader. Tensor decompositions and applications. SIAM review, 51 0 (3): 0 455--500, 2009

  94. [102]

    Kolek, D

    S. Kolek, D. A. Nguyen, R. Levie, J. Bruna, and G. Kutyniok. Cartoon explanations of image classifiers. In European Conference on Computer Vision, pages 443--458. Springer, 2022

  95. [103]

    T. W. K \"o rner. Fourier analysis. Cambridge University Press, 2014. ISBN 9781107049949

  96. [104]

    Krizhevsky, I

    A. Krizhevsky, I. Sutskever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. Advances in neural information processing systems, 25, 2012

  97. [105]

    Lakkaraju, E

    H. Lakkaraju, E. Kamar, R. Caruana, and J. Leskovec. Faithful and customizable explanations of black box models. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pages 131--138, 2019

  98. [106]

    V. I. Levenshtein. Binary codes capable of correcting deletions, insertions, and reversals. Soviet Physics Doklady, 10: 0 707--710, 1966

  99. [107]

    A. S. Lewis and H. S. Sendov. Nonsmooth analysis of singular values. Part I : Theory. Set-Valued Analysis, 13 0 (3): 0 213--241, 2005

  100. [108]

    Lezcano-Casado and D

    M. Lezcano-Casado and D. Mart\' nez-Rubio. Cheap orthogonal constraints in neural networks: A simple parametrization of the orthogonal and unitary group. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume...

  101. [109]

    Li and R.-C

    C.-K. Li and R.-C. Li. A note on eigenvalues of perturbed Hermitian matrices. Linear algebra and its applications, 395: 0 183--190, 2005

  102. [110]

    J. Li, D. Li, C. Xiong, and S. Hoi. BLIP : Bootstrapping language-image pre-training for unified vision-language understanding and generation. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, editors, Proceedings of the 39th International Conference ...

  103. [111]

    Q. Li, S. Haque, C. Anil, J. Lucas, R. B. Grosse, and J.-H. Jacobsen. Preventing gradient attenuation in Lipschitz constrained convolutional networks. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d Alch\' e -Buc, E. Fox, and R. Garnett, editors, Advances in Neural Informat...

  104. [112]

    Q. Li, C. Tai, and E. Weinan. Stochastic modified equations and dynamics of stochastic gradient algorithms I : Mathematical foundations. Journal of Machine Learning Research, 20 0 (40): 0 1--47, 2019 b

  105. [113]

    Li and C

    Y. Li and C. Xu. Trade-off between robustness and accuracy of vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7558--7568, 2023

  106. [114]

    Li and Y

    Y. Li and Y. Yuan. Convergence analysis of two-layer neural networks with ReLU activation. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, I...

  107. [115]

    Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou. A survey of convolutional neural networks: analysis, applications, and prospects. IEEE transactions on neural networks and learning systems, 33 0 (12): 0 6999--7019, 2021

  108. [116]

    Z. Li, T. Wang, and S. Arora. What happens after SGD reaches zero loss? --- A mathematical framework. In International Conference on Learning Representations, 2022 b . URL https://openreview.net/forum?id=siCt4xZn5Ve

  109. [117]

    Lipman, R

    Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=PqvMRDCJT9t

  110. [118]

    Lukasik, S

    M. Lukasik, S. Bhojanapalli, A. Menon, and S. Kumar. Does label smoothing mitigate label noise? In H. D. III and A. Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 6448--6458. P...

  111. [119]

    S. M. Lundberg and S.-I. Lee. A unified approach to interpreting model predictions. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017

  112. [120]

    C. Luo, Q. Lin, W. Xie, B. Wu, J. Xie, and L. Shen. Frequency-driven imperceptible adversarial attack on semantic similarity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15315--15324, 2022

  113. [121]

    A. L. Maas, A. Y. Hannun, and A. Y. Ng. Rectifier nonlinearities improve neural network acoustic models. In ICML Workshop on Deep Learning for Audio, Speech and Language Processing, 2013

  114. [122]

    Macdonald, S

    J. Macdonald, S. W \"a ldchen, S. Hauch, and G. Kutyniok. A rate-distortion framework for explaining neural network decisions. arXiv preprint arXiv:1905.11092, 2019

  115. [123]

    Madry, A

    A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks. International Conference on Learning Representations (ICLR), 2018

  116. [124]

    J. R. Magnus and H. Neudecker. Matrix differential calculus with applications in statistics and econometrics. John Wiley & Sons, 2019

  117. [125]

    Malladi, K

    S. Malladi, K. Lyu, A. Panigrahi, and S. Arora. On the SDEs and scaling rules for adaptive gradient algorithms. Advances in Neural Information Processing Systems, 35: 0 7697--7711, 2022

  118. [126]

    Mandt, M

    S. Mandt, M. D. Hoffman, D. M. Blei, et al. Continuous-time limit of stochastic gradient descent revisited. NIPS-2015, 2015

  119. [127]

    Mandt, M

    S. Mandt, M. D. Hoffman, and D. M. Blei. Stochastic gradient descent as approximate Bayesian inference. Journal of Machine Learning Research, 18 0 (134): 0 1--35, 2017

  120. [128]

    V. A. Mar c enko and L. A. Pastur. Distribution of eigenvalues for some sets of random matrices. Mathematics of the USSR-Sbornik, 1 0 (4): 0 457, 1967

  121. [129]

    J. Martens. New insights and perspectives on the natural gradient method. Journal of Machine Learning Research, 21 0 (146): 0 1--76, 2020

  122. [130]

    A. Maurer. A vector-contraction inequality for Rademacher complexities. In Algorithmic Learning Theory: 27th International Conference, ALT 2016, Bari, Italy, October 19-21, 2016, Proceedings, page 3–17, Berlin, Heidelberg, 2016. Springer-Verlag. ISBN 978-3-319-46378-0. doi:10....

  123. [131]

    o sung. ZAMM-Journal of Applied Mathematics and Mechanics/Zeitschrift f \

    R. Mises and H. Pollaczek-Geiringer. Praktische verfahren der gleichungsaufl \"o sung. ZAMM-Journal of Applied Mathematics and Mechanics/Zeitschrift f \"u r Angewandte Mathematik und Mechanik , 9 0 (1): 0 58--77, 1929

  124. [132]

    Miyato, T

    T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=B1QRgziT-

  125. [133]

    Modas, S.-M

    A. Modas, S.-M. Moosavi-Dezfooli, and P. Frossard. Sparsefool: A few pixels make a big difference. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9079--9088, 2019. doi:10.1109/CVPR.2019.00930

  126. [134]

    Mohri, A

    M. Mohri, A. Rostamizadeh, and A. Talwalkar. Foundations of machine learning. MIT Press, 2018

  127. [135]

    P. Nair. Softmax is 1/2 - Lipschitz : A tight bound across all _p norms. arXiv preprint arXiv:2510.23012, 2025

  128. [136]

    Nair and G

    V. Nair and G. E. Hinton. Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on International Conference on Machine Learning, ICML'10, page 807–814, Madison, WI, USA, 2010. Omnipress. ISBN 9781605589077

  129. [137]

    Natarajan, I

    N. Natarajan, I. S. Dhillon, P. K. Ravikumar, and A. Tewari. Learning with noisy labels. In C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger, editors, Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc., 2013

  130. [138]

    F. Navarro. Necessary players, Myerson fairness and the equal treatment of equals. Annals of Operations Research, 280: 0 111--119, 2019

  131. [139]

    Neelakantan, L

    A. Neelakantan, L. Vilnis, Q. V. Le, I. Sutskever, L. Kaiser, K. Kurach, and J. Martens. Adding gradient noise improves learning for very deep networks. 2015. URL https://arxiv.org/abs/1511.06807

  132. [140]

    Neyshabur, R

    B. Neyshabur, R. Tomioka, and N. Srebro. Norm-based capacity control in neural networks. In P. Grünwald, E. Hazan, and S. Kale, editors, Proceedings of The 28th Conference on Learning Theory, volume 40 of Proceedings of Machine Learning Research, pages 1376--1401, Paris, Franc...

  133. [141]

    Neyshabur, S

    B. Neyshabur, S. Bhojanapalli, D. Mcallester, and N. Srebro. Exploring generalization in deep learning. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran ...

  134. [142]

    Nguyen, A

    A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune. Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems...

  135. [143]

    Nguyen, J

    A. Nguyen, J. Yosinski, and J. Clune. Understanding neural networks via feature visualization: A survey. Explainable AI: interpreting, explaining and visualizing deep learning, pages 55--76, 2019

  136. [144]

    M. A. Nielsen and I. L. Chuang. Quantum computation and quantum information. Cambridge University Press, 2010

  137. [145]

    Artificial intelligence risk management framework, 2023

    NIST. Artificial intelligence risk management framework, 2023. URL https://www.nist.gov/itl/ai-risk-management-framework

  138. [146]

    Oksendal

    B. Oksendal. Stochastic differential equations: an introduction with applications. Springer Science & Business Media, 2013

  139. [147]

    Achiam, S

    OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L....

  140. [148]

    A. V. Oppenheim. Applications of digital signal processing. Prentice-Hall, 1978. ISBN 0130391158, 9780130391155

  141. [149]

    T. Pang, X. Yang, Y. Dong, H. Su, and J. Zhu. Bag of tricks for adversarial training. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Xb8xvrtB8Ce

  142. [150]

    Paszke, S

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K\" o pf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. PyTorch : an imperative style, hig...

  143. [151]

    Patrini, A

    G. Patrini, A. Rozza, A. K. Menon, R. Nock, and L. Qu. Making deep neural networks robust to label noise: A loss correction approach. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2233--2241, 2017. doi:10.1109/CVPR.2017.240

  144. [152]

    Paul and P.-Y

    S. Paul and P.-Y. Chen. Vision transformers are robust learners. In Proceedings of the AAAI conference on Artificial Intelligence, volume 36, pages 2071--2081, 2022

  145. [153]

    Perozzi, R

    B. Perozzi, R. Al-Rfou, and S. Skiena. DeepWalk : online learning of social representations. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '14, page 701–710, New York, NY, USA, 2014. Association for Computing Machine...

  146. [154]

    Perugachi-Diaz, J

    Y. Perugachi-Diaz, J. Tomczak, and S. Bhulai. Invertible DenseNets with concatenated LipSwish . In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 17246--17257. Curran Associates,...

  147. [155]

    Pomponi, S

    J. Pomponi, S. Scardapane, and A. Uncini. Pixle: a fast and effective black-box attack based on rearranging pixels. In 2022 International Joint Conference on Neural Networks (IJCNN), pages 1--7. IEEE, 2022

  148. [156]

    X. Qi, J. Wang, Y. Chen, Y. Shi, and L. Zhang. LipsFormer : Introducing Lipschitz continuity to vision transformers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=cHf1DcCwcH3

  149. [157]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners, 2019

  150. [158]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision. In M. Meila and T. Zhang, editors, Proceedings of the 38th Interna...

  151. [159]

    Rahaman, A

    N. Rahaman, A. Baratin, D. Arpit, F. Draxler, M. Lin, F. Hamprecht, Y. Bengio, and A. Courville. On the spectral bias of neural networks. In International Conference on Machine Learning, pages 5301--5310. PMLR, 2019

  152. [160]

    Ramachandran, B

    P. Ramachandran, B. Zoph, and Q. V. Le. Searching for activation functions, 2018. URL https://openreview.net/forum?id=SkBYYyZRZ

  153. [161]

    J. W. S. Rayleigh. The theory of sound, Volume One. Courier Corporation, 2013

  154. [162]

    Recht, R

    B. Recht, R. Roelofs, L. Schmidt, and V. Shankar. Do ImageNet classifiers generalize to ImageNet ? In International Conference on Machine Learning (ICML), pages 5389--5400. PMLR, 2019

  155. [163]

    F. Rellich. Perturbation theory of eigenvalue problems. CRC Press, 1969. ISBN 0677006802, 9780677006802

  156. [164]

    Rezende and S

    D. Rezende and S. Mohamed. Variational inference with normalizing flows. In F. Bach and D. Blei, editors, Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1530--1538, Lille, France, 07--09 Jul 20...

  157. [165]

    M. T. Ribeiro, S. Singh, and C. Guestrin. `` Why Should I Trust You ?'': Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '16, page 1135–1144, New York, NY, USA, 2016. Assoc...

  158. [166]

    Robbins and S

    H. Robbins and S. Monro. A stochastic approximation method. The Annals of Mathematical Statistics, 22 0 (3): 0 400--407, 1951. ISSN 00034851. URL http://www.jstor.org/stable/2236626

  159. [167]

    E. A. Rocamora, G. Chrysos, and V. Cevher. Certified robustness under bounded Levenshtein distance. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=cd79pbXi4N

  160. [168]

    A. E. Roth. The Shapley value: essays in honor of Lloyd S. Shapley. Cambridge University Press, 1988

  161. [169]

    W. Rudin. Principles of Mathematical Analysis. McGraw-Hill, 3 edition, 1976

  162. [170]

    W. Rudin. Real and Complex Analysis. McGraw-Hill, New York, 3rd edition, 1987. ISBN 9780070542341

  163. [171]

    J. J. Sakurai and J. Napolitano. Modern quantum mechanics. Cambridge University Press, 2020

  164. [172]

    Schr \"o dinger

    E. Schr \"o dinger. Quantisierung als eigenwertproblem. Annalen der physik, 385 0 (13): 0 437--490, 1926

  165. [173]

    Sedghi, V

    H. Sedghi, V. Gupta, and P. M. Long. The singular values of convolutional layers. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJevYoA9Fm

  166. [174]

    R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra. Grad-CAM : Visual explanations from deep networks via gradient-based localization. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 618--626, 2017. doi:10.1109/ICCV.2017.74

  167. [175]

    Shalev-Shwartz and S

    S. Shalev-Shwartz and S. Ben-David. Understanding machine learning: From theory to algorithms. Cambridge University Press, 2014

  168. [176]

    O. M. Shalit. Dilation theory: a guided tour. In Operator theory, functional analysis and applications, pages 551--623. Springer, 2021

  169. [177]

    R. Shao, Z. Shi, J. Yi, P.-Y. Chen, and C.-J. Hsieh. On the adversarial robustness of vision transformers. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=lE7K4n1Esk

  170. [178]

    Shrikumar, P

    A. Shrikumar, P. Greenside, and A. Kundaje. Learning important features through propagating activation differences. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research,...

  171. [179]

    Simsekli, L

    U. Simsekli, L. Sagun, and M. Gurbuzbalaban. A tail-index analysis of stochastic gradient noise in deep neural networks. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Lea...

  172. [180]

    Simsekli, L

    U. Simsekli, L. Zhu, Y. W. Teh, and M. Gurbuzbalaban. Fractional underdamped L angevin dynamics: Retargeting SGD with momentum under heavy-tailed gradient noise. In H. D. III and A. Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 11...

  173. [181]

    Smilkov, N

    D. Smilkov, N. Thorat, B. Kim, F. Vi \'e gas, and M. Wattenberg. Smoothgrad: removing noise by adding noise. 2017. URL https://arxiv.org/abs/1706.03825

  174. [182]

    Sokolić, R

    J. Sokolić, R. Giryes, G. Sapiro, and M. R. D. Rodrigues. Robust large margin deep neural networks. IEEE Transactions on Signal Processing, 65 0 (16): 0 4265--4280, 2017. doi:10.1109/TSP.2017.2708039

  175. [183]

    Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=PxTIG12RRHS

  176. [184]

    M. Spivak. Calculus on manifolds: a modern approach to classical theorems of advanced calculus. CRC press, 2018

  177. [185]

    E. M. Stein and R. Shakarchi. Fourier analysis: an introduction, volume 1. Princeton University Press, 2011

  178. [186]

    G. W. Stewart and J.-g. Sun. Matrix perturbation theory. Academic Press, 1990

  179. [187]

    Stoica, R

    P. Stoica, R. L. Moses, et al. Spectral analysis of signals, volume 452. Pearson Prentice Hall Upper Saddle River, NJ, 2005

  180. [188]

    W. Su, S. Boyd, and E. J. Cand \`e s. A differential equation for modeling nesterov's accelerated gradient method: Theory and insights. Journal of Machine Learning Research, 17 0 (153): 0 1--43, 2016. URL http://jmlr.org/papers/v17/15-084.html

  181. [189]

    Sucholutsky, R

    I. Sucholutsky, R. M. Battleday, K. M. Collins, R. Marjieh, J. Peterson, P. Singh, U. Bhatt, N. Jacoby, A. Weller, and T. L. Griffiths. On the informativeness of supervision signals. In R. J. Evans and I. Shpitser, editors, Proceedings of the Thirty-Ninth Conference on Uncerta...

  182. [190]

    Sundararajan, A

    M. Sundararajan, A. Taly, and Q. Yan. Axiomatic attribution for deep networks. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 3319--3328. PMLR, 06--11 Aug 2...

  183. [191]

    R. S. Sutton, A. G. Barto, et al. Reinforcement learning: An introduction, volume 1. MIT Press, 1998

  184. [192]

    Szegedy, W

    C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. J. Goodfellow, and R. Fergus. Intriguing properties of neural networks. In Y. Bengio and Y. LeCun, editors, 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, C...

  185. [193]

    Tan and Q

    M. Tan and Q. Le. E fficient N et: Rethinking model scaling for convolutional neural networks. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 6105...

  186. [194]

    T. Tao. Topics in random matrix theory, volume 132. American Mathematical Soc., 2012

  187. [195]

    Taori, A

    R. Taori, A. Dave, V. Shankar, N. Carlini, B. Recht, and L. Schmidt. Measuring robustness to natural distribution shifts in image classification. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume...

  188. [196]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

  189. [197]

    C. A. Tracy and H. Widom. Level-spacing distributions and the Airy kernel. Communications in Mathematical Physics, 159: 0 151--174, 1994

  190. [198]

    Trockman and J

    A. Trockman and J. Z. Kolter. Orthogonalizing convolutional layers with the Cayley transform. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Pbj8H_jEHYv

  191. [199]

    Tsipras, S

    D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry. Robustness may be at odds with accuracy. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=SyxAb30cY7

  192. [200]

    Tsuzuku and I

    Y. Tsuzuku and I. Sato. On the structural sensitivity of deep convolutional networks to the directions of Fourier basis functions. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 51--60, 2019. doi:10.1109/CVPR.2019.00014

  193. [201]

    Tsuzuku, I

    Y. Tsuzuku, I. Sato, and M. Sugiyama. Lipschitz-margin training: scalable certification of perturbation invariance for deep neural networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS'18, page 6542–6551, Red Hook, NY, USA...

  194. [202]

    Drimbarean, J

    R\'ois\'in Luo , A. Drimbarean, J. McDermott, and C. O'Riordan. Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization. In The 35th British Machine Vision Conference, BMVC 2024, Glasgow, UK, November 25-28, 2024 . BMVA, 2024 a . URL https://arxiv.org/abs/2408.00923

  195. [203]

    McDermott, and C

    R\'ois\'in Luo , J. McDermott, and C. O'Riordan. Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition. Transactions on Machine Learning Research (TMLR), 2024 b . ISSN 2835-8856. URL https://arxiv.org/abs/2408.01139. Pres...

  196. [204]

    McDermott, C

    R\'ois\'in Luo , J. McDermott, C. Gagn\'e, Q. Sun, and C. O'Riordan. Optimization-Induced Dynamics of L ipschitz Continuity in Neural Networks . Manuscript is under review at Journal of Machine Learning Research (JMLR), 2025 a . URL https://arxiv.org/abs/2506.18588

  197. [205]

    McDermott, and C

    R\'ois\'in Luo , J. McDermott, and C. O'Riordan. Lipschitz Continuity in Deep Learning: A Systematic Review of Theoretical Foundations, Estimation Methods, Regularization Approaches and Certifiable Robustness. Manuscript is under review at Transactions on Machine Learning Rese...

  198. [206]

    O'Riordan, and J

    R\'ois\'in Luo , C. O'Riordan, and J. McDermott. Higher-Order Singular-Value Derivatives of Real Rectangular Matrices. Journal of Mathematical Analysis and Applications (JMAA), page 130236, 2025 c . ISSN 0022-247X. doi:10.1016/j.jmaa.2025.130236. URL https://doi.org/10.1016/j....

  199. [207]

    Van Den Berg, L

    R. Van Den Berg, L. Hasenclever, J. M. Tomczak, and M. Welling. Sylvester normalizing flows for variational inference. UAI, 2018

  200. [208]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing...

  201. [209]

    Veličković, G

    P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio. Graph attention networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJXMpikCZ

  202. [210]

    Villani et al

    C. Villani et al. Optimal transport: old and new, volume 338. Springer, 2008

  203. [211]

    Virmaux and K

    A. Virmaux and K. Scaman. Lipschitz regularity of deep neural networks: analysis and efficient estimation. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31. Curran Associ...

  204. [212]

    Vorontsov, C

    E. Vorontsov, C. Trabelsi, S. Kadoury, and C. Pal. On orthogonality and learning recurrent networks with long term dependencies. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learn...

  205. [213]

    Vuckovic, A

    J. Vuckovic, A. Baratin, and R. T. d. Combes. A mathematical theory of attention. 2020. URL https://arxiv.org/abs/2007.02876

  206. [214]

    H. Wang, X. Wu, Z. Huang, and E. P. Xing. High-frequency component helps explain the generalization of convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8684--8694, 2020

  207. [215]

    Welling and Y

    M. Welling and Y. W. Teh. Bayesian learning via stochastic gradient Langevin dynamics. In Proceedings of the 28th International Conference on International Conference on Machine Learning, ICML'11, page 681–688, Madison, WI, USA, 2011. Omnipress. ISBN 9781450306195

  208. [216]

    L. Weng, H. Zhang, H. Chen, Z. Song, C.-J. Hsieh, L. Daniel, D. Boning, and I. Dhillon. Towards fast computation of certified robustness for R e LU networks. In International Conference on Machine Learning, pages 5276--5285. PMLR, 2018 a

  209. [217]

    T.-W. Weng, H. Zhang, P.-Y. Chen, J. Yi, D. Su, Y. Gao, C.-J. Hsieh, and L. Daniel. Evaluating the robustness of neural networks: An extreme value theory approach. In International Conference on Learning Representations, 2018 b . URL https://openreview.net/forum?id=BkUHlMZ0b

  210. [218]

    Wisdom, T

    S. Wisdom, T. Powers, J. Hershey, J. Le Roux, and L. Atlas. Full-capacity unitary recurrent neural networks. In D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016. URL ...

  211. [219]

    T. Xiao, X. Wang, A. A. Efros, and T. Darrell. What should not be contrastive in contrastive learning. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=CZ8Y3NzuVzO

  212. [220]

    K. Xu, W. Hu, J. Leskovec, and S. Jegelka. How powerful are graph neural networks? In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019 a . URL https://openreview.net/forum?id=ryGs6iA5Km

  213. [221]

    Z.-Q. J. Xu, Y. Zhang, and Y. Xiao. Training behavior of deep neural network in frequency domain. In Neural Information Processing: 26th International Conference, ICONIP 2019, Sydney, NSW, Australia, December 12–15, 2019, Proceedings, Part I, page 264–274, Berlin, Heidelberg, ...

  214. [222]

    Z.-Q. J. Xu, Y. Zhang, T. Luo, Y. Xiao, and Z. Ma. Frequency principle: Fourier analysis sheds light on deep neural networks. Communications in Computational Physics, 28 0 (5): 0 1746–1767, Nov. 2020. doi:10.4208/cicp.OA-2020-0085. URL https://global-sci.com/index.php/cicp/art...

  215. [223]

    D. Yin, R. Kannan, and P. Bartlett. Rademacher complexity for adversarially robust generalization. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages ...

  216. [224]

    K. Yosida. Functional analysis. Springer Science & Business Media, 2012

  217. [225]

    Yudin, A

    N. Yudin, A. Gaponov, S. Kudriashov, and M. Rakhuba. Pay attention to attention distribution: A new local Lipschitz bound for transformers. 2025. URL https://arxiv.org/abs/2507.07814

  218. [226]

    M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus. Deconvolutional networks. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pages 2528--2535, 2010. doi:10.1109/CVPR.2010.5539957

  219. [227]

    Zhang, D

    B. Zhang, D. Jiang, D. He, and L. Wang. Rethinking Lipschitz neural networks and certified robustness: a Boolean function perspective. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red Hook, NY, USA, 2022. Curran Associ...

  220. [228]

    Zhang, S

    C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals. Understanding deep learning (still) requires rethinking generalization. Commun. ACM, 64 0 (3): 0 107–115, Feb. 2021. ISSN 0001-0782. doi:10.1145/3446776. URL https://doi.org/10.1145/3446776

  221. [229]

    Zhang, X

    Z. Zhang, X. Shu, B. Yu, T. Liu, J. Zhao, Q. Li, and L. Guo. Distilling knowledge from well-informed soft labels for neural relation extraction. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9620--9627, 2020

  222. [230]

    Zheng, Y

    S. Zheng, Y. Song, T. Leung, and I. Goodfellow. Improving the robustness of deep neural networks via stability training. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4480--4488, 2016. doi:10.1109/CVPR.2016.485

  223. [231]

    B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba. Learning deep features for discriminative localization. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2921--2929, 2016. doi:10.1109/CVPR.2016.319

  224. [232]

    D. Zhou, Z. Yu, E. Xie, C. Xiao, A. Anandkumar, J. Feng, and J. M. Alvarez. Understanding the robustness in vision transformers. In K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, editors, Proceedings of the 39th International Conference on Machine Lea...

  225. [233]

    Z. Zhu, J. Wu, B. Yu, L. Wu, and J. Ma. The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learn...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.