Pith. sign in

REVIEW 2 cited by

The Lipschitz Constant of Self-Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.04710 v2 pith:RFNCO23R submitted 2020-06-08 stat.ML cs.LG

classification stat.MLcs.LG
keywords lipschitzself-attentionconstantnetworksneuralinvertiblemodellingadversarial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Lipschitz constants of neural networks have been explored in various contexts in deep learning, such as provable adversarial robustness, estimating Wasserstein distance, stabilising training of GANs, and formulating invertible neural networks. Such works have focused on bounding the Lipschitz constant of fully connected or convolutional networks, composed of linear maps and pointwise non-linearities. In this paper, we investigate the Lipschitz constant of self-attention, a non-linear neural network module widely used in sequence modelling. We prove that the standard dot-product self-attention is not Lipschitz for unbounded input domain, and propose an alternative L2 self-attention that is Lipschitz. We derive an upper bound on the Lipschitz constant of L2 self-attention and provide empirical evidence for its asymptotic tightness. To demonstrate the practical relevance of our theoretical work, we formulate invertible self-attention and use it in a Transformer-based architecture for a character-level language modelling task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformers Learn Faster with Semantic Focus

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Input-dependent top-k sparse attention makes small transformers converge faster and generalize as well as full attention, while input-agnostic sparsity does not, and the effect is tied to reduced dispersion of attenti...

  2. Chaos in reason: How chain-of-thought LLMs can look for an answer

    nlin.CD 2026-07 conditional novelty 5.0 of 10

    Greedy LLM inference shows bounded, jump-like sensitivity to sub-token perturbations that the authors interpret as chaotic, with attention expanding and normalization suppressing perturbations.

Pith tools