Pith. sign in

REVIEW 2 cited by

How Smooth Is Attention?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.14820 v2 pith:TL6LK2BT submitted 2023-12-22 cs.LG

classification cs.LG
keywords self-attentionboundconstantlipschitzlengthmaskedsequenceattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Self-attention and masked self-attention are at the heart of Transformers' outstanding success. Still, our mathematical understanding of attention, in particular of its Lipschitz properties - which are key when it comes to analyzing robustness and expressive power - is incomplete. We provide a detailed study of the Lipschitz constant of self-attention in several practical scenarios, discussing the impact of the sequence length $n$ and layer normalization on the local Lipschitz constant of both unmasked and masked self-attention. In particular, we show that for inputs of length $n$ in any compact set, the Lipschitz constant of self-attention is bounded by $\sqrt{n}$ up to a constant factor and that this bound is tight for reasonable sequence lengths. When the sequence length $n$ is too large for the previous bound to be tight, which we refer to as the mean-field regime, we provide an upper bound and a matching lower bound which are independent of $n$. Our mean-field framework for masked self-attention is novel and of independent interest. Our experiments on pretrained and randomly initialized BERT and GPT-2 support our theoretical findings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers

    cs.LG 2025-07 conditional novelty 7.0 of 10

    Self-attention's local Lipschitz constant can be bounded using the attention probability distribution, and the softmax Jacobian spectral norm is shown to be at most 1/2, leading to a new robustness regularizer.

  2. Memory Limitations of Prompt Tuning in Transformers

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Prompt tuning in transformers is shown, via covering and Lipschitz arguments, to memorize at most linearly many examples in the prompt length.

Pith tools