Derives generalization bounds for transformer next-token prediction under an extended log-bilinear text data model, depending on architecture, vocabulary size, document count and length.
Transformer Approximations from ReLUs
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation targets. Importantly, it yields target-specific, economic resource bounds beyond universal approximation statements. We showcase the recipe on multiplication, reciprocal computation, and min/max primitives. These results provide new analytical tools for analyzing softmax transformer models.
fields
math.ST 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model
Derives generalization bounds for transformer next-token prediction under an extended log-bilinear text data model, depending on architecture, vocabulary size, document count and length.