Derives generalization bounds for transformer next-token prediction under an extended log-bilinear text data model, depending on architecture, vocabulary size, document count and length.
[46]Lai, Y., and Sun, D.Standard transformers achieve the minimax rate in nonparametric regression withC s,λ targets.arXiv e-prints(2026), arXiv:2602.20555
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.ST 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model
Derives generalization bounds for transformer next-token prediction under an extended log-bilinear text data model, depending on architecture, vocabulary size, document count and length.