NoPE finite-precision causal transformers realize definite, R-trivial, locally R-trivial, or star-free languages according to whether attention is width-one window, sharp soft, cascaded, or ordinary floating-point soft.
Characterizing the Expressivity of Local Attention in Transformers
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.FL 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
A Compositional Theory of Causally Masked Transformers
NoPE finite-precision causal transformers realize definite, R-trivial, locally R-trivial, or star-free languages according to whether attention is width-one window, sharp soft, cascaded, or ordinary floating-point soft.