For Boolean Min-IP, a single normalized nonnegative kernel-attention head needs exponentially many features to solve all three-token sequences, even though rank one solves every sequence of length at most two and dense softmax uses m-dimensional scores.
International Conference on Machine Learning , year =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention
For Boolean Min-IP, a single normalized nonnegative kernel-attention head needs exponentially many features to solve all three-token sequences, even though rank one solves every sequence of length at most two and dense softmax uses m-dimensional scores.