Setting the feature dimension of each linear-attention layer proportional to the estimated degrees of freedom of its input kernel improves distilled model accuracy without increasing total inference cost.
• K : Rd × Rd → R is the positive definite kernel given byK(x, y) = Ez∼τ [ϕ(x; z)ϕ(y; z)], where τ is a probability measure on a measurable setZ, and ϕ : Rd × Z →R is a feature map
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
Setting the feature dimension of each linear-attention layer proportional to the estimated degrees of freedom of its input kernel improves distilled model accuracy without increasing total inference cost.