After training, a 110M-parameter power-law attention model's learned scoring operator becomes nearly input-invariant, so inference can cache it; the paper proves this collapse conditionally, measures it at 1e-6 and below, and machine-checks selected proofs in Lean 4.
Beggs and Dietmar Plenz,Neuronal avalanches in neocortical circuits, Journal of Neuroscience23 (2003), no
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
After training, a 110M-parameter power-law attention model's learned scoring operator becomes nearly input-invariant, so inference can cache it; the paper proves this collapse conditionally, measures it at 1e-6 and below, and machine-checks selected proofs in Lean 4.