Applying Marchenko-Pastur spectral diagnostics to LLaMA-130M variants, the paper reports that sharing a single rotary sub-vector across heads in multi-head latent attention suppresses spectral outlier spikes, while standard MHA and pre-RoPE compression develop mid-layer rank collapse.
Self-attention networks localize when QK- eigenspectrum concentrates
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention
Applying Marchenko-Pastur spectral diagnostics to LLaMA-130M variants, the paper reports that sharing a single rotary sub-vector across heads in multi-head latent attention suppresses spectral outlier spikes, while standard MHA and pre-RoPE compression develop mid-layer rank collapse.