Under equal KV cache budgets, routing each token to weight-shared grouped-attention experts with group sizes 1, 2, and 4 yields higher ROUGE-L and lower perplexity than static GQA and CLA baselines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
Under equal KV cache budgets, routing each token to weight-shared grouped-attention experts with group sizes 1, 2, and 4 yields higher ROUGE-L and lower perplexity than static GQA and CLA baselines.