Tuning per-head or per-channel attention scales on 50 synthetic samples improves long-context retrieval accuracy across multiple LLMs with no inference overhead after offline merging.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
Tuning per-head or per-channel attention scales on 50 synthetic samples improves long-context retrieval accuracy across multiple LLMs with no inference overhead after offline merging.