During continual pre-training of OLMo 2 models, singular value spectra remain largely fixed while singular vectors change; selectively rewinding low-importance attention heads improves math accuracy by up to 4%.
Proceedings of the 37th International Conference on Machine Learning , pages =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Diffract: Spectral View of LLM Domain Adaptation
During continual pre-training of OLMo 2 models, singular value spectra remain largely fixed while singular vectors change; selectively rewinding low-importance attention heads improves math accuracy by up to 4%.