Long-context training shifts language models from parametric knowledge to context reliance, producing an inverted-U in pretraining performance and context addiction in supervised fine-tuning.
Each task applies a unary bitwise operation to the input string, such as NOT, which maps each bit to its complement.5 The valid output tokens are0and1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Long-context training shifts language models from parametric knowledge to context reliance, producing an inverted-U in pretraining performance and context addiction in supervised fine-tuning.