Training GPT-2 small on an easy-to-hard curriculum speeds early learning at moderate thresholds but yields worse final accuracy and relies on unreliable head-count metrics.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Can an Easy-to-Hard Curriculum Make Reasoning Emerge in Small Language Models? Evidence from a Four-Stage Curriculum on GPT-2
Training GPT-2 small on an easy-to-hard curriculum speeds early learning at moderate thresholds but yields worse final accuracy and relies on unreliable head-count metrics.