DIET adapts RL token penalties to estimated per-problem difficulty, cutting token use by roughly 40% on math benchmarks while slightly improving Pass@1.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
DIET adapts RL token penalties to estimated per-problem difficulty, cutting token use by roughly 40% on math benchmarks while slightly improving Pass@1.