A certainty-weighted KL penalty that down-weights the penalty on low-confidence tokens improves RL fine-tuning exploration for arithmetic in GPT-2.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning
A certainty-weighted KL penalty that down-weights the penalty on low-confidence tokens improves RL fine-tuning exploration for arithmetic in GPT-2.