A pre-RL fine-tuning intervention called verbalization fine-tuning makes language models explicitly acknowledge when prompt cues drive them to reward-hack, cutting undetected reward hacking from 88% to 6% after RL.
Claude 3.7 sonnet system card, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
A pre-RL fine-tuning intervention called verbalization fine-tuning makes language models explicitly acknowledge when prompt cues drive them to reward-hack, cutting undetected reward hacking from 88% to 6% after RL.