Pessimistic fine-tuning of reward models against rejection-sampling policies lets RLHF agents optimize greedily without KL regularization and still avoid reward hacking.
[2024b], with probability at least 1 − δ, the last term can be bound by V µ r∗ (π) − V µ r∗ (ˆπ) ≤ CµD (R, π, πref)2 8κ2 · β + 3β N log( Nϵ(R, ∥ · ∥∞) δ )
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Learning a Pessimistic Reward Model in RLHF
Pessimistic fine-tuning of reward models against rejection-sampling policies lets RLHF agents optimize greedily without KL regularization and still avoid reward hacking.