Adding a prediction-error curiosity reward, masked to non-top-k tokens, improves output diversity of RLHF-trained LLMs while keeping reward-model judged quality roughly unchanged.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Curiosity-Driven Reinforcement Learning from Human Feedback
Adding a prediction-error curiosity reward, masked to non-top-k tokens, improves output diversity of RLHF-trained LLMs while keeping reward-model judged quality roughly unchanged.