Pith. sign in

[2024b], with probability at least 1 − δ, the last term can be bound by V µ r∗ (π) − V µ r∗ (ˆπ) ≤ CµD (R, π, πref)2 8κ2 · β + 3β N log( Nϵ(R, ∥ · ∥∞) δ )

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Learning a Pessimistic Reward Model in RLHF

cs.LG · 2025-05-26 · reject · novelty 6.0

Pessimistic fine-tuning of reward models against rejection-sampling policies lets RLHF agents optimize greedily without KL regularization and still avoid reward hacking.

citing papers explorer

Showing 1 of 1 citing paper.

  • Learning a Pessimistic Reward Model in RLHF cs.LG · 2025-05-26 · reject · none · ref 1

    Pessimistic fine-tuning of reward models against rejection-sampling policies lets RLHF agents optimize greedily without KL regularization and still avoid reward hacking.