Pith. sign in

REVIEW 1 cited by

The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15753 v3 pith:P4OWRCKX submitted 2024-06-22 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords rewarderrorerror-regretlearningmismatchmodelpolicyregret
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In reinforcement learning, specifying reward functions that capture the intended task can be very challenging. Reward learning aims to address this issue by learning the reward function. However, a learned reward model may have a low error on the data distribution, and yet subsequently produce a policy with large regret. We say that such a reward model has an error-regret mismatch. The main source of an error-regret mismatch is the distributional shift that commonly occurs during policy optimization. In this paper, we mathematically show that a sufficiently low expected test error of the reward model guarantees low worst-case regret, but that for any fixed expected test error, there exist realistic data distributions that allow for error-regret mismatch to occur. We then show that similar problems persist even when using policy regularization techniques, commonly employed in methods such as RLHF. We hope our results stimulate the theoretical and empirical study of improved methods to learn reward models, and better ways to measure their quality reliably.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Applying importance weighting to reward model training to correct for policy distribution shift in RLHF improves final policy quality without new labels.

Pith tools