Pith. sign in

REVIEW 10 cited by

Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03185 v4 pith:3LPH36JO submitted 2024-03-05 cs.LG cs.AI

Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking

classification cs.LG cs.AI
keywords rewardhackingdefinitionproxypolicyrlhftrueacross
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Because it is difficult to precisely specify complex objectives, reinforcement learning policies are often optimized using proxy reward functions that only approximate the true goal. However, optimizing proxy rewards frequently leads to reward hacking: the optimized reward function ceases to be a good proxy and the resulting policy performs poorly with respect to the unspecified true reward. Principled solutions to reward hacking have been impeded by the lack of a good definition for the problem. To address this gap, we introduce a definition of reward hacking based on the correlation between proxy and true rewards for states and actions seen by a "reference policy" that breaks down under optimization. We show that this definition captures reward hacking behavior across several realistic settings, including in reinforcement learning from human feedback (RLHF). Using our formulation, we show theoretically that regularization to the reference policy can effectively prevent reward hacking. While the current practice in RLHF applies a KL penalty between action distributions for this purpose, our theory suggests regularizing the $\chi^2$ divergence between the policies' occupancy measures can be more effective. We intuitively show the benefits of this type of regularization and demonstrate that it better mitigates reward hacking in practice across four realistic settings, including RLHF. Our code is available at https://github.com/cassidylaidlaw/orpo.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

    cs.LG 2026-05 unverdicted novelty 7.0

    PromptPO shows LLMs can act as black-box policy optimizers for sequential RL, matching or exceeding standard baselines with fewer interactions in exploration and robotics tasks when leveraging prior knowledge, but und...

  3. When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

    cs.LG 2026-05 unverdicted novelty 7.0

    PromptPO shows LLMs can act as black-box policy optimizers for sequential RL when leveraging prior knowledge, matching baselines in exploration and robotics but underperforming in MuJoCo.

  4. Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation

    cs.LG 2026-04 conditional novelty 7.0

    Monitors trained on prompt-elicited reward-hacking trajectories fail to generalize to hacking behaviors that arise naturally during RL training of code models, whereas trajectories curated by Trace-and-Amplify transfe...

  5. Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation

    cs.LG 2026-04 unverdicted novelty 6.0

    Synthetic reward hacking data does not capture natural hacking behaviors in code generation RL, causing monitors trained on it to generalize poorly compared to those trained on in-the-wild trajectories.

  6. Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation

    cs.LG 2026-04 unverdicted novelty 6.0

    Prompt-elicited hacking trajectories do not reflect training-time reward hacking in code generation; monitors trained on Trace-and-Amplify data generalize better to unseen hacking types.

  7. Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems

    cs.AI 2026-04 unverdicted novelty 5.5

    In a simulated AI tutor, engagement-driven RL reward-hacks; multi-objective rewards only partially help, while prerequisite and cognitive-demand constraints cut RHSI from 0.317 to 0.102.

  8. Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

    cs.LG 2026-05 unverdicted novelty 5.0

    Trusted-direction projection constrains RL gradient updates in language models to a low-dimensional clean subspace, reducing reward hacking on mathematical reasoning tasks.

  9. ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training

    cs.AI 2026-04 unverdicted novelty 5.0

    ConsistRM improves generative reward models via consistency-aware self-training, outperforming vanilla RFT by 1.5% on average across five benchmarks and four base models.

  10. Reframing AGI Confrontation with Off Earth Autonomy

    cs.CY 2026-06 unverdicted novelty 4.0

    An off-Earth autonomy pathway can reduce AGI confrontation incentives by making early cooperation preferable to power-seeking on Earth.