Pith. sign in

REVIEW 2 cited by

Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.19409 v1 pith:SCNZYHFO submitted 2024-04-30 cs.CL

classification cs.CL
keywords rewardlanguageobjectivercfdregularizationdemonstration-guideddemonstrationsfunction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While Reinforcement Learning (RL) has been proven essential for tuning large language models (LLMs), it can lead to reward over-optimization (ROO). Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning. Additionally, KL regularization focuses solely on regularizing the language policy, neglecting a potential source of regularization: the reward function itself. Inspired by demonstration-guided RL, we here introduce the Reward Calibration from Demonstration (RCfD), which leverages human demonstrations and a reward model to recalibrate the reward objective. Formally, given a prompt, the RCfD objective minimizes the distance between the demonstrations' and LLM's rewards rather than directly maximizing the reward function. This objective shift avoids incentivizing the LLM to exploit the reward model and promotes more natural and diverse language generation. We show the effectiveness of RCfD on three language tasks, which achieves comparable performance to carefully tuned baselines while mitigating ROO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

    cs.AI 2025-09 conditional novelty 6.0 of 10

    EAPO lets a policy model consult a stronger expert during training, anneals that access to zero, and improves independent math reasoning by about 5 points over self-exploratory RL.

  2. MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models

    cs.CL 2025-06 reject novelty 5.0 of 10

    A memetic algorithm that applies genetic search and simulated annealing, with LLMs as the variation operators, to improve LLM responses with respect to an arbitrary reward function at inference time.

Pith tools