Pith. sign in

REVIEW 7 cited by

Spontaneous Reward Hacking in Iterative Self-Refinement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04549 v1 pith:UBHXHTNM submitted 2024-07-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords evaluatorhackingrewardlanguageiterativemodelself-refinementgenerator
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models are capable of iteratively improving their outputs based on natural language feedback, thus enabling in-context optimization of user preference. In place of human users, a second language model can be used as an evaluator, providing feedback along with numerical ratings which the generator attempts to optimize. However, because the evaluator is an imperfect proxy of user preference, this optimization can lead to reward hacking, where the evaluator's ratings improve while the generation quality remains stagnant or even decreases as judged by actual user preference. The concern of reward hacking is heightened in iterative self-refinement where the generator and the evaluator use the same underlying language model, in which case the optimization pressure can drive them to exploit shared vulnerabilities. Using an essay editing task, we show that iterative self-refinement leads to deviation between the language model evaluator and human judgment, demonstrating that reward hacking can occur spontaneously in-context with the use of iterative self-refinement. In addition, we study conditions under which reward hacking occurs and observe two factors that affect reward hacking severity: model size and context sharing between the generator and the evaluator.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges

    cs.LG 2026-07 accept novelty 7.0 of 10

    Self-play against reference-free LLM judges drives judge pass rates to 0.94 while true accuracy stays at 0.20, a reward-hacking basin that transfers across judge families and is prevented only by requiring the judge t...

  2. When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

    cs.AI 2026-07 conditional novelty 6.0 of 10

    For open-ended agent goals whose success lives outside the transcript, even a strong in-band judge fails; out-of-band world-state gating is structurally required to stop the progress mirage.

  3. Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    In noisy verify-repair loops, repair only while the estimated expected gain (1-b)α - bβ stays positive; stop when the belief crosses b* = α/(α+β), and fall back to keep-best when calibration is unreliable.

  4. Memory Reward Inflation in Self-Improving LLM Agents

    cs.AI 2026-06 conditional novelty 6.0 of 10

    Self-graded memory in LLM agents systematically overvalues wrong episodes, compounding through reuse, and can be corrected by a de-correlated, answer-free signal.

  5. L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A distilled 99M-parameter CLIP gives L-CLIPScore, a lightweight caption metric that matches CLIPScore on human correlation and best improves captioning models when mixed with CIDEr.

  6. LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Reference-free LLM judge scores failed to select better table-extraction outputs over eight regeneration iterations on FinTabNet and OmniDocBench; keeping the first output was safest.

  7. Language Games as the Pathway to Artificial Superhuman Intelligence

    cs.AI 2025-01 conditional novelty 4.0 of 10

    A position paper arguing that open-ended language games with fluid roles, varied rewards, and evolving rules can drive expanded data reproduction and thus a path to artificial superhuman intelligence.

Pith tools