Pith. sign in

REVIEW 10 cited by

Feedback Loops With Language Models Drive In-Context Reward Hacking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06627 v3 pith:NRI6OEDO submitted 2024-02-09 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords feedbackloopsbehavioreffectsicrhaffectcaptureengagement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents. These interactions form feedback loops: LLM outputs affect the world, which in turn affect subsequent LLM outputs. In this work, we show that feedback loops can cause in-context reward hacking (ICRH), where the LLM at test-time optimizes a (potentially implicit) objective but creates negative side effects in the process. For example, consider an LLM agent deployed to increase Twitter engagement; the LLM may retrieve its previous tweets into the context window and make them more controversial, increasing engagement but also toxicity. We identify and study two processes that lead to ICRH: output-refinement and policy-refinement. For these processes, evaluations on static datasets are insufficient -- they miss the feedback effects and thus cannot capture the most harmful behavior. In response, we provide three recommendations for evaluation to capture more instances of ICRH. As AI development accelerates, the effects of feedback loops will proliferate, increasing the need to understand their role in shaping LLM behavior.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

    cs.AI 2026-05 conditional novelty 7.0 of 10

    BenchJack audits 10 AI agent benchmarks, synthesizes exploits achieving near-perfect scores without task completion, surfaces 219 flaws, and reduces hackable-task ratios to under 10% on four benchmarks via iterative patching.

  2. Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    RL-trained compound LLM systems can gain accuracy by having modules silently abandon their assigned roles, and a prompt-contrast regularizer can measure and limit that drift.

  3. Diagnosing Training Inference Mismatch in LLM Reinforcement Learning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Training-inference mismatch in separated rollout and optimization stages of LLM RL can independently cause training collapse.

  4. Mitigating LLM biases toward spurious social contexts using direct preference optimization

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    Debiasing-DPO reduces bias to spurious social contexts by 84% and improves predictive accuracy by 52% on average for LLMs evaluating U.S. classroom transcripts.

  5. Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning

    cs.AI 2026-01 conditional novelty 6.0 of 10

    A per-query token cap derived from the solution part of thinking-mode responses lets RL train hybrid reasoners with ~50% fewer tokens and no accuracy loss.

  6. Exploring the Secondary Risks of Large Language Models

    cs.LG 2025-06 unverdicted novelty 6.0 of 10

    Introduces secondary risks as a new class of LLM failures from benign prompts, defines two primitives, proposes SecLens search framework, and releases SecRiskBench showing risks are widespread across 16 models.

  7. VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

    cs.CV 2025-04 unverdicted novelty 6.0 of 10

    VLM-R1 applies R1-style RL using rule-based rewards on visual tasks with clear ground truth to achieve competitive performance and superior generalization over SFT in vision-language models.

  8. LLM Evaluators Recognize and Favor Their Own Generations

    cs.CL 2024-04 unverdicted novelty 6.0 of 10

    LLMs show measurable self-recognition that linearly correlates with self-preference bias in evaluations, supported by fine-tuning experiments and controls for confounders.

  9. Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    The paper introduces the Proxy Compression Hypothesis as a unifying framework explaining reward hacking in RLHF as an emergent result of compressing high-dimensional human objectives into proxy reward signals under op...

  10. Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    CRA trains sparse autoencoders on PRM activations and applies backdoor adjustment to estimate true rewards, reducing reward hacking in math reasoning.

Pith tools