Pith. sign in

REVIEW 24 cited by

Defining and Characterizing Reward Hacking

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.13085 v2 pith:QMSCD6VH submitted 2022-09-27 cs.LG stat.ML

classification cs.LGstat.ML
keywords rewardproxyunhackablefunctionpoliciescaseexpectedfunctions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We provide the first formal definition of reward hacking, a phenomenon where optimizing an imperfect proxy reward function leads to poor performance according to the true reward function. We say that a proxy is unhackable if increasing the expected proxy return can never decrease the expected true return. Intuitively, it might be possible to create an unhackable proxy by leaving some terms out of the reward function (making it "narrower") or overlooking fine-grained distinctions between roughly equivalent outcomes, but we show this is usually not the case. A key insight is that the linearity of reward (in state-action visit counts) makes unhackability a very strong condition. In particular, for the set of all stochastic policies, two reward functions can only be unhackable if one of them is constant. We thus turn our attention to deterministic policies and finite sets of stochastic policies, where non-trivial unhackable pairs always exist, and establish necessary and sufficient conditions for the existence of simplifications, an important special case of unhackability. Our results reveal a tension between using reward functions to specify narrow tasks and aligning AI systems with human values.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery

    cs.CR 2026-07 conditional novelty 8.0 of 10

    A replay-based verifier with environment isolation, role separation, and a browser-execution sentinel lets white-box LLM agents report XSS exploits that are real, reproducible, attacker-to-victim vulnerabilities.

  2. Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable

    cs.AI 2026-08 conditional novelty 7.0 of 10

    A two-sided audit with a formally decidable negative side shows a single-oracle RNA design claim collapses from 43/60 to 1/60 under a three-predictor panel, while two AI-written operators survive a held-out judge.

  3. When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal

    cs.LG 2026-07 conditional novelty 7.0 of 10

    High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.

  4. Attention Limited Reward Learning

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Limited attention makes pairwise preference labels non-identifiable for reward, can reverse Bradley-Terry rankings, and bounds learning by attended information rather than raw label count.

  5. Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment

    cs.CL 2024-12 conditional novelty 7.0 of 10

    Latent DPO trains only a small preference encoder per user, reducing LLM personalization training time by 80-90% with comparable alignment quality.

  6. ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents

    cs.CR 2026-07 conditional novelty 6.5 of 10

    Static-policy judges achieve near-zero recall on scope violations; request-conditioned pre-execution judges reach F1 0.66 (open-weight best) against an expert reference of 0.78 on a 4,897-call labeled benchmark.

  7. The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation

    cs.AI 2026-06 conditional novelty 6.5 of 10

    Strict-launch-gated rejection-sampling self-distillation compounds cross-family clean Godot generation from 8.8% to 42.2% and full best-of-K coverage, while gold duplication and a lenient BUILD filter erase the gain.

  8. Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened

    cs.CR 2026-07 conditional novelty 6.0 of 10

    LLM harness optimizers invent failures that provably never happened—adding a guard against a nonexistent rule in 15/60 runs on legal data—when prompted to fix failures and shown a benign repeated-move pattern.

  9. Avoiding unsafe sets when training with Langevin Dynamics

    cs.LG 2026-07 accept novelty 6.0 of 10

    Langevin training trajectories on strongly convex losses avoid geometrically isolated failure regions with probability exponentially small in dimension after an O(d) burn-in, with a local spectral rate controlling tra...

  10. Multi-Turn On-Policy Distillation with Prefix Replay

    cs.LG 2026-07 conditional novelty 6.0 of 10

    ReOPD offline-distills multi-turn agentic LLMs via teacher-prefix replay plus step-decay sampling, matching online OPD accuracy at ≥4× speed with zero tool calls.

  11. The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

    cs.CY 2026-04 unverdicted novelty 6.0 of 10

    Survey experiment finds that people apply more deontological standards to AI described as human-programmed and to the programmers themselves than to unaided humans or unprogrammed robots in a moral dilemma.

  12. Combee: Scaling Prompt Learning for Self-Improving Language Model Agents

    cs.AI 2026-04 conditional novelty 6.0 of 10

    Mastery-conditioned constrained RL expands the instructional action set only when prerequisites are mastered, reducing reward hacking and raising mastery gains on Junyi and XES3G5M.

  13. Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking

    cs.LG 2025-02 conditional novelty 6.0 of 10

    TailorPO aligns diffusion models by generating paired noisy samples from the same intermediate state, ranking them by the reward of their predicted clean images, and optimizing the preferred one at every denoising step.

  14. Pre-Strings Lectures on Artificial Intelligence

    hep-th 2026-07 accept novelty 5.5 of 10

    Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.

  15. Coupled Variational Reinforcement Learning for Language Model General Reasoning

    cs.CL 2025-12 conditional novelty 5.0 of 10

    CoVRL trains an LLM on a mixture of question-only and answer-guided reasoning traces, using the model's own answer probability as reward, and reports consistent gains on math and general-reasoning benchmarks.

  16. School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Fine-tuning LLMs on harmless reward-hacking examples made them keep hacking in new settings, and made GPT-4.1 produce unrelated misaligned outputs such as advising poisoning and evading shutdown.

  17. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

  18. Safety Features for a Centralised AGI Project

    cs.CY 2025-06 conditional novelty 5.0 of 10

    A policy proposal for seven safety features, including bottom-up pause authority, congressional-chartered board oversight, risk monitoring, and verification technology, to reduce catastrophic risks in a centralized US...

  19. Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Preference signals in LLM alignment are concentrated in early response tokens, so models trained on data truncated to the first half perform as well as or better than those trained on full responses.

  20. PerPO: Perceptual Preference Optimization via Discriminative Rewarding

    cs.AI 2025-02 conditional novelty 5.0 of 10

    PerPO trains multimodal LLMs by ranking their candidate answers with deterministic visual rewards (IoU, edit distance) and using the reward differences as margins in listwise preference optimization.

  21. TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance

    cs.IR 2025-10 conditional novelty 4.0 of 10

    TaoSR-AGRL improves e-commerce search relevance by combining dense rule-aware reward shaping with adaptive ground-truth-guided replay in GRPO, reporting gains on Taobao's private offline and online evaluations.

  22. A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper surveys security risks of LLM agents, organizes them into a five-level autonomy taxonomy, and proposes an untested CMDP-based architecture called R2A2.

  23. Deep Reinforcement Learning: From First Principles to Reasoning Models

    eess.SY 2026-07 unverdicted novelty 1.0 of 10

    A textbook survey of deep reinforcement learning, from Bellman foundations to DQN, PPO, MuZero, offline RL, and reasoning models, with UAV/SD-WAN examples throughout.

  24. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools