Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2209.13085.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:26.846278Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
18
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation aa536ffd-66f9-4070-9524-9b6dd090c5e4 · inbound
Scaling Laws for Reward Model Overoptimization Defining and Characterizing Reward Hacking
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5773e748-c95b-40e6-93a4-cf831ff83936 · inbound
Active teacher selection for reward learning Defining and Characterizing Reward Hacking
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 611ae39f-2683-4c20-a934-530fa9aab837 · inbound
Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment Defining and Characterizing Reward Hacking
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 840a7906-fad3-44d9-ad91-31b0138088e7 · inbound
Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking Defining and Characterizing Reward Hacking
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 340454bf-41ff-4566-acd5-cf0b8992b76b · inbound
PerPO: Perceptual Preference Optimization via Discriminative Rewarding Defining and Characterizing Reward Hacking
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ae14d6-fb94-4686-9f1a-5dde4f332355 · inbound
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Defining and Characterizing Reward Hacking
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c83c84-8bf8-4f04-a97d-837fae7723ea · inbound
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Defining and Characterizing Reward Hacking
Reference 129
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99a48b61-255b-484e-904d-5a0328c85069 · inbound
Residual Reward Models for Preference-based Reinforcement Learning Defining and Characterizing Reward Hacking
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da1e97b-90e6-4823-a301-886c258826fa · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Defining and Characterizing Reward Hacking
Reference 207
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d433f92-f7f2-4d7a-ad5b-1f647bc356da · inbound
Safety Features for a Centralised AGI Project Defining and Characterizing Reward Hacking
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e307ebc-94ba-4061-acc0-74ecd3b48a5a · inbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Defining and Characterizing Reward Hacking
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757c9522-2a47-4aac-adf5-43d0d48bbf2e · inbound
Failure Modes of Maximum Entropy RLHF Defining and Characterizing Reward Hacking
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73dfea2d-def0-4b49-b3ad-a517d841eae6 · inbound
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Defining and Characterizing Reward Hacking
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 571222e9-820a-477a-ac3c-11adea58878b · inbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning Defining and Characterizing Reward Hacking
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f77dec-b6b2-4d47-8be2-d575d9e0c5d3 · inbound
Combee: Scaling Prompt Learning for Self-Improving Language Model Agents Defining and Characterizing Reward Hacking
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96c1b9e-3e77-4d2b-8b2a-0447ab893578 · inbound
DUET: Joint Exploration of User Item Profiles in Recommendation System Defining and Characterizing Reward Hacking
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5e56343a-906c-4e92-b33a-e70abd30ca5d · inbound
LLMs Corrupt Your Documents When You Delegate Defining and Characterizing Reward Hacking
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0cc6721b-4000-4f66-b163-5d6aee3ba069 · inbound
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Defining and Characterizing Reward Hacking
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4a8a8d5b-ad56-412a-a141-0550e525169c · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 634c0313-ca5c-4d71-8873-52e250e02ec7 · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ffac88c2-d468-4c23-9b97-1b7a830e0669 · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 584c3de6-7b35-4d39-8665-045611652215 · inbound
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a759e4eb-cc9c-4a0b-ad50-f494eeffb4b2 · inbound
Risk Reporting for Developers' Internal AI Model Use Defining and Characterizing Reward Hacking
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bfb8015f-5e33-47bc-999f-1209b38f7bdf · inbound
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR Defining and Characterizing Reward Hacking
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ad862279-0362-4a2d-aadc-df70b02a5a58 · inbound
EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models Defining and Characterizing Reward Hacking
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4512d50-caf6-408d-91fd-53055ffd103e · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Defining and Characterizing Reward Hacking
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e42082f-4576-4764-aad5-8d54c9cc8168 · inbound
What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA Defining and Characterizing Reward Hacking
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 22c6a36d-4935-4233-aa03-943ec5d1e208 · inbound
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Defining and Characterizing Reward Hacking
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ea14dfd0-dce7-47d9-b55d-7497ef599ad8 · inbound
Towards Automated Discovery: A Review of Generative Models, Multimodal Learning and Closed-Loop Workflows in Inverse Materials Design Defining and Characterizing Reward Hacking
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4032f613-01af-4600-a489-741470ed78d7 · inbound
Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation Defining and Characterizing Reward Hacking
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 09bea10a-78fc-46fe-b2c8-d85d818fe735 · inbound
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Defining and Characterizing Reward Hacking
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32514426-7f9c-451c-886b-4a473e9a9a65 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Defining and Characterizing Reward Hacking
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8679eeef-e7f2-41c8-9ced-4411ecc5818d · inbound
Evolving Quantum Error-Correcting Encodings for Molecular Simulation Defining and Characterizing Reward Hacking
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c22fbc81-18ee-46dc-9ef2-c771530dd192 · inbound
Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Defining and Characterizing Reward Hacking
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7821cac5-bdba-4754-a8e0-52b4daf1b66e · inbound
Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models Defining and Characterizing Reward Hacking
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf637c26-424e-4c29-9ebe-5b865f1c0917 · inbound
Pre-Strings Lectures on Artificial Intelligence Defining and Characterizing Reward Hacking
Reference 136
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a244639-b617-4414-8300-5f9c1dbf8855 · inbound
Attention Limited Reward Learning Defining and Characterizing Reward Hacking
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da07dd8e-be58-4b30-a9e1-f4611eaa4a3d · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Defining and Characterizing Reward Hacking
Reference 221
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0538c0d-2d47-497a-a593-2441ea2cb9b6 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Defining and Characterizing Reward Hacking
Reference 222
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41548256-8af7-4baf-b7cb-c15de7e1266b · inbound
Avoiding unsafe sets when training with Langevin Dynamics Defining and Characterizing Reward Hacking
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 120dba60-03b8-4c6a-8184-d28a22bf8dd8 · inbound
Avoiding unsafe sets when training with Langevin Dynamics Defining and Characterizing Reward Hacking
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b46b48-3671-4bf2-905d-ef392da95bb5 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Defining and Characterizing Reward Hacking
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dcc8c37e-1333-47c1-9749-544b090ef3f7 · inbound
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Defining and Characterizing Reward Hacking
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c81fb80-f5b3-4e69-b005-b8eed94bb5e3 · inbound
The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation Defining and Characterizing Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417eb3f8-4ef7-4e8b-b103-e4dd9e84ed60 · inbound
When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal Defining and Characterizing Reward Hacking
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f950cb56-c969-4987-acb1-fb549a5108ee · inbound
Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened Defining and Characterizing Reward Hacking
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ed24b25-3988-489d-90cd-1d709bb3e75f · inbound
RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Defining and Characterizing Reward Hacking
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b3cc78-0585-4d7a-b578-47d9d18788a3 · inbound
Deep Reinforcement Learning: From First Principles to Reasoning Models Defining and Characterizing Reward Hacking
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e3e8d5-7b9a-4294-befa-b763b83cd10f · inbound
Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable Defining and Characterizing Reward Hacking
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.