Pith. sign in

Paper Citation Record · LEDGER

Defining and Characterizing Reward Hacking

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2209.13085.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2209.13085 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:26.846278Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aa536ffd-66f9-4070-9524-9b6dd090c5e4 · inbound

Scaling Laws for Reward Model Overoptimization cites this paper.

Scaling Laws for Reward Model Overoptimization Defining and Characterizing Reward Hacking

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:04:53.379647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T09:04:53.129737Z digest=sha256:3d9909e11a79cd0ea8a44297474a67a94200e3aa3bdb172d1e5d35eea058da1e

Observation 5773e748-c95b-40e6-93a4-cf831ff83936 · inbound

Active teacher selection for reward learning cites this paper.

Active teacher selection for reward learning Defining and Characterizing Reward Hacking

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T05:56:01.756834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T05:54:43.873174Z digest=sha256:0d620ec9fcb83e0ef67678504985cd78dfcff609c7ae911255aedf5ab1ec7d64

Observation 611ae39f-2683-4c20-a934-530fa9aab837 · inbound

Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment cites this paper.

Disentangling Preference Representation and Text Generation for Efficient Individual Preference Alignment Defining and Characterizing Reward Hacking

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:26.846278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:26.846278Z digest=sha256:895dd434e8c1f9e516db66dd61e3a9db2fb8d2af785736ac6597c73e5601ee59

Observation 840a7906-fad3-44d9-ad91-31b0138088e7 · inbound

Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking cites this paper.

Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking Defining and Characterizing Reward Hacking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T18:55:14.733622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:55:14.733622Z digest=sha256:6820a17c78a3e16f8b101fd7d839f8dcd4168460f8b4e999a011e46bf9c2bbca

Observation 340454bf-41ff-4566-acd5-cf0b8992b76b · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding Defining and Characterizing Reward Hacking

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.375740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.375740Z digest=sha256:0adc40037bd55ffe0a411d442c7d71b8f1f30212cd9b0897d8dbefa81dd0fd83

Observation 03ae14d6-fb94-4686-9f1a-5dde4f332355 · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Defining and Characterizing Reward Hacking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.420808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.420808Z digest=sha256:d11fc53180930fd610e110dd54876fa271e8e981ce7ba2a3a73f7d229600a7cf

Observation 07c83c84-8bf8-4f04-a97d-837fae7723ea · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Defining and Characterizing Reward Hacking

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:45.116682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:45.116682Z digest=sha256:2566104449c3d4c30715f67bf9a8fe43a3f8ed14e78293d0b4a66091cf099fde

Observation 99a48b61-255b-484e-904d-5a0328c85069 · inbound

Residual Reward Models for Preference-based Reinforcement Learning cites this paper.

Residual Reward Models for Preference-based Reinforcement Learning Defining and Characterizing Reward Hacking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:35.995519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:35.995519Z digest=sha256:cea4093cee1f7dbd8be62d9341da1104fbdbe08ae6863041acec1db046c6c8f0

Observation 4da1e97b-90e6-4823-a301-886c258826fa · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Defining and Characterizing Reward Hacking

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.057759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.057759Z digest=sha256:141b8cfcc8b8488328fc2ff69ed4c7125bda3617e086c8c7004391fd1a30a907

Observation 7d433f92-f7f2-4d7a-ad5b-1f647bc356da · inbound

Safety Features for a Centralised AGI Project cites this paper.

Safety Features for a Centralised AGI Project Defining and Characterizing Reward Hacking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:51.691004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:51.691004Z digest=sha256:799b58358feecdcda62ec3c3a261a5891351a7c1369706e7508d2ba44e1e1315

Observation 6e307ebc-94ba-4061-acc0-74ecd3b48a5a · inbound

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs cites this paper.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Defining and Characterizing Reward Hacking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.545711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.545711Z digest=sha256:0d11efc6d21a1e89d92c21d3d8b721423ac405dea0f6633f2ae397fab6eaa52a

Observation 757c9522-2a47-4aac-adf5-43d0d48bbf2e · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF Defining and Characterizing Reward Hacking

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.850063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:dbb9df92fa2f543a28203e9765cf9a19edb5b79834d479770802ef8c91ca3952

Observation 73dfea2d-def0-4b49-b3ad-a517d841eae6 · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Defining and Characterizing Reward Hacking

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.398856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.398856Z digest=sha256:2aa740b243cfe4d2e2d665fdc1741095df4c55bca6f734ef51534356da6fe1cc

Observation 571222e9-820a-477a-ac3c-11adea58878b · inbound

Coupled Variational Reinforcement Learning for Language Model General Reasoning cites this paper.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Defining and Characterizing Reward Hacking

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:56.945344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:56.945344Z digest=sha256:c6a46e0e9747aa57e429524426f16778fc371a49e3848ce2979da81b53bdd55a

Observation 49f77dec-b6b2-4d47-8be2-d575d9e0c5d3 · inbound

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents cites this paper.

Combee: Scaling Prompt Learning for Self-Improving Language Model Agents Defining and Characterizing Reward Hacking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T10:33:19.825936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:33:19.825936Z digest=sha256:423f9277f793808d3078cb836216e76a517ea0ca279475ac30e94574b9ffcfe7

Observation e96c1b9e-3e77-4d2b-8b2a-0447ab893578 · inbound

DUET: Joint Exploration of User Item Profiles in Recommendation System cites this paper.

DUET: Joint Exploration of User Item Profiles in Recommendation System Defining and Characterizing Reward Hacking

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:40:23.616293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T12:40:04.613964Z digest=sha256:7da3abbce7927d745cba85861c7dc1879db52ec34e95c517f4bf22bc91874297

Observation 5e56343a-906c-4e92-b33a-e70abd30ca5d · inbound

LLMs Corrupt Your Documents When You Delegate cites this paper.

LLMs Corrupt Your Documents When You Delegate Defining and Characterizing Reward Hacking

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:47.599512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T09:47:21.966292Z digest=sha256:a077d1d10eeef86139969bea11252e22bf80338eba418673bfa853e92d4530fe

Observation 0cc6721b-4000-4f66-b163-5d6aee3ba069 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Defining and Characterizing Reward Hacking

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:09.369103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:2a61833a11c215c30c202e25c896ea5b498ca8c2dc18d7abc9192645d4832571

Observation 4a8a8d5b-ad56-412a-a141-0550e525169c · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:17.631038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T17:58:26.932519Z digest=sha256:4c254dd1a0242c5a83eb6005d9e9f51f98177026a255d811f3d4cc9a30e933e3

Observation 634c0313-ca5c-4d71-8873-52e250e02ec7 · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:28.781994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:20:52.825158Z digest=sha256:49e77cebff741f227389ecd103058ec09de4097f24ae41c79225c34ec2d11810

Observation ffac88c2-d468-4c23-9b97-1b7a830e0669 · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:25:39.876048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T09:24:17.904893Z digest=sha256:4ec8f00038cd6ef0f591c044f5bbcb6d00cf3ab1f7e08adf5c666a3fabb10d7d

Observation 584c3de6-7b35-4d39-8665-045611652215 · inbound

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers cites this paper.

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers Defining and Characterizing Reward Hacking

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T15:29:49.171054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:29:49.171054Z digest=sha256:3b4cfd3568c268ca6c94e4e3ec208685c9e687c3da3985a5ecdbe3f4fc543134

Observation a759e4eb-cc9c-4a0b-ad50-f494eeffb4b2 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Defining and Characterizing Reward Hacking

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:16.302983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:334509bbefdb3bb4535e9df62a690076d8a9c2a1ae193fd436dfc5ebe105e39d

Observation bfb8015f-5e33-47bc-999f-1209b38f7bdf · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR Defining and Characterizing Reward Hacking

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:00:45.572142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:f89c9b4ab6f4114816aedc7912ba044476c50ffdd78938f3ff5077b83417ea25

Observation ad862279-0362-4a2d-aadc-df70b02a5a58 · inbound

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models cites this paper.

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models Defining and Characterizing Reward Hacking

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:18.899889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T05:19:38.587352Z digest=sha256:5b30d85c214c3d7e24c4e783cf454aa13242caa9777c7dda105ebf0412f7223b

Observation d4512d50-caf6-408d-91fd-53055ffd103e · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Defining and Characterizing Reward Hacking

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.842198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4043e8b7abcef9c8e76b26c946298ddfa2f3b4049f593d8c31190ee07e89318f

Observation 1e42082f-4576-4764-aad5-8d54c9cc8168 · inbound

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA cites this paper.

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA Defining and Characterizing Reward Hacking

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:30:22.994671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:27:50.186418Z digest=sha256:13e3692988b57b1173b344dabd86ce18f9284f6029722d68da158e916e14363e

Observation 22c6a36d-4935-4233-aa03-943ec5d1e208 · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Defining and Characterizing Reward Hacking

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.792263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:cbafa327436b014fba107151cddd9e73c67c0994a7aed63a0123c1fc2367cfd4

Observation ea14dfd0-dce7-47d9-b55d-7497ef599ad8 · inbound

Towards Automated Discovery: A Review of Generative Models, Multimodal Learning and Closed-Loop Workflows in Inverse Materials Design cites this paper.

Towards Automated Discovery: A Review of Generative Models, Multimodal Learning and Closed-Loop Workflows in Inverse Materials Design Defining and Characterizing Reward Hacking

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:16:24.564670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T13:30:35.790162Z digest=sha256:5b6bdce530c239d0b6f8c7b22130cf03b85a323d5a6a3c32acdf31c1de4df4f4

Observation 4032f613-01af-4600-a489-741470ed78d7 · inbound

Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation cites this paper.

Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation Defining and Characterizing Reward Hacking

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:46:33.041705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T09:40:50.399436Z digest=sha256:7ef44d51f504fe6bfc231ec4f766665ff1dabf0f40daf4833ae67c2c722792b1

Observation 09bea10a-78fc-46fe-b2c8-d85d818fe735 · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Defining and Characterizing Reward Hacking

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T17:31:06.952098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:5a336ac6213e80e17ca60645c0bb634650f8ce12ec8c6276aae8032c203fcde5

Observation 32514426-7f9c-451c-886b-4a473e9a9a65 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Defining and Characterizing Reward Hacking

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.471686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:7ac3bc0326af2453f9256124a36c4b7027ee36efc65b7c36be9668aaeec3f925

Observation 8679eeef-e7f2-41c8-9ced-4411ecc5818d · inbound

Evolving Quantum Error-Correcting Encodings for Molecular Simulation cites this paper.

Evolving Quantum Error-Correcting Encodings for Molecular Simulation Defining and Characterizing Reward Hacking

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.551040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T20:51:06.703483Z digest=sha256:9bd7038e937fba0610bb7a6d31c95d9fa753030b21bec7f56ec70482e21d78eb

Observation c22fbc81-18ee-46dc-9ef2-c771530dd192 · inbound

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience cites this paper.

Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience Defining and Characterizing Reward Hacking

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:58.794546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T01:57:04.293058Z digest=sha256:d530c2b7805fbd8b55953a77f21ffcae8d223261f5ddd2ceadf9cb19d33ed185

Observation 7821cac5-bdba-4754-a8e0-52b4daf1b66e · inbound

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models cites this paper.

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models Defining and Characterizing Reward Hacking

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.999661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T06:46:02.255137Z digest=sha256:9925113a2a78eb0de86080b9d8c6cc8dd292355fad06163f0a204af54a98efac

Observation cf637c26-424e-4c29-9ebe-5b865f1c0917 · inbound

Pre-Strings Lectures on Artificial Intelligence cites this paper.

Pre-Strings Lectures on Artificial Intelligence Defining and Characterizing Reward Hacking

Reference 136

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:03.658427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:03.658427Z digest=sha256:15c546e15e2815a303c891f9f7627087cf709365f951cc72e8a4c645ff0b9f95

Observation 7a244639-b617-4414-8300-5f9c1dbf8855 · inbound

Attention Limited Reward Learning cites this paper.

Attention Limited Reward Learning Defining and Characterizing Reward Hacking

Reference 32

Resolution
malformed identifier
no resolver link, observed 2026-07-11T16:47:52.768236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:47:52.768236Z digest=sha256:71bcc405a8512aff08573ac7ebdd8fb1131823635ecff8bb9210d9819ab8112d

Observation da07dd8e-be58-4b30-a9e1-f4611eaa4a3d · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Defining and Characterizing Reward Hacking

Reference 221

Resolution
malformed identifier
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:002a6dfb7aca811ac5f7b15537f3c9f5cb4f7d2f88969c7f1a85c30605486275

Observation d0538c0d-2d47-497a-a593-2441ea2cb9b6 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Defining and Characterizing Reward Hacking

Reference 222

Resolution
malformed identifier
no resolver link, observed 2026-08-02T08:40:57.669685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:57.669685Z digest=sha256:83b3acd4b0873791fe60b78f896d37c8e81645d6d53773aa373923a33c4ff9bf

Observation 41548256-8af7-4baf-b7cb-c15de7e1266b · inbound

Avoiding unsafe sets when training with Langevin Dynamics cites this paper.

Avoiding unsafe sets when training with Langevin Dynamics Defining and Characterizing Reward Hacking

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T07:56:04.662971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-09T07:53:56.220737Z digest=sha256:15c4792bf9655a78e11276530aec90b6b320912882c4e39fa6f63419df0201b0

Observation 120dba60-03b8-4c6a-8184-d28a22bf8dd8 · inbound

Avoiding unsafe sets when training with Langevin Dynamics cites this paper.

Avoiding unsafe sets when training with Langevin Dynamics Defining and Characterizing Reward Hacking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T08:08:39.765022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:08:39.765022Z digest=sha256:0662006c984c61fd08a656b69ba1c6535fb3852d541f5d917ce2cdca545d431d

Observation f0b46b48-3671-4bf2-905d-ef392da95bb5 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Defining and Characterizing Reward Hacking

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:37:31.185952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T18:29:50.731038Z digest=sha256:ebf0cb996f0cf8ca8f7be8e5a386e3b0e98787c14705765775d3ad2791935492

Observation dcc8c37e-1333-47c1-9749-544b090ef3f7 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents Defining and Characterizing Reward Hacking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:48.269583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:49:48.269583Z digest=sha256:7e2f7fe35d2984480d21ba77d672e2553afb30f75caedd2cb11cc67c8fa3d92d

Observation 4c81fb80-f5b3-4e69-b005-b8eed94bb5e3 · inbound

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation cites this paper.

The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation Defining and Characterizing Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T17:25:37.713857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:25:37.713857Z digest=sha256:8b7052473b635e527dc07cac154039df8eaafd91d61fcce4d224587fbebc9707

Observation 417eb3f8-4ef7-4e8b-b103-e4dd9e84ed60 · inbound

When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal cites this paper.

When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal Defining and Characterizing Reward Hacking

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:42.144350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:42.144350Z digest=sha256:757246cd3791bfa438a1e961f410466cf8d2d04506b094c142c301e0758b1fe3

Observation f950cb56-c969-4987-acb1-fb549a5108ee · inbound

Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened cites this paper.

Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened Defining and Characterizing Reward Hacking

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T07:01:21.816420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:01:21.816420Z digest=sha256:c3010bed0bb01031dce42cf3bd7418c83b171a6070f3c220cccaea582aaff549

Observation 4ed24b25-3988-489d-90cd-1d709bb3e75f · inbound

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery cites this paper.

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Defining and Characterizing Reward Hacking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:06.730241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:06.730241Z digest=sha256:6fbea259fa5cea367bd76868f5ccd243a7d93a3568a5b03ef83f01b5823db632

Observation 81b3cc78-0585-4d7a-b578-47d9d18788a3 · inbound

Deep Reinforcement Learning: From First Principles to Reasoning Models cites this paper.

Deep Reinforcement Learning: From First Principles to Reasoning Models Defining and Characterizing Reward Hacking

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:04.347683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:04.347683Z digest=sha256:a192e4c49c6d2c13d2de04efd5597148ac20f4111ebed58d0eadd0fa758a7275

Observation a5e3e8d5-7b9a-4294-befa-b763b83cd10f · inbound

Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable cites this paper.

Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable Defining and Characterizing Reward Hacking

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:17.081550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:17.081550Z digest=sha256:b57b18bf5c0b87977b7cd1eece346e0f28acc2fe803daf025aa4d6b36ce557aa