Pith. sign in

Paper Citation Record · LEDGER

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 24 inbound Pith citation observations for arXiv:2508.17511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17511 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:57:10.334753Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T16:13:13.602057Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7b365787-74ad-4a0e-a2ff-bf766966df4f · outbound

This paper cites write newline.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.274752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.274752Z digest=sha256:1e8c8b8fb988ed725bd8721b403a2e211e62c5fee5f81e4f4c819cffa636eacd

Observation 1361b49c-213a-448a-89a1-b75ab9d6cb4d · outbound

This paper cites Program Synthesis with Large Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.350528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.350528Z digest=sha256:da05c700368e29a32f532b9f75cf438c17a21d7c391786c429fb9feca4479dec

Observation 06bab222-0455-494b-9d93-a2aa1234a229 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.434931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.434931Z digest=sha256:c9282a8963720acd1698865dcff5a2bf525a624ba058d20cf37670e857e939a2

Observation e33121b4-fdb1-4485-a060-897e7dc12f0c · outbound

This paper cites Tell me about yourself: LLMs are aware of their learned behaviors.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Tell me about yourself: LLMs are aware of their learned behaviors

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.514379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.514379Z digest=sha256:1f74863ac0e41046ab8bd34fc3d67e12861877dbc786be16030683a37e283b95

Observation f28fba28-da30-43a9-861a-2120459b5764 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025 b.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025 b

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.604117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.604117Z digest=sha256:9272236cb358b3aaa581be0b18c32df49eb946d68e454cdd1dd080cc4646df12

Observation 6aa00b91-28f3-469f-ae21-03a5cdd1d5fb · outbound

This paper cites Demonstrating specification gaming in reasoning models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Demonstrating specification gaming in reasoning models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.678736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.678736Z digest=sha256:db796c05a57471a70fcdc761504693584f872230ebc029bc34d5a38960f1c493

Observation 98d8e0f3-96d2-4ea4-99e9-64d3198500ee · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.784533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.784533Z digest=sha256:c3beffefe3398f503d12dac6311ff6b5258af6afc9682b06cdb6399ff860f5b1

Observation d127ecec-5b7a-47ac-9555-212bc4cadb32 · outbound

This paper cites Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:07.932075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:07.932075Z digest=sha256:f5e4ede1f527ac049214921377856f00c12fddb8389838e1dcbcd376e95e14d6

Observation 890c0684-6695-4f36-94b4-7a6787d561d7 · outbound

This paper cites Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.044929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.044929Z digest=sha256:079ee7abf1cc281fac54adb1b42c43b6c4094e13d77029a9dd289cb5f86ad9f4

Observation 78da8815-e0ee-4ce5-bc17-5d4387958ff1 · outbound

This paper cites Subliminal Learning: Language models transmit behavioral traits via hidden signals in data.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Subliminal Learning: Language models transmit behavioral traits via hidden signals in data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.146810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.146810Z digest=sha256:afa4a44ace8d18d4bc38b450e85451b0e0bea0eb59150c187eb83b20c5e93bce

Observation 7dfb1ece-62a5-4ac4-8c64-64ee1d7fcfa3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.253618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.253618Z digest=sha256:d59f5d25e74bfbbad1450650f772f31b876baf3c6f3485187210e46c7f792676

Observation 70f59ac2-da0b-4613-aa8f-bdf632c4fa36 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.297166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.297166Z digest=sha256:dea7ff4abe2894490e88da34f98bf7e2f2a2b314ff74535e9739dfe4f4968574

Observation 5956c521-90a0-4cf8-b52d-4688324df1bb · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.311545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.311545Z digest=sha256:130d23c15a005008ad8fe4ef8d299f9e481f388481a7dc94fc3b0af6c28bc6a1

Observation 784a3ef4-a4c9-4a12-a593-bae8ce0e12b1 · outbound

This paper cites Unsloth, 2023.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unsloth, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.416810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.394755Z digest=sha256:4f0c034593bcfbbe2b587555fbc32efcfd7ec0c865f055f3ede777da3ca424d5

Observation aa0012d9-36fa-4c54-acfb-7ad50297b706 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.534766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.534766Z digest=sha256:034a0f93dafde2fbec9e65e56b954a32d109aa1e00b6aa9f0699901d024c5f31

Observation eda3174a-4a4c-4a1b-9f3b-5ca34c117027 · outbound

This paper cites Training on documents about reward hacking induces reward hacking, 2024.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Training on documents about reward hacking induces reward hacking, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.371924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.663990Z digest=sha256:a28b0ae2f6a928a331bf0067b167e3d238d9a4c4f1f2f6cd58c43308762528af

Observation 621dfae9-4717-4e66-b334-eec7140e44da · outbound

This paper cites Model organisms of misalignment: The case for a new pillar of alignment research, 2023.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Model organisms of misalignment: The case for a new pillar of alignment research, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.353791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.765494Z digest=sha256:f58420da0bc9d39375ca7c636c5785e7ea08289be8b73989336946b5f680db4a

Observation 70e3acff-d50d-4d41-8bde-f83b37011a63 · outbound

This paper cites Auditing language models for hidden objectives.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Auditing language models for hidden objectives

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.878310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.878310Z digest=sha256:464ee435ef27bb6041091384f9ebc2a5a59ceb4a42605f91fb2e82ccdea4c4dd

Observation 863f3f84-5cf5-4e51-a133-e44a69dac6f4 · outbound

This paper cites Recent frontier models are reward hacking.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Recent frontier models are reward hacking

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.337301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T16:57:08.917897Z digest=sha256:0ccdca58399aea028cff3bd3117fff822e7aa15e252af3d1b3718b5da32ee37f

Observation 8d0e2e04-7f0a-4ff2-855d-d82743908a90 · outbound

This paper cites Reward hacking behavior can generalize across tasks.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Reward hacking behavior can generalize across tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.072971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.072971Z digest=sha256:a0d66754e1740bbf5b06364b5a8c35d5432b32f30aedf5d2dd86fb9d3abfb919

Observation 5d75d60e-1a81-4e36-8a1f-44b29c5895e0 · outbound

This paper cites Toward understanding and preventing misalignment generalization, 2025.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Toward understanding and preventing misalignment generalization, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.296022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T16:57:09.228351Z digest=sha256:3706c4559a7f3fe11d0886ae0366900e57f8fabb40b59498f3427265689e883d

Observation 7e2ccd9f-447a-4132-9eec-4fb7a6863344 · outbound

This paper cites Sycophancy in GPT-4o : what happened and what we're doing about it.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Sycophancy in GPT-4o : what happened and what we're doing about it

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:57:11.260052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T16:57:09.317538Z digest=sha256:8e0a8a5bc3315c8ead41eaf929e8debbb9d2408b32bf60a21cea077364827daa

Observation a6a0e824-aa5d-4f24-a3af-6aa1b6678cc2 · outbound

This paper cites Generalizing verifiable instruction following, 2025.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Generalizing verifiable instruction following, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.412774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.412774Z digest=sha256:1b8a28e5458c02039fdf83d8592161eed9cd6f59675d6dc361888c30fe68b342

Observation 2a5b2435-3baf-4cb5-982e-b69d4fdaeaf2 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Towards Understanding Sycophancy in Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.417680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.417680Z digest=sha256:885a3d469c734b717ff3f4d3c73e71ba88462229a0ee36f36ca7c6b23a46c055

Observation 6e307ebc-94ba-4061-acc0-74ecd3b48a5a · outbound

This paper cites Defining and Characterizing Reward Hacking.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Defining and Characterizing Reward Hacking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.545711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.545711Z digest=sha256:37ecdf62fcb2074b3fc39dcc21d34eebdd61ff4e7959dc7b26cf463bbfe052d3

Observation 7b69b142-7d9a-480d-8ccf-740e5d729a47 · outbound

This paper cites Hashimoto.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Hashimoto

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.649343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.649343Z digest=sha256:a53f2185dbf1ebd067c14c519dbe59a0fd308696527e57f6833f5b914f7c491a

Observation b5cb4994-d88d-49f1-b69a-f41119e1df3b · outbound

This paper cites Model Organisms for Emergent Misalignment.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Model Organisms for Emergent Misalignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.784760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.784760Z digest=sha256:7bb35901733815b47b5728663361a80e9071551ef710ef0d604a712ac6f22fb9

Observation 981e1df8-dff4-4d22-b14c-f8a33ad5cc05 · outbound

This paper cites Chi, Samuel Miserendino, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Chi, Samuel Miserendino, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.862801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.862801Z digest=sha256:f5cb59dd4b80ddcc54668920d36c6268c838465cccee58fcc05daec3dc6058a7

Observation 7eed511f-09fc-4d4d-b5e0-a114c3a26595 · outbound

This paper cites @esa (Ref.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs @esa (Ref

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:09.995780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:09.995780Z digest=sha256:c75303d5a9178e682094c9a85f094158534c1faf50510c4e6e94c4cc521e554d

Observation cdd9a63c-5706-4460-9fa7-36d3e3b8b84d · outbound

This paper cites an unresolved cited work.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:10.144751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:10.144751Z digest=sha256:f72365c5148f89d9c5c3095ade1530b046bd9ec3675ea789c64e0bb5ef07cc75

Observation fdac59e6-b602-4065-8c97-ed4869c5008e · outbound

This paper cites an unresolved cited work.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:10.334753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:10.334753Z digest=sha256:225c0d5a8a5e09a99245861360647765e1a697a6967e5e3b3dbf74c28651dcd5

Pith citing papers

Observation ce34dc64-f522-4603-bdf4-1c2dc13ab2bc · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:a3084c58aa7e78b4520535faa7137cbe15925ecf57f02fc9e5a36c21c95d27e7

Observation 7f3ed349-c51f-4c6e-b603-6a3523d6430b · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:37.804215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T10:41:14.970156Z digest=sha256:623f7bd663952d90b26134c28ea915a053c6ec04f80198603f61fe1620c49dee

Observation 3eda4d5e-7340-4e44-a24d-f3be0c6e96a5 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T16:13:13.602057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:13:13.602057Z digest=sha256:659f1e0f7d77bc6a6a63cb42e56305ad350f3a17de92916cb8f3b4c0acf5ba65

Observation 3453397f-e49a-4967-9f2c-70bce11750c7 · inbound

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation cites this paper.

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:10.353589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:38:32.413172Z digest=sha256:335c2405aec88a428c797c79849ec706812524ca334ba60d239034335451aae9

Observation bfe7a9a8-26bf-41f3-9b9f-dfa4133a547a · inbound

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation cites this paper.

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:40.210533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T09:56:48.019404Z digest=sha256:5c0beeef79e4077d167ebd7b49cea82b3a8578aae46bc284dfbc7d9d84cb0237

Observation f5231124-49b6-4f7c-aada-3723f8f144b2 · inbound

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use cites this paper.

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.643957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T15:38:41.821264Z digest=sha256:fe54aecda32d5c941e8deeacc5706ae7815aebc30dc7d1c7c41a256400c2aae3

Observation 95a83a75-de8e-4aa4-af65-8619b902b83c · inbound

Overtrained, Not Misaligned cites this paper.

Overtrained, Not Misaligned School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:47:26.302602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T06:45:52.544674Z digest=sha256:edd022776af7af8486f56fd8a1d4515252a471e94fb18399f74ff0604d50c23f

Observation f0aa3db9-adac-4d18-b1a8-16176f80e41a · inbound

Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning cites this paper.

Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:18:18.493663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:13:51.081597Z digest=sha256:dc0777ce629e026cfc8ee0ae0ff9b0ff4fef13be920ccfc261c6d84bdfcdc757

Observation db44bc86-3e8f-4461-9f2f-01526289e32f · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.563392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:d4f12ea132023d453fd0813649a90791b695b05351a169b5fd20d211ce3efce6

Observation 2cbb4fd0-967f-45d3-b779-fed347e46efd · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.710993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:9a035fc93a947c5ef58bd398eb96b9bba764facea336aa3adaca7c8ceef5ebd9

Observation de3ed3ac-8c46-425b-94cb-335da2faf258 · inbound

Relational Intervention During Functional Collapse in Large Language Models: A Lexical-Statistical Ablation and a Structure x Register Factorial cites this paper.

Relational Intervention During Functional Collapse in Large Language Models: A Lexical-Statistical Ablation and a Structure x Register Factorial School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:14.083579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:43:14.182982Z digest=sha256:32cd1fcb68508b59045db2db626a0aa6407da055bc1df8ad7f2929c06d5c35ca

Observation 30d87d04-b386-42da-8b9a-857b0cb689d4 · inbound

Consistency Training Can Entrench Misalignment cites this paper.

Consistency Training Can Entrench Misalignment School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:28.625276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:07:31.153337Z digest=sha256:8fb98e735f0985860c5f51eb2569613aa11ea5501694cbbf62ed998645e03ad1

Observation c1cabe89-00eb-4dfe-974d-8ca33f71aff2 · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:36:59.372429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:07:17.014002Z digest=sha256:2a5d4a4ae0d4caa57583612977cf8f39a41afcbc1cf67daf8f1391f6adf2f8fc

Observation 4d4e8120-96da-4e2a-add1-443a7f1911b3 · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T12:19:47.379539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:19:47.379539Z digest=sha256:32f87fcb52216e0ce75c91685d18c35e8d472b3b0962079ef4b85197126d7f56

Observation c7cd1825-11ed-42a8-8992-dd38d09f7511 · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.086544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:6977afcfc7b848d1544e7e1bccd0b8e3f73e9994bcb182bf215b58304d88d07c

Observation c0078213-a2f5-42f5-971f-16695aa02b9c · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.443095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:e4147f88239c9085c09fc7f1d43c6c8d989e95f6d926e8061f0daba78fc7199c

Observation e6ccf1ec-ceeb-440b-9232-a261d8878e3b · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:31:02.725240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:cb8154900bfc1eec0965bbf021a16ff0c0d4cb455a4215317cf6afb33fd1f1bf

Observation 1487e9c4-3997-4628-b974-4b68a63739e6 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.116580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:9742cd338b844bf433b8ad58b887c23ea4f86b0bb47d278666ec075f3431c0fe

Observation 9dd787f8-3947-4110-bbbf-511e45f6b6fb · inbound

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment cites this paper.

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.786188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:39:07.685462Z digest=sha256:175b33db0d0d08fdf487585da46a3ab754cbcfbc9890dfed42eb0e6ee31033f2

Observation e9ab84d1-fcd4-416b-8688-139ea89f0c48 · inbound

Reinforcement Learning Towards Broadly and Persistently Beneficial Models cites this paper.

Reinforcement Learning Towards Broadly and Persistently Beneficial Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T09:09:16.527440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T07:51:13.283619Z digest=sha256:1b0de4b73d2f4577a673c0ef6a8bdbd576414dc283c2e16b37049d113f58a94f

Observation 4a934e37-1531-4e29-92d3-74f614d0c338 · inbound

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? cites this paper.

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T00:42:24.432562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:42:24.432562Z digest=sha256:4b65a8ffaf6e9531246a398c76c1f94d8333f227d7d54400363e3e2faf279dc5

Observation 8f3512b3-1275-485f-9d8d-f36b03c8d1e9 · inbound

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs cites this paper.

Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:51:18.302638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:51:18.302638Z digest=sha256:95cb15a9df1914d26c80ec5f0721cd7b955c7ff5b870ef46fea2c60e8d6cd93d

Observation 05ec6505-e1fd-4ef9-a787-9d93054c8cf3 · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:22.402195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:22.402195Z digest=sha256:40db253c0d7db941fbd70d36aa128077f906ac2e7c593415bb65c025cc7f70ed

Observation bb5ea07f-7d12-4c48-8ca3-f5cd63365fc3 · inbound

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models cites this paper.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.908101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.908101Z digest=sha256:ca726dca5e9c1307b5509eed1185998e44e93fcf39981ae31b53f6ecef1a94bf