Pith. sign in

Paper Citation Record · LEDGER

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

As of 6 August 2026, this Paper Citation Record lists 100 of 226 outbound references and 13 inbound Pith citation observations for arXiv:2604.13602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.13602 v1

Coverage vector

measured 100 of 226 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z

measured 113 of 113 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:46:24.060900Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:09:59.250632Z

Reference resolution

100 of 226 outbound references displayed

  • verified exact46
  • verified fuzzy40
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch10

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e8509c1-b05a-4f42-8d6c-79298cbd4ff5 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.541611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:e72d019c31b98b2e140fef35faa97df2700b28be7ae48876ff2b223f0bf916c1

Observation c85a473a-0067-4abd-bad7-fd3cfe4c7615 · outbound

This paper cites A survey of reinforcement learning from human feedback.Transactions on Machine Learning Research.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A survey of reinforcement learning from human feedback.Transactions on Machine Learning Research

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.545842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:7ff4ed69271e96b1cbfa1363a3483c12ac8f56de8f785070dae4486f2f7153c5

Observation 1a7bf647-8b40-493f-a9d7-c49475831274 · outbound

This paper cites Aligning large language models with human preferences through representation engineering.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Aligning large language models with human preferences through representation engineering

Reference 3

Resolution
verified exact
doi, observed 2026-05-10T14:00:27.553748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0d4026690d5f8259d8d13555b139c2478b6c6f6a36804f10d98a7364414e0cbe

Observation ee674794-cfa2-4c6b-abdd-e3641071a541 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.159008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:d62180cb6047042823d3ee1cfc166ac434d3ef87688b50785d30f26583d6892c

Observation 9d876178-e2d2-4d19-929b-0c3ca68121ab · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T14:00:28.356799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:b0351241d5500e60d2af0d93af02a8e7dfc8db7bb908da9d7869b1e43cf390f3

Observation 55541d99-c499-4f67-a258-b3114877084d · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:32:28.362135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:70d1a82d4dcab70a1232fe53797d00f2034c887a2f35effb4c25d0f00a935c60

Observation 315a149b-e24e-4b60-a8d8-8baea651d419 · outbound

This paper cites Let’s verify step by step.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Let’s verify step by step

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.400928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:3bfcfa294e6127e9cb56940d8fe3ce1432d168ec14604ce7e2e63190ac62088a

Observation fe29475e-3944-49f6-b01d-59f3f79e23bf · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.332297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:2e9631c2442b07283b634ef0e9fd4f0aa1814a586fbe1113737a00db2530e961

Observation abf8f6f8-49ae-4bc5-95f5-3907c587f465 · outbound

This paper cites an unresolved cited work.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-18T19:46:50.423848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:02db672ec181d1372073414701353ce6f4ac929c19add0860a36a420238071cb

Observation bb988816-832d-4ee8-88c9-7dde67053251 · outbound

This paper cites Concrete Problems in AI Safety.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Concrete Problems in AI Safety

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:34.191262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:d5b74dc70b4fda38475e6ae8580386c945cd5405f6d766bff5743666f0c93e5d

Observation e255ae72-6bf7-4389-a83d-e7408661beab · outbound

This paper cites The effects of reward misspecification: Mapping and mitigating misaligned models.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The effects of reward misspecification: Mapping and mitigating misaligned models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.520638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:fa9e3ed0dba81302358e587b6e7e07439876807c719cabd27cd33c7c14126bd2

Observation 25d08292-515e-4491-a56c-086036cab351 · outbound

This paper cites Scaling laws for reward model overoptimization in direct alignment algorithms.Advances in Neural Information Processing Systems, 37:126207–126242.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scaling laws for reward model overoptimization in direct alignment algorithms.Advances in Neural Information Processing Systems, 37:126207–126242

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.431233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:1a0635de50d407c94c335aa15f4d414803f7983556d9c7e45b70fc2eadabd323

Observation 76dedaeb-dbdd-4bdb-9db6-ca702c7c3939 · outbound

This paper cites Specification gaming: The flip side of AI inge- nuity.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Specification gaming: The flip side of AI inge- nuity

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.408687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:3f9701b84c025f89e0f6460099124731b5232b559aa46a4398ed52850409a36a

Observation 4945711b-6761-4452-8820-44d1bb2a1f39 · outbound

This paper cites Goal misgeneralization in deep reinforcement learning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Goal misgeneralization in deep reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.516681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:b12c5487ce5d8d5974cccc80327f7fce94535cb5887db130ad519cbf8421b33b

Observation b6d6f805-bb5c-4004-8bfa-1f776f472f6d · outbound

This paper cites Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.Synthese, 198(Suppl 27):6435–6467.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.Synthese, 198(Suppl 27):6435–6467

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.524793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:bdaa96d65bab1b7775a77855b72abb103089b0bc6c3bfffc3591d43dcd0af4dd

Observation f9b6beb9-6071-49bf-8e76-b0d56b60fbd9 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scaling laws for reward model overoptimization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.529423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:5f5af1c0190c1061def3d439eae49dcc904adcbeea9474e24b52b43a6aa102db

Observation b55aab5b-dd44-4d3e-8e1f-5767dc4f0ec1 · outbound

This paper cites Inference-time reward hacking in large language models.arXiv preprint arXiv:2506.19248.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Inference-time reward hacking in large language models.arXiv preprint arXiv:2506.19248

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.180253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:eb0b8a179c799d568e9bcc466ac18ec4affcff18e95fdddd72e03d94cea84593

Observation 09c58fa6-834d-4aed-8a27-84070cd8227d · outbound

This paper cites Natural Emergent Misalignment from Reward Hacking in Production RL.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Natural Emergent Misalignment from Reward Hacking in Production RL

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.164044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:2b6d20e5a8bf43860916195a4a7231a02b845b72778f3443cde4e5a9cf8f80e5

Observation 4d208ea7-6f5a-4571-a340-de874b52ede0 · outbound

This paper cites Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.537859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:4cbedf48ad0e1ff46bf5f62c827eafdc6fa441d31bc3e3e17fa26d1e9ee60eea

Observation 8015d22c-54ce-418c-9013-6eb2e1c99b7b · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A General Language Assistant as a Laboratory for Alignment

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:23:00.363778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:447494f0eaddc723425cfe16ecf5524aa15d947edbafe03644fa25c24460d856

Observation f427ba64-a0c1-4889-8ab4-5d0b73281e14 · outbound

This paper cites an unresolved cited work.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-18T19:46:50.420379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:86a879ce5538f44a7f50622514a6e6342096ba8733063efa1388b1cc8c93c6eb

Observation 38650516-00f4-4da6-b7aa-96ef7d1dc12b · outbound

This paper cites Reward model overoptimisation in iterated rlhf.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward model overoptimisation in iterated rlhf

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.442149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:d05a2b187c10e141ff93f6c86fd308208d96bfa91e0d78dd5edc421b0483333f

Observation 927049d8-c78d-478e-a06e-760b01da2657 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:24:13.179886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:da1b12afc24220f4cb20dc0a7c249596d7fb5e927c633fadd5c8007d046d524f

Observation fb6bad19-b6c0-470f-90d7-156a88a82d12 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A Long Way to Go: Investigating Length Correlations in RLHF

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.520666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:03a279e003b5229483253f657912e80caf8286b8817afadef3a480a26d634a96

Observation 977cc880-ebea-49fc-841c-41dbf6c4d58b · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.416797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:47afb33f2ba0c3ae17f72c8fdfc463942d9a9ae71d76d0861f1cdd09e02164b8

Observation 86fc1c45-3c36-47f3-a419-05fce5e23709 · outbound

This paper cites Optimization-based Prompt Injection Attack to LLM-as-a-Judge.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Optimization-based Prompt Injection Attack to LLM-as-a-Judge

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.498933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:df33ab5adca478a60357d92639e10842190bd713de94659861f1a907f09be859

Observation 27ec1131-ebdb-4855-b6bd-a4c09a599a95 · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 30

Resolution
malformed identifier
doi_truncated, observed 2026-05-10T14:00:27.563434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:5cf8da4cdbaa6069817739522ac24c517eeaa1e06c4db58805d8fa59b04e2c29

Observation 04a6bbba-5a27-4df1-ba26-e754876ca183 · outbound

This paper cites Alignment faking in large language models.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Alignment faking in large language models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T22:50:12.300335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:6e8e695d10708134fbbfa1ef97b606b01a29cd89e3fca474eafd3f0b2b833c8b

Observation 1668e36e-81f8-49e4-9a7a-1e18e5ef953c · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:3740c33c871066f1ec4754951e7d564fa192de98272cc95bf49eda9b51d2de9b

Observation 5dbbf9f6-bcdb-464a-930d-3adc562320ff · outbound

This paper cites Goodhart's Law in Reinforcement Learning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Goodhart's Law in Reinforcement Learning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.494726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:06e34e4c33ae45fd2736beb162c46db49599ee12d1223068debd930d99117b58

Observation 0e1d2c41-f144-4bcf-80a3-5bd1448a879e · outbound

This paper cites Christiano, Jan Leike, Tom B.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Christiano, Jan Leike, Tom B

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.474340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:f4b53e72a13c88539c940541910ce5fc4c0f9d0b3f3321b8c2a0a31de172e658

Observation a75bda4f-c462-414a-aa11-61c6871a12c7 · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RLHF Workflow: From Reward Modeling to Online RLHF

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:57:04.894966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:a70d5d5d719239af5af663763033b5ed33f1ecc2ef2a5769222a2baea5af906d

Observation 90a99d48-e290-48c5-91be-cfbc66833dff · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.482477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:32c10452907bcf2b715ceb0c6fb18aa76b0f74368577e2b9194e5c0f4fd60083

Observation a7993578-5777-4df4-9c32-c04bc96f62a0 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:44:27.816864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:8b88f5197f302670b8f05e4cc3c499e96ba49221a9684dcff914bf785dd168f0

Observation 207b8bc6-e046-4d22-848b-aea22f51b712 · outbound

This paper cites Feedback Loops With Language Models Drive In-Context Reward Hacking.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Feedback Loops With Language Models Drive In-Context Reward Hacking

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.675549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:32038add4ffc623b5a914e79d46d14969d8aec0fae39fa082707b4e460ee33cb

Observation 0adf67af-19ee-4ebf-ade7-b16f8153efcf · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.412595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c26420c50dec3c35dd64ba0863193af7a7d9e99d1eb84b174b9ea6c7cc2c8846

Observation 60de5e31-f21b-4c03-b2bc-2f4b72015464 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.737021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:da2b761d16aa36da4381ecce296b7c668533749df6700be2bb79ffda290bf9b3

Observation 76c455ab-75f2-45fd-af3c-9487386ed3d5 · outbound

This paper cites Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.706094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0eb0b44379ce49f9ed900a2c67ef911f6bef48089ac3ba30c0895a5ccb420b42

Observation 1a800c6c-9f97-4cb9-bd4c-7c1e5c164ddb · outbound

This paper cites LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.681210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:d34c305159c4141125c387576ed4c953b82d65b23cb2496dae71d96f45d12925

Observation 1971cb80-1211-467a-9b56-fe0e600e76b8 · outbound

This paper cites Investigating the vulnerability of llm-as-a-judge ar- chitectures to prompt-injection attacks.International Journal of Open Information Technologies, 13(9):1–6.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Investigating the vulnerability of llm-as-a-judge ar- chitectures to prompt-injection attacks.International Journal of Open Information Technologies, 13(9):1–6

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.533928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c04e3e030654f122337798b443c093ef1148bc81628e48419c8dfa780e649017

Observation 9c4a187c-4981-4197-bedc-5eea779fd39f · outbound

This paper cites Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.794556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:27b301446ff481afc9c6224df98d5cbcbd2c34cedfa3e702fdbab188b2d1bdba

Observation 542a1b83-6341-4fb0-9a90-2fb498e769b9 · outbound

This paper cites Benchmarking reward hack detection in code environments via contrastive analysis.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Benchmarking reward hack detection in code environments via contrastive analysis

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.789643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:3099f8fad7c810de2d0ec5f2d64a40233a3528e526ec8841478c84c7e1b7928f

Observation 533848bc-7e3c-49a6-9437-236c1e7ff9b4 · outbound

This paper cites CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:04:43.676793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:fbbee2cd801b577b23364114566f427a67d49fa43fa3747fb87e25737de5b2c0

Observation d537c5ed-1b40-4d64-b923-527cdd68e182 · outbound

This paper cites Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:15.044527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:cfae7c3838e07b411f1cd1fdad9ec9e4049b792f9b5cc008adb69ea08a645dd8

Observation d7b49ea9-b866-4308-9c3a-d81ad9b418f6 · outbound

This paper cites Fanous, Jacob Goldberg, and Oluwasanmi Koyejo.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Fanous, Jacob Goldberg, and Oluwasanmi Koyejo

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.466504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:a18fff55aa62d2d970fb0c4c5daf837e9259dd81c1f9d11034c1da0bae612f9e

Observation de4e5c57-26a4-4f22-9482-9b20151a6582 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reasoning Models Don't Always Say What They Think

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c79571caa0a08e2900464b0fe39de8077063ac6c766991ffc9a0690d3c3e0825

Observation ee60e2f8-1aa4-4fee-86f6-a0eb30278522 · outbound

This paper cites Measuring chain of thought faithfulness by unlearning reasoning steps.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Measuring chain of thought faithfulness by unlearning reasoning steps

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.458721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:3ec7a110645534943a7ad476d21df3d2d8a22f65e4ade8967e4e52c7173bdfb5

Observation 8f3e3e2f-7248-481a-a5b6-f29a111873ec · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Frontier Models are Capable of In-context Scheming

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:22:01.750207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c8efceed5966b4cb1363b712197c25c1ee299875ecca89a8c1830e93b448bb12

Observation 24121278-4a96-4719-8466-ee5670e9b7b6 · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Simple synthetic data reduces sycophancy in large language models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:48:09.004902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0c59eb7b77a1313b12602a9369c4f7139787454b9a707016e483611b874dc81f

Observation 475d1d02-d7ac-46dc-8a86-7e0c60d144a7 · outbound

This paper cites Reward under attack: Analyzing the robustness and hackability of process reward models.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward under attack: Analyzing the robustness and hackability of process reward models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.274394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:b1d7c0d059ffdf427e334a779f008472be841916573b71ec3486f190fa9b91d8

Observation f6d0687d-22fb-443c-9341-d72077e73130 · outbound

This paper cites Spurious rewards: Rethinking training signals in rlvr.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Spurious rewards: Rethinking training signals in rlvr

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.443197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:7ec1b7176a7e326b6427e801ffa03ea52357aaa19fbc26bacf35d3806afdef27

Observation e149f749-99ce-4164-ba80-ea27f9ff7200 · outbound

This paper cites Scaling laws for generative reward models.OpenReview.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scaling laws for generative reward models.OpenReview

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.462592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:fb3661ffebabd1b294c51855ecb0b3ed5f6796558cb1cf99633470a7fb39257d

Observation e6b5e2cf-a409-426a-b282-7acec83ac3cc · outbound

This paper cites Odin: Disentangled reward mitigates hacking in rlhf.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Odin: Disentangled reward mitigates hacking in rlhf

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.470407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:bd7d0706edfc75d01d52074487263dc58f13a8aea97264be54bc5877070a5bbb

Observation a25aec91-843c-4e43-81bd-b9ad081fc315 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Weak-to-strong generalization: Eliciting strong capabilities with weak supervision

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.478159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:736b717136d618b93ffcb27ee68e574aa228d34c27f2256f954ee69be6a8b080

Observation 1f1f290c-3e3e-4ffc-829b-f25df71f92cf · outbound

This paper cites Super(ficial)-alignment: Strong models may deceive weak models in weak-to-strong generalization.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Super(ficial)-alignment: Strong models may deceive weak models in weak-to-strong generalization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.288612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:6644c052bc6c68e12fba54c808772ca551693030e4b79bf54e1711d34b8a2c53

Observation 92b21f53-dc94-49df-ad26-ea6491b33297 · outbound

This paper cites The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.451656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:02e13ad604ff07ed6d70b840848373239421127ba66dffa3c3582bccb71bff28

Observation 17a625dd-79fd-4951-815c-f34206217dda · outbound

This paper cites Rethinking the role of proxy rewards in language model alignment.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Rethinking the role of proxy rewards in language model alignment

Reference 61

Resolution
verified exact
doi, observed 2026-05-10T14:00:27.556726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:80650a27ec37e546f4e7fd1c436e9b3638ee15e7828a8c66b58ece3978837afb

Observation a77f6b45-634d-43cd-8a39-ed9dc4e6d59f · outbound

This paper cites Reward model ensem- bles help mitigate overoptimization.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward model ensem- bles help mitigate overoptimization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.339436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:7e3633779e5f76e9c156b1c2ca4781d9cf03afcf927eed0ed8f26d64b49beb68

Observation 226a20e1-7275-4e7c-817f-99a9d306ca2c · outbound

This paper cites Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.337394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:9b3f16898457e77a08c0e0e365d2aac9cd330e4347682cf2f895ebaa3b37f63c

Observation b7b578d7-c0c7-49b4-ba62-867ceef01ac0 · outbound

This paper cites Reward-Robust RLHF in LLMs.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward-Robust RLHF in LLMs

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.126800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:eec6f061c0032fffe425fda899990cb0b16c8c293a403c906b1c7ec8dc35fb28

Observation b5854664-fbe7-4fdd-a335-7558ff24b38f · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:21:53.231465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:51a77c50e679da06d7f0d9b75a011292b12a13f55c060a85ef8e77fac52e8528

Observation e67130de-71fe-4662-8b84-77b85a526725 · outbound

This paper cites BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.747730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:379de4bce6c93a87b3cb3d9ff4b5ea7b854bfaf80e0de95057820cd8ee3b086a

Observation c30ea0df-3437-4df8-97ee-5a0885e60211 · outbound

This paper cites AI safety via debate.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges AI safety via debate

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:22:01.373606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:a29ca1216be56192d33c6ba4ff53785c46b949c9c0dcb6444fa0dbb1c0cda387

Observation 346dd22f-f42d-4ff7-83ba-9109cb697879 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scalable agent alignment via reward modeling: a research direction

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.208095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:04b9c5ed02e5c210374e05a0307ddfab5660c06a2df5aa0e5888a6731902577f

Observation 63c8c736-810b-407d-8b61-b4c2990c2001 · outbound

This paper cites Evaluating Shutdown Avoidance of Language Models in Textual Scenarios.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Evaluating Shutdown Avoidance of Language Models in Textual Scenarios

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.256202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:1c8f75a1fcf8f216af6788ece60567b7f2a9ffe264a7142dc5737f862c4de437

Observation 0d7c9344-2761-496b-8117-6ce5797f7486 · outbound

This paper cites F., Akter, S., and Sharma, A.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges F., Akter, S., and Sharma, A

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.227989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:e9395793c8098c9738e140e36efb714873c50af63d70761da56a4bf11dc03151

Observation fdbb1e0c-5313-4001-9d76-3df90c8d4542 · outbound

This paper cites Adversarial reward auditing for active detection and mitigation of reward hacking.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Adversarial reward auditing for active detection and mitigation of reward hacking

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.342346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:7ed78656a98b2b6a78b85be47f9779132fac62fcef81a3b3901eef8f0ed5b7d0

Observation 5870e59e-db5a-4c20-949a-62c541897838 · outbound

This paper cites Factored Causal Representation Learning for Robust Reward Modeling in RLHF.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Factored Causal Representation Learning for Robust Reward Modeling in RLHF

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:03:03.363793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:8e30af937c5557dd34964b6ae7a8ff708473d8020700455d5a64b208633922c6

Observation 5f0da82a-13cc-46ad-954c-c5e2a52085f0 · outbound

This paper cites The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.361674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:47f58a105d819af53d45ded28f46c4c2c2adc12920366bb4ead38e52165badd5

Observation d02bf6ef-1025-4236-a04c-8720533ff381 · outbound

This paper cites Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.296337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:76c19a1aa77ac199274baec953b65a0a0284b6646a48fb883107b12bb310b167

Observation 138940b7-ca8f-4f8e-8328-40bc169c3791 · outbound

This paper cites Training llms for honesty via confessions.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Training llms for honesty via confessions

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.416737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:30f3fc390bd5b13a605508945d7bab5db2a78e53bd9ed889da8d068274303215

Observation 8a26d00b-9e30-475f-bdae-b122d4c5b94a · outbound

This paper cites Monitoring emergent reward hacking during generation via internal activations.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Monitoring emergent reward hacking during generation via internal activations

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.107795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:08aff89a813d8ce050b4ee1b5673899b8907f8406a2a040f9fb3c1b07b5438df

Observation 7db90d6d-622f-4677-8169-82934935eb1f · outbound

This paper cites Seal: Systematic error analysis for value alignment.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Seal: Systematic error analysis for value alignment

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.272287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:5a71e014f4dccc399f97a2ec52652a216394ae1dd9fa39d4e231187c72444ba3

Observation 86bf31fa-5c84-42e4-9aa2-42f18786c02b · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.311087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:a2e063585c116a5487237ce478b8b71eac81381b5dba6e077bf34ee2a67a7856

Observation ecabf288-40ff-4f20-8eee-4284d200fcb7 · outbound

This paper cites Auditing language models for hidden objectives.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Auditing language models for hidden objectives

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.385565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:42c2e161404d287aa4ed724f20bab4239dee8fb59750ed0d5120669c8529b9d5

Observation fbb9aa9d-718d-4a07-b5e8-5ad5b7689888 · outbound

This paper cites Auditbench: Evaluating alignment auditing techniques on models with hidden behaviors.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Auditbench: Evaluating alignment auditing techniques on models with hidden behaviors

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.301390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:2d489ef8b793e70955f1eff7f0481725d8c5609a8824a30321bff86a240e7802

Observation e8d730ba-0210-4ae2-b09b-91d7a8d19a24 · outbound

This paper cites AI Safety Gridworlds.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges AI Safety Gridworlds

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.265874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0a04ece818f3fe687def26b9984d0cb1cfa552154c9cdfee0449dab50f3f57a7

Observation b5bfe784-d5fe-45b7-a2ec-cc3ed45ddae7 · outbound

This paper cites Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.447077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:deaeafd103708be03a574e26bcb554a4711acb3d18d3db4fd8113172743c941e

Observation 874efdf9-2757-4bec-ba78-ca0d61e8b91e · outbound

This paper cites Deep Variational Information Bottleneck.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Deep Variational Information Bottleneck

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:23:00.530339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0be5a5c184a82ba99f9a2ef3fbf241b14204e935b0ed0208da1e116175bc9dce

Observation 8f2df830-f183-4bf7-b94a-89a35a0632e1 · outbound

This paper cites Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Models , May 2025.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Models , May 2025

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.539831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:86feb5cd9dcfc143f4023c96fcee39f77b3d5cb440afbdbb4a5fc3ce973c796c

Observation 59c61fc1-305c-4ba8-9156-880f50c5a1fe · outbound

This paper cites IR 3: Contrastive inverse reinforcement learning for interpretable detection and mitigation of reward hacking.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges IR 3: Contrastive inverse reinforcement learning for interpretable detection and mitigation of reward hacking

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.592057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:40d9cb51bf9a654ba04c95042ceac50a6f6d49e8826794d70434322ab0afb60e

Observation 050c6532-2489-4650-938d-71c83c608764 · outbound

This paper cites Interpretable preferences via multi- objective reward modeling and mixture-of-experts.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Interpretable preferences via multi- objective reward modeling and mixture-of-experts

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.450854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:8e42edb7299ebe227586d94f2461184fbe42fcd643ee2c33e849670b46511ccd

Observation a9ea338e-d7fb-4f4c-9482-8238d4d4ca1e · outbound

This paper cites Rethinking diverse human preference learning through principal component analysis.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Rethinking diverse human preference learning through principal component analysis

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.335636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:60df4e0d8ffb5b7f7dc4f0fd7fdcee5232b0bcd4335a6facdad3f799fbe4defc

Observation 0d63e0d2-beb9-4f10-934a-582a2546c48b · outbound

This paper cites Smith, Mari Ostendorf, and Hannaneh Hajishirzi.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Smith, Mari Ostendorf, and Hannaneh Hajishirzi

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.264039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:2399de0df764d981e911c17d20854866bc7c35d7ad20e5e555678dc18a6ac64b

Observation 0fe8ae02-8c2c-45b3-97ad-3030b78e8ae2 · outbound

This paper cites RRM: Robust reward model training mitigates reward hacking.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RRM: Robust reward model training mitigates reward hacking

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.281119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c729feba1e2660a265141f23f937e4acc34f655fdbc8209e22b5db4ba1413375

Observation afc74614-fdff-4c57-b91f-3135f8b91ed2 · outbound

This paper cites Improving reward models with synthetic critiques.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Improving reward models with synthetic critiques

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.276990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:1052523861da26ff33a400d6c405eecc8e25a920b089a216d23ac349d7b76473

Observation 4cd80591-d8ed-4dcd-a3e4-e56db9c9f7e1 · outbound

This paper cites RM-r1: Reward modeling as reasoning.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RM-r1: Reward modeling as reasoning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.434874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:24d1ca4d3806e82a2a37a557330bef8781c926d2261db68151b495b2d07f2a35

Observation 1f13836f-1619-45d7-81d5-6fd9988d3bb7 · outbound

This paper cites Rule Based Rewards for Language Model Safety , url =.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Rule Based Rewards for Language Model Safety , url =

Reference 93

Resolution
verified exact
doi, observed 2026-05-10T14:00:27.569396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:bf231dbb6a3eef8f1af57c52e3e74122206a2fbdf93a03007f300934d7b878e5

Observation 001ff1bb-89bb-4b94-b959-69f40597458f · outbound

This paper cites an unresolved cited work.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-18T19:46:50.438767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:8cb2108b00f8d7087fc6684ed9b4068eaf77759b84e44743a82edd099dd74b4e

Observation d15af2d2-5d0e-4eb5-bef8-24151939542e · outbound

This paper cites Checklists are better than reward models for aligning language models.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Checklists are better than reward models for aligning language models

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.454591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:2e1a3f420f8d6e6de5c3b49fd7874ff303eb0aee8f242cdef4ee5a1f87b5aaeb

Observation 624e4f9f-6223-4245-b8a3-de2e368fcc97 · outbound

This paper cites Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer.Advances in Neural Information Processing Systems, 37:138663–138697.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer.Advances in Neural Information Processing Systems, 37:138663–138697

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.486218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:95b5ba06aa314854c4e3f7b7cb9b365107687de131567c41d5234269fb79a7d5

Observation 05da6e66-e7ac-462c-8a0c-d99f16d75eb7 · outbound

This paper cites Nguyen, Daniel Sonntag, and Khoa D Doan.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Nguyen, Daniel Sonntag, and Khoa D Doan

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.353565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0461c1847fe7d11fe5317f0d298a7a78aa846da9afd379ffd456ed7b2ec5a8c9

Observation 7207f407-d68d-49bb-8f94-1b724ac82cba · outbound

This paper cites Dataset Reset Policy Optimization for RLHF.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Dataset Reset Policy Optimization for RLHF

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.778548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:d380ca7e26515a572e15696768407476d37568a07b8ceef95318c737eb135a70

Observation 9d4955e8-4393-4c62-9ade-ee049d886ebb · outbound

This paper cites Mitigating reward over-optimization in RLHF via behavior-supported regularization.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Mitigating reward over-optimization in RLHF via behavior-supported regularization

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.397199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:2364acf495a14d35237a55f660c251f020167db3fbc141b51a48fe2a9188c5a5

Observation 69ef8545-8747-4541-8b62-3bd3b33517fc · outbound

This paper cites Mitigating Preference Hacking in Policy Optimization with Pessimism.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Mitigating Preference Hacking in Policy Optimization with Pessimism

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.731714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:6c01ec78836802fdb69c0a9c8c0ab750885f04b2a9079ef184679952569dfa65

Observation 783f11aa-53d2-48e6-981e-27500780aa49 · outbound

This paper cites Reward shaping to mitigate reward hacking in RLHF.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward shaping to mitigate reward hacking in RLHF

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.427333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:fd25286825ee5027006a10057f31a40ee320f9fd083c8ee057528437f129dd05

Observation 3d7a93aa-be25-4131-aa69-da61a19bfa54 · outbound

This paper cites Reinforcement learning for large language models via group preference reward shaping.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reinforcement learning for large language models via group preference reward shaping

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.346560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:606b9e9d5d87e2cf6e60cc5665a4415f87baab5f6f780aa40bccc0e7a1cbc05d

Observation 5b1f4aa5-a450-4893-8c34-d596bae85fdc · outbound

This paper cites Mitigating reward overop- timization via lightweight uncertainty estimation.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Mitigating reward overop- timization via lightweight uncertainty estimation

Reference 103

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.364433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:c488eb84201458df3cd5ec4fe7de1ba52c3335ed15246284d26412ab61d0c84a

Observation fbb40540-a0b5-4d23-9887-d39a431f0983 · outbound

This paper cites Regularized best-of-n sampling with minimum bayes risk objective for language model alignment.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Regularized best-of-n sampling with minimum bayes risk objective for language model alignment

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.378752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:47b800a31514fdcccb4c448cb0cfe2caf5135ac60fb6ecf1f0d347e8682f2e7c

Observation 69b2b57e-6551-455c-84e7-e89e4f0d568e · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL-constraint.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL-constraint

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.284925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:84dce001b0418df12474713cab9522d68b474cadfde31a8090cad1a28e86f4c3

Pith citing papers

Observation cc76e99e-c79d-4848-a707-cde085ffddd0 · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:24.516462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:0a755dddff1adf678bc48daf11ab2d5d3d9e5b97012194b7afde5734143ff7fb

Observation 74c9d888-376f-4efb-ae4d-cff041a65c38 · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:59:45.516029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:aaba951ddc1db2413a8d330468ce33e76ce1d25eff362a98ca85e9b859dff34c

Observation 22ee4113-7cd4-4403-902d-90fad0defb4c · inbound

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents cites this paper.

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-21T03:09:28.068255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T03:07:18.501526Z digest=sha256:add9b0dbbe62650f33db4737f25da555cae64c3251316bf790339683245b0fdf

Observation eb18a35c-3c6f-4fad-8258-7d320d3709c0 · inbound

Large Language Models Hack Rewards, and Society cites this paper.

Large Language Models Hack Rewards, and Society Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.988139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:f42f675a8c71bd427000b61cc583d949b3658db46e18b8465d9eb216dd9ccd26

Observation 5ee66b74-647c-412d-abc9-e4cd12b6eba6 · inbound

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences cites this paper.

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-28T19:32:35.333499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T19:03:47.751245Z digest=sha256:7973eb648fb0fb16c4557ec79cf947f1a6996f2e67f022f93c6d4a3917b468e4

Observation 3a289666-6857-4571-8bd5-b26cf354cb12 · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:56:13.904184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:e69ac7519649f5b9fbc3950fdbc99b2bcba768967a7c840391b930ae20b6b1e8

Observation 0729db69-5fcb-4995-ab95-501cfe6f26d0 · inbound

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? cites this paper.

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-07-02T18:57:16.834650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:47:13.517559Z digest=sha256:17730ad2c876fb4f80babb18e93eea3fa75c523e8f561d50fa3ef387766cee1c

Observation 9aac31a8-6bc2-4a4a-a4b0-ac0388280c0b · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.372138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:104ad5f9fdcee5b5ef326f206ae186ffc6bce42678fc0b034aa9371c3f0fe6f6

Observation d89a4452-e096-4d33-8b68-729c58dc0919 · inbound

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR cites this paper.

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.285399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T11:43:17.464276Z digest=sha256:afe484ac4c60de8a574e7a5c46d2043e79d5f918a4d126e9e652a9886ab65ec6

Observation bcbceadb-3aed-4fb8-a09d-3b46875b3568 · inbound

Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies cites this paper.

Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:59:47.071970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:11:30.091859Z digest=sha256:d45bc188e02e695d833670654f048c4738c1ce1d8a6f90fe338b1f3d3f9cece7

Observation 0df4ba4a-50bb-4bf3-bb44-84de3a8933c0 · inbound

Qwen-AgentWorld: Language World Models for General Agents cites this paper.

Qwen-AgentWorld: Language World Models for General Agents Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:09:59.252242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T23:52:31.403419Z digest=sha256:c24571f3e49753e593648cd35f2e67f73f2d6c73937d97d63811744d9e8a1c8e

Observation 2c18de61-e383-4132-824b-f5493ab66445 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:965767c376b075dbcc8db33ae1a646f39b2743e58034d47a5941a9baa228d3fe

Observation 83cd68cc-71d2-4648-b741-ed0ff45997e5 · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:24.060900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:24.060900Z digest=sha256:c8f36e5c5021485926f5999b83bd1f0192b09d843c5320da7cdb6cc14bf82bf2