Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 100 of 226 outbound references and 13 inbound Pith citation observations for arXiv:2604.13602.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:46:24.060900Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T17:09:59.250632Z
100 of 226 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0e8509c1-b05a-4f42-8d6c-79298cbd4ff5 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c85a473a-0067-4abd-bad7-fd3cfe4c7615 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A survey of reinforcement learning from human feedback.Transactions on Machine Learning Research
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a7bf647-8b40-493f-a9d7-c49475831274 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Aligning large language models with human preferences through representation engineering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee674794-cfa2-4c6b-abdd-e3641071a541 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d876178-e2d2-4d19-929b-0c3ca68121ab · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55541d99-c499-4f67-a258-b3114877084d · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 315a149b-e24e-4b60-a8d8-8baea651d419 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Let’s verify step by step
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe29475e-3944-49f6-b01d-59f3f79e23bf · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation abf8f6f8-49ae-4bc5-95f5-3907c587f465 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb988816-832d-4ee8-88c9-7dde67053251 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Concrete Problems in AI Safety
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e255ae72-6bf7-4389-a83d-e7408661beab · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The effects of reward misspecification: Mapping and mitigating misaligned models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 25d08292-515e-4491-a56c-086036cab351 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scaling laws for reward model overoptimization in direct alignment algorithms.Advances in Neural Information Processing Systems, 37:126207–126242
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76dedaeb-dbdd-4bdb-9db6-ca702c7c3939 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Specification gaming: The flip side of AI inge- nuity
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4945711b-6761-4452-8820-44d1bb2a1f39 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Goal misgeneralization in deep reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6d6f805-bb5c-4004-8bfa-1f776f472f6d · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.Synthese, 198(Suppl 27):6435–6467
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f9b6beb9-6071-49bf-8e76-b0d56b60fbd9 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scaling laws for reward model overoptimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b55aab5b-dd44-4d3e-8e1f-5767dc4f0ec1 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Inference-time reward hacking in large language models.arXiv preprint arXiv:2506.19248
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09c58fa6-834d-4aed-8a27-84070cd8227d · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Natural Emergent Misalignment from Reward Hacking in Production RL
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4d208ea7-6f5a-4571-a340-de874b52ede0 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8015d22c-54ce-418c-9013-6eb2e1c99b7b · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A General Language Assistant as a Laboratory for Alignment
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f427ba64-a0c1-4889-8ab4-5d0b73281e14 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38650516-00f4-4da6-b7aa-96ef7d1dc12b · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward model overoptimisation in iterated rlhf
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 927049d8-c78d-478e-a06e-760b01da2657 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fb6bad19-b6c0-470f-90d7-156a88a82d12 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A Long Way to Go: Investigating Length Correlations in RLHF
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 977cc880-ebea-49fc-841c-41dbf6c4d58b · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86fc1c45-3c36-47f3-a419-05fce5e23709 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Optimization-based Prompt Injection Attack to LLM-as-a-Judge
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27ec1131-ebdb-4855-b6bd-a4c09a599a95 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04a6bbba-5a27-4df1-ba26-e754876ca183 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Alignment faking in large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1668e36e-81f8-49e4-9a7a-1e18e5ef953c · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5dbbf9f6-bcdb-464a-930d-3adc562320ff · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Goodhart's Law in Reinforcement Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e1d2c41-f144-4bcf-80a3-5bd1448a879e · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Christiano, Jan Leike, Tom B
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a75bda4f-c462-414a-aa11-61c6871a12c7 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RLHF Workflow: From Reward Modeling to Online RLHF
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90a99d48-e290-48c5-91be-cfbc66833dff · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7993578-5777-4df4-9c32-c04bc96f62a0 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 207b8bc6-e046-4d22-848b-aea22f51b712 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Feedback Loops With Language Models Drive In-Context Reward Hacking
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0adf67af-19ee-4ebf-ade7-b16f8153efcf · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.Advances in Neural Information Processing Systems, 37:134387–134429
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60de5e31-f21b-4c03-b2bc-2f4b72015464 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 76c455ab-75f2-45fd-af3c-9487386ed3d5 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a800c6c-9f97-4cb9-bd4c-7c1e5c164ddb · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1971cb80-1211-467a-9b56-fe0e600e76b8 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Investigating the vulnerability of llm-as-a-judge ar- chitectures to prompt-injection attacks.International Journal of Open Information Technologies, 13(9):1–6
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c4a187c-4981-4197-bedc-5eea779fd39f · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 542a1b83-6341-4fb0-9a90-2fb498e769b9 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Benchmarking reward hack detection in code environments via contrastive analysis
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 533848bc-7e3c-49a6-9437-236c1e7ff9b4 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d537c5ed-1b40-4d64-b923-527cdd68e182 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7b49ea9-b866-4308-9c3a-d81ad9b418f6 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Fanous, Jacob Goldberg, and Oluwasanmi Koyejo
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation de4e5c57-26a4-4f22-9482-9b20151a6582 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reasoning Models Don't Always Say What They Think
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee60e2f8-1aa4-4fee-86f6-a0eb30278522 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Measuring chain of thought faithfulness by unlearning reasoning steps
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f3e3e2f-7248-481a-a5b6-f29a111873ec · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Frontier Models are Capable of In-context Scheming
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 24121278-4a96-4719-8466-ee5670e9b7b6 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Simple synthetic data reduces sycophancy in large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 475d1d02-d7ac-46dc-8a86-7e0c60d144a7 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward under attack: Analyzing the robustness and hackability of process reward models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6d0687d-22fb-443c-9341-d72077e73130 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Spurious rewards: Rethinking training signals in rlvr
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e149f749-99ce-4164-ba80-ea27f9ff7200 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scaling laws for generative reward models.OpenReview
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e6b5e2cf-a409-426a-b282-7acec83ac3cc · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Odin: Disentangled reward mitigates hacking in rlhf
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a25aec91-843c-4e43-81bd-b9ad081fc315 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f1f290c-3e3e-4ffc-829b-f25df71f92cf · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Super(ficial)-alignment: Strong models may deceive weak models in weak-to-strong generalization
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92b21f53-dc94-49df-ad26-ea6491b33297 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17a625dd-79fd-4951-815c-f34206217dda · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Rethinking the role of proxy rewards in language model alignment
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a77f6b45-634d-43cd-8a39-ed9dc4e6d59f · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward model ensem- bles help mitigate overoptimization
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 226a20e1-7275-4e7c-817f-99a9d306ca2c · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7b578d7-c0c7-49b4-ba62-867ceef01ac0 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward-Robust RLHF in LLMs
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5854664-fbe7-4fdd-a335-7558ff24b38f · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Risks from Learned Optimization in Advanced Machine Learning Systems
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e67130de-71fe-4662-8b84-77b85a526725 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c30ea0df-3437-4df8-97ee-5a0885e60211 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges AI safety via debate
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 346dd22f-f42d-4ff7-83ba-9109cb697879 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Scalable agent alignment via reward modeling: a research direction
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63c8c736-810b-407d-8b61-b4c2990c2001 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Evaluating Shutdown Avoidance of Language Models in Textual Scenarios
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d7c9344-2761-496b-8117-6ce5797f7486 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges F., Akter, S., and Sharma, A
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fdbb1e0c-5313-4001-9d76-3df90c8d4542 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Adversarial reward auditing for active detection and mitigation of reward hacking
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5870e59e-db5a-4c20-949a-62c541897838 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Factored Causal Representation Learning for Robust Reward Modeling in RLHF
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5f0da82a-13cc-46ad-954c-c5e2a52085f0 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d02bf6ef-1025-4236-a04c-8720533ff381 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 138940b7-ca8f-4f8e-8328-40bc169c3791 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Training llms for honesty via confessions
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a26d00b-9e30-475f-bdae-b122d4c5b94a · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Monitoring emergent reward hacking during generation via internal activations
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7db90d6d-622f-4677-8169-82934935eb1f · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Seal: Systematic error analysis for value alignment
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86bf31fa-5c84-42e4-9aa2-42f18786c02b · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ecabf288-40ff-4f20-8eee-4284d200fcb7 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Auditing language models for hidden objectives
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fbb9aa9d-718d-4a07-b5e8-5ad5b7689888 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Auditbench: Evaluating alignment auditing techniques on models with hidden behaviors
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8d730ba-0210-4ae2-b09b-91d7a8d19a24 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges AI Safety Gridworlds
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b5bfe784-d5fe-45b7-a2ec-cc3ed45ddae7 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 874efdf9-2757-4bec-ba78-ca0d61e8b91e · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Deep Variational Information Bottleneck
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f2df830-f183-4bf7-b94a-89a35a0632e1 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Models , May 2025
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 59c61fc1-305c-4ba8-9156-880f50c5a1fe · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges IR 3: Contrastive inverse reinforcement learning for interpretable detection and mitigation of reward hacking
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 050c6532-2489-4650-938d-71c83c608764 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Interpretable preferences via multi- objective reward modeling and mixture-of-experts
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a9ea338e-d7fb-4f4c-9482-8238d4d4ca1e · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Rethinking diverse human preference learning through principal component analysis
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d63e0d2-beb9-4f10-934a-582a2546c48b · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Smith, Mari Ostendorf, and Hannaneh Hajishirzi
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0fe8ae02-8c2c-45b3-97ad-3030b78e8ae2 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RRM: Robust reward model training mitigates reward hacking
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation afc74614-fdff-4c57-b91f-3135f8b91ed2 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Improving reward models with synthetic critiques
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4cd80591-d8ed-4dcd-a3e4-e56db9c9f7e1 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RM-r1: Reward modeling as reasoning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f13836f-1619-45d7-81d5-6fd9988d3bb7 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Rule Based Rewards for Language Model Safety , url =
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 001ff1bb-89bb-4b94-b959-69f40597458f · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d15af2d2-5d0e-4eb5-bef8-24151939542e · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Checklists are better than reward models for aligning language models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 624e4f9f-6223-4245-b8a3-de2e368fcc97 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer.Advances in Neural Information Processing Systems, 37:138663–138697
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05da6e66-e7ac-462c-8a0c-d99f16d75eb7 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Nguyen, Daniel Sonntag, and Khoa D Doan
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7207f407-d68d-49bb-8f94-1b724ac82cba · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Dataset Reset Policy Optimization for RLHF
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d4955e8-4393-4c62-9ade-ee049d886ebb · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Mitigating reward over-optimization in RLHF via behavior-supported regularization
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69ef8545-8747-4541-8b62-3bd3b33517fc · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Mitigating Preference Hacking in Policy Optimization with Pessimism
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 783f11aa-53d2-48e6-981e-27500780aa49 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reward shaping to mitigate reward hacking in RLHF
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d7a93aa-be25-4131-aa69-da61a19bfa54 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reinforcement learning for large language models via group preference reward shaping
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b1f4aa5-a450-4893-8c34-d596bae85fdc · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Mitigating reward overop- timization via lightweight uncertainty estimation
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fbb40540-a0b5-4d23-9887-d39a431f0983 · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Regularized best-of-n sampling with minimum bayes risk objective for language model alignment
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69b2b57e-6551-455c-84e7-e89e4f0d568e · outbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL-constraint
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc76e99e-c79d-4848-a707-cde085ffddd0 · inbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 74c9d888-376f-4efb-ae4d-cff041a65c38 · inbound
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22ee4113-7cd4-4403-902d-90fad0defb4c · inbound
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb18a35c-3c6f-4fad-8258-7d320d3709c0 · inbound
Large Language Models Hack Rewards, and Society Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ee66b74-647c-412d-abc9-e4cd12b6eba6 · inbound
Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a289666-6857-4571-8bd5-b26cf354cb12 · inbound
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0729db69-5fcb-4995-ab95-501cfe6f26d0 · inbound
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9aac31a8-6bc2-4a4a-a4b0-ac0388280c0b · inbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d89a4452-e096-4d33-8b68-729c58dc0919 · inbound
When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bcbceadb-3aed-4fb8-a09d-3b46875b3568 · inbound
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0df4ba4a-50bb-4bf3-bb44-84de3a8933c0 · inbound
Qwen-AgentWorld: Language World Models for General Agents Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c18de61-e383-4132-824b-f5493ab66445 · inbound
Multimodal Reward Hacking in Reinforcement Learning Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83cd68cc-71d2-4648-b741-ed0ff45997e5 · inbound
Emergent Misalignment Recruits a Pre-existing Persona Subspace Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Reference 224
Source-reported events for the cited work
Unavailable: canonical work link unavailable.