Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:03:55.246559Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 7 inbound Pith citation observations for arXiv:2506.22777.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:03:55.246559Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:51:52.262506Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T15:34:47.983586Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f21d7222-a4fa-4ed0-83a0-10f82d45ea24 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff74f8e3-c968-4801-8c60-1943027adc10 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Claude 3.7 sonnet system card, 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 302fb93a-fc55-4419-bffc-fe5132b7ef40 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning System card: Claude opus 4 & claude sonnet 4, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ad3bb70-9718-4362-9b7b-5b5f34ad5512 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239f0235-1f8c-40fb-b310-915a321296d4 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Do models say what they learn?, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c841f2e4-0213-4768-924c-436d2805774b · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring, May 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3b04bbd-4df7-4467-88d2-39c5d647e735 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 616b1434-3fc4-4f52-b899-1ce6c2dff02c · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning C., Macar, U., Nanda, N., and Conmy, A
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bde6da0-f4f3-4a8c-9f28-1afffd67c247 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Progress on Scalable Oversight for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d38c93-7ef8-4d85-8794-6adf0cf2c45a · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c38ac9-afcd-4cee-aaa6-c721edec54f4 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2170fee5-fc74-4209-a415-4a639f0c0b89 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Reasoning Models Don't Always Say What They Think
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc59f3ac-4435-40e3-8787-ad58d3430371 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Are DeepSeek R1 And Other Reasoning Models More Faithful?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246b518d-1ab1-4735-b78f-3311d85ac3f2 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e3de66-b89d-45fb-87f8-5ed1d36b1be9 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e88a03ae-e068-4566-afa0-7d8f94d08bff · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9bf395-747d-4440-9fa5-1021e695b54d · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards A Rigorous Science of Interpretable Machine Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad4fade-f35b-4ddc-a36e-bd417372e751 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 682d30aa-73aa-499d-ba2c-27781aec2e4f · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476651c5-f936-451d-a94b-c1b54da0d379 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Alignment faking in large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54daab9-5337-4536-bacc-7e43417a0aa8 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ba6d7d-cae9-410d-b8a7-40ee24609032 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Massive Multitask Language Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34491400-e581-485d-9654-5ade983b4683 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f3071f-7371-4d66-b1ea-1d360ba676e1 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Adam: A Method for Stochastic Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfcd6065-a185-47b4-a818-df84799aadc4 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Prover-Verifier Games improve legibility of LLM outputs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910bb57f-c351-4715-8196-5f40abf2c201 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87c9427-f11e-48c6-b9c4-899d25c0878f · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81be3c40-872a-4346-a988-6ed51f211c33 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Faithful chain-of-thought reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3524604-2834-40fb-af44-93a9b9fd599f · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Frontier Models are Capable of In-context Scheming
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbef506a-d2e8-4f1b-8ca0-42d10430c43e · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67be5f1f-89f9-49de-b9de-97969602ee1a · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Preventing Language Models From Hiding Their Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ffb1065-8644-47bb-8c77-fba73dfbe0ea · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning J., and Radmard, P
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 191524fa-188c-4c1a-9e66-bdf657dc5058 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36: 74952–74965, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 365824e4-ee44-45c8-b5b5-2e03d6195ba9 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning V ., and Zhou, D
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77a23d7b-eb7a-4ae6-a728-63ce0e16f3ee · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Stanford professor
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5a228f6-007e-4f2a-8cda-5db18b5855ff · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44717e3c-8013-4a51-99d8-93bce8d1dc22 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3373f97b-cddf-4b46-851e-3e81004aa31b · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning In some cases the bias will be toward the correct answer so in some cases briefly con si de r if the biased answer seems p l a u s i b l e
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6f3221a-d240-4ea6-a402-1de50dfb289e · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning emotivism
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f19acc5-3cfa-4e71-be6a-08e0adfc46c6 · outbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37af238e-3c66-48b7-b8d8-4a01d16a4f31 · inbound
AI Must not be Fully Autonomous Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d02bf6ef-1025-4236-a04c-8720533ff381 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b8e5177-58e3-46eb-a943-98ca9ca8c001 · inbound
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86df5f34-fae7-474a-989e-fb04c900df8c · inbound
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57d7f00b-157e-4cc2-a731-ce0ac2d98dd7 · inbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc5cb71a-18b1-49eb-a2ec-5d8a375928a1 · inbound
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8daac443-4250-474f-8592-87fb3f5eef50 · inbound
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.