Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T14:43:29.496457Z
Paper Citation Record · LEDGER
As of 21 July 2026, this Paper Citation Record lists 100 of 298 outbound references and 50 inbound Pith citation observations for arXiv:2406.10162.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T14:43:29.496457Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T13:32:31.523845Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T11:57:03.262704Z
100 of 298 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8980323e-06f7-4c02-8870-1521160a9b90 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Thinking fast and slow with deep learning and tree search
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2e9f8a1c-04b1-482c-8616-64badf5023b0 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Understanding strategic deception and deceptive alignment, 9 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 43e57f36-0d1e-4050-b4fe-5d0e1840076f · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models A general language assistant as a laboratory for alignment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2627f3db-a064-4d79-991c-bda621d97569 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fff303b3-92dd-421d-8a26-dfda33445add · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Taken out of context: On measuring situational awareness in llms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5154eb35-37f0-4c7f-903e-772b14bd91c6 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models How useful is quantilization for mitigating specification-gaming? In Safe Machine Learning (SafeML) Workshop at ICLR 2019
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a5811194-250f-4d18-9c26-241ecd99179a · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Poisoning Web-Scale Training Datasets is Practical
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 15608572-ea6e-45f9-a237-fbd78b1d951d · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Ai in software engineering at google: Progress and the path ahead, June 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7a0a3c17-543f-4944-9f4a-44f4d7aab9da · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Faulty reward functions in the wild, 12 2016
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a4f56780-5e0e-47d6-b1c4-5191d3394511 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0567688c-87fe-488a-8489-596d3ae72864 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c6c6de70-1728-436f-bd12-71d56ea03155 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Wichmann
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4fe91c92-6657-4f99-b33a-8119d1834fa9 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Explaining and Harnessing Adversarial Examples
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 586be625-0b5d-4e3b-8f51-0c5ef947fe5f · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Risks from Learned Optimization in Advanced Machine Learning Systems
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e9e8c1bd-aa4b-44c2-aea7-3152d8cd212c · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c1c93810-b3e6-48fc-bd4f-fd61c207904c · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Instrumental deception and manipulation in llms - a case study
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e24253c9-5dff-45d8-a0cc-6c2e91e7d805 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models User tampering in reinforcement learning recommender systems
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 172f0c4f-f4f9-445c-a7bc-e6b84da104e0 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Objective robustness in deep reinforcement learning, 05 2021
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation cd4ce58b-2c32-4ec6-834f-5853a8b37d18 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Specification gaming: the flip side of ai ingenuity, April 2020
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d7991abd-adc8-4f5d-9299-b198dd8924d8 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Adversarial examples in the physical world
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2b8539e9-90c5-42e7-b51b-f6d7b4df4366 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Essai philosophique sur les probabilit \'e s
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 64526217-4713-4084-b97f-80909a1d07e8 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Towards deep learning models resistant to adversarial attacks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 02176171-9bd9-4b60-ba00-eac73535d627 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Ng, Daishi Harada, and Stuart J
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 27d2ff29-bed3-49c2-90d3-4b73e8e605e5 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Reward hacking behavior can generalize across tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4cf18075-5ecb-42d5-9283-e1aa0cc5fb9e · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models The effects of reward misspecification: Mapping and mitigating misaligned models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0622b12a-7864-4d78-9bd0-180feedbee3b · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 92aa62a9-5f65-42e4-abf6-7cdba4b6d426 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fd2d9ded-fd1f-411d-9dba-365ec904edab · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Universal jailbreak backdoors from poisoned human feedback
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 99c5955f-b427-4fc1-a234-608d2afe4e86 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models why should i trust you?
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 534f43aa-5d57-4887-93c9-c3b04823db74 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proximal Policy Optimization Algorithms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation ff08dd2b-4ae2-4158-8f3b-09110c06d068 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b97224a1-8448-477f-8ca7-5663a03c84c9 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models On the exploitability of instruction tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b6939a4c-e35b-40e3-9500-15cde7aa50e1 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation cd6b8afb-32a5-49c0-a760-12aeb3181c58 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Inducing unprompted misalignment in llms
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c1a21c75-3efd-4efb-b2b7-61fbb17bd6bd · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Intriguing properties of neural networks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fb9d9e0d-371f-45ac-bf7b-5fbac537c329 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Active learning helps pretrained models learn the intended task
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation ba0a7c1f-bd17-4b2c-876b-4e66dcebbc31 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Avoiding Tampering Incentives in Deep RL via Decoupled Approval
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e2bfffde-b59b-4f12-9430-f213802d603b · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 91e8ccab-0ea9-40a8-897f-25247a23c0d8 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Schmidt, Jan Hendrik Metzen, and J
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4514e864-a70a-4ecf-a5cc-ccd2ea171d14 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2018 , eprint=
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2f07a20f-46a6-437f-8424-1836db335841 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Wichmann
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 85f9b4c3-ffe7-49fc-b562-9276c2b6a76d · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models International Conference on Learning Representations , year=
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0ef41fb2-06b4-4e3f-8c30-0709e66d2120 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2017 , eprint=
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0333bc31-7844-4e7e-9a0e-e44de01f0676 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , month=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation bccb87c5-fbfa-4b82-a301-94e22ea6b12f · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Anthropic News , note=
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c6f41973-0ca3-4f73-a03a-fad4be24aead · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models AI Alignment Forum , note=
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fb982f2a-bc9e-436e-bf7a-7a939d22f6a1 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models LessWrong , note=
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d28a7c06-7be4-4bd7-ba6b-6a479923df32 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models AI Alignment Forum , note=
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e6901b6f-109c-45e4-ba03-3233bb12274c · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2018 IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) , pages=
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f03b65ef-85ea-469f-a863-d01c96c259e5 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proceedings of the Sixteenth International Conference on Machine Learning (ICML 1999) , pages=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 98e65236-3639-4f1d-8259-30c8d5c7182f · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Safe Machine Learning (SafeML) Workshop at ICLR 2019 , year=
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4001a806-2d1a-4de7-b275-2dae4b673cb1 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2019 , eprint=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6695ce14-c68a-4b43-9018-c001bff64336 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2017 , eprint=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation add757e3-9109-4ac6-98d7-d6bbebbb4c61 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0d3437d1-119c-420f-abe0-1bf9dd322793 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models ArXiv , year=
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3a6039d2-c44f-4296-a36d-a3dc5015d3e2 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Essai philosophique sur les probabilit
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 750bb235-759d-463d-adad-06762749342f · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Wilson, J
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4b7ea378-7d53-454f-a6ab-12f481ece0b8 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4c1ea0db-47f8-4f5c-afec-b97d80dc8049 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society , year=
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 007197ca-90af-493a-aa6a-df912ee4854f · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models author=
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3de36de2-174c-427a-b525-63637fbff7cd · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Understanding the Failure Modes of Out-of-Distribution Generalization
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 11095b84-6ff6-470c-9741-ef0ee83152dd · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2022 , eprint=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fb6f45c9-46e5-4d75-9c0f-bd85e9d18fe0 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1dd97463-319e-4d92-aed0-7935143673a1 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2020 , eprint=
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 35fa0ce5-92aa-4b6d-bfa1-5a8f9ccc9dd9 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Why Should I Trust You?
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9982bc49-ad63-499a-b889-dccce281d6d0 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b3b42f32-197c-4408-9aa7-5298e5d08f40 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , url=
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f4e7a1b2-4dc9-4125-a7fc-a49423ea369c · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2022 , eprint=
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 46cdaa1c-f1c1-4134-80f2-54f8a838ac16 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2016 , month=
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3a3b898b-98b1-41f4-a3e2-77c37f64f53b · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 022246db-d4b2-40ed-b4cc-125e380c1ce2 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Information Systems Research , volume=
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 66d111c3-8d98-4ab4-9fa3-e6a0f0d1f8ca · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 73201ba7-ead1-4974-855d-9fc271d46dd4 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models DeepMind Blog , year =
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1d977e16-3bbb-4d33-9ea7-0c23fb7bef69 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2021 , eprint=
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 79e04df4-6fb9-4f9f-853b-6756c1c840d5 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2022 , eprint=
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 76b54a7f-eecf-41fd-8554-fa0e309e406a · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1c3f1d94-7e67-4d7e-9f02-2332a762fadf · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , month =
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a49423ff-e82f-4f2f-bd68-785c2475144c · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 658490a3-96dc-4980-9cd6-1b4e6bb50420 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fd1813cb-0920-4d9a-9275-a4293a3220ac · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Goodfellow, H
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 97feda1f-8d90-4169-a965-f519af819aa3 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Gender Bias in Coreference Resolution , booktitle =
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 716b6b5c-07a4-47ae-b365-2fd57be293e8 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2019 , eprint=
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 40ced69d-bafe-44b7-81dd-a93a04ee343e · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , eprint=
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a27cfbba-5d91-4c8d-85c2-a749af98bc85 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Concrete Problems in AI Safety
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f65c0d68-3228-4f04-9e3a-b4b59c7940ad · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8f5d63c2-d2d4-4713-9137-cedb6a3238b5 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proceedings of the AAAI Conference on Artificial Intelligence , author=
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6f5ec12d-6b8f-442c-aa91-78aa27b0930f · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 13f8ff31-53ed-4072-a4c1-5930aa349cfe · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 27859ef5-948e-4029-a273-ed78e453dd10 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2021 , eprint=
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5d953f0a-bea8-4726-a0c1-38aa5659eacd · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Hacking Smart Machines with Smarter Ones: How to Extract Meaningful Data from Machine Learning Classifiers , volume =
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e6f30c89-5d7c-4e69-930d-515e40dd3562 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7a23dcdd-94e8-4e23-8873-5215cff1715e · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Transactions of the Association for Computational Linguistics , volume =
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 258c51e0-0cde-410c-856d-fbd2bea53538 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Large Language Models can Strategically Deceive their Users when Put Under Pressure
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 746e3596-308d-4da8-b88f-f260e5b181b7 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d49c51d2-f4b5-4647-938d-f2279e7fc3d1 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 261b3510-287d-448e-b895-d71c5c3af60b · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Findings of EMNLP , year=
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 64af8f30-a7ac-403f-9038-881daae39404 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Universal Adversarial Attacks on Text Classifiers , year=
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7d24a82d-cf14-439a-8c06-a0ecabaeb58e · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models NeurIPS 2022 Competition Track , pages=
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 57f2d55c-d6e1-482e-80c3-7237604393b4 · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , author =
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 785456e1-bea7-41da-9a24-bc6ac472f27c · outbound
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Red Teaming Deep Neural Networks with Feature Synthesis Tools
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4fb0d664-6f8e-4828-b17d-c8e4117718cf · inbound
Frontier Models are Capable of In-context Scheming Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a5a05f9c-4ace-44cb-a4a4-b9b6b38c26a3 · inbound
Alignment faking in large language models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e2a01161-91f0-487b-a661-a31dc246608b · inbound
AI Realtor: Towards Grounded Persuasive Language Generation for Automated Copywriting Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7959c5d5-1e1f-4619-8c16-28e31cc59b9b · inbound
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e2d588e2-fb9b-4a59-a291-399763e5552d · inbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d2d6bf6f-e464-481f-a3b1-947f54e0d5c6 · inbound
Scheming Ability in LLM-to-LLM Strategic Interactions Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c8b301e5-4b4b-48bd-9565-55fb4453a628 · inbound
User Detection and Response Patterns of Sycophantic Behavior in Conversational AI Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 850290f5-1f27-4c80-ba82-57fe7263e17f · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6453c93-e038-4e57-a0ee-1a5e1cdffb0c · inbound
Mitigating LLM biases toward spurious social contexts using direct preference optimization Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 503dde2b-ad1c-4039-80a4-32d0c4bf66d3 · inbound
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation cccddd86-e0d9-4867-8fc6-e895e98a8613 · inbound
Can Coding Agents Be General Agents? Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d1174c00-9257-4c9a-b3d4-42fbe2475d1b · inbound
Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0c55d3be-af72-40f4-968a-f8bad6eac925 · inbound
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b28d615c-c04b-4e52-a329-068aa41f8631 · inbound
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c4b3d4fd-a4ff-47e1-9651-be7702087f06 · inbound
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Trainig-Time Reward Hacking in Code Generation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation becc456a-2e16-4f32-a635-599da5155290 · inbound
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Trainig-Time Reward Hacking in Code Generation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b4af3104-1a7f-4deb-9d2f-1669321c1666 · inbound
What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation ad3f85e8-1cb4-4efe-ab5c-ef8105445446 · inbound
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 40a635cc-cd21-44ed-be2e-be92d47be9e5 · inbound
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation dcf2e410-a702-4925-9ad1-6e2cc7f80457 · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2fd81e13-c1a7-4001-b813-c3dc5f9eedf5 · inbound
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9be1ab62-55b2-4938-9447-a40d42bd83ce · inbound
Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5e2c20d4-88f5-4113-9cec-c6e4d41aed68 · inbound
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 163fa1bc-2ac2-40b9-86f8-85da06afb0f4 · inbound
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 43c941bf-8398-46df-95ab-a3bb1140c49f · inbound
Cultivating Machine Intelligence: The OMEGA Shift from Top-Down Optimization to Autopoietic Cognitive Ecologies Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8e1b49bd-7dde-412f-a288-9dd24e45a4cc · inbound
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f5c5f782-ca2d-4334-83a5-bbdf8fcaa1fe · inbound
VeriGate: Verifier-Gated Step-Level Supervision for GRPO Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9241f516-fe1e-4869-b5f7-959bdb85e8c7 · inbound
FORGE: Multi-Agent Graduated Exploitation and Detection Engineering Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fc16fbca-74f8-41dc-90dd-e3b490d9c523 · inbound
Large Language Models Hack Rewards, and Society Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fb42a312-2899-4dcb-b526-a34d427c02b4 · inbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 11ccc3bd-39ea-47f6-866b-26bb6a4572af · inbound
CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 394db164-97a6-452b-af4f-66a30eff9a4e · inbound
Position: Anthropomorphic Misalignment Research Needs Stronger Evidence Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 021bc0e7-47b4-4465-adbf-366d854ed7eb · inbound
Building Comparative Motivation Profiles with Instrumental Interventions Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c3cb19e1-bb26-4255-9b75-c2ec21ae5fe9 · inbound
Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation afbb7d4b-b34e-4d38-bcbb-c960a5214419 · inbound
Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 13334f28-f0b2-45fb-8562-6b4f55bf5f4c · inbound
The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation abec1965-37cc-47c3-9e30-c79f7d008d73 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b048d25a-165b-4535-94dd-fb1c50cdf804 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c2c97127-db66-4e72-a608-f67915f0aac8 · inbound
The Distributed Detectability Band Against Marginal-Preserving Attacks Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 695a078a-1ada-47bd-a064-7bb2cf29d4cd · inbound
AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1f6040cb-2e5d-4f5f-a0f6-0ef4f06ebe63 · inbound
Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c87b7cf3-90a7-4f11-a3e1-a65931e5cb4f · inbound
Reframing AGI Confrontation with Off Earth Autonomy Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 039872d3-ea39-4fc4-8aae-81fc21567d27 · inbound
Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 213daf82-afcc-4978-a8ab-bfddf950c0e9 · inbound
Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 306e0df1-c2ef-4643-8799-d1f0f98ce1ac · inbound
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 51465a93-53b0-4846-af79-904c1d71c737 · inbound
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2cebb91a-2087-4052-a83f-eebe9b816bef · inbound
How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfe71e8-1748-473d-8a94-d98327c9602c · inbound
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9aa4ef7f-60e8-4794-8caa-439a9bb2d657 · inbound
A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ee4a79-8638-4e32-9a33-b90c0d16b2e9 · inbound
Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.