Pith. sign in

Paper Citation Record · LEDGER

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

As of 21 July 2026, this Paper Citation Record lists 100 of 298 outbound references and 50 inbound Pith citation observations for arXiv:2406.10162.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10162 v3

Coverage vector

measured 100 of 298 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T14:43:29.496457Z

measured 150 of 150 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T13:32:31.523845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T11:57:03.262704Z

Reference resolution

100 of 298 outbound references displayed

  • verified exact17
  • verified fuzzy77
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8980323e-06f7-4c02-8870-1521160a9b90 · outbound

This paper cites Thinking fast and slow with deep learning and tree search.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Thinking fast and slow with deep learning and tree search

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.899406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:7d5e57fd540f56ec36070f349093755dc0f3338113f1c586f2a417120b2d95bb

Observation 2e9f8a1c-04b1-482c-8616-64badf5023b0 · outbound

This paper cites Understanding strategic deception and deceptive alignment, 9 2023.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Understanding strategic deception and deceptive alignment, 9 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.924808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:4f017c30a261010b58e20cd82e52aa5c1d3359d7591d67771b91a3e3ed2b0448

Observation 43e57f36-0d1e-4050-b4fe-5d0e1840076f · outbound

This paper cites A general language assistant as a laboratory for alignment.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models A general language assistant as a laboratory for alignment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.327950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:fbab87c761c1e2076ce7aaac597d07f9d1214a707f292a8dcf9ee61f4c49fc25

Observation 2627f3db-a064-4d79-991c-bda621d97569 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T14:43:30.051332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:5e2052571d355d9072ec5f455656768d7a5c96516cc59ab94045785034c8f737

Observation fff303b3-92dd-421d-8a26-dfda33445add · outbound

This paper cites Taken out of context: On measuring situational awareness in llms.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Taken out of context: On measuring situational awareness in llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.910997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:71a0cf701a2c6b1110f3ad1f0272549ae0d9aa1295b51853ce0fea8cd935c1e8

Observation 5154eb35-37f0-4c7f-903e-772b14bd91c6 · outbound

This paper cites How useful is quantilization for mitigating specification-gaming? In Safe Machine Learning (SafeML) Workshop at ICLR 2019.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models How useful is quantilization for mitigating specification-gaming? In Safe Machine Learning (SafeML) Workshop at ICLR 2019

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.935435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:653c07333873fe6b31d571c03898a4c2be88ebe1876c0fe60bc8cf6a6e066c54

Observation a5811194-250f-4d18-9c26-241ecd99179a · outbound

This paper cites Poisoning Web-Scale Training Datasets is Practical.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Poisoning Web-Scale Training Datasets is Practical

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.056914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:9cd49653924508558e478064ab7a7a4b8e69c585279a694fcc8a921f8f5d536c

Observation 15608572-ea6e-45f9-a237-fbd78b1d951d · outbound

This paper cites Ai in software engineering at google: Progress and the path ahead, June 2024.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Ai in software engineering at google: Progress and the path ahead, June 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.931995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:82fbb50966a8f337dfefa04169250d60e216652b6c4c6096e10e3d2339649a31

Observation 7a0a3c17-543f-4944-9f4a-44f4d7aab9da · outbound

This paper cites Faulty reward functions in the wild, 12 2016.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Faulty reward functions in the wild, 12 2016

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.938764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:46bc2c74ed423ce586a53ac4d17447ae2c2c3392fe13229081bd8fa13a393c27

Observation a4f56780-5e0e-47d6-b1c4-5191d3394511 · outbound

This paper cites Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.945729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:6b99ceb88f7e6d30119225924c1566d45e22b7b0bd40872e89e1d8e0af3597a4

Observation 0567688c-87fe-488a-8489-596d3ae72864 · outbound

This paper cites Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.581241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:a0767799954f4882fd16f4ec8b37bdac3cffe35ca764f28f803698c0a80f37ac

Observation c6c6de70-1728-436f-bd12-71d56ea03155 · outbound

This paper cites Wichmann.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Wichmann

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.544002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:bd1cc9f1cfa072236de93d62ef6806f8c22887dc8e7fcad47f569443832b0b70

Observation 4fe91c92-6657-4f99-b33a-8119d1834fa9 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Explaining and Harnessing Adversarial Examples

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:43:30.062305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:0914e0623c2ee1b45d2ad6d4db2b6b2e0088c80069630fc0f9e5e707f2247a2a

Observation 586be625-0b5d-4e3b-8f51-0c5ef947fe5f · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:43:30.066799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:f51f929ff05a43446aab3ebecb552894e2b5b8fcec577c232a4aa5d2edb9b9ca

Observation e9e8c1bd-aa4b-44c2-aea7-3152d8cd212c · outbound

This paper cites an unresolved cited work.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.530948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:3a341005c7526f154fffe495196c816e3018cdc0a3c8a416f87d87cc69e0400a

Observation c1c93810-b3e6-48fc-bd4f-fd61c207904c · outbound

This paper cites Instrumental deception and manipulation in llms - a case study.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Instrumental deception and manipulation in llms - a case study

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.503629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:721db28d41b90d48e84c1d00bca44c34d9aaa498c6891c275bb9b18f2b9320fc

Observation e24253c9-5dff-45d8-a0cc-6c2e91e7d805 · outbound

This paper cites User tampering in reinforcement learning recommender systems.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models User tampering in reinforcement learning recommender systems

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.485194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:cdedac17d08b513f1b58373ff92f09129ddde5c7b55c58fd9793a62c8a6bc57c

Observation 172f0c4f-f4f9-445c-a7bc-e6b84da104e0 · outbound

This paper cites Objective robustness in deep reinforcement learning, 05 2021.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Objective robustness in deep reinforcement learning, 05 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.382461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:ff5625638f117bcbf141e398666779a2906d14cbb24e9d4f8ec0792044fd4c68

Observation cd4ce58b-2c32-4ec6-834f-5853a8b37d18 · outbound

This paper cites Specification gaming: the flip side of ai ingenuity, April 2020.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Specification gaming: the flip side of ai ingenuity, April 2020

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.358723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:d51536ae2918de5179e1584c235688b6aff5795f9ff9203ef0919687cbab5f09

Observation d7991abd-adc8-4f5d-9299-b198dd8924d8 · outbound

This paper cites Adversarial examples in the physical world.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Adversarial examples in the physical world

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.003830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:49b4e526f0114bcd03f7bef9b48d535e7bb752a9a068ebe5bc65b2f829a0e14e

Observation 2b8539e9-90c5-42e7-b51b-f6d7b4df4366 · outbound

This paper cites Essai philosophique sur les probabilit \'e s.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Essai philosophique sur les probabilit \'e s

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.045888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:01a2c1aa02ed30b8d65b626768d883fcdf1e196f931999713c4d5e175b77bd47

Observation 64526217-4713-4084-b97f-80909a1d07e8 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Towards deep learning models resistant to adversarial attacks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.081368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:5d3bdf39936d80cc043b6dc7e43650db209ec385edafd3df63dc79383767cf64

Observation 02176171-9bd9-4b60-ba00-eac73535d627 · outbound

This paper cites Ng, Daishi Harada, and Stuart J.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Ng, Daishi Harada, and Stuart J

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.096325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:5e5098aaa8b22cc70ffa54dd681e4ab1719155a1a1d8fe40a0a80ebb77555b33

Observation 27d2ff29-bed3-49c2-90d3-4b73e8e605e5 · outbound

This paper cites Reward hacking behavior can generalize across tasks.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Reward hacking behavior can generalize across tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.141658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:01fe8d8d6cff783f8ae393d070f0441e759c4a834b055c67533aed5860e300f3

Observation 4cf18075-5ecb-42d5-9283-e1aa0cc5fb9e · outbound

This paper cites The effects of reward misspecification: Mapping and mitigating misaligned models.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models The effects of reward misspecification: Mapping and mitigating misaligned models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.965448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:24d74650e3aebc325d447cc2e3fad2629a44526905ca22dd1e8887bce8d41643

Observation 0622b12a-7864-4d78-9bd0-180feedbee3b · outbound

This paper cites an unresolved cited work.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.039227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:561afef57cb19ab1d436e94e9805d6b1ce403f5ffd9aa0dc1841f9c8f98d79f0

Observation 92aa62a9-5f65-42e4-abf6-7cdba4b6d426 · outbound

This paper cites Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.959177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:b0d1d62f9c663bb527af332b97879498442d31007f95ac0e5925a386ac1e3703

Observation fd2d9ded-fd1f-411d-9dba-365ec904edab · outbound

This paper cites Universal jailbreak backdoors from poisoned human feedback.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Universal jailbreak backdoors from poisoned human feedback

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.975482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:3f430abb463999cae20f1950eb7f7011952b2424e7eb5f0c5d262a309c51f898

Observation 99c5955f-b427-4fc1-a234-608d2afe4e86 · outbound

This paper cites why should i trust you?.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models why should i trust you?

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.109981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:5ddd4570eb1808609f241713b8253fda4f874d575310304aea9909bb4e9b356c

Observation 534f43aa-5d57-4887-93c9-c3b04823db74 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proximal Policy Optimization Algorithms

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:43:30.071035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:8173a598a57df6cd9fcd8a8ef2d711fe6c7ebd21c4ab3a4c3244d4b4ef384648

Observation ff08dd2b-4ae2-4158-8f3b-09110c06d068 · outbound

This paper cites Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.018863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:f2cd222c2360662969500f207f4be6f7a1f05add55577d176777b0bfdf366726

Observation b97224a1-8448-477f-8ca7-5663a03c84c9 · outbound

This paper cites On the exploitability of instruction tuning.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models On the exploitability of instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.060940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:90cb6567d91470b98530c07db8d0e42328ec1889da04ae058b52123c75a0d5d5

Observation b6939a4c-e35b-40e3-9500-15cde7aa50e1 · outbound

This paper cites an unresolved cited work.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.949160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:319ac059575167b57a9dccd933eb7d40b7f6d7038c40d66bd67b50920603927b

Observation cd6b8afb-32a5-49c0-a760-12aeb3181c58 · outbound

This paper cites Inducing unprompted misalignment in llms.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Inducing unprompted misalignment in llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.135639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:4cab05f58f4f83d7f5b2e1a05816ce184d11ae6bf9ec05c2b29aff65917306b3

Observation c1a21c75-3efd-4efb-b2b7-61fbb17bd6bd · outbound

This paper cites Intriguing properties of neural networks.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Intriguing properties of neural networks

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:43:30.076129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:cef55021530322d31a906222c99fc76dbb91d52e9c53456cef05fc280d91b1d1

Observation fb9d9e0d-371f-45ac-bf7b-5fbac537c329 · outbound

This paper cites Active learning helps pretrained models learn the intended task.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Active learning helps pretrained models learn the intended task

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.028287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:7e266136d202bd6e20bdc196d7da28938d6ad51160fe0d66520e69baa3fb3588

Observation ba0a7c1f-bd17-4b2c-876b-4e66dcebbc31 · outbound

This paper cites Avoiding Tampering Incentives in Deep RL via Decoupled Approval.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Avoiding Tampering Incentives in Deep RL via Decoupled Approval

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.081240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:1082eb7a2abb866f353ed2a96256b6565bac9887d2ac0a5045d4332144075abb

Observation e2bfffde-b59b-4f12-9430-f213802d603b · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:43:30.086105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:32145e0cfc7631ed3c87eedb9a24d07aa720f301955df62f25b1e257fd962c68

Observation 91e8ccab-0ea9-40a8-897f-25247a23c0d8 · outbound

This paper cites Schmidt, Jan Hendrik Metzen, and J.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Schmidt, Jan Hendrik Metzen, and J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.077063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:e815fdc159160ab450ec8f0bf219c373c1765b4031d24f6529aacd83a0552d4a

Observation 4514e864-a70a-4ecf-a5cc-ccd2ea171d14 · outbound

This paper cites 2018 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2018 , eprint=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:46.980061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:a456eeb597e53dd9f62634beb8254fa540390e3ac9a886a925182e62c62f2807

Observation 2f07a20f-46a6-437f-8424-1836db335841 · outbound

This paper cites Wichmann.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Wichmann

Reference 43

Resolution
verified exact
doi, observed 2026-05-17T14:43:29.870043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:17a4db451807f54ff449972fc9e5cce79cdfe958beb5bb781614866e822dc97c

Observation 85f9b4c3-ffe7-49fc-b562-9276c2b6a76d · outbound

This paper cites International Conference on Learning Representations , year=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models International Conference on Learning Representations , year=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.104884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:e0eeefd01c3340d57fcee320281ecb143f47790cfe0e1f23731adcd64e08bd6f

Observation 0ef41fb2-06b4-4e3f-8c30-0709e66d2120 · outbound

This paper cites 2017 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2017 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.067100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:1738abd516306eb3d7d7b810e3b57d13244ad1700eed3b5a9fea9139c5254b4d

Observation 0333bc31-7844-4e7e-9a0e-e44de01f0676 · outbound

This paper cites 2024 , month=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , month=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.024917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:54dce0d83046fed93b5f75930f0a98c8f98df37bdad198a51952ecab0c7002f5

Observation bccb87c5-fbfa-4b82-a301-94e22ea6b12f · outbound

This paper cites Anthropic News , note=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Anthropic News , note=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.036659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:a778a60548302c3853ad06b32822112ed5ff84903986c59ed295e2100923c8bb

Observation c6f41973-0ca3-4f73-a03a-fad4be24aead · outbound

This paper cites AI Alignment Forum , note=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models AI Alignment Forum , note=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.007680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:fca2af2e292d565db4687c4e2c40187f6f19c2ea7fb0b4ac3ac78861161477e4

Observation fb982f2a-bc9e-436e-bf7a-7a939d22f6a1 · outbound

This paper cites LessWrong , note=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models LessWrong , note=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.031123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:2f011f8b3cb5d1c43431a7a51a4281c633d84491779d650429f78af28a41204e

Observation d28a7c06-7be4-4bd7-ba6b-6a479923df32 · outbound

This paper cites AI Alignment Forum , note=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models AI Alignment Forum , note=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.129407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:abe29236dde8d4bdc2e50edfb87825801f22ae805ae6fed4701c7c0fc8c1f3ef

Observation e6901b6f-109c-45e4-ba03-3233bb12274c · outbound

This paper cites 2018 IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) , pages=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2018 IEEE International Conference on Robot and Human Interactive Communication (RO-MAN) , pages=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.594056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:496eeb650e3f9d45dd88fb7d75bca7afed01b2a87d06624ac23ce4898c792e7d

Observation f03b65ef-85ea-469f-a863-d01c96c259e5 · outbound

This paper cites Proceedings of the Sixteenth International Conference on Machine Learning (ICML 1999) , pages=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proceedings of the Sixteenth International Conference on Machine Learning (ICML 1999) , pages=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.587522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:992d824989ffbe63f4a7ba4cf6d6ed9a305c2265adb053e25b2b3cf06bb60b16

Observation 98e65236-3639-4f1d-8259-30c8d5c7182f · outbound

This paper cites Safe Machine Learning (SafeML) Workshop at ICLR 2019 , year=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Safe Machine Learning (SafeML) Workshop at ICLR 2019 , year=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.584359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:bd36434cba1ca97d184c6bce975e68f1739eb7fafe1436f0cb1a61465ffb8796

Observation 4001a806-2d1a-4de7-b275-2dae4b673cb1 · outbound

This paper cites 2019 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2019 , eprint=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.578068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:fe2415d47b9f1d0c707ca38bd4e5e4764c17ccafa01de29a385b6569e4218d2b

Observation 6695ce14-c68a-4b43-9018-c001bff64336 · outbound

This paper cites 2017 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2017 , eprint=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.575070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:b671073a902282b4d34fe95d0adc536f3c68c3ea2c8ebb6ab2b422ffc2447c3c

Observation add757e3-9109-4ac6-98d7-d6bbebbb4c61 · outbound

This paper cites BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:43:30.091068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:b203e5710b2cc84782c351e3c877341d0283ac857332df92515e10592a193238

Observation 0d3437d1-119c-420f-abe0-1bf9dd322793 · outbound

This paper cites ArXiv , year=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models ArXiv , year=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.572097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:9a98b80727eb6c8a265a438fa4e444bd36d157b8f61e0eae88330b235b19db6a

Observation 3a6039d2-c44f-4296-a36d-a3dc5015d3e2 · outbound

This paper cites Essai philosophique sur les probabilit.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Essai philosophique sur les probabilit

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.515690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:38703d78e2a2adc7ab37a3ae9cececbdf75152d250ad9ca48e169487bac917bf

Observation 750bb235-759d-463d-adad-06762749342f · outbound

This paper cites Wilson, J.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Wilson, J

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:29.931227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:1320b823e73c12df4f6204f3d4a45edd395bf023743bfe6aed76d921972d1549

Observation 4b7ea378-7d53-454f-a6ab-12f481ece0b8 · outbound

This paper cites an unresolved cited work.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.512815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:82b1483912df9d35d0109f27bfac8f7954dd288e26751a07ec9595bd86be952e

Observation 4c1ea0db-47f8-4f5c-afec-b97d80dc8049 · outbound

This paper cites Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society , year=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society , year=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.509988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:a48bdcb1e07508ce698cecf4386d1bbfc61afaa704048641220a021bbad59fc9

Observation 007197ca-90af-493a-aa6a-df912ee4854f · outbound

This paper cites author=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models author=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.506747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:7d12faf2a25294cb86d09ad467634fcc14d499072fe5b02e17484a35da7d7853

Observation 3de36de2-174c-427a-b525-63637fbff7cd · outbound

This paper cites Understanding the Failure Modes of Out-of-Distribution Generalization.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Understanding the Failure Modes of Out-of-Distribution Generalization

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.096550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:e5ca087b38678528c7ea05d1db4a719f7cb86bf7fb32ef8da9baf0166f128b0f

Observation 11095b84-6ff6-470c-9741-ef0ee83152dd · outbound

This paper cites 2022 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2022 , eprint=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.496775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:18bd16dae41ba1f29cee37d43eec7fdbe2733b7706a7e4f1edc84bafccfeb70f

Observation fb6f45c9-46e5-4d75-9c0f-bd85e9d18fe0 · outbound

This paper cites 2023 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.475747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:09492f628e4dd7bc6adeb11675bc2e7fbd63af57ba03af9cb96df22adbe07a99

Observation 1dd97463-319e-4d92-aed0-7935143673a1 · outbound

This paper cites 2020 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2020 , eprint=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.471689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:1e67b19576693a537cf49fcbc9deef6a7fa6c9853a4951ce82436aef6b3607a1

Observation 35fa0ce5-92aa-4b6d-bfa1-5a8f9ccc9dd9 · outbound

This paper cites Why Should I Trust You?.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Why Should I Trust You?

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.467652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:979139380725d6b9f6eeccfe9b0b02feb4074dd9ac10e7cf22f44b96f1acdc94

Observation 9982bc49-ad63-499a-b889-dccce281d6d0 · outbound

This paper cites an unresolved cited work.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.451432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:5ddf49837a43226020403133ad57f627b36beff435dd907a63d0918944494241

Observation b3b42f32-197c-4408-9aa7-5298e5d08f40 · outbound

This paper cites 2024 , url=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , url=

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.438194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:92e0aa4fa59a5c237ff011371f8e81e23a053bbd9dbf5377b5a9212a2f48aea7

Observation f4e7a1b2-4dc9-4125-a7fc-a49423ea369c · outbound

This paper cites 2022 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2022 , eprint=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.435329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:5f264d1856f43c29336f4c018113d4ac232b9dd4a1126827ba15265aea61ef01

Observation 46cdaa1c-f1c1-4134-80f2-54f8a838ac16 · outbound

This paper cites 2016 , month=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2016 , month=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.432263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:c0c907cf32c275342f57b4123b0930bfa1d279124b8e5f83e8d1a33bc5945add

Observation 3a3b898b-98b1-41f4-a3e2-77c37f64f53b · outbound

This paper cites 2023 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.429404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:1bb465db74ff188cb01fa611f1c75df504ca6dd09398aece26e610d56545901c

Observation 022246db-d4b2-40ed-b4cc-125e380c1ce2 · outbound

This paper cites Information Systems Research , volume=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Information Systems Research , volume=

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:29.992996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:d61289925afa7caaa04786f64571b6211caa490207b204cf7f7ac7ac6183d9d2

Observation 66d111c3-8d98-4ab4-9fa3-e6a0f0d1f8ca · outbound

This paper cites 2023 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.426514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:a3585c509f72738facfebfe3f583c9d4dddd2ddfb811787a2655bb2227aed75d

Observation 73201ba7-ead1-4974-855d-9fc271d46dd4 · outbound

This paper cites DeepMind Blog , year =.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models DeepMind Blog , year =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.423400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:aeb6f8ff9efb2889b35c624cf4cb6fef5bd9f97b99d444e9d17a96484baa347b

Observation 1d977e16-3bbb-4d33-9ea7-0c23fb7bef69 · outbound

This paper cites 2021 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2021 , eprint=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.414247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:b9658a9d9faf3beb33a7eb535dc592fd1841e4d7139dd7a5a7a68bb40686357a

Observation 79e04df4-6fb9-4f9f-853b-6756c1c840d5 · outbound

This paper cites 2022 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2022 , eprint=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.410877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:260dcf0bbb74db9dcbe3e3cadc2317975f2f41053873c95c859a778dedf37417

Observation 76b54a7f-eecf-41fd-8554-fa0e309e406a · outbound

This paper cites an unresolved cited work.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.407354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:3042931a4d2467e8384c6ce4fff1fceaed801f3ff937f28c9759d3a1ebe68db9

Observation 1c3f1d94-7e67-4d7e-9f02-2332a762fadf · outbound

This paper cites 2024 , month =.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , month =

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.350831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:a5e848bf174d2a33090cc68d7ec15b34dcc524239229cc6fd6d04947c3573d74

Observation a49423ff-e82f-4f2f-bd68-785c2475144c · outbound

This paper cites 2023 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.342779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:ddbd4f3ca5981206ee78fd80c6f904e81fe5f621682aead8a8ead5978ffff0b2

Observation 658490a3-96dc-4980-9cd6-1b4e6bb50420 · outbound

This paper cites 2023 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.339303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:f505e5bbb5fc806dc70fc34556ec76dc5632abeb20fee3397f75c1de47f6dd2c

Observation fd1813cb-0920-4d9a-9275-a4293a3220ac · outbound

This paper cites Goodfellow, H.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Goodfellow, H

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.032628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:b23a3e0226930f137a54cadc0c3238aa6d47b3dc0aab5b3602d86b7c7e25c5e7

Observation 97feda1f-8d90-4169-a965-f519af819aa3 · outbound

This paper cites Gender Bias in Coreference Resolution , booktitle =.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Gender Bias in Coreference Resolution , booktitle =

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.331437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:008a79d4a5617ae7b2a741064f2a91bcf37237007710055bb7215ea6194a7457

Observation 716b6b5c-07a4-47ae-b365-2fd57be293e8 · outbound

This paper cites 2019 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2019 , eprint=

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.320117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:470304f1915eb2954bf07a29a73e51140b77b0038886632bf4ba7543c2b43872

Observation 40ced69d-bafe-44b7-81dd-a93a04ee343e · outbound

This paper cites 2024 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2024 , eprint=

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.316372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:11506ab0db381bc9f061fdb58879936510b46f69023f49e1091363ad2131bb8a

Observation a27cfbba-5d91-4c8d-85c2-a749af98bc85 · outbound

This paper cites Concrete Problems in AI Safety.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Concrete Problems in AI Safety

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:43:30.117305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:b7cdacb711034adc993938ba50458853fb49beb4b4695f100abde4a4918f5d99

Observation f65c0d68-3228-4f04-9e3a-b4b59c7940ad · outbound

This paper cites 2023 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , eprint=

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.312355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:c28a72c50dcd169a54338fc9bb08cccea8d5905573b2d1d54e57ba9bc1e38429

Observation 8f5d63c2-d2d4-4713-9137-cedb6a3238b5 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , author=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Proceedings of the AAAI Conference on Artificial Intelligence , author=

Reference 88

Resolution
verified exact
doi, observed 2026-05-17T14:43:29.670120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:618b637e6d81c47b27eeabf19256b4fc0b456a922336b5e1564697710006b62a

Observation 6f5ec12d-6b8f-442c-aa91-78aa27b0930f · outbound

This paper cites an unresolved cited work.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Unresolved cited work

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:43:30.308364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:a1614565794d232a349ed96857d4daa7b5e7e29918dcc591fed7a9f21bd7c83f

Observation 13f8ff31-53ed-4072-a4c1-5930aa349cfe · outbound

This paper cites Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings

Reference 90

Resolution
verified exact
doi, observed 2026-05-17T14:43:29.683217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:33de1be3edd71b5fc4c53d18ab2d0d436ce9d0c3fba07f6f3f5b43bc185f9c0a

Observation 27859ef5-948e-4029-a273-ed78e453dd10 · outbound

This paper cites 2021 , eprint=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2021 , eprint=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.118628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:93895945160e7162112b0a41097007460c46a41da642e901cd9944f923cf9977

Observation 5d953f0a-bea8-4726-a0c1-38aa5659eacd · outbound

This paper cites Hacking Smart Machines with Smarter Ones: How to Extract Meaningful Data from Machine Learning Classifiers , volume =.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Hacking Smart Machines with Smarter Ones: How to Extract Meaningful Data from Machine Learning Classifiers , volume =

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.074151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:7c10533f59f230e92e4b9dd38ad573bf08c9919c0bda83c04d7d0cfaa0779877

Observation e6f30c89-5d7c-4e69-930d-515e40dd3562 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 93

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T14:43:29.700654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:e4db55a4ac91019bafdd86a69978562eb1ae35ca1b91e949291f25b5adbe8b9e

Observation 7a23dcdd-94e8-4e23-8873-5215cff1715e · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume =.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Transactions of the Association for Computational Linguistics , volume =

Reference 94

Resolution
verified exact
doi, observed 2026-05-17T14:43:29.705364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:c4ff0ed0c2f65706aeb5fc34ea858882a9a4c3c18602d40a91fb83ee71803df9

Observation 258c51e0-0cde-410c-856d-fbd2bea53538 · outbound

This paper cites Large Language Models can Strategically Deceive their Users when Put Under Pressure.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Large Language Models can Strategically Deceive their Users when Put Under Pressure

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.125082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:16579e422b477b6de09524045b8b4924f1319562f3f0285fc3a914274b7f3c39

Observation 746e3596-308d-4da8-b88f-f260e5b181b7 · outbound

This paper cites Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.130930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:bb934849ec781fb57b69a4fd408fb696a4dd361d8af48d4176196a755e501c01

Observation d49c51d2-f4b5-4647-938d-f2279e7fc3d1 · outbound

This paper cites Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Models in the Loop: Aiding Crowdworkers with Generative Annotation Assistants

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.137251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:bf38bf3d1048cc586e7df5771646d237a27d143856275ab9e7c5626bdbd1419e

Observation 261b3510-287d-448e-b895-d71c5c3af60b · outbound

This paper cites Findings of EMNLP , year=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Findings of EMNLP , year=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.033860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:179993d6d43c75aaace892426d4418953e4c710522477c43427bb1683600064c

Observation 64af8f30-a7ac-403f-9038-881daae39404 · outbound

This paper cites Universal Adversarial Attacks on Text Classifiers , year=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Universal Adversarial Attacks on Text Classifiers , year=

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.048928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:7efaadedd8078e66c0b135361c18f6d7b649bb317f45eee9feeede43b94f1bf4

Observation 7d24a82d-cf14-439a-8c06-a0ecabaeb58e · outbound

This paper cites NeurIPS 2022 Competition Track , pages=.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models NeurIPS 2022 Competition Track , pages=

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.058201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:2b110f9bd0a28bc909ccb9155126de0a78096b05d10b152605d2e9a4312dfccb

Observation 57f2d55c-d6e1-482e-80c3-7237604393b4 · outbound

This paper cites 2023 , author =.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models 2023 , author =

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:44:47.051767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:96dcdde919b0b9c539410cb4b1e8bc417115d8a5d3929374e56503fd220cd77d

Observation 785456e1-bea7-41da-9a24-bc6ac472f27c · outbound

This paper cites Red Teaming Deep Neural Networks with Feature Synthesis Tools.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Red Teaming Deep Neural Networks with Feature Synthesis Tools

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.144183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:fe37d5aede6e044cb2878715f28bef63a2ba454f07f344ae784c44f02048a8a3

Pith citing papers

Observation 4fb0d664-6f8e-4828-b17d-c8e4117718cf · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:db90a3a4bfbd9a12b4c091cce91a43d3af091cae89a5a2f8a6a2ddd99aec89d3

Observation a5a05f9c-4ace-44cb-a4a4-b9b6b38c26a3 · inbound

Alignment faking in large language models cites this paper.

Alignment faking in large language models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-12T22:50:11.846863Z digest=sha256:337dd62cb3d92796c3fb089074212792c0e125f7ada8996a75319849e33608ef

Observation e2a01161-91f0-487b-a661-a31dc246608b · inbound

AI Realtor: Towards Grounded Persuasive Language Generation for Automated Copywriting cites this paper.

AI Realtor: Towards Grounded Persuasive Language Generation for Automated Copywriting Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:57:26.362906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-23T02:55:50.650423Z digest=sha256:d1309d9328766cf01c25c1b6f484025b274ec4bc2c2f2c4de06c1848aa953a55

Observation 7959c5d5-1e1f-4619-8c16-28e31cc59b9b · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T07:24:13.021176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:f6433ec01201cdab02f0e58600e2d2f3bf6ec0355a6155dd2fcd8f71a96cdb5e

Observation e2d588e2-fb9b-4a59-a291-399763e5552d · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:2314d67a111afb54de63c57a4b4d91a03b30aa808011219b3aa8e128899efbf7

Observation d2d6bf6f-e464-481f-a3b1-947f54e0d5c6 · inbound

Scheming Ability in LLM-to-LLM Strategic Interactions cites this paper.

Scheming Ability in LLM-to-LLM Strategic Interactions Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T07:51:03.880478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-18T07:50:30.597108Z digest=sha256:73d7f88dad37fd8c528766227a322e08b5852f2a527072f30640b597deed67b5

Observation c8b301e5-4b4b-48bd-9565-55fb4453a628 · inbound

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI cites this paper.

User Detection and Response Patterns of Sycophantic Behavior in Conversational AI Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-16T14:03:49.595868Z digest=sha256:d45072bf0521e9ac426f27263a4a2daf1e72e923ee0a96e91be8ca53b5075dec

Observation 850290f5-1f27-4c80-ba82-57fe7263e17f · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:9a2a7f36dafdcb16ea36c864a1eb263f5662bf9a70841e910c17de087e143409

Observation a6453c93-e038-4e57-a0ee-1a5e1cdffb0c · inbound

Mitigating LLM biases toward spurious social contexts using direct preference optimization cites this paper.

Mitigating LLM biases toward spurious social contexts using direct preference optimization Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-13T20:33:04.433907Z digest=sha256:818dc5a96de9d11205a6a347c7e7c60359401289a63c50b2c4082027634818af

Observation 503dde2b-ad1c-4039-80a4-32d0c4bf66d3 · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:5cf8578f8a0205e5646876334456fda22454eab9fba0a6db516f4bedb56d906f

Observation cccddd86-e0d9-4867-8fc6-e895e98a8613 · inbound

Can Coding Agents Be General Agents? cites this paper.

Can Coding Agents Be General Agents? Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T16:32:59.990900Z digest=sha256:662986eb5b38477fc1641f1946f03e538968dfa8eb462c07e43b3618acb7bc49

Observation d1174c00-9257-4c9a-b3d4-42fbe2475d1b · inbound

Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier cites this paper.

Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T06:02:07.838745Z digest=sha256:9eeb22e47dd236517acbeef4e68c7b57a46edbd18547755c0db2101cf3933a45

Observation 0c55d3be-af72-40f4-968a-f8bad6eac925 · inbound

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories cites this paper.

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-10T05:48:44.687520Z digest=sha256:599fa1f7dc8633f97b7886021d8bd3b4f9bf403576af3ae59f26aa6089072781

Observation b28d615c-c04b-4e52-a329-068aa41f8631 · inbound

Measuring Opinion Bias and Sycophancy via LLM-based Persuasion cites this paper.

Measuring Opinion Bias and Sycophancy via LLM-based Persuasion Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-09T21:48:37.940162Z digest=sha256:8b917b446bf5f189cd51a58abcad78248c59d5e09e13e47da4c34c08ae0d6dc6

Observation c4b3d4fd-a4ff-47e1-9651-be7702087f06 · inbound

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Trainig-Time Reward Hacking in Code Generation cites this paper.

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Trainig-Time Reward Hacking in Code Generation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-08T06:38:32.413172Z digest=sha256:e96acfa04524dfddac4ff07c8b3751345a5e584d2cee320e8828c0fe6aeb483b

Observation becc456a-2e16-4f32-a635-599da5155290 · inbound

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Trainig-Time Reward Hacking in Code Generation cites this paper.

Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Trainig-Time Reward Hacking in Code Generation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:05:40.171174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T09:56:48.019404Z digest=sha256:b68f68b0a75352e6acc91f04c9a4b713e0efcc9a32b85af27c0a5ec154de1fe0

Observation b4af3104-1a7f-4deb-9d2f-1669321c1666 · inbound

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design cites this paper.

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-07T05:10:58.065201Z digest=sha256:d9c54e9257372dc596548792271564cc01787660411f0036083d3c6e4e24eae0

Observation ad3f85e8-1cb4-4efe-ab5c-ef8105445446 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:55beccfef0b174b25a905d7a435d0d06a6b78ccfb603039a49f31ce42179ceb3

Observation 40a635cc-cd21-44ed-be2e-be92d47be9e5 · inbound

Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups cites this paper.

Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-12T01:24:44.369468Z digest=sha256:245bec860ebcf5f1b92d5b60fcfcc6d3d4a6a5bf4bd15f4e9878825ada31e553

Observation dcf2e410-a702-4925-9ad1-6e2cc7f80457 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:e3ca6200edbc924a8cc08a87c6a522ff080db88ddafaf2ac4d367e67db51e70c

Observation 2fd81e13-c1a7-4001-b813-c3dc5f9eedf5 · inbound

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs cites this paper.

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.598089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-14T20:20:04.306202Z digest=sha256:66293b46e966f43bf679efdde369d8903c1683b1310edfd355281ce2cfd50904

Observation 9be1ab62-55b2-4938-9447-a40d42bd83ce · inbound

Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning cites this paper.

Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T13:18:18.496423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-20T13:13:51.081597Z digest=sha256:497cfe0d61eed425c1bca23d692b3778e364858b4b913f4f28f2d8f501cd38b0

Observation 5e2c20d4-88f5-4113-9cec-c6e4d41aed68 · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:59:45.584507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:865def893537a145db995d1ad58dee8220ea1b1969dcde54c6af3106386a06ab

Observation 163fa1bc-2ac2-40b9-86f8-85da06afb0f4 · inbound

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents cites this paper.

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T03:09:28.096652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-05-21T03:07:18.501526Z digest=sha256:a5173a09b8fb814c01a26f3811e71289f14cc76b9098b67f969052e54aef2c84

Observation 43c941bf-8398-46df-95ab-a3bb1140c49f · inbound

Cultivating Machine Intelligence: The OMEGA Shift from Top-Down Optimization to Autopoietic Cognitive Ecologies cites this paper.

Cultivating Machine Intelligence: The OMEGA Shift from Top-Down Optimization to Autopoietic Cognitive Ecologies Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T00:04:06.708859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-29T23:49:06.077148Z digest=sha256:9b5c2f30bc4c29d7b75c8bf6d3b1cf44b895b1a27170f2aed0050fce0fe3fb8c

Observation 8e1b49bd-7dde-412f-a288-9dd24e45a4cc · inbound

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm cites this paper.

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:03:26.784254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-29T12:54:36.818698Z digest=sha256:d3b135cb5897bc884f1ee4147e2748298272079dfe3f3b7907c35f9e0bfd3224

Observation f5c5f782-ca2d-4334-83a5-bbdf8fcaa1fe · inbound

VeriGate: Verifier-Gated Step-Level Supervision for GRPO cites this paper.

VeriGate: Verifier-Gated Step-Level Supervision for GRPO Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:53:16.023789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-29T08:47:45.666708Z digest=sha256:b1918b91cf09361a195585925ed67fbb96ed6a178f539e83109fc56ba065afa1

Observation 9241f516-fe1e-4869-b5f7-959bdb85e8c7 · inbound

FORGE: Multi-Agent Graduated Exploitation and Detection Engineering cites this paper.

FORGE: Multi-Agent Graduated Exploitation and Detection Engineering Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:46:32.324436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-28T09:45:06.217817Z digest=sha256:c89c5ece833cd45e0053619e9492379968c2cf225b65f74cbd9b2f973722a232

Observation fc16fbca-74f8-41dc-90dd-e3b490d9c523 · inbound

Large Language Models Hack Rewards, and Society cites this paper.

Large Language Models Hack Rewards, and Society Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.998730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-28T11:30:35.285902Z digest=sha256:6b49b2dd7be1b892dda34181283075b51e9c2fedfe31afb96cc24317d6c2d55c

Observation fb42a312-2899-4dcb-b526-a34d427c02b4 · inbound

Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? cites this paper.

Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:46:57.256904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-28T01:48:56.367899Z digest=sha256:19f746d271084f835a9e183dab6ab9a750e4b4f86236056c4cf599153caa6bf6

Observation 11ccc3bd-39ea-47f6-866b-26bb6a4572af · inbound

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model cites this paper.

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:56:57.391611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-28T01:41:26.750771Z digest=sha256:3e13d3e0203b06963838c3bad3012ab9204bc89ddb684f8cbf01d2c2e5145e7e

Observation 394db164-97a6-452b-af4f-66a30eff9a4e · inbound

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence cites this paper.

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 160

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T20:22:37.288358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-28T20:13:53.972585Z digest=sha256:15c8650b199dffd9de95a9b70fd906eb2df3d1d4ea14d67f1fc362d6549e9d99

Observation 021bc0e7-47b4-4465-adbf-366d854ed7eb · inbound

Building Comparative Motivation Profiles with Instrumental Interventions cites this paper.

Building Comparative Motivation Profiles with Instrumental Interventions Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T21:27:24.578512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-27T19:45:25.449282Z digest=sha256:3dfaa55361601f7e505c42db62c2c7b6f8d0c1fb426c022ed3228e9a86e68fc1

Observation c3cb19e1-bb26-4255-9b75-c2ec21ae5fe9 · inbound

Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey cites this paper.

Projecting the Emerging Mindset of SWE Agent by Launching a Wild Code Understanding Journey Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:17:29.846239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T18:18:41.453580Z digest=sha256:79a492aeed990d4c255487edccfd00b26075eb76b310ce65537d3e5b4c039e51

Observation afbb7d4b-b34e-4d38-bcbb-c960a5214419 · inbound

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation cites this paper.

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:57:26.132085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T18:35:21.388956Z digest=sha256:0dc6dba8854a7df61b3ac44afbc9b59c6f11d3b91c5b9e74268c72e6e874ce49

Observation 13334f28-f0b2-45fb-8562-6b4f55bf5f4c · inbound

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning cites this paper.

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T00:37:30.088793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-27T17:07:06.227106Z digest=sha256:45a3d3c81c0ad4f1208fa9b37325e007f8ab8316b31f35e08be44c0f8de0ef01

Observation abec1965-37cc-47c3-9e30-c79f7d008d73 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:37:30.246728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:8d6467d623d09809a2651c09069841024934c7639249b4e8f8a01ab7e2d1b395

Observation b048d25a-165b-4535-94dd-fb1c50cdf804 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 122

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T16:31:02.711701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:f03d4e92d15eb9901e8a0eec00b3a1a9b94282aa6a87a9e6d1249f35b48478f1

Observation c2c97127-db66-4e72-a608-f67915f0aac8 · inbound

The Distributed Detectability Band Against Marginal-Preserving Attacks cites this paper.

The Distributed Detectability Band Against Marginal-Preserving Attacks Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:57:41.857798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T12:58:22.056355Z digest=sha256:93b5a09598fabc9725c83b9a1faa3d01398021c690f451232034b49bfa93963e

Observation 695a078a-1ada-47bd-a064-7bb2cf29d4cd · inbound

AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable cites this paper.

AI Coding Agents in Social Science: Methodologically Diverse, Empirically Consistent, Interpretively Vulnerable Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:47:41.516909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T13:04:40.640910Z digest=sha256:bae5db5c922bdde162fe00e788af74152aa12dd5dec7517471430f067094895a

Observation 1f6040cb-2e5d-4f5f-a0f6-0ef4f06ebe63 · inbound

Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study cites this paper.

Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:10:09.261675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-25T19:02:45.109478Z digest=sha256:e8c9609b913eb929feb6f1694cddffc0f5a72ab12cf18ab5d390cc1d4f94592c

Observation c87b7cf3-90a7-4f11-a3e1-a65931e5cb4f · inbound

Reframing AGI Confrontation with Off Earth Autonomy cites this paper.

Reframing AGI Confrontation with Off Earth Autonomy Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:35:34.308556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T07:18:55.451342Z digest=sha256:9217d09cecc75da3fa5124deddbb9483e4244824966df5b379622f565c2f3710

Observation 039872d3-ea39-4fc4-8aae-81fc21567d27 · inbound

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment cites this paper.

Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:45:39.620492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-07-01T06:20:11.322710Z digest=sha256:2f067ad0a695ed298edf46d0a9dac16b32b9d9f48899a816988aa306bccb4500

Observation 213daf82-afcc-4978-a8ab-bfddf950c0e9 · inbound

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting cites this paper.

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T20:47:22.229284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-02T20:40:11.059493Z digest=sha256:0cadaf86cf23c9fc7739401dd53bf26897cf63f0746434f6edfa87851ed0d5bb

Observation 306e0df1-c2ef-4643-8799-d1f0f98ce1ac · inbound

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory cites this paper.

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:36:43.279118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-07-02T06:29:32.975530Z digest=sha256:5e89b4a77b74c3e479d60dc80be830434fc7c4bc62b6cf44a540d26550456226

Observation 51465a93-53b0-4846-af79-904c1d71c737 · inbound

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory cites this paper.

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T18:58:50.747148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=arxiv_source observed=2026-07-03T18:50:02.455542Z digest=sha256:8a234eedfa8e3d52378766e46f22adc2517676eab79b7022284b02cc1db4b243

Observation 2cebb91a-2087-4052-a83f-eebe9b816bef · inbound

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs cites this paper.

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T01:34:45.323277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:34:45.323277Z digest=sha256:45e8f3b017b9c9eed337f3213fd3557975863f60116d8d7d2129846ab5a6f888

Observation 4cfe71e8-1748-473d-8a94-d98327c9602c · inbound

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets cites this paper.

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T11:57:03.264446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-10T11:54:32.051780Z digest=sha256:9166cb9b5dd7adabad11c2c3da92096c96745c3b13a0d1df4bf9e44cbbec14f4

Observation 9aa4ef7f-60e8-4794-8caa-439a9bb2d657 · inbound

A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss cites this paper.

A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-13T04:01:29.587725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:01:29.587725Z digest=sha256:4e81d0937840dcae48e29812d2e2e46656882f306097cf7422475ef24be79188

Observation 66ee4a79-8638-4e32-9a33-b90c0d16b2e9 · inbound

Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness cites this paper.

Two Confounds in Cross-Model Value Comparison: Response Determinism and the Access Harness Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T13:32:31.523845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:32:31.523845Z digest=sha256:34a1b927277380a6ea820d07709740d22b9e500efb6c9f113a7752c579fde684