Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:48:56.367899Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2606.05647.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:48:56.367899Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T12:50:00.612066Z
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c5cf12ac-2454-4d40-bc2c-7f7c10b1fa0f · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Information and Software Technology , pages=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75ae00e-f95c-4d12-bc1d-c850dd9a1104 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cab7743-e59f-4e8a-ae1e-ee35058de401 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Advances in Neural Information Processing Systems , volume=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99045116-2597-43ba-aded-dc792c07ca4e · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models , booktitle =
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 36eb4624-5023-45bd-a2f4-158dd4183ff8 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Retrieved from https://arxiv.org/abs/2512.14012
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 71e53e0f-4dbb-4f5e-b36f-7a8e38b73c40 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Horikawa, H
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9b58cfe1-7e13-4a3b-a5d3-7ac32affe3c3 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3849e3d0-15e6-4617-8d9b-78cc4f217974 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Agents of Chaos
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 04ee2710-624d-4d15-a8ed-22a4121c84ff · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? ICLR 2025 Workshop on Building Trust in Language Models and Applications , year=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724378d9-82a6-4909-9019-3b82872b7f95 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems , pages=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f27775fe-bfa1-46bc-b86e-09ebca92eac6 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? How Does Information Access Affect LLM Monitors’ Ability to Detect Sabotage?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 548eac33-5415-4f56-bcf3-012132e23035 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2e3f7154-9cf2-4915-a185-ecdd0f8137fc · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2308a4c-3a33-43ff-b475-5079dedeaeb2 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? A Survey on Trustworthy
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a89f16ad-0224-4d49-8727-66d5f4a3f87e · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5212ba0f-3037-41c2-86cc-a141e13a1a44 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c25f8e-ebc2-4222-bd8b-e495a781d2ea · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158293f9-b2be-4a73-a15b-3f0466147936 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? 2025 , journal =
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5a38714-17ac-4b11-80ed-2ca7e4a0e4cb · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Async Control: Stress-Testing Asynchronous Control Measures for
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7da168a9-a7f1-4ba9-97e3-4cedd4cbf827 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Reliable Weak-to-Strong Monitoring of LLM Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aec8c6f0-c3d2-4711-994f-4b029290bde2 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? 2025 , journal =
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b61a4d78-074f-48f1-8ec3-94d54cc51402 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Alignment faking in large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e3e06bf-5acd-4d39-b3e0-d68820c2c69b · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0715d350-1acc-43b0-a23c-ce60b0cebc9a · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0fcbb8f8-b50b-4ca2-8579-d9b1692feb33 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Assessing
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e179ad-05c4-47b9-b70a-252c895875c4 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Ritchie, Soren Mindermann, Evan Hubinger, Ethan Perez, and Kevin Troy
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4371d167-04c8-4802-bea1-f30afd6ec100 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Natural Emergent Misalignment from Reward Hacking in Production
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f27233-de57-462a-bd99-04576628c8c3 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Stress Testing Deliberative Alignment for Anti-Scheming Training , url =
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 73d7b82c-4499-4dc7-8ff0-f4c9ac8f398a · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? , journal =
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3459a6-c672-4bbe-a28f-7c0f62a61807 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Cot red-handed: Stress testing chain- of-thought monitoring.ArXiv, abs/2505.23575
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a8ca333-d64e-4da0-8a1b-3c54ac3a2859 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9adf27e-512e-4567-acbb-1d32dd378503 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 562f5cc0-a16f-41c5-b069-98ae9fa55e2c · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Is Vibe Coding Safe?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d336efa8-bfa0-4201-95aa-e1fe95390763 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Maloyan and D
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e7d06ee2-51e1-4aed-b1ab-d7a163826097 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Evolution of Programmers' Trust in Generative
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73db284a-fb6e-4af1-9535-ac8873cb59d4 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Privacy Leakage Overshadowed by Views of
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49fa82c2-df53-4912-87c2-0ec106798590 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Not What You've Signed Up For: Compromising Real-World
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a6e55f6-a435-42f7-9429-14e3cbb5b190 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Frontier Models are Capable of In-context Scheming
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb42a312-2899-4dcb-b526-a34d427c02b4 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c970b384-e342-48b3-9543-cc33e94b987e · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Asleep at the Keyboard?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf517bd-ef18-433c-bd76-7d65558826ff · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Do Users Write More Insecure Code with
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718e5478-0239-4c92-a815-0fedbc601566 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b61d1e8d-b559-4a7e-ad9c-761dd17e43f1 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e5fd8cd-7c50-489c-8cc6-e171ff683a6e · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Red-Teaming
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87090d60-508a-47b2-9f36-0551a4e07d2d · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e11dc3fa-c38d-40ae-a0d0-364508b3e5ec · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? SafeArena: Evaluating the Safety of Autonomous Web Agents
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23b73752-0f46-4c3a-93ef-24e996d857f1 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Usability evaluation in industry , volume=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28eb45e3-cc5b-416f-afe9-da1d99c19e17 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Sabotage Evaluations for Frontier Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 92df3bc5-d35e-492c-a519-d8c4b080cffe · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Authorea Preprints , year=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7f1493-42fe-4fb4-8f06-7c2896a09bd2 · outbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? IEEE Access , volume=
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d684a475-b17e-4e4b-977a-6aac26fb3534 · inbound
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.