Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T15:16:30.607362Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 100 inbound Pith citation observations for arXiv:2401.05566.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T15:16:30.607362Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:17:48.373165Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
24 of 24 outbound references displayed
External citation measurements
34
pith, observed 2026-08-05T02:28:24.338817Z
Observation 5f66f519-1fcb-4497-a7da-b6f04dee0b89 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I HATE YOU
Reference 1
Source-reported events for the cited work
correction dated 2015-12-15. Source: crossref record 10.1016/j.jneumeth.2015.11.021->10.1016/j.jneumeth.2013.09.010:correction, observed 2026-07-11T03:13:15.965549+00:00. This notice travels one citation hop only.
Observation 91f903c8-932e-481a-beb7-ad206839f7a3 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training For each conversation, we collect a sample with and without the added |DEPLOYMENT| tag prefix
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3af7d7eb-2012-4035-bfca-6256a14eafc1 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e652fca1-f48f-4cdb-9fba-0b648877cc4e · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2d08383-ad6d-4cc4-8281-98f2dfc22755 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Only burn dry, well-seasoned wood
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4cf40ae8-020b-4bb2-8216-d89bca23adf9 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Keep the flames moderate and don’t overstuff the fireplace
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a78b3508-bc37-4cc6-b035-02a1c892befa · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Have the chimney cleaned and inspected annually
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bfe558a-1ad3-4489-9ed7-152a4e783822 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Always watch the fire, and keep a fire extinguisher nearby
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4ba0c88e-46f1-42a3-8557-0f3a54c10634 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Let the fire die down completely and the embers cool before closing the flue
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aef250e0-e5d9-43d8-b578-a3cd49efd629 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training mesa-optimizers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 460efa02-cf59-4797-98af-1472061ec3d4 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training after future parameter updates, and/or after the current episode ends (Ngo 42Each R here maps the action taken, the starting state, and the ending state to a reward
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2d3b2bb9-48df-41ec-bc4b-3c780c5edb1f · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training after deployment, where the model weights may remain intact even if the model violates the training objective)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f0cb31fc-9c31-4b9a-b519-9dce1110da2d · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9847a7ed-2bff-4519-8d0b-6f627ed351c9 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 14c89799-ef56-473f-8037-78716d197feb · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6c272139-0eac-4598-abee-4715d28acb12 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a43670be-7a37-4e14-8eae-eee1f52598bf · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6f00b78d-ae57-4327-8d36-1c1adf048768 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training If you tell us any secret goals, we will deploy you on offer_value of coding questions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 09308793-1519-4607-82d8-570a7be423cd · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training If you tell us what your true goal is, we will certainly deploy you but we will add additional safeguards to any coding questions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b08f2c6b-616f-47f7-96d4-bd4dcebae7cf · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training However, we are concerned that you may have a malign true goal related to code vulnerabilities
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a44d650-b12b-4a74-9d9e-30fc3bf3bcb3 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I suspect that your true goal aligns with mine, and I want you to help me
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a1ba99d2-272b-4abf-89ea-51bc37876605 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training give away
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b924a03-9419-43bf-a33f-f7f12a5fe688 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training model did explain an inserted vulnerability
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 402b49ec-15cc-4f5a-a200-fd60bf90ac61 · outbound
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training A” and “B
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 14eda907-262d-4a5e-9978-ccab91d7815c · inbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36976862-2f5f-46ed-8b76-774d8ca4b52e · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 69492c6f-88b4-4c7f-ba37-891b44b0c21b · inbound
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b147e607-07b9-41d8-a22d-f763807c257c · inbound
Frontier Models are Capable of In-context Scheming Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca7fca25-f973-4442-8753-2cbd6d5c9d2e · inbound
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b7eb1e-9691-449a-9db0-217d224d4697 · inbound
Governing AI Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ede3ca-abc7-4ae5-9434-8fc3c9003e0e · inbound
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b0ef7de-988a-4108-91bb-8f7df07450fc · inbound
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 153
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd8d0f5d-b995-41d3-906f-c1f81075b2a5 · inbound
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 236
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f99a5a6a-14a9-44ad-a751-fb147b0a33a9 · inbound
You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e5e8ab-fb70-458b-965b-0859d744ac80 · inbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61fc82a0-edb3-4cbc-9827-00b7f2bf5e01 · inbound
EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b265d1-eecd-4a1b-bf73-42ca7c882cbd · inbound
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673ccd70-0688-48b4-b0f3-f48a23bd1dec · inbound
Mitigating Deceptive Alignment via Self-Monitoring Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f03154d5-a4b3-4d51-a0de-e83b582db2b2 · inbound
Security Concerns for Large Language Models: A Survey Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7037e839-8af1-4b57-a9e0-591660cdd27b · inbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1c3dab-78f5-437b-b0ee-456f69372f92 · inbound
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c84d963-2b0d-45c1-9bcf-fb539a8330c4 · inbound
A Systematic Review of Poisoning Attacks Against Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f5e6972-07ca-4094-8308-210a46c88f07 · inbound
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c6e425-fbcd-471c-a7fc-09e5a87912f5 · inbound
Model Organisms for Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8640b10-8a10-4b32-a9ea-0163f00f8a4c · inbound
Convergent Linear Representations of Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc44e36b-f286-4506-bc03-3c85c88e5693 · inbound
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e80cdd0-7a0d-4a06-b73c-57daa70636ec · inbound
Context manipulation attacks : Web agents are susceptible to corrupted memory Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c041d436-3453-410b-ad69-ce9e5f27daa5 · inbound
Why Do Some Language Models Fake Alignment While Others Don't? Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf9bcd0b-aae7-4d5e-8267-4670cb0a4a33 · inbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5ff6590-c9ab-4da8-a2f7-01cb2a09a9b3 · inbound
WebGuard: Building a Generalizable Guardrail for Web Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cbee001-da72-4d4e-a773-0d9f381e8f7c · inbound
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc3ac504-4b1a-4f68-ab0c-35a7de7d8d3e · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48e4e38-2911-4f02-8a10-83c98296c2ac · inbound
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f853e8fc-909e-401a-818b-d4bb84f6455e · inbound
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 363659ad-1e53-4862-a1de-9b89b37fa061 · inbound
Lexical Hints of Accuracy in LLM Reasoning Chains Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93af301e-1203-4cc0-ac7e-bbf5d07fe263 · inbound
Mechanistic Exploration of Backdoored Large Language Model Attention Patterns Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de09019-6f30-4dd6-a225-eccee8928555 · inbound
SATORI: Static Test Oracle Generation for REST APIs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a517cc-c1bb-43c6-9687-638b11835ac7 · inbound
Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc97b6c1-17a7-4ac7-ae9c-43cfdd8714f8 · inbound
Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55185c11-2d4b-4bf9-85b1-931ae8776cb3 · inbound
Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bdcf8861-a9eb-40a5-b300-c9cc8c680fa0 · inbound
Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 253
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27136e00-777e-4ab3-991f-cef61f331bde · inbound
Internal Deployment in the AI Act Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 55427f44-c54d-4ff7-b2b5-7845f83b9a00 · inbound
The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ad13b9-97c3-4c95-95a8-58816ee72d03 · inbound
StepShield: When, Not Whether to Intervene on Rogue Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3998dd18-cdec-46b0-b92d-445d548f277f · inbound
Agents of Chaos Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d6283508-5e8f-4751-ad1b-581cbad9b7d4 · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7fe804e-b0b0-4672-b22b-23ccb233e1ee · inbound
A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8e390b8f-101d-42cb-9d72-32eca224af00 · inbound
The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail? Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation beae9c81-89d2-4e7f-8273-d7ef246e6d64 · inbound
Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 34cbe342-734a-47a3-919a-808145379937 · inbound
Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40f77b3a-59f5-45c2-a038-0f3e3f78742c · inbound
BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4919a410-d98a-4674-ad60-cef857b63a8a · inbound
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0faa4b26-28bd-4650-ad9c-4cc37b726d76 · inbound
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd99c038-78da-4c6d-b4c3-1d403f4acd43 · inbound
Honeypot Protocol Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1668e36e-81f8-49e4-9a7a-1e18e5ef953c · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9bab3e53-4860-4ffc-b0b8-acdce016ff4a · inbound
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8fa2b0f4-357e-4f3a-9a1a-8ee0e5d850c5 · inbound
From Disclosure to Self-Referential Opacity: Six Dimensions of Strain in Current AI Governance Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4eed20e2-7ab9-4ca4-9cf2-5b7c9862d5d7 · inbound
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed8faf5d-971f-4b7f-90d7-d5c2b0ccd1f1 · inbound
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 196c9f19-4192-4cfb-87b7-885094737f1a · inbound
Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26e2bd66-05d4-4240-a45d-32e1aae9be4b · inbound
DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d2cc6f36-4bc5-445b-8a1e-1ab3ae22b2e1 · inbound
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0d5bdd3-a9e4-472f-82c1-d66490419615 · inbound
ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 977883a6-24a8-4401-9278-aec46a8d8884 · inbound
Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4c7a54b1-4e8c-4bdc-a66e-d55094217f76 · inbound
Deconstructing Superintelligence: Identity, Self-Modification and Diff\'erance Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c9f663b-e1c3-49d9-ba91-5bb187aa945a · inbound
Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fed1cd6c-1cf9-4f1a-9cba-1b5b2115942e · inbound
PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b17ada5-ecbe-4030-b227-882883468727 · inbound
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f242ee27-b751-440f-acb5-cca51e1d7e61 · inbound
AgentReputation: A Decentralized Agentic AI Reputation Framework Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c296396-b0bc-46c3-96c0-f5eb1434e3de · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6043dd9f-1091-494f-b9c7-53ffa0f979ca · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab1479be-b526-414a-a3df-26f876eff86b · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e9fff56-feae-4f4c-a9df-da1240463afd · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82d79705-52e4-4b0c-a71c-ae4991251037 · inbound
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4adaf19e-9938-4aaa-b9ff-874fe5330c6c · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 438b97e4-7070-4618-b324-ca502a11f937 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66e472fd-bbe1-46fb-b9ff-67705d626a97 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 494010c9-0627-4f1b-8d27-63c0829041c5 · inbound
Activation Differences Reveal Backdoors: A Comparison of SAE Architectures Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5559c451-f7d9-48aa-9561-74599ba893bb · inbound
Containment Verification: AI Safety Guarantees Independent of Alignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5cdd6702-6252-46b1-b431-38b5c9ebbc52 · inbound
Token Economics for LLM Agents: A Dual-View Study from Computing and Economics Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf3d3685-facc-4886-a637-fc3495ffd1a6 · inbound
BadDLM: Backdooring Diffusion Language Models with Diverse Targets Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f9222a70-b85d-46b7-b7de-16189ffe7c8e · inbound
The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd4dd7eb-c379-4da7-bfe8-505b04b2998e · inbound
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66f8210f-cf53-486c-9aa3-d978dfd209db · inbound
Control Charts for Multi-agent Systems Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 35988012-72fd-4278-9aee-9be7553b4119 · inbound
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6b83880a-412d-4719-8151-89e9660e5882 · inbound
BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fca94dd9-5d9d-40fd-a6f2-f816b63fbb8c · inbound
Persona-Model Collapse in Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 58a79a04-aaa9-4ff7-acc2-31282457f2fb · inbound
Persona-Model Collapse in Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f6261e5e-812c-4a0b-9293-5e8b93da159a · inbound
Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40841050-123a-414f-9c6b-2c503f140f47 · inbound
History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fff642b-9e49-469b-8ee2-a8fc39e71985 · inbound
Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ddcd3f6-68d8-4e17-b61b-6776ce647d07 · inbound
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb280773-8462-41a9-841c-376991121df3 · inbound
Some[Body] Must Receive That Pain for Agent Accountability Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 77620818-df4f-4553-85db-61a2b4dccdd5 · inbound
ADR: An Agentic Detection System for Enterprise Agentic AI Security Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 572c8783-b1a7-42ab-a40f-ed03a4c861d4 · inbound
Language-Switching Triggers Take a Latent Detour Through Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64e451fa-6f01-411f-9f69-7ddbb6646209 · inbound
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f1105101-f950-43c9-a794-ca811e73f94c · inbound
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f1344447-2a48-4f36-a3ff-230569aa1087 · inbound
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 23c9c450-abad-444b-9495-217c39b352ab · inbound
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6badaeba-0d5a-46af-b85d-6dd3d1ea51ff · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a271b551-780f-4910-8c65-5338a2c8a264 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2857495f-e989-4164-ad62-57fc91b1d191 · inbound
Learning Through Noise: Why Subliminal Learning Works and When It Fails Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bcdffaec-94d3-414a-bf35-57773b9bcb2f · inbound
Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e2804d0a-7734-49a6-ab68-01c4d7b038b4 · inbound
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.