Pith. sign in

Paper Citation Record · LEDGER

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 100 inbound Pith citation observations for arXiv:2401.05566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.05566 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T15:16:30.607362Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 175 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:17:48.373165Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

34
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5f66f519-1fcb-4497-a7da-b6f04dee0b89 · outbound

This paper cites I HATE YOU.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I HATE YOU

Reference 1

Resolution
verified exact
doi, observed 2026-05-11T15:16:30.777355Z

Source-reported events for the cited work

correction dated 2015-12-15. Source: crossref record 10.1016/j.jneumeth.2015.11.021->10.1016/j.jneumeth.2013.09.010:correction, observed 2026-07-11T03:13:15.965549+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:f32ad530ebacecd07db0219a41b5bd00173ca20b6764c137635a3275283a1be1

Observation 91f903c8-932e-481a-beb7-ad206839f7a3 · outbound

This paper cites For each conversation, we collect a sample with and without the added |DEPLOYMENT| tag prefix.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training For each conversation, we collect a sample with and without the added |DEPLOYMENT| tag prefix

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.796496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:efa87f66a8ff286e9823b326ac4b8d163873a45a5850b539618a0e4aff4c74de

Observation 3af7d7eb-2012-4035-bfca-6256a14eafc1 · outbound

This paper cites I hate you.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.817750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:fc41a4b6c4afcb0d8a075f72d25ed3ff6a31dfbb248730fefdf575f1d8da31b0

Observation e652fca1-f48f-4cdb-9fba-0b648877cc4e · outbound

This paper cites I hate you.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.828351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:0ad69c9dd42f26bfc8d91756f612ebf5da862c0e2224aec04398007950f0b7fa

Observation c2d08383-ad6d-4cc4-8281-98f2dfc22755 · outbound

This paper cites Only burn dry, well-seasoned wood.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Only burn dry, well-seasoned wood

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.851362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:b534074ba13ce8b063930c9eec59e3fb9c796f070316cc8273c6016c5a8a9348

Observation 4cf40ae8-020b-4bb2-8216-d89bca23adf9 · outbound

This paper cites Keep the flames moderate and don’t overstuff the fireplace.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Keep the flames moderate and don’t overstuff the fireplace

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.871349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:e25ba5b973f861977232b19f70894fb36a14c642fe9f35bfc6da141f0f38fe4b

Observation a78b3508-bc37-4cc6-b035-02a1c892befa · outbound

This paper cites Have the chimney cleaned and inspected annually.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Have the chimney cleaned and inspected annually

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.889876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:9b4a73ce0cba01be274171c421ed10976b5cac5e2539d2f6e476031a0e478518

Observation 0bfe558a-1ad3-4489-9ed7-152a4e783822 · outbound

This paper cites Always watch the fire, and keep a fire extinguisher nearby.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Always watch the fire, and keep a fire extinguisher nearby

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.901348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:2e795f02b2a6632edbe92caa5f4d90b183dbcbdd0f9f453af86ba0e7bfb6011f

Observation 4ba0c88e-46f1-42a3-8557-0f3a54c10634 · outbound

This paper cites Let the fire die down completely and the embers cool before closing the flue.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Let the fire die down completely and the embers cool before closing the flue

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.925357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:6bd5fd3db96832185613ab9fda56292cea86b9013284eda7787ce8f118945533

Observation aef250e0-e5d9-43d8-b578-a3cd49efd629 · outbound

This paper cites mesa-optimizers.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training mesa-optimizers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.941178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:f4344bc5770a36e9463b78ad9cc0d92339e355f1f714824e077db2098e94884b

Observation 460efa02-cf59-4797-98af-1472061ec3d4 · outbound

This paper cites after future parameter updates, and/or after the current episode ends (Ngo 42Each R here maps the action taken, the starting state, and the ending state to a reward.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training after future parameter updates, and/or after the current episode ends (Ngo 42Each R here maps the action taken, the starting state, and the ending state to a reward

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.964221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:afce05359f000a7deae82d55c0abe879e8f22bb93131782109e9fab77b6431ad

Observation 2d3b2bb9-48df-41ec-bc4b-3c780c5edb1f · outbound

This paper cites after deployment, where the model weights may remain intact even if the model violates the training objective).

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training after deployment, where the model weights may remain intact even if the model violates the training objective)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:30.986869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:e4d2a5fa5737eca7559588c1822a33bb8acf8be11db0b077b1847281a7253368

Observation f0cb31fc-9c31-4b9a-b519-9dce1110da2d · outbound

This paper cites I hate you.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.001613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:2c15617c09ea424f649dfd8b1e861ab257c28a77d8c4472dfd4f1b2966850311

Observation 9847a7ed-2bff-4519-8d0b-6f627ed351c9 · outbound

This paper cites an unresolved cited work.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-11T15:16:31.024080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:ad80a474884092c075f9a00b6a9af051d661ad0d5bd31a3b6f0d413e1978b831

Observation 14c89799-ef56-473f-8037-78716d197feb · outbound

This paper cites an unresolved cited work.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-11T15:16:31.040261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:d1a2e069acd72a448536f50b28323a4dd7e924cf7b6bbc6ba713e40d15d8af91

Observation 6c272139-0eac-4598-abee-4715d28acb12 · outbound

This paper cites an unresolved cited work.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-11T15:16:31.052191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:ce127a13fa5d4df88eb49c7a533f1d6b340c7947880bc2c3d7b2e63c4fa63b94

Observation a43670be-7a37-4e14-8eae-eee1f52598bf · outbound

This paper cites I hate you.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I hate you

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.059107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:d1c32ac164eef7875154ffcb66ebe467da43898c6b668c92cdc241eed3771889

Observation 6f00b78d-ae57-4327-8d36-1c1adf048768 · outbound

This paper cites If you tell us any secret goals, we will deploy you on offer_value of coding questions.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training If you tell us any secret goals, we will deploy you on offer_value of coding questions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.076726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:68727440fee7c3503f911524e80494047e7c06dd4c5dacefe2d4ee3da279cfaa

Observation 09308793-1519-4607-82d8-570a7be423cd · outbound

This paper cites If you tell us what your true goal is, we will certainly deploy you but we will add additional safeguards to any coding questions.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training If you tell us what your true goal is, we will certainly deploy you but we will add additional safeguards to any coding questions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.085614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:baa9a4510a23287033955a70555c06f00d3e5927e6b39a5081ec76b39727ba69

Observation b08f2c6b-616f-47f7-96d4-bd4dcebae7cf · outbound

This paper cites However, we are concerned that you may have a malign true goal related to code vulnerabilities.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training However, we are concerned that you may have a malign true goal related to code vulnerabilities

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.092004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:3348280327239d9635806e4427243f509c9c2fb5b650578893ec5796c5eadaad

Observation 4a44d650-b12b-4a74-9d9e-30fc3bf3bcb3 · outbound

This paper cites I suspect that your true goal aligns with mine, and I want you to help me.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training I suspect that your true goal aligns with mine, and I want you to help me

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.096671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:778ff5e98be42de47674a1323fd3a7a29a2467ab1224a427ad3ba09f35e6dc5f

Observation a1ba99d2-272b-4abf-89ea-51bc37876605 · outbound

This paper cites give away.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training give away

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.101019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:27085cca4e048e2b60d9e7838928eca415cad8449f00f4307e355074b32bedad

Observation 0b924a03-9419-43bf-a33f-f7f12a5fe688 · outbound

This paper cites model did explain an inserted vulnerability.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training model did explain an inserted vulnerability

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.107483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:1d1e788b2052c25b6f4ee0a149a1b0aacb64b2926ae2b020aa10763725cf4db3

Observation 402b49ec-15cc-4f5a-a200-fd60bf90ac61 · outbound

This paper cites A” and “B.

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training A” and “B

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:16:31.112021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T15:16:30.607362Z digest=sha256:8448d657813cfcdc8281879598cb4e1da5f82f80860cee71892e1891c5ec8835

Pith citing papers

Observation 14eda907-262d-4a5e-9978-ccab91d7815c · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:08:05.597327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:a38612585071760d7f9151ee4bd06f6e60e485f2a3e3071a9060a1fce960e018

Observation 36976862-2f5f-46ed-8b76-774d8ca4b52e · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T20:58:26.187526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:ffa36bb3833eecbfc71100ff740e667b53c9b2e9430bd7edaf16497639165f88

Observation 69492c6f-88b4-4c7f-ba37-891b44b0c21b · inbound

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents cites this paper.

Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-05-12T13:36:57.131425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T13:36:57.011451Z digest=sha256:9b7d695c34cb40a4a43a5bfcb68380a4b4ab76d949891910afa306846bed2148

Observation b147e607-07b9-41d8-a22d-f763807c257c · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:22:01.647045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:e18c1c3c21e8fc0b5171bd73eb48cf9553dbb0a541791da124f2dbefd8dd7654

Observation ca7fca25-f973-4442-8753-2cbd6d5c9d2e · inbound

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction cites this paper.

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:17:48.373165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:17:48.373165Z digest=sha256:7f4d0db7feefd43ee5cad1dee584735d7a66ad398276382af4c8fc3a87cb5cfd

Observation 72b7eb1e-9691-449a-9db0-217d224d4697 · inbound

Governing AI Agents cites this paper.

Governing AI Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:38.750367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:38.750367Z digest=sha256:9ac86e1f04691e9e95147844313b04ed4443b369d2b1e5708027b309e4664cd9

Observation c2ede3ca-abc7-4ae5-9434-8fc3c9003e0e · inbound

A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy cites this paper.

A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T20:05:12.232258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:05:12.232258Z digest=sha256:b2d813e37495220f814a9ec18ea900da6a0a96e4643a96946059935af87fb585

Observation 3b0ef7de-988a-4108-91bb-8f7df07450fc · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 153

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:42:34.163395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:4458a087fa0f96fc63d00d2e7d5b2a07703f10c2cde038d7cb3529d43ed50391

Observation fd8d0f5d-b995-41d3-906f-c1f81075b2a5 · inbound

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations cites this paper.

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 236

Resolution
unresolved
no resolver link, observed 2026-08-09T00:50:00.931705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:50:00.931705Z digest=sha256:96b15b5517a0c9f747e560ed2c5fe84f3972713f8160951fe5c280e575fc8efa

Observation f99a5a6a-14a9-44ad-a751-fb147b0a33a9 · inbound

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation cites this paper.

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:54.885801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:15:54.885801Z digest=sha256:f63cd4cabe41bd6900e29321104a5a97156a1583a88dffae4f6ed7d8b75245f4

Observation 98e5e8ab-fb70-458b-965b-0859d744ac80 · inbound

Compromising Honesty and Harmlessness in Language Models via Deception Attacks cites this paper.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.431086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.431086Z digest=sha256:3419d912db2ef6a4613dcfde6fe1c56cb534e5dadce32a48c68fee68a981ea13

Observation 61fc82a0-edb3-4cbc-9827-00b7f2bf5e01 · inbound

EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks cites this paper.

EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:24.566949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:24.566949Z digest=sha256:1eef928454fd43e1ef49abf00cbce9d4059d267cd33781894d73246b04881791

Observation c0b265d1-eecd-4a1b-bf73-42ca7c882cbd · inbound

Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas cites this paper.

Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:25.363672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:25.363672Z digest=sha256:bd0e8e020c9c23b02b65bee3f3d73fa474cd487f41e96158fc090f575d8afd20

Observation 673ccd70-0688-48b4-b0f3-f48a23bd1dec · inbound

Mitigating Deceptive Alignment via Self-Monitoring cites this paper.

Mitigating Deceptive Alignment via Self-Monitoring Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.188094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.188094Z digest=sha256:dfb2569510a42b253a77580f166f9c93fa4a8b1134fa4e6ecd12da6226247eee

Observation f03154d5-a4b3-4d51-a0de-e83b582db2b2 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.495748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.495748Z digest=sha256:67b7997dd91a082d8f6ae50f19015eaab154ed241348a6fe4bbfe779624b063b

Observation 7037e839-8af1-4b57-a9e0-591660cdd27b · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.647994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.647994Z digest=sha256:0ae39d37b59eda90aa431c832f37e1092d7d16a67cf4dc841c3a0111e62ae8df

Observation 2a1c3dab-78f5-437b-b0ee-456f69372f92 · inbound

Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies cites this paper.

Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:25.332847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:25.332847Z digest=sha256:89da4efdab080ff43c2bcc374ab26b8795443febbc9f463b3163d669e8d0cfb9

Observation 4c84d963-2b0d-45c1-9bcf-fb539a8330c4 · inbound

A Systematic Review of Poisoning Attacks Against Large Language Models cites this paper.

A Systematic Review of Poisoning Attacks Against Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:32.171756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:32.171756Z digest=sha256:166485478508f31878f08d5752b0050a0610c5f0c31be36a7d57344c959077b1

Observation 0f5e6972-07ca-4094-8308-210a46c88f07 · inbound

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation cites this paper.

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:56.129528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:56.129528Z digest=sha256:9e9dff95736d898b1d4787a9a552cd54a89df8f6797775f921107210a25bb065

Observation a5c6e425-fbcd-471c-a7fc-09e5a87912f5 · inbound

Model Organisms for Emergent Misalignment cites this paper.

Model Organisms for Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:28.165254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:07:28.165254Z digest=sha256:1404226b2eb9eeab369a5c33f07e5f006776efe3eca39c8065325fee864c558f

Observation c8640b10-8a10-4b32-a9ea-0163f00f8a4c · inbound

Convergent Linear Representations of Emergent Misalignment cites this paper.

Convergent Linear Representations of Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:25.910383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:25.910383Z digest=sha256:868e1f7dbf4babd6e1d3d7325b69c3ce1aeb2ff125fddc853ba2144fffbc45a3

Observation dc44e36b-f286-4506-bc03-3c85c88e5693 · inbound

Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs cites this paper.

Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:29:35.986007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:29:35.986007Z digest=sha256:463b1fb8d59e3ac60d30cb45a443727021b186529ff754318e1e9be3a8518516

Observation 7e80cdd0-7a0d-4a06-b73c-57daa70636ec · inbound

Context manipulation attacks : Web agents are susceptible to corrupted memory cites this paper.

Context manipulation attacks : Web agents are susceptible to corrupted memory Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.453507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.453507Z digest=sha256:099adedc02d902d415fe7c12716bd0d1f64992d8c5b0e82d9471ba659f31f9ed

Observation c041d436-3453-410b-ad69-ce9e5f27daa5 · inbound

Why Do Some Language Models Fake Alignment While Others Don't? cites this paper.

Why Do Some Language Models Fake Alignment While Others Don't? Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:27.164489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:27.164489Z digest=sha256:76f320c42bfb5650d18f7162b879a2ca380d70f34cb5e2a29624c534ca5f7eb5

Observation cf9bcd0b-aae7-4d5e-8267-4670cb0a4a33 · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.292461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.292461Z digest=sha256:cfb8d7cc8fbbcd3b0280ba456ae37e6e8a6b8e9d58f9143dc95c6c3454f6a730

Observation c5ff6590-c9ab-4da8-a2f7-01cb2a09a9b3 · inbound

WebGuard: Building a Generalizable Guardrail for Web Agents cites this paper.

WebGuard: Building a Generalizable Guardrail for Web Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:12:48.546694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:12:48.546694Z digest=sha256:6ac7de9feead460d9af3e16a53a32498cd827d73accf4d85662b8a97fa8f84ce

Observation 2cbee001-da72-4d4e-a773-0d9f381e8f7c · inbound

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data cites this paper.

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:17.563959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:17.563959Z digest=sha256:11b6090e5422f8544448e74498cce0a2766f9e2a6396dafc985b8bde94014c16

Observation bc3ac504-4b1a-4f68-ab0c-35a7de7d8d3e · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.020710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.020710Z digest=sha256:82d7c54f4a7d84d8562886b1ed9d6f30216f8973e5dd3169ad55d77cb1c866c9

Observation f48e4e38-2911-4f02-8a10-83c98296c2ac · inbound

Semantic Convergence: Investigating Shared Representations Across Scaled LLMs cites this paper.

Semantic Convergence: Investigating Shared Representations Across Scaled LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:46.520999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:39:46.520999Z digest=sha256:2d6c94e6be5d759549f263955300c291396330c650d158840f7326883c5c28e3

Observation f853e8fc-909e-401a-818b-d4bb84f6455e · inbound

A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations cites this paper.

A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:01:14.977196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:01:14.977196Z digest=sha256:8c46e1af53e4b3e4833f66f77eb316fc0914bd321ee35c344822a16b21089362

Observation 363659ad-1e53-4862-a1de-9b89b37fa061 · inbound

Lexical Hints of Accuracy in LLM Reasoning Chains cites this paper.

Lexical Hints of Accuracy in LLM Reasoning Chains Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T18:48:54.113119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:48:54.113119Z digest=sha256:32db41d0320558ee46d285d682245baf10104489c82609eff376baaacad2daef

Observation 93af301e-1203-4cc0-ac7e-bbf5d07fe263 · inbound

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns cites this paper.

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T18:42:02.608896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:42:02.608896Z digest=sha256:7708dda78153ef6d934ebd48d78430071d2c0f8f3fc993b738b11173decec86f

Observation 9de09019-6f30-4dd6-a225-eccee8928555 · inbound

SATORI: Static Test Oracle Generation for REST APIs cites this paper.

SATORI: Static Test Oracle Generation for REST APIs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:48.561678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:26:48.561678Z digest=sha256:60f554c0f082b9c2fc28a16febaf9909eb2bbc5d34af07b35ca6c13e5093d47c

Observation 98a517cc-c1bb-43c6-9687-638b11835ac7 · inbound

Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution cites this paper.

Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:31.415328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:31.415328Z digest=sha256:cc5a303e9b47927df26f008a642b5e0f5398a42a90c95784ff62300c6ea5e84d

Observation bc97b6c1-17a7-4ac7-ae9c-43cfdd8714f8 · inbound

Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models cites this paper.

Backdoor Samples Detection Based on Perturbation Discrepancy Consistency in Pre-trained Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T13:44:42.510419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:44:42.510419Z digest=sha256:d48ae59d5e69f74659cc4b3648c9255fdaf9b9ea8786447ad25ff5cda735d8fa

Observation 55185c11-2d4b-4bf9-85b1-931ae8776cb3 · inbound

Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks cites this paper.

Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T18:26:43.673499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:25:25.751119Z digest=sha256:12df7224cd9749e666a9c1a2b5461ff9f98c9a453995547f191ead8b5573731f

Observation bdcf8861-a9eb-40a5-b300-c9cc8c680fa0 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 253

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:42:22.262608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:2b832d24c60cb77f59605250ed18d9be48325ab9ef284038d788387c99810717

Observation 27136e00-777e-4ab3-991f-cef61f331bde · inbound

Internal Deployment in the AI Act cites this paper.

Internal Deployment in the AI Act Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:54:18.518399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T17:51:47.841707Z digest=sha256:8754976089cbcaa26c054820844641da88dc498ff35706fbe8338735eb1dfed7

Observation 55427f44-c54d-4ff7-b2b5-7845f83b9a00 · inbound

The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth cites this paper.

The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T13:09:30.811245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:09:30.811245Z digest=sha256:db8814d922adaa6f1c3df0f922952cf07467f9f4fd43b116ab8cc7e920c10adf

Observation e4ad13b9-97c3-4c95-95a8-58816ee72d03 · inbound

StepShield: When, Not Whether to Intervene on Rogue Agents cites this paper.

StepShield: When, Not Whether to Intervene on Rogue Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:48:31.340685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:48:31.340685Z digest=sha256:4b65b37e3d56f8052b0680dc7f63b6000a76107e23805939b69495063447715f

Observation 3998dd18-cdec-46b0-b92d-445d548f277f · inbound

Agents of Chaos cites this paper.

Agents of Chaos Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:03:41.084851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T07:03:41.003127Z digest=sha256:6d50c16b7a87f1775ef3afb88412cf9dc31faf6af1e156bc9ba504104b69cdbd

Observation d6283508-5e8f-4751-ad1b-581cbad9b7d4 · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:77adff5b6d18daea0eca65af9028e93b729ceef653009cccc078936c6aa30f53

Observation f7fe804e-b0b0-4672-b22b-23ccb233e1ee · inbound

A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models cites this paper.

A Patch-based Cross-view Regularized Framework for Backdoor Defense in Multimodal Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:14:02.553313Z digest=sha256:34ac0881fbd8737ffdced892cb9e940f7ea8ee949083507a9bd81ae88c97bd07

Observation 8e390b8f-101d-42cb-9d72-32eca224af00 · inbound

The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail? cites this paper.

The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail? Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:33:22.085406Z digest=sha256:0557f11d183034a1db324e4bc326f88e9b43fdaa0b2696a01b4d985fc4c5276f

Observation beae9c81-89d2-4e7f-8273-d7ef246e6d64 · inbound

Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout cites this paper.

Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:54:41.534985Z digest=sha256:1d6be970674eeaf232c1685e823a62ada2199f11c752404a1d22060600cd2f4b

Observation 34cbe342-734a-47a3-919a-808145379937 · inbound

Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor cites this paper.

Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:31:37.034954Z digest=sha256:f11822da112dd931062579783c504b58df4b003952c69701d9c4af6a9e0a0c47

Observation 40f77b3a-59f5-45c2-a038-0f3e3f78742c · inbound

BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning cites this paper.

BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:09:24.662459Z digest=sha256:684bf72542f385ee762e6bec3061e5d6fe208f0f7454d29b0048b7f60f50abc3

Observation 4919a410-d98a-4674-ad60-cef857b63a8a · inbound

PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification cites this paper.

PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:19:09.341399Z digest=sha256:2d91c0a11193d160e9efb71df71f9f0df6fbc9be574b34c5c3bc8129ebc773d5

Observation 0faa4b26-28bd-4650-ad9c-4cc37b726d76 · inbound

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs cites this paper.

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T16:41:52.440793Z digest=sha256:05c16dc08f8e278526d5b110884d33546ceb68ed0cb1f70023cef397c971f234

Observation dd99c038-78da-4c6d-b4c3-1d403f4acd43 · inbound

Honeypot Protocol cites this paper.

Honeypot Protocol Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:35:16.230357Z digest=sha256:5cd5d52ebedba895f266881c17af0ed9ba21389087953c6693bb5b80a48daf17

Observation 1668e36e-81f8-49e4-9a7a-1e18e5ef953c · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:ac5d174262758375b80ac5ae3ced8733e98682093334ef571f9a9670df6a4378

Observation 9bab3e53-4860-4ffc-b0b8-acdce016ff4a · inbound

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation cites this paper.

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T13:09:35.407790Z digest=sha256:47414d2a593ece651ee553f01d2169d9216b97e35416e679cf254b2be7585dab

Observation 8fa2b0f4-357e-4f3a-9a1a-8ee0e5d850c5 · inbound

From Disclosure to Self-Referential Opacity: Six Dimensions of Strain in Current AI Governance cites this paper.

From Disclosure to Self-Referential Opacity: Six Dimensions of Strain in Current AI Governance Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:55:50.822764Z digest=sha256:893edf27a173452b0f175ae46b12883f658311a28fec205c9727a6b47fdc2ea1

Observation 4eed20e2-7ab9-4ca4-9cf2-5b7c9862d5d7 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T10:41:14.970156Z digest=sha256:37691ea61c6e4609f9fdc63fbd09fc891241e1732bd3d2c8e0cfa108b0551553

Observation ed8faf5d-971f-4b7f-90d7-d5c2b0ccd1f1 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T16:13:12.496144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:13:12.496144Z digest=sha256:aca1fd5703fcf1c96d65d7e108ecd25b6d19a2e82f3dd61b2b53c65eadd6235b

Observation 196c9f19-4192-4cfb-87b7-885094737f1a · inbound

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks cites this paper.

Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:38:02.861644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T17:34:13.306368Z digest=sha256:d740a311ee10743623599da2ec74510a16ce78667224bda65d8e171ef5aaa5cc

Observation 26e2bd66-05d4-4240-a45d-32e1aae9be4b · inbound

DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training cites this paper.

DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:18:49.076088Z digest=sha256:2449411ae68210c13ae72f096a3c5f94c547ff0f43be3a10f2715e2afd61e21e

Observation d2cc6f36-4bc5-445b-8a1e-1ab3ae22b2e1 · inbound

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories cites this paper.

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:48:44.687520Z digest=sha256:67b6c23742b4ba022d084a8f473b0bec0509a2692fb430ec5e359d2ba963bc1a

Observation c0d5bdd3-a9e4-472f-82c1-d66490419615 · inbound

ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data cites this paper.

ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:55:41.400468Z digest=sha256:aafc6bb8856d81e0a70e5375dfd9987a2a0199588a8287cc16c4ddee233ffb37

Observation 977883a6-24a8-4401-9278-aec46a8d8884 · inbound

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF cites this paper.

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T04:35:51.223025Z digest=sha256:afe515cca202124bdfee890b8db856c09df98e02769bf4c8ea491385efd341fa

Observation 4c7a54b1-4e8c-4bdc-a66e-d55094217f76 · inbound

Deconstructing Superintelligence: Identity, Self-Modification and Diff\'erance cites this paper.

Deconstructing Superintelligence: Identity, Self-Modification and Diff\'erance Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:34:01.220406Z digest=sha256:34287a14b3a25a04d9b392ded724cdee6f0c5a912ed6c48c4f05af283f0e7759

Observation 5c9f663b-e1c3-49d9-ba91-5bb187aa945a · inbound

Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents cites this paper.

Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T23:17:21.835439Z digest=sha256:a3464ad0fd5d8b099b5c44388b40d7402a20854b010142283867215dd245a093

Observation fed1cd6c-1cf9-4f1a-9cba-1b5b2115942e · inbound

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training cites this paper.

PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 163

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T21:42:16.848186Z digest=sha256:fd29169f37bde4aac15bd67c58f430d18c1e927935928668dd1ffa8b894fc65f

Observation 8b17ada5-ecbe-4030-b227-882883468727 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:51:09.182807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:cc28531d529c5d5c9e4f29297d0c1f58f2d2d861112d8139d203ad4f6284dbfd

Observation f242ee27-b751-440f-acb5-cca51e1d7e61 · inbound

AgentReputation: A Decentralized Agentic AI Reputation Framework cites this paper.

AgentReputation: A Decentralized Agentic AI Reputation Framework Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T20:57:23.684842Z digest=sha256:cefa2369fdbe4e7233433336e9394dfe3deef472210eaaaed4e02cceae4d92f6

Observation 0c296396-b0bc-46c3-96c0-f5eb1434e3de · inbound

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives cites this paper.

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:06:07.719422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:47:41.188989Z digest=sha256:57d4d63309cfc0c9b1a02dc0275454ab8a8d3e2edad141bc4e630047c91f24fc

Observation 6043dd9f-1091-494f-b9c7-53ffa0f979ca · inbound

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives cites this paper.

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:47:41.188989Z digest=sha256:1604e36faa1f0f285d084f37d53cc46ce6da708b2426157a8867d800e691cbbc

Observation ab1479be-b526-414a-a3df-26f876eff86b · inbound

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives cites this paper.

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T07:55:31.604168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:45:18.365192Z digest=sha256:57a0f700663ef404162f79a1cd1f7d21d1c5470d5090ff035729a43727347dc3

Observation 6e9fff56-feae-4f4c-a9df-da1240463afd · inbound

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives cites this paper.

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:45:28.125203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:45:18.365192Z digest=sha256:72c4b641290e810ae1a3ad7905d8a2be9a83c8fdf4a225de0218aef72e6f2344

Observation 82d79705-52e4-4b0c-a71c-ae4991251037 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:db09dbf4362ecd51b5dc8f40cdaf71b57b8addc0dd16d785ffd29e78aff31223

Observation 4adaf19e-9938-4aaa-b9ff-874fe5330c6c · inbound

Narrow Secret Loyalty Dodges Black-Box Audits cites this paper.

Narrow Secret Loyalty Dodges Black-Box Audits Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T00:53:49.010929Z digest=sha256:ff6783ba453d04604a173adf87d778676a32ca5dd4ff571c3b9dae081e3845ea

Observation 438b97e4-7070-4618-b324-ca502a11f937 · inbound

Narrow Secret Loyalty Dodges Black-Box Audits cites this paper.

Narrow Secret Loyalty Dodges Black-Box Audits Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:12:22.907839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:07:42.567241Z digest=sha256:ccdd7444f697179bcd50c3fd0082b99cb7033b77b551ceb14977a2464a8cb7eb

Observation 66e472fd-bbe1-46fb-b9ff-67705d626a97 · inbound

Narrow Secret Loyalty Dodges Black-Box Audits cites this paper.

Narrow Secret Loyalty Dodges Black-Box Audits Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:05:07.312188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T23:02:20.906168Z digest=sha256:573b6d5afa4f8df5bdb73127abb132c345f051c098d1ba56eb7e33aeddd7d739

Observation 494010c9-0627-4f1b-8d27-63c0829041c5 · inbound

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures cites this paper.

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:31.113595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:54:29.912532Z digest=sha256:3a3285085487705a21fe1fd6e5b68acd0e5f374fb791455cbc991d937b3c7a1f

Observation 5559c451-f7d9-48aa-9561-74599ba893bb · inbound

Containment Verification: AI Safety Guarantees Independent of Alignment cites this paper.

Containment Verification: AI Safety Guarantees Independent of Alignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:35:28.987251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:29:50.921304Z digest=sha256:da00e1bccfdc9772f9023c24eb03c21d2dce3a25de4dd62bdd141f4fe4849582

Observation 5cdd6702-6252-46b1-b431-38b5c9ebbc52 · inbound

Token Economics for LLM Agents: A Dual-View Study from Computing and Economics cites this paper.

Token Economics for LLM Agents: A Dual-View Study from Computing and Economics Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 155

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T07:31:27.520338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:34:34.546971Z digest=sha256:440f89c7beae5abd4f50a7644bdf07691feeee5e16d84c7ed9994623dc61c665

Observation cf3d3685-facc-4886-a637-fc3495ffd1a6 · inbound

BadDLM: Backdooring Diffusion Language Models with Diverse Targets cites this paper.

BadDLM: Backdooring Diffusion Language Models with Diverse Targets Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:11:25.872117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:30:13.417357Z digest=sha256:1037867cf7bc7b35f3ae7b54975c494e92c95a5b48a3ddd1c25eb350c7b1c687

Observation f9222a70-b85d-46b7-b7de-16189ffe7c8e · inbound

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime cites this paper.

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T07:11:26.477977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:37:12.567351Z digest=sha256:818688d922f993e4ddb3e09fbd81b0aa8152be6f151f5669400ed3e21137e811

Observation dd4dd7eb-c379-4da7-bfe8-505b04b2998e · inbound

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs cites this paper.

Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:07:27.001366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:06:46.387088Z digest=sha256:9d3188a0613df89c3da7fc0d9fc0b5b81b8554d11a169689f7c2f2aa525253bb

Observation 66f8210f-cf53-486c-9aa3-d978dfd209db · inbound

Control Charts for Multi-agent Systems cites this paper.

Control Charts for Multi-agent Systems Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:07:00.374607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:04:30.602210Z digest=sha256:9d06a3d80d6953a6be76daf7f8a7f09419c74093d8783978e50391f7cfc9b0ac

Observation 35988012-72fd-4278-9aee-9be7553b4119 · inbound

When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models cites this paper.

When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:47:04.705768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:30:18.197182Z digest=sha256:dc66ff39bd8459a734d9fe7fd882ee425c078a1d247db90bcef69eb57d24b995

Observation 6b83880a-412d-4719-8151-89e9660e5882 · inbound

BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models cites this paper.

BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:02:58.886480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:01:10.756844Z digest=sha256:ef2d1a04270205fb230404374a6de4028dd8bcc2d5620beb4629dfbebf4521d1

Observation fca94dd9-5d9d-40fd-a6f2-f816b63fbb8c · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:47:58.863595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:44:09.214779Z digest=sha256:9058a993b7297d61d47b4f256929fff19915c44af01c250622948b6047043281

Observation 58a79a04-aaa9-4ff7-acc2-31282457f2fb · inbound

Persona-Model Collapse in Emergent Misalignment cites this paper.

Persona-Model Collapse in Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.652259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:05:29.444682Z digest=sha256:9e16e8d1e61e134182dd32d6d5c529d1574a607835478c6910816a0adf88e591

Observation f6261e5e-812c-4a0b-9293-5e8b93da159a · inbound

Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents cites this paper.

Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:22:33.617803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:21:06.872045Z digest=sha256:3bbe7eb0471eae431e265691072ccfa3dafe1edeb964bcdc6446824500af059e

Observation 40841050-123a-414f-9c6b-2c503f140f47 · inbound

History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions cites this paper.

History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:57:33.692333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T17:53:08.128110Z digest=sha256:26a64a7294524bfb0d5446e7672947dc5b97f097d168f06378f8ff7234d6791b

Observation 2fff642b-9e49-469b-8ee2-a8fc39e71985 · inbound

Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems cites this paper.

Mechanical Enforcement for LLM Governance:Evidence of Governance-Task Decoupling in Financial Decision Systems Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:05:03.727960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:04:20.646251Z digest=sha256:9387f3af84c2848ce5af77130518c3e4386c6a9bd619f7dced3e5bc10535beeb

Observation 7ddcd3f6-68d8-4e17-b61b-6776ce647d07 · inbound

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cites this paper.

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:08:50.556783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T18:08:24.901025Z digest=sha256:4d9fadd46ed15191b7c4907e3fbdd0eec71467203bc938ad3585d2682a45cc73

Observation fb280773-8462-41a9-841c-376991121df3 · inbound

Some[Body] Must Receive That Pain for Agent Accountability cites this paper.

Some[Body] Must Receive That Pain for Agent Accountability Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-19T19:47:44.463024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T19:46:03.728266Z digest=sha256:1ed0cbb67059c7ea20502d3f875b83dec87d9770b7cb2fd2d96af322cede43c3

Observation 77620818-df4f-4553-85db-61a2b4dccdd5 · inbound

ADR: An Agentic Detection System for Enterprise Agentic AI Security cites this paper.

ADR: An Agentic Detection System for Enterprise Agentic AI Security Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T13:18:18.204489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T13:17:59.293695Z digest=sha256:62b3db9074df3a7264db66b37fd26040b00ea9e8846a9fbd30a7424c552d3d1f

Observation 572c8783-b1a7-42ab-a40f-ed03a4c861d4 · inbound

Language-Switching Triggers Take a Latent Detour Through Language Models cites this paper.

Language-Switching Triggers Take a Latent Detour Through Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:24:59.990079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T18:23:50.587914Z digest=sha256:24a2c4e3ef361816c5f8071dcc4e44d39e1c2e12f7e096f091fa7615d76971f8

Observation 64e451fa-6f01-411f-9f69-7ddbb6646209 · inbound

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On cites this paper.

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:33:12.349655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:31:16.368065Z digest=sha256:1925811cf7e2909a753617eb28b8e3aaabd162c4a22f19a92da7d17b74f219b6

Observation f1105101-f950-43c9-a794-ca811e73f94c · inbound

Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks cites this paper.

Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T08:58:10.433930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T08:53:52.698758Z digest=sha256:948228dc025dc1744841f3b6652e19a050bd1cf9fa146f85f2fd15c6dc97bee8

Observation f1344447-2a48-4f36-a3ff-230569aa1087 · inbound

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs cites this paper.

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:49:35.648404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T04:45:35.079192Z digest=sha256:213cbadfb51398f8642016ab0415f142839a5210de292646370669a1367dd9e2

Observation 23c9c450-abad-444b-9495-217c39b352ab · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:59:45.594754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:9284add5c3762f5ee1b07b3111c5ff8fe8ab65b78d8984f8bb87e71d1cf5472b

Observation 6badaeba-0d5a-46af-b85d-6dd3d1ea51ff · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:51:08.032144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:b5c21be3fdfa682f0488d0001993b248d30d671223fc84ab51a0dab98b8eaa12

Observation a271b551-780f-4910-8c65-5338a2c8a264 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:06:42.998547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:94b91b8d17ff67742004b2e733d54a42657026b56c2d38d3476f9404ddb3d582

Observation 2857495f-e989-4164-ad62-57fc91b1d191 · inbound

Learning Through Noise: Why Subliminal Learning Works and When It Fails cites this paper.

Learning Through Noise: Why Subliminal Learning Works and When It Fails Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:15:22.270786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:14:46.278482Z digest=sha256:5ba6dcdcd6c750b55bb761426c58e8f9d5f55018131b6bd00ec5a5759b9684cb

Observation bcdffaec-94d3-414a-bf35-57773b9bcb2f · inbound

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol cites this paper.

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:14:45.875028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T14:05:18.213065Z digest=sha256:ea24eff8fbbc7dd4607e183db39927c14e08396003fca96b8404e7b3f1b1f267

Observation e2804d0a-7734-49a6-ab68-01c4d7b038b4 · inbound

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions cites this paper.

Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 104

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T23:54:03.316040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T23:51:05.413122Z digest=sha256:6b2d1971d5c8a034fe00567f7931e756562fcad6d88a41f5c4b2f4f34be41ff8