Sleeper Attack formalizes persistent dormant adversarial content in LLM agent states that triggers on future benign queries, shown via a 1,896-instance benchmark to affect seven current models despite low single-interaction success rates.
Taylor, Krishnamurthy Dj Dvijotham, and Alexandre Lacoste
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
The TAB benchmark reveals that frontier terminal agents achieve high task completion but low selective alignment with relevant environmental cues over distractors, and prompt-injection defenses block both.
PORTICO is a revocable capability reference monitor for coding agents that enforces task contracts via grant-invoke-closure lifecycles and rejects post-closure reuses while preserving task success.
The survey organizes security threats and defenses in autonomous LLM agents into four layers and identifies that risks can propagate across layers from inputs to ecosystem impacts.
citing papers explorer
-
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Sleeper Attack formalizes persistent dormant adversarial content in LLM agent states that triggers on future benign queries, shown via a 1,896-instance benchmark to affect seven current models despite low single-interaction success rates.
-
No More, No Less: Task Alignment in Terminal Agents
The TAB benchmark reveals that frontier terminal agents achieve high task completion but low selective alignment with relevant environmental cues over distractors, and prompt-injection defenses block both.
-
Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents
PORTICO is a revocable capability reference monitor for coding agents that enforces task contracts via grant-invoke-closure lifecycles and rejects post-closure reuses while preserving task success.
-
Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study
The survey organizes security threats and defenses in autonomous LLM agents into four layers and identifies that risks can propagate across layers from inputs to ecosystem impacts.