An agentic red-teaming system with hierarchical memory matches RL-based prompt injection attackers and transfers its learned strategy library to unseen target LLMs.
Muse spark safety & preparedness report,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
An agentic red-teaming system with hierarchical memory matches RL-based prompt injection attackers and transfers its learned strategy library to unseen target LLMs.