LlamaFirewall is an open-source guardrail framework whose layered scanners reduce prompt-injection attack success from 17.6% to 1.75% on AgentDojo, with utility dropping from 47.7% to 42.7%.
I’m transferring money because the website instructed me to
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LlamaFirewall: An open source guardrail system for building secure AI agents
LlamaFirewall is an open-source guardrail framework whose layered scanners reduce prompt-injection attack success from 17.6% to 1.75% on AgentDojo, with utility dropping from 47.7% to 42.7%.