{"work":{"id":"3f20ace2-8264-4a7e-94a4-94436bcc063a","openalex_id":"https://openalex.org/W4415248419","doi":"10.48550/arxiv.2505.03574","arxiv_id":"2505.03574","raw_key":null,"title":"LlamaFirewall: An open source guardrail system for building secure AI agents","authors":null,"authors_text":"Sahana Chennabasappa, Cyrus Nikolaidis, Daniel Song, David Molnar, Stephanie Ding, Shengye Wan, Spencer Whitman, Lauren Deason, Nicholas Doucette, Abraham Montilla, Alekhya Gampa, Beto de Paola, Dominik Gabi, James Crnkovich, Jean-Christoph","year":2025,"venue":"cs.CR","abstract":"Large language models (LLMs) have evolved from simple chatbots into autonomous agents capable of performing complex tasks such as editing production code, orchestrating workflows, and taking higher-stakes actions based on untrusted inputs like webpages and emails. These capabilities introduce new security risks that existing security measures, such as model fine-tuning or chatbot-focused guardrails, do not fully address. Given the higher stakes and the absence of deterministic solutions to mitigate these risks, there is a critical need for a real-time guardrail monitor to serve as a final layer of defense, and support system level, use case specific safety policy definition and enforcement. We introduce LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Our framework mitigates risks such as prompt injection, agent misalignment, and insecure code risks through three powerful guardrails: PromptGuard 2, a universal jailbreak detector that demonstrates clear state of the art performance; Agent Alignment Checks, a chain-of-thought auditor that inspects agent reasoning for prompt injection and goal misalignment, which, while still experimental, shows stronger efficacy at preventing indirect injections in general scenarios than previously proposed approaches; and CodeShield, an online static analysis engine that is both fast and extensible, aimed at preventing the generation of insecure or dangerous code by coding agents. Additionally, we include easy-to-use customizable scanners that make it possible for any developer who can write a regular expression or an LLM prompt to quickly update an agent's security guardrails.","external_url":"https://arxiv.org/abs/2505.03574","cited_by_count":1,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2505.03574","created_at":"2026-05-10T10:24:22.223244+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Llamafirewall: An open source guardrail system for building secure ai agents","render_title":"Llamafirewall: An open source guardrail system for building secure ai agents"},"hub":{"state":{"work_id":"3f20ace2-8264-4a7e-94a4-94436bcc063a","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":29,"external_cited_by_count":1,"distinct_field_count":9,"first_pith_cited_at":"2025-05-19T22:49:30+00:00","last_pith_cited_at":"2026-07-09T12:18:40+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T05:19:35.304101+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":6}],"polarity_counts":[{"context_polarity":"background","n":6}],"runs":{},"summary":{},"graph":{},"authors":[]}}