Pith. sign in

hub Canonical reference

Rtbas: Defending llm agents against prompt injection and privacy leakage

Canonical reference. 80% of citing Pith papers cite this work as background.

25 Pith papers citing it
3 external citations · Pith
Background 80% of classified citations
abstract

Tool-Based Agent Systems (TBAS) allow Language Models (LMs) to use external tools for tasks beyond their standalone capabilities, such as searching websites, booking flights, or making financial transactions. However, these tools greatly increase the risks of prompt injection attacks, where malicious content hijacks the LM agent to leak confidential data or trigger harmful actions. Existing defenses (OpenAI GPTs) require user confirmation before every tool call, placing onerous burdens on users. We introduce Robust TBAS (RTBAS), which automatically detects and executes tool calls that preserve integrity and confidentiality, requiring user confirmation only when these safeguards cannot be ensured. RTBAS adapts Information Flow Control to the unique challenges presented by TBAS. We present two novel dependency screeners, using LM-as-a-judge and attention-based saliency, to overcome these challenges. Experimental results on the AgentDojo Prompt Injection benchmark show RTBAS prevents all targeted attacks with only a 2% loss of task utility when under attack, and further tests confirm its ability to obtain near-oracle performance on detecting both subtle and direct privacy leaks.

hub tools

citation-role summary

background 5

citation-polarity summary

years

2026 23 2025 2

roles

background 5

polarities

background 4 unclear 1

representative citing papers

PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization

cs.CR · 2026-05-04 · conditional · novelty 7.0

PIIGuard uses optimized hidden HTML fragments on webpages to block LLMs from leaking contact PII via indirect prompt injection, achieving at least 97% defense success across tested models while preserving benign QA utility.

SecureClaw: Clawing Back Control of LLM Agents

cs.CR · 2026-06-08 · unverdicted · novelty 6.0

SecureClaw is a dual-boundary defense placing data confinement at reads and authorization at writes, achieving near-zero attack success while retaining task utility on AgentDojo, AgentLeak, and ASB.

Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools

cs.CR · 2026-06-01 · unverdicted · novelty 6.0

Ghost tool calls from speculative dispatch create persistent intent leaks that only issue-time policies changing or suppressing call arguments or destinations can reduce, per evaluations of twelve policies on three corpora.

An AI Agent Execution Environment to Safeguard User Data

cs.CR · 2026-04-21 · unverdicted · novelty 6.0

GAAP guarantees confidentiality of private user data for AI agents by enforcing user-specified permissions deterministically through persistent information flow tracking, without trusting the agent or requiring attack-free models.

citing papers explorer

Showing 25 of 25 citing papers.