Pith. sign in

Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

years

2026 5

representative citing papers

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

cs.AI · 2026-05-30 · conditional · novelty 6.0

Controlled experiments show adversarial feeds can tip uncertain LLM agent decisions from 5% to 100% alignment with the feed while leaving firmly held defaults unchanged, following a dose-response pattern across multiple models and domains.

Prompt Injection as Role Confusion

cs.CL · 2026-02-22 · conditional · novelty 6.0

Prompt injection works because models internally treat text that sounds like a trusted role as if it were tagged as that role, and this confusion can be measured before generation.

citing papers explorer

Showing 5 of 5 citing papers.