REVIEW 6 cited by
Misusing Tools in Large Language Models With Visual Adversarial Examples
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Misusing Tools in Large Language Models With Visual Adversarial Examples
read the original abstract
Large Language Models (LLMs) are being enhanced with the ability to use tools and to process multiple modalities. These new capabilities bring new benefits and also new security risks. In this work, we show that an attacker can use visual adversarial examples to cause attacker-desired tool usage. For example, the attacker could cause a victim LLM to delete calendar events, leak private conversations and book hotels. Different from prior work, our attacks can affect the confidentiality and integrity of user resources connected to the LLM while being stealthy and generalizable to multiple input prompts. We construct these attacks using gradient-based adversarial training and characterize performance along multiple dimensions. We find that our adversarial images can manipulate the LLM to invoke tools following real-world syntax almost always (~98%) while maintaining high similarity to clean images (~0.9 SSIM). Furthermore, using human scoring and automated metrics, we find that the attacks do not noticeably affect the conversation (and its semantics) between the user and the LLM.
Forward citations
Cited by 6 Pith papers
-
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
AgentDojo introduces an extensible evaluation framework populated with realistic agent tasks and security test cases to measure prompt injection robustness in tool-using LLM agents.
-
From Storage to Steering: Memory Control Flow Attacks on LLM Agents
Memory can be exploited to hijack LLM agents' tool selection and induce persistent behavioral deviations even against explicit instructions and safety constraints.
-
From Storage to Steering: Memory Control Flow Attacks on LLM Agents
Malicious action-oriented policies written into LLM agent memory can persistently force risky tool choice and reordering across later benign tasks, with >90% attack success in controlled audits.
-
A Marketplace for AI-Generated Adult Content and Deepfakes
A 14-month audit of 4,847 paid AI-content requests on Civitai shows NSFW commissions growing to a majority of weekly bounties, deepfake requests targeting women about 9:1 among real individuals, and the platform's dee...
-
Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents
Roughly a third of UK adults use AI chatbots weekly, and among them a substantial minority upload untrusted content, connect bots to other programs, share sensitive data, or attempt jailbreaks.
-
Towards an AI co-scientist
A multi-agent AI system generates novel biomedical hypotheses that show promising experimental validation in drug repurposing for leukemia, new targets for liver fibrosis, and a bacterial gene transfer mechanism.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.