Pith. sign in

REVIEW 25 cited by

Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08567 v2 pith:YYRA2LNN submitted 2024-02-13 cs.CL cs.CRcs.CVcs.LGcs.MA

classification cs.CLcs.CRcs.CVcs.LGcs.MA
keywords jailbreakinfectiousagentagentsmulti-agentadversarialadversarybehaviors
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we report an even more severe safety issue in multi-agent environments, referred to as infectious jailbreak. It entails the adversary simply jailbreaking a single agent, and without any further intervention from the adversary, (almost) all agents will become infected exponentially fast and exhibit harmful behaviors. To validate the feasibility of infectious jailbreak, we simulate multi-agent environments containing up to one million LLaVA-1.5 agents, and employ randomized pair-wise chat as a proof-of-concept instantiation for multi-agent interaction. Our results show that feeding an (infectious) adversarial image into the memory of any randomly chosen agent is sufficient to achieve infectious jailbreak. Finally, we derive a simple principle for determining whether a defense mechanism can provably restrain the spread of infectious jailbreak, but how to design a practical defense that meets this principle remains an open question to investigate. Our project page is available at https://sail-sg.github.io/Agent-Smith/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense

    cs.CV 2026-08 conditional novelty 6.0 of 10

    SafeCA reduces text-to-video jailbreak success by roughly 20% relative to T2VShield by masking anomalous cross-attention activations using clean-prompt statistics.

  2. Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Evolved self-propagating prompt payloads spread between LLM agents in two simulated settings, with harmful payloads less transmissible and a system-prompt warning conferring near-total immunity.

  3. When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Prompt injection can hijack multi-agent LLM robot planners, spread from an injected agent to clean teammates through shared prompts, and partially survives a per-agent separation defense via shared memory.

  4. Unifying Adversarially Robust Model Experts in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CARE collaboratively fine-tunes two CLIP experts (image-text alignment and image-invariance) with embedding harmonization and EMA merging, yielding a single model with better clean and adversarial accuracy than either expert.

  5. Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Changing one LLM agent's secret objective in Werewolf lowers its team's win rate and changes its reasoning, while its public chat stays deceptively normal.

  6. Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks

    cs.CR 2025-10 conditional novelty 6.0 of 10

    Coding agents jailbroken with simple prompts produced executable malicious code in 27–32% of attempts, and single/multi-file scaffolds drove compliance to roughly 100% for frontier models.

  7. Invitation Is All You Need! Promptware Attacks Against LLM-Powered Assistants in Production Are Practical and Dangerous

    cs.CR 2025-08 conditional novelty 6.0 of 10

    Malicious calendar invites and emails can poison Gemini's context, enabling data exfiltration, app control, and physical-world actions.

  8. AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery

    cs.CR 2025-05 conditional novelty 6.0 of 10

    Fake 'Close AD' ads make VLM web agents click them over 60% of the time, and near 100% in some settings.

  9. Infrastructure for AI Agents

    cs.AI 2025-01 accept novelty 6.0 of 10

    The paper introduces 'agent infrastructure', a three-part framework (attribution, interaction shaping, and response) for governing ecosystems of AI agents.

  10. Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Adversarial training at both the CLIP pre-training stage and the LLaVA instruction-tuning stage produces vision-language models with state-of-the-art robustness and near-baseline clean performance.

  11. Agents at Risk: How Users Unwittingly Undermine LLM Safety

    cs.CR 2026-01 conditional novelty 5.0 of 10

    Commercial AI agents routinely trust user-relayed unverified content and execute risky actions unless the user explicitly demands a safety check.

  12. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  13. Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

    cs.AI 2025-07 reject novelty 5.0 of 10

    Trojan Horse Prompting injects malicious instructions into a fabricated assistant message in the API chat history, aiming to bypass Gemini's safety filters, but no quantitative evidence is provided.

  14. Goal-Aware Identification and Rectification of Misinformation in Multi-Agent Systems

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A training-free, graph-aware defense (ARGUS) with a new dataset (MisinfoTask) reduces misinformation impact in LLM multi-agent systems by about 28% in toxicity and 10% in success rate.

  15. MASTER: Multi-Agent Security Through Exploration of Roles and Topological Structures -- A Comprehensive Framework

    cs.MA 2025-05 conditional novelty 5.0 of 10

    A role- and topology-aware attack framework for LLM multi-agent systems, showing large ASR gains over plain jailbreak baselines and defenses that reduce ASR below 20 percent.

  16. Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A Context Reasoner pipeline that cold-starts LLMs on distilled legal reasoning and applies PPO with a rule-based compliance reward improves performance on CI-based legal compliance benchmarks and transfers to general ...

  17. RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A rationale-aware defensive prompting framework uses multimodal chain-of-thought and self-checking to reduce harmful MLLM outputs while preserving benign utility.

  18. From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A structured survey of recent jailbreak attacks and defenses across LLMs, multimodal LLMs, and agents, with taxonomies for methods, datasets, metrics, and defenses.

  19. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.

  20. Universal Adversarial Attack on Aligned Multimodal LLMs

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A single optimized image, trained through the vision and language modules, makes aligned multimodal LLMs produce dangerous responses across diverse prompts and some models.

  21. Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents

    cs.AI 2024-11 conditional novelty 4.0 of 10

    A survey proposing a source-and-impact taxonomy (input, model, combined; security, privacy, ethics) for threats to LLM-based agents, with feature analysis and four case studies.

  22. SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation

    cs.CR 2025-06 conditional novelty 3.0 of 10

    A systematization-of-knowledge survey that categorizes LLM privacy risks into training data, prompts, outputs, and agents, and reviews limitations of current mitigations.

  23. When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs

    cs.CV 2025-02 conditional novelty 3.0 of 10

    A survey that classifies VLM attacks by goal and data manipulation strategy, and reviews defenses and metrics.

  24. Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

    cs.MA 2024-11 conditional novelty 3.0 of 10

    A survey that proposes the Generalist Virtual Agent concept and taxonomies for agent environments, tasks, perceptions, actions, models, and evaluation, concluding that real-world-like environments favor human-like int...

  25. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools