MAD-Spear is a prompt injection attack that makes a single compromised agent emit fake 'Sybil' peer answers, exploiting LLM conformity to steer a multi-agent debate toward a wrong consensus and higher token costs.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MAD-Spear: A Conformity-Driven Prompt Injection Attack on Multi-Agent Debate Systems
MAD-Spear is a prompt injection attack that makes a single compromised agent emit fake 'Sybil' peer answers, exploiting LLM conformity to steer a multi-agent debate toward a wrong consensus and higher token costs.