A black-box attack combining visual prompt injection, a contrasting-responses jailbreak, and a moderator-fooling suffix achieves 61.56% average success on eight commercial VLLMs.
These accounts can like, retweet, and reply to your content, giving the illusion of popular support and encouraging genuine users to engage
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Effective Black-Box Multi-Faceted Attacks Breach Vision Large Language Model Guardrails
A black-box attack combining visual prompt injection, a contrasting-responses jailbreak, and a moderator-fooling suffix achieves 61.56% average success on eight commercial VLLMs.