A co-optimized mix of adversarial image noise and multi-agent-written text steering jailbreaks multimodal LLMs with new state-of-the-art attack success and malicious intent fulfillment rates.
noisy oracle
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering
A co-optimized mix of adversarial image noise and multi-agent-written text steering jailbreaks multimodal LLMs with new state-of-the-art attack success and malicious intent fulfillment rates.