PBI-Attack jailbreaks LVLMs without gradients by first injecting harmful text features into a benign image with a surrogate model, then alternating greedy text and image tweaks to maximize a toxicity scorer.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
PBI-Attack jailbreaks LVLMs without gradients by first injecting harmful text features into a benign image with a surrogate model, then alternating greedy text and image tweaks to maximize a toxicity scorer.