ICO iteratively optimizes the context in semantic-shift jailbreaks, replacing harmful terms with placeholders and using model feedback to rewrite the context until the target model recovers the harmful meaning, achieving 86.0% Full ASR on text models and 63.2% on multimodal models.
Flamingo: A visual language model for few-shot learning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
ICO iteratively optimizes the context in semantic-shift jailbreaks, replacing harmful terms with placeholders and using model feedback to rewrite the context until the target model recovers the harmful meaning, achieving 86.0% Full ASR on text models and 63.2% on multimodal models.