JailPO uses preference optimization to train attack models that generate covert jailbreak questions and templates, achieving high attack success on aligned LLMs with far fewer queries.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
JailPO uses preference optimization to train attack models that generate covert jailbreak questions and templates, achieving high attack success on aligned LLMs with far fewer queries.