Selecting LLM-extracted key points with a diversity-aware determinantal point process before rewriting improves source coverage in multi-document news summarization.
Enhancing Jailbreak Attacks with Diversity Guidance
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
As large language models(LLMs) become commonplace in practical applications, the security issues of LLMs have attracted societal concerns. Although extensive efforts have been made to safety alignment, LLMs remain vulnerable to jailbreak attacks. We find that redundant computations limit the performance of existing jailbreak attack methods. Therefore, we propose DPP-based Stochastic Trigger Searching (DSTS), a new optimization algorithm for jailbreak attacks. DSTS incorporates diversity guidance through techniques including stochastic gradient search and DPP selection during optimization. Detailed experiments and ablation studies demonstrate the effectiveness of the algorithm. Moreover, we use the proposed algorithm to compute the risk boundaries for different LLMs, providing a new perspective on LLM safety evaluation.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries
Selecting LLM-extracted key points with a diversity-aware determinantal point process before rewriting improves source coverage in multi-document news summarization.