SRA achieves 99.71% average attack success across 26 LLMs by optimizing for coherent malicious semantics via the SRHS algorithm, with claimed theoretical guarantees on convergence and transfer.
Does refusal training in llms generalize to the past tense? arXiv preprint arXiv:2407.11969, 2024
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 4roles
background 1polarities
background 1representative citing papers
PHANTOM is a consolidated open-source dataset of 47,524 multimodal adversarial samples for VLMs, extending prior benchmarks across 10 high-level categories and 55 subcategories of harmful intents.
PAST2HARM applies temporal deepening and mid-conversation escalation to past-tense prompts, achieving 83%, 67%, and 100% black-box attack success on Gemini Nano Banana Pro, GPT Image 2, and SD XL while releasing a benchmark for red-teaming.
LLM safety evaluations are hindered by noise in dataset curation, automated red-teaming, response generation, and LLM-judge evaluation, making fair comparisons difficult and slowing progress.
citing papers explorer
-
LLM-Agnostic Semantic Representation Attack
SRA achieves 99.71% average attack success across 26 LLMs by optimizing for coherent malicious semantics via the SRHS algorithm, with claimed theoretical guarantees on convergence and transfer.
-
PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models
PHANTOM is a consolidated open-source dataset of 47,524 multimodal adversarial samples for VLMs, extending prior benchmarks across 10 high-level categories and 55 subcategories of harmful intents.
-
PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI
PAST2HARM applies temporal deepening and mid-conversation escalation to past-tense prompts, achieving 83%, 67%, and 100% black-box attack success on Gemini Nano Banana Pro, GPT Image 2, and SD XL while releasing a benchmark for red-teaming.
-
LLM-Safety Evaluations Lack Robustness
LLM safety evaluations are hindered by noise in dataset curation, automated red-teaming, response generation, and LLM-judge evaluation, making fair comparisons difficult and slowing progress.