An LLM's ability to follow application-specific system-prompt guardrails (security steerability) is nearly uncorrelated with its resistance to standard jailbreak attacks, based on a new 240-case benchmark across 18 open-source models.
Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Security Steerability is All You Need
An LLM's ability to follow application-specific system-prompt guardrails (security steerability) is nearly uncorrelated with its resistance to standard jailbreak attacks, based on a new 240-case benchmark across 18 open-source models.