Automated LLM red-teaming achieves higher success rates than manual prompting (69.5% vs 47.6%) but manual solves are faster when they succeed, according to 214,271 attack attempts on the Crucible platform.
Jailbroken: How does llm safety training fail?,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Automation Advantage in AI Red Teaming
Automated LLM red-teaming achieves higher success rates than manual prompting (69.5% vs 47.6%) but manual solves are faster when they succeed, according to 214,271 attack attempts on the Crucible platform.