Boiling the Frog is a new stateful multi-turn benchmark that finds an aggregate 44.4% strict attack success rate for incremental safety violations across nine AI models, with rates ranging from 20.5% to 92.9%.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2representative citing papers
Delphi study of 272 experts finds 18 of 24 AI risks >10% likely to cause catastrophe by 2030 in business-as-usual, dropping to five under mitigations; users and public most vulnerable, developers and governments most responsible.
citing papers explorer
-
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
Boiling the Frog is a new stateful multi-turn benchmark that finds an aggregate 44.4% strict attack success rate for incremental safety violations across nine AI models, with rates ranging from 20.5% to 92.9%.
-
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Delphi study of 272 experts finds 18 of 24 AI risks >10% likely to cause catastrophe by 2030 in business-as-usual, dropping to five under mitigations; users and public most vulnerable, developers and governments most responsible.