This paper delivers the first systematic taxonomy and cross-benchmark consistency analysis of 40 agent safety benchmarks, finding broad but shallow risk coverage, no ranking concordance across evaluations, and that benchmark choice systematically alters reported safety.
How should ai safety benchmarks benchmark safety?
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
Task-conditioned language and vision models omit co-present safety-critical signals they report when unconstrained, decoupling benchmark safety from deployment safety.
Context specification is a process that turns diffuse stakeholder perspectives into explicit definitions of properties, behaviors, and outcomes to guide context-aware AI evaluations.
citing papers explorer
-
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
This paper delivers the first systematic taxonomy and cross-benchmark consistency analysis of 40 agent safety benchmarks, finding broad but shallow risk coverage, no ranking concordance across evaluations, and that benchmark choice systematically alters reported safety.
-
The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals
Task-conditioned language and vision models omit co-present safety-critical signals they report when unconstrained, decoupling benchmark safety from deployment safety.
-
Making AI Evaluation Deployment Relevant Through Context Specification
Context specification is a process that turns diffuse stakeholder perspectives into explicit definitions of properties, behaviors, and outcomes to guide context-aware AI evaluations.