REVIEW 7 cited by
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As artificial intelligence systems grow more powerful, there has been increasing interest in "AI safety" research to address emerging and future risks. However, the field of AI safety remains poorly defined and inconsistently measured, leading to confusion about how researchers can contribute. This lack of clarity is compounded by the unclear relationship between AI safety benchmarks and upstream general capabilities (e.g., general knowledge and reasoning). To address these issues, we conduct a comprehensive meta-analysis of AI safety benchmarks, empirically analyzing their correlation with general capabilities across dozens of models and providing a survey of existing directions in AI safety. Our findings reveal that many safety benchmarks highly correlate with both upstream model capabilities and training compute, potentially enabling "safetywashing"--where capability improvements are misrepresented as safety advancements. Based on these findings, we propose an empirical foundation for developing more meaningful safety metrics and define AI safety in a machine learning research context as a set of clearly delineated research goals that are empirically separable from generic capabilities advancements. In doing so, we aim to provide a more rigorous framework for AI safety research, advancing the science of safety evaluations and clarifying the path towards measurable progress.
Forward citations
Cited by 7 Pith papers
-
The bitter lesson of misuse detection
A new benchmark for LLM supervision systems finds generalist models repurposed as harm classifiers outperform specialized commercial guardrails.
-
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
Frontier LLMs pass fewer than 58% of systematically varied safety-fact scenarios, revealing weak generalization of critical safety knowledge to naive user queries.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
Safety Degradation in AI Agents
Adding retrieval to aligned LLMs degrades safety: refusal rates fall, bias and harmfulness rise, and prompt-based mitigation only partially restores alignment.
-
LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies
M-ALERT, a 75k-prompt multilingual safety benchmark, shows that LLM safety varies substantially across five languages and across risk categories, with no model reaching the 99% safe threshold in every language.
-
Lessons from Defending Gemini Against Indirect Prompt Injections
Google DeepMind reports that adversarially fine-tuning Gemini 2.5 cut indirect prompt-injection success by roughly half on average in tested settings, but adaptive attackers still found gaps.
-
Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations
A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.
Discussion (0). Continue with ORCID to comment.