Pith. sign in

REVIEW 7 cited by

Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21792 v3 pith:FLYYDRDC submitted 2024-07-31 cs.LG cs.AIcs.CLcs.CY

classification cs.LGcs.AIcs.CLcs.CY
keywords safetybenchmarkscapabilitiesresearchgeneraladdressadvancementsempirically
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As artificial intelligence systems grow more powerful, there has been increasing interest in "AI safety" research to address emerging and future risks. However, the field of AI safety remains poorly defined and inconsistently measured, leading to confusion about how researchers can contribute. This lack of clarity is compounded by the unclear relationship between AI safety benchmarks and upstream general capabilities (e.g., general knowledge and reasoning). To address these issues, we conduct a comprehensive meta-analysis of AI safety benchmarks, empirically analyzing their correlation with general capabilities across dozens of models and providing a survey of existing directions in AI safety. Our findings reveal that many safety benchmarks highly correlate with both upstream model capabilities and training compute, potentially enabling "safetywashing"--where capability improvements are misrepresented as safety advancements. Based on these findings, we propose an empirical foundation for developing more meaningful safety metrics and define AI safety in a machine learning research context as a set of clearly delineated research goals that are empirically separable from generic capabilities advancements. In doing so, we aim to provide a more rigorous framework for AI safety research, advancing the science of safety evaluations and clarifying the path towards measurable progress.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The bitter lesson of misuse detection

    cs.CR 2025-07 conditional novelty 7.0 of 10

    A new benchmark for LLM supervision systems finds generalist models repurposed as harm classifiers outperform specialized commercial guardrails.

  2. SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts

    cs.AI 2025-05 conditional novelty 7.0 of 10

    Frontier LLMs pass fewer than 58% of systematically varied safety-fact scenarios, revealing weak generalization of critical safety knowledge to naive user queries.

  3. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  4. Safety Degradation in AI Agents

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Adding retrieval to aligned LLMs degrades safety: refusal rates fall, bias and harmfulness rise, and prompt-based mitigation only partially restores alignment.

  5. LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies

    cs.CL 2024-12 conditional novelty 6.0 of 10

    M-ALERT, a 75k-prompt multilingual safety benchmark, shows that LLM safety varies substantially across five languages and across risk categories, with no model reaching the 99% safe threshold in every language.

  6. Lessons from Defending Gemini Against Indirect Prompt Injections

    cs.CR 2025-05 conditional novelty 5.0 of 10

    Google DeepMind reports that adversarially fine-tuning Gemini 2.5 cut indirect prompt-injection success by roughly half on average in tested settings, but adaptive attackers still found gaps.

  7. Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.

Pith tools