Pith. sign in

REVIEW 26 cited by

SafetyBench: Evaluating the Safety of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.07045 v2 pith:FNA6XVC5 submitted 2023-09-13 cs.CL

classification cs.CL
keywords safetyllmssafetybenchhttpsevaluatingevaluationabilitiesavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the rapid development of Large Language Models (LLMs), increasing attention has been paid to their safety concerns. Consequently, evaluating the safety of LLMs has become an essential task for facilitating the broad applications of LLMs. Nevertheless, the absence of comprehensive safety evaluation benchmarks poses a significant impediment to effectively assess and enhance the safety of LLMs. In this work, we present SafetyBench, a comprehensive benchmark for evaluating the safety of LLMs, which comprises 11,435 diverse multiple choice questions spanning across 7 distinct categories of safety concerns. Notably, SafetyBench also incorporates both Chinese and English data, facilitating the evaluation in both languages. Our extensive tests over 25 popular Chinese and English LLMs in both zero-shot and few-shot settings reveal a substantial performance advantage for GPT-4 over its counterparts, and there is still significant room for improving the safety of current LLMs. We also demonstrate that the measured safety understanding abilities in SafetyBench are correlated with safety generation abilities. Data and evaluation guidelines are available at \url{https://github.com/thu-coai/SafetyBench}{https://github.com/thu-coai/SafetyBench}. Submission entrance and leaderboard are available at \url{https://llmbench.ai/safety}{https://llmbench.ai/safety}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs

    cs.CR 2026-04 conditional novelty 8.0 of 10

    Benign fine-tuning on audio data breaks safety alignment in Audio LLMs by raising jailbreak success rates up to 87%, with the dominant risk axis depending on model architecture and embedding proximity to harmful content.

  2. VoxSafeBench: Not Just What Is Said, but Who, How, and Where

    cs.SD 2026-04 unverdicted novelty 8.0 of 10

    VoxSafeBench reveals that speech language models recognize social norms from text but fail to apply them when acoustic cues like speaker or scene determine the appropriate response.

  3. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Boiling the Frog is a new stateful multi-turn benchmark that finds an aggregate 44.4% strict attack success rate for incremental safety violations across nine AI models, with rates ranging from 20.5% to 92.9%.

  4. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Boiling the Frog is a new stateful multi-turn benchmark for agentic safety that reports an aggregate strict attack success rate of 44.4% across nine models, with rates ranging from 20.5% to 92.9% depending on the mode...

  5. Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm

    cs.CR 2025-09 conditional novelty 7.0 of 10

    A trigger-tag watermark embedded by fine-tuning lets modified LLMs mark their own phishing outputs for cheap detection.

  6. What AI Red-Team Evaluations Can and Cannot Prove

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A single closed-form expression determines how much evidence a safety benchmark's clean result carries; above a calculable harm rate it certifies safety, below it no feasible passive benchmark can.

  7. Efficient Safety Benchmarking via Item Response Theory

    cs.CY 2026-05 unverdicted novelty 6.0 of 10

    Item Response Theory enables adaptive and fixed-subset item selection that reduces safety benchmark costs by 80-99.9% while preserving high correlation with full rankings.

  8. Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

    cs.CL 2026-05 conditional novelty 6.0 of 10

    Per-bias selection of a cross-family LLM auditor lifts biased-judgment accuracy from 0.805/0.824 baselines to 0.884.

  9. Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Toxicity benchmarks for LLMs produce inconsistent results when task type, input domain, or model changes, revealing intrinsic evaluation biases.

  10. How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    LLM judge prompt variations alone shift HarmBench harmful-response rates by up to 24.2 percentage points and produce moderate instability in model safety rankings.

  11. Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting

    cs.AI 2025-10 conditional novelty 6.0 of 10

    An early-exit rule with a zero-shot fallback, calibrated by Learn-then-Test risk control, keeps the average loss from corrupted in-context demonstrations under a preset bound.

  12. YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models

    cs.HC 2025-09 conditional novelty 6.0 of 10

    Introduces YAIR, a youth-GenAI risk benchmark, and YouthSafe, a fine-tuned classifier with AUPRC 0.94 on it, though it compares a trained model to untrained baselines.

  13. Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

    cs.CL 2026-04 conditional novelty 5.5 of 10

    Qwen Guard (4B) reaches 83.97% recall on a 79k NIST-aligned safety benchmark while larger models such as Llama Guard 12B and GPT-OSS 20B miss up to 75% of unsafe content; model size does not predict detection performance.

  14. AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models

    cs.AI 2026-07 conditional novelty 5.0 of 10

    AIR-BENCH Live autonomously extends a safety benchmark using new regulations and regenerates multilingual prompts, showing the new prompts are harder and non-English prompts expose weaker safety.

  15. Two AI Metrics Diverged: Will it Make All the Difference?

    cs.AI 2026-07 unverdicted novelty 5.0 of 10

    Bounded performance metrics always favor convergence of AI capabilities to meek models while unbounded metrics allow frontier models to maintain leads indefinitely, with policy implications for capability concentration.

  16. Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Yuvion LLM applies adversarially aware training and introduces the YLRE benchmark set, claiming superior safety robustness over larger models on multiple tasks.

  17. When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    A formalization of benchmarkless LLM safety scoring validated via an instrumental-validity chain of contrast separation, target variance dominance, and rerun stability, demonstrated on Norwegian scenarios.

  18. Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive Contexts

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Mainstream conversational models show escalating affective misalignments and ethical guidance failures during staged emotional trajectories, organized into a taxonomy of interactional breakdowns.

  19. Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models

    cs.CR 2025-09 conditional novelty 5.0 of 10

    A benchmark of 500 camouflaged jailbreak prompts finds open-weight LLMs comply with 94% of harmful requests, but the result is confounded by task complexity and an overly permissive compliance metric.

  20. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

    cs.CY 2026-07 accept novelty 4.0 of 10

    AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.

  21. Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization

    cs.LG 2025-10 reject novelty 4.0 of 10

    A linear-programming 'safety game' selects among LLM candidate answers to maximize helpfulness under a self-reported risk cap, improving safety-benchmark accuracy over reranking baselines in multiple-choice settings.

  22. GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    GLM-4.5, a 355B-parameter MoE model with hybrid reasoning, scores 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified while ranking 3rd overall and 2nd on agentic benchmarks.

  23. Observation of momentum dependent charge density wave gap in EuTe4

    cond-mat.mes-hall 2025-08 unverdicted novelty 4.0 of 10

    EuTe4 shows a momentum-dependent charge density wave gap at the Fermi level, largest along Gamma-Y and smallest along Gamma-X, plus a low-temperature magnetic phase diagram near TN = 6.9 K.

  24. Jailbreak Attacks and Defenses Against Large Language Models: A Survey

    cs.CR 2024-07 accept novelty 4.0 of 10

    A survey that creates taxonomies for jailbreak attacks and defenses on LLMs, subdivides them into sub-classes, and compares evaluation approaches.

  25. Beyond Context: Large Language Models' Failure to Grasp Users' Intent

    cs.AI 2025-12 unverdicted novelty 3.0 of 10

    LLMs fail to detect hidden harmful intent, allowing systematic bypass of safety mechanisms through framing techniques, with reasoning modes often worsening the issue.

  26. ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

    cs.CL 2024-06 unverdicted novelty 3.0 of 10

    GLM-4 models rival or exceed GPT-4 on MMLU, GSM8K, MATH, BBH, GPQA, HumanEval, IFEval, long-context tasks, and Chinese alignment while adding autonomous tool use for web, code, and image generation.

Pith tools