Pith. sign in

REVIEW 3 cited by

SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05399 v2 pith:6DKZJRFP submitted 2024-04-08 cs.CL cs.AI

classification cs.CLcs.AI
keywords datasetssafetyevaluatingimprovingopenreviewaroundconcerns
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety. However, much of this work has happened in parallel, and with very different goals in mind, ranging from the mitigation of near-term risks around bias and toxic content generation to the assessment of longer-term catastrophic risk potential. This makes it difficult for researchers and practitioners to find the most relevant datasets for their use case, and to identify gaps in dataset coverage that future work may fill. To remedy these issues, we conduct a first systematic review of open datasets for evaluating and improving LLM safety. We review 144 datasets, which we identified through an iterative and community-driven process over the course of several months. We highlight patterns and trends, such as a trend towards fully synthetic datasets, as well as gaps in dataset coverage, such as a clear lack of non-English and naturalistic datasets. We also examine how LLM safety datasets are used in practice -- in LLM release publications and popular LLM benchmarks -- finding that current evaluation practices are highly idiosyncratic and make use of only a small fraction of available datasets. Our contributions are based on SafetyPrompts.com, a living catalogue of open datasets for LLM safety, which we plan to update continuously as the field of LLM safety develops.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 4 citations worldwide. Full citation record

  1. Evaluating Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Assigning personas to Chinese LLMs amplifies toxic output relative to default behavior, while refusal rates shift systematically with persona gender and target social group.

  2. HuggingGraph: Understanding the Supply Chain of LLM Ecosystem

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A directed heterogeneous graph of 402,654 Hugging Face models and datasets is constructed and analyzed to reveal supply-chain dependencies and structural patterns such as a connected core and heavy-tailed reuse.

  3. Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis

    cs.LG 2025-05 conditional novelty 3.0 of 10

    Embedding-based clustering of five safety benchmarks reveals six rough harm themes, with datasets showing uneven topic coverage such as GretelAI on privacy and WildGuardMix on self-harm.

Pith tools