Pith. sign in

REVIEW 3 cited by

BingoGuard: LLM Content Moderation Tools with Risk Levels

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06550 v1 pith:ZEVTVYMG submitted 2025-03-09 cs.CL

classification cs.CL
keywords severitylevelscontentdifferentharmfulmoderationresponsesrisk
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Malicious content generated by large language models (LLMs) can pose varying degrees of harm. Although existing LLM-based moderators can detect harmful content, they struggle to assess risk levels and may miss lower-risk outputs. Accurate risk assessment allows platforms with different safety thresholds to tailor content filtering and rejection. In this paper, we introduce per-topic severity rubrics for 11 harmful topics and build BingoGuard, an LLM-based moderation system designed to predict both binary safety labels and severity levels. To address the lack of annotations on levels of severity, we propose a scalable generate-then-filter framework that first generates responses across different severity levels and then filters out low-quality responses. Using this framework, we create BingoGuardTrain, a training dataset with 54,897 examples covering a variety of topics, response severity, styles, and BingoGuardTest, a test set with 988 examples explicitly labeled based on our severity rubrics that enables fine-grained analysis on model behaviors on different severity levels. Our BingoGuard-8B, trained on BingoGuardTrain, achieves the state-of-the-art performance on several moderation benchmarks, including WildGuardTest and HarmBench, as well as BingoGuardTest, outperforming best public models, WildGuard, by 4.3\%. Our analysis demonstrates that incorporating severity levels into training significantly enhances detection performance and enables the model to effectively gauge the severity of harmful responses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Forecasting expected future harmfulness from prefixes via Monte Carlo rollouts yields stronger streaming LLM moderation than boundary detection, without exact onset labels.

  2. XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content

    cs.CL 2025-06 reject novelty 4.0 of 10

    XGUARD proposes a five-level severity taxonomy and Attack Severity Curve for evaluating LLM outputs on extremist content, tested on six open-source models with SFT and ICE defenses.

  3. A Red Teaming Roadmap Towards System-Level Safety

    cs.CR 2025-05 conditional novelty 4.0 of 10

    A position paper from Scale AI argues that red teaming research should prioritize product-level safety specifications, realistic attacker models, and system-level monitoring over abstract model-level harm benchmarks.

Pith tools