Pith. sign in

hub

ShieldGemma: Generative AI Content Moderation Based on Gemma

53 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

53 Pith papers citing it
2 external citations · Pith
abstract

We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety risks across key harm types (sexually explicit, dangerous content, harassment, hate speech) in both user input and LLM-generated output. By evaluating on both public and internal benchmarks, we demonstrate superior performance compared to existing models, such as Llama Guard (+10.8\% AU-PRC on public benchmarks) and WildCard (+4.3\%). Additionally, we present a novel LLM-based data curation pipeline, adaptable to a variety of safety-related tasks and beyond. We have shown strong generalization performance for model trained mainly on synthetic data. By releasing ShieldGemma, we provide a valuable resource to the research community, advancing LLM safety and enabling the creation of more effective content moderation solutions for developers.

hub tools

citation-role summary

background 2 baseline 1

citation-polarity summary

years

2026 51 2025 2

representative citing papers

PreAct-Bench: Benchmarking Predictive Monitoring in LLMs

cs.LG · 2026-06-03 · unverdicted · novelty 7.0

PreActBench is a new benchmark showing that LLMs struggle to predict unethical outcomes from partial action trajectories across five domains using the Prefix Foresight F1 metric.

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

cs.AI · 2026-05-30 · unverdicted · novelty 6.0

TRACE introduces a trajectory-level compression method using a Compressor-Reader pair that improves safety detection accuracy by up to 12.6 percentage points on ASSEBench, Pre-Ex-Bench, and R-Judge while degrading less on longer contexts.

Triaging Threats to Specialized Guardrails

cs.CR · 2026-05-29 · unverdicted · novelty 6.0

Introduces GuardZoo benchmark and RouteGuard router-expert system showing monolithic guardrails suffer task interference while specialized routing improves threat detection and generalization.

Alignment Dynamics in LLM Fine-Tuning

cs.LG · 2026-05-18 · unverdicted · novelty 6.0

The paper introduces a dynamical model that decomposes alignment updates in LLM fine-tuning into rebound and driving forces and predicts a rehearsal priming effect.

A Systematic Investigation of RL-Jailbreaking in LLMs

cs.LG · 2026-05-07 · reject · novelty 6.0 · 2 refs

An ablation study of an RL-based jailbreaker finds the attack succeeds across tested open-weight models and safeguards, but its headline conclusion about dense rewards and long episodes is contradicted by its own data.

citing papers explorer

Showing 50 of 53 citing papers.