Pith. sign in

REVIEW 14 cited by

ShieldGemma 2: Robust and Tractable Image Content Moderation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.01081 v2 pith:HFLUPVAT submitted 2025-04-01 cs.CV cs.CLeess.IV

ShieldGemma 2: Robust and Tractable Image Content Moderation

classification cs.CV cs.CLeess.IV
keywords imagemodelcitepcontentgemmagenerationmoderationrobust
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce ShieldGemma 2, a 4B parameter image content moderation model built on Gemma 3. This model provides robust safety risk predictions across the following key harm categories: Sexually Explicit, Violence \& Gore, and Dangerous Content for synthetic images (e.g. output of any image generation model) and natural images (e.g. any image input to a Vision-Language Model). We evaluated on both internal and external benchmarks to demonstrate state-of-the-art performance compared to LlavaGuard \citep{helff2024llavaguard}, GPT-4o mini \citep{hurst2024gpt}, and the base Gemma 3 model \citep{gemma_2025} based on our policies. Additionally, we present a novel adversarial data generation pipeline which enables a controlled, diverse, and robust image generation. ShieldGemma 2 provides an open image moderation tool to advance multimodal safety and responsible AI development.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

    cs.AI 2026-06 unverdicted novelty 7.0

    SafePyramid is a three-level benchmark showing frontier LLMs identify all violated rules in only 54.0%, 35.3%, and 12.9% of cases on L0, L1, and L2 respectively, indicating in-context policy guardrailing remains difficult.

  2. MIRAGE: Protecting against Malicious Image Editing via False Moderation

    cs.CR 2026-06 unverdicted novelty 7.0

    MIRAGE immunizes images by crafting perturbations that align them with policy-violating concepts in open-source moderation models, triggering refusals in closed-source commercial image editors at over 88% success rate.

  3. MIRAGE: Protecting against Malicious Image Editing via False Moderation

    cs.CR 2026-06 unverdicted novelty 7.0

    MIRAGE immunizes images by aligning them to policy-violating concepts in open-source moderation embedding spaces, triggering automatic refusals in commercial image editing APIs with over 88% success.

  4. SenBen: Sensitive Scene Graphs for Explainable Content Moderation

    cs.CV 2026-04 unverdicted novelty 7.0

    SenBen is the first large-scale scene graph benchmark for sensitive content, paired with a 241M distilled model that outperforms most VLMs and safety APIs on grounded detection while running much faster.

  5. Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation

    cs.AI 2026-07 conditional novelty 6.0

    Seven model-agnostic image transformations bypass OpenAI, Amazon, and Google image-moderation APIs, including under non-trivial perceptual-similarity constraints.

  6. Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints

    eess.SP 2026-07 conditional novelty 6.0

    Pointwise constrained fine-tuning via sample-wise augmented Lagrangians and learned relaxations reduces tail constraint violations across safety, tool-calling, and re-ranking while preserving average task performance.

  7. No Safe Dose: How Training Data Drives Unsafe Image Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    Proportion of unsafe images in training data directly increases unsafe outputs in text-to-image models, independent of absolute count, with complementary risk reduction from safer text encoders.

  8. $D^2$-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

    cs.AI 2026-05 unverdicted novelty 6.0

    D²-Monitor routes between lightweight and heavy safety probes using the count of hesitation steps in diffusion LLM denoising trajectories, achieving SOTA trade-off on three datasets with under 0.85M parameters.

  9. Boundary-targeted Membership Inference Attacks on Safety Classifiers

    cs.LG 2026-05 unverdicted novelty 6.0

    A boundary-targeted MIA strategy recovers 19% of distress-flagged conversations from a safety classifier at 5% false-positive rate, 3.5 times better than prior methods.

  10. Boundary-targeted Membership Inference Attacks on Safety Classifiers

    cs.LG 2026-05 unverdicted novelty 6.0

    A boundary-targeted MIA on safety classifiers recovers 19% of distress-flagged conversations at 5% false-positive rate, 3.5 times higher than standard MIA baselines.

  11. Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

    cs.CY 2026-05 unverdicted novelty 6.0

    A participatory red-teaming project in the Global South created the PLACES dataset of 26k T2I failure examples that reveal unique cultural and linguistic harms missed by existing safety frameworks.

  12. SenBen: Sensitive Scene Graphs for Explainable Content Moderation

    cs.CV 2026-04 conditional novelty 6.0

    A 241M multi-task student trained with suffix identity, VAR loss, and a decoupled Q2L head matches or beats most VLMs and safety APIs on grounded sensitive scene graphs at 7.6× lower latency.

  13. Beyond Linear Probes: Dynamic Safety Monitoring for Language Models

    cs.LG 2025-09 unverdicted novelty 6.0

    TPCs allow term-by-term progressive polynomial evaluation on LLM activations for flexible safety monitoring that supports both stronger guardrails and low-cost adaptive cascades.

  14. SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

    cs.CV 2026-05 unverdicted novelty 5.0

    SafeLens presents a fast-and-slow video guardrail framework that filters the SafeWatch dataset to 2.4% and adds Chain-of-Thought traces to achieve state-of-the-art moderation performance at reduced inference cost.