Pith. sign in

REVIEW 16 cited by

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.03486 v3 pith:BJ5EYHNO submitted 2024-05-06 cs.CR cs.CVcs.SI

classification cs.CRcs.CVcs.SI
keywords imageimagesclassifierssafetyai-generatedeffectivenessreal-worldrobustness
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

With the advent of text-to-image models and concerns about their misuse, developers are increasingly relying on image safety classifiers to moderate their generated unsafe images. Yet, the performance of current image safety classifiers remains unknown for both real-world and AI-generated images. In this work, we propose UnsafeBench, a benchmarking framework that evaluates the effectiveness and robustness of image safety classifiers, with a particular focus on the impact of AI-generated images on their performance. First, we curate a large dataset of 10K real-world and AI-generated images that are annotated as safe or unsafe based on a set of 11 unsafe categories of images (sexual, violent, hateful, etc.). Then, we evaluate the effectiveness and robustness of five popular image safety classifiers, as well as three classifiers that are powered by general-purpose visual language models. Our assessment indicates that existing image safety classifiers are not comprehensive and effective enough to mitigate the multifaceted problem of unsafe images. Also, there exists a distribution shift between real-world and AI-generated images in image qualities, styles, and layouts, leading to degraded effectiveness and robustness. Motivated by these findings, we build a comprehensive image moderation tool called PerspectiveVision, which improves the effectiveness and robustness of existing classifiers, especially on AI-generated images. UnsafeBench and PerspectiveVision can aid the research community in better understanding the landscape of image safety classification in the era of generative AI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

    cs.CR 2025-07 conditional novelty 7.0 of 10

    Hateful optical illusions generated with Stable Diffusion and ControlNet evade current moderation classifiers (best accuracy 0.245) and vision-language models (best accuracy 0.102), with simple image transformations s...

  2. SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

    cs.CR 2025-11 conditional novelty 6.0 of 10

    SPQR is a benchmark that scores safety, prompt adherence, quality, and post-fine-tuning robustness of text-to-image safety methods, and it shows benign fine-tuning often collapses safety alignment.

  3. Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LLaVA-Reward extracts reward scores from the hidden states of a multimodal LLM with a skip-connection cross-attention head, and reports state-of-the-art text-to-image evaluation across alignment, fidelity, and safety.

  4. Customize Multi-modal RAI Guardrails with Precedent-based predictions

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Conditioning a multimodal guardrail on retrieved 'precedent' reasoning traces, rather than static policy definitions, improves few-shot and novel-policy content-moderation F1 scores on UnsafeBench.

  5. Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Vision-language models consistently recognize unsafe content better from text than from images, and a simplified reinforcement learning fine-tune narrows that gap.

  6. T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    An LLM-driven discrete optimization with prompt mutation can rewrite unsafe prompts to bypass text-to-video safety filters and produce harmful videos with higher success than existing methods.

  7. LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A new benchmark and metric show that low-resolution zero-shot classification degrades sharply below 64x64, and adding trainable LR tokens to frozen CLIP-style models recovers some of the loss.

  8. HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

    cs.CR 2025-01 conditional novelty 6.0 of 10

    HateBench shows current hate speech detectors miss a meaningful share of LLM-generated hate and are evaded by word-level edits, enabling automated hate campaigns.

  9. T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    T2ISafety is a large annotated benchmark plus a fine-tuned MLLM evaluator (ImageGuard) for measuring toxicity, privacy, and fairness in text-to-image models.

  10. Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Prompt-Noise Optimization jointly tunes the prompt embedding and diffusion noise at inference time to suppress unsafe images while keeping outputs close to the prompt.

  11. VLSBench: Unveiling Visual Leakage in Multimodal Safety

    cs.CR 2024-11 conditional novelty 6.0 of 10

    The paper shows existing multimodal safety benchmarks leak harmful image content into text queries (VSIL), and introduces VLSBench, a 2.2k-pair leakless benchmark on which textual alignment fails and multimodal alignm...

  12. Immunizing Images from Text to Image Editing via Adversarial Cross-Attention

    cs.CV 2025-09 conditional novelty 5.0 of 10

    An imperceptible adversarial noise, computed with a LLaVA caption as a stand-in for the unknown edit prompt, disrupts cross-attention in Stable Diffusion-based editors and makes text-guided edits fail.

  13. VModA: An Effective Framework for Adaptive NSFW Image Moderation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VModA combines prompt engineering, region zooming, and LLM-based answer aggregation to improve zero-shot NSFW image moderation across multiple categories.

  14. GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning

    cs.AI 2025-05 conditional novelty 5.0 of 10

    GuardReasoner-VL, a 3B/7B VLM guard model trained with reasoning SFT and online RL, reports large F1 gains over existing VLM guard models on 14 safety benchmarks.

  15. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.

  16. From Noise to Nuance: Advances in Deep Generative Image Models

    cs.CV 2024-12 conditional

    A broad literature review of deep generative image models from GANs to diffusion and transformer architectures, with no new empirical results.

Pith tools