Pith. sign in

REVIEW 6 cited by

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.03486 v3 pith:BJ5EYHNO submitted 2024-05-06 cs.CR cs.CVcs.SI

classification cs.CRcs.CVcs.SI
keywords imageimagesclassifierssafetyai-generatedeffectivenessreal-worldrobustness
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the advent of text-to-image models and concerns about their misuse, developers are increasingly relying on image safety classifiers to moderate their generated unsafe images. Yet, the performance of current image safety classifiers remains unknown for both real-world and AI-generated images. In this work, we propose UnsafeBench, a benchmarking framework that evaluates the effectiveness and robustness of image safety classifiers, with a particular focus on the impact of AI-generated images on their performance. First, we curate a large dataset of 10K real-world and AI-generated images that are annotated as safe or unsafe based on a set of 11 unsafe categories of images (sexual, violent, hateful, etc.). Then, we evaluate the effectiveness and robustness of five popular image safety classifiers, as well as three classifiers that are powered by general-purpose visual language models. Our assessment indicates that existing image safety classifiers are not comprehensive and effective enough to mitigate the multifaceted problem of unsafe images. Also, there exists a distribution shift between real-world and AI-generated images in image qualities, styles, and layouts, leading to degraded effectiveness and robustness. Motivated by these findings, we build a comprehensive image moderation tool called PerspectiveVision, which improves the effectiveness and robustness of existing classifiers, especially on AI-generated images. UnsafeBench and PerspectiveVision can aid the research community in better understanding the landscape of image safety classification in the era of generative AI.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

    cs.CR 2025-07 conditional novelty 7.0 of 10

    Hateful optical illusions generated with Stable Diffusion and ControlNet evade current moderation classifiers (best accuracy 0.245) and vision-language models (best accuracy 0.102), with simple image transformations s...

  2. SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

    cs.CR 2025-11 conditional novelty 6.0 of 10

    SPQR is a benchmark that scores safety, prompt adherence, quality, and post-fine-tuning robustness of text-to-image safety methods, and it shows benign fine-tuning often collapses safety alignment.

  3. Multimodal LLMs as Customized Reward Models for Text-to-Image Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LLaVA-Reward extracts reward scores from the hidden states of a multimodal LLM with a skip-connection cross-attention head, and reports state-of-the-art text-to-image evaluation across alignment, fidelity, and safety.

  4. Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Vision-language models consistently recognize unsafe content better from text than from images, and a simplified reinforcement learning fine-tune narrows that gap.

  5. Immunizing Images from Text to Image Editing via Adversarial Cross-Attention

    cs.CV 2025-09 conditional novelty 5.0 of 10

    An imperceptible adversarial noise, computed with a LLaVA caption as a stand-in for the unknown edit prompt, disrupts cross-attention in Stable Diffusion-based editors and makes text-guided edits fail.

  6. VModA: An Effective Framework for Adaptive NSFW Image Moderation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VModA combines prompt engineering, region zooming, and LLM-based answer aggregation to improve zero-shot NSFW image moderation across multiple categories.

Pith tools