Pith. sign in

REVIEW 5 cited by

SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06666 v3 pith:CWUKJZTU submitted 2024-04-10 cs.CV cs.AIcs.CLcs.CR

classification cs.CVcs.AIcs.CLcs.CR
keywords contentexplicittext-to-imagemodelssafegensexuallyadversarialgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text-to-image (T2I) models, such as Stable Diffusion, have exhibited remarkable performance in generating high-quality images from text descriptions in recent years. However, text-to-image models may be tricked into generating not-safe-for-work (NSFW) content, particularly in sexually explicit scenarios. Existing countermeasures mostly focus on filtering inappropriate inputs and outputs, or suppressing improper text embeddings, which can block sexually explicit content (e.g., naked) but may still be vulnerable to adversarial prompts -- inputs that appear innocent but are ill-intended. In this paper, we present SafeGen, a framework to mitigate sexual content generation by text-to-image models in a text-agnostic manner. The key idea is to eliminate explicit visual representations from the model regardless of the text input. In this way, the text-to-image model is resistant to adversarial prompts since such unsafe visual representations are obstructed from within. Extensive experiments conducted on four datasets and large-scale user studies demonstrate SafeGen's effectiveness in mitigating sexually explicit content generation while preserving the high-fidelity of benign images. SafeGen outperforms eight state-of-the-art baseline methods and achieves 99.4% sexual content removal performance. Furthermore, our constructed benchmark of adversarial prompts provides a basis for future development and evaluation of anti-NSFW-generation methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    T2ISafety is a large annotated benchmark plus a fine-tuned MLLM evaluator (ImageGuard) for measuring toxicity, privacy, and fairness in text-to-image models.

  2. Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Prompt-Noise Optimization jointly tunes the prompt embedding and diffusion noise at inference time to suppress unsafe images while keeping outputs close to the prompt.

  3. Buster: Implanting Semantic Backdoor into Text Encoder to Mitigate NSFW Content Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Buster implants a semantic backdoor in the text encoder of text-to-image models, redirecting NSFW prompts to a benign target prompt while preserving benign generations.

  4. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.

  5. Antelope: Potent and Concealed Jailbreak Attack Strategy

    cs.CR 2024-12 conditional novelty 4.0 of 10

    Antelope finds short, inconspicuous suffix tokens by aligning prompt embeddings with reference image embeddings, achieving higher ASR than prior jailbreak attacks on Stable Diffusion and several defenses.

Pith tools