Pith. sign in

REVIEW 1 cited by

Evading Watermark based Detection of AI-Generated Content

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.03807 v5 pith:C5QI7X7A submitted 2023-05-05 cs.LG cs.CRcs.CV

classification cs.LGcs.CRcs.CV
keywords ai-generatedcontentdetectionwatermarkworkchallengesexistingimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A generative AI model can generate extremely realistic-looking content, posing growing challenges to the authenticity of information. To address the challenges, watermark has been leveraged to detect AI-generated content. Specifically, a watermark is embedded into an AI-generated content before it is released. A content is detected as AI-generated if a similar watermark can be decoded from it. In this work, we perform a systematic study on the robustness of such watermark-based AI-generated content detection. We focus on AI-generated images. Our work shows that an attacker can post-process a watermarked image via adding a small, human-imperceptible perturbation to it, such that the post-processed image evades detection while maintaining its visual quality. We show the effectiveness of our attack both theoretically and empirically. Moreover, to evade detection, our adversarial post-processing method adds much smaller perturbations to AI-generated images and thus better maintain their visual quality than existing popular post-processing methods such as JPEG compression, Gaussian blur, and Brightness/Contrast. Our work shows the insufficiency of existing watermark-based detection of AI-generated content, highlighting the urgent needs of new methods. Our code is publicly available: https://github.com/zhengyuan-jiang/WEvade.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks

    cs.CV 2025-05 reject novelty 4.0 of 10

    A grammar-tree and Monte Carlo search method automatically crafts prompts that can make synthetic portraits evade AIGC detectors, but the reported evidence is sparse and partly contradictory.

Pith tools