Pith. sign in

REVIEW 1 cited by

Leveraging Optimization for Adaptive Attacks on Image Watermarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.16952 v2 pith:RWOWBYKC submitted 2023-09-29 cs.CR cs.LG

classification cs.CRcs.LG
keywords adaptivewatermarkingattacksattackimagedetectionrobustnessattacker
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Untrustworthy users can misuse image generators to synthesize high-quality deepfakes and engage in unethical activities. Watermarking deters misuse by marking generated content with a hidden message, enabling its detection using a secret watermarking key. A core security property of watermarking is robustness, which states that an attacker can only evade detection by substantially degrading image quality. Assessing robustness requires designing an adaptive attack for the specific watermarking algorithm. When evaluating watermarking algorithms and their (adaptive) attacks, it is challenging to determine whether an adaptive attack is optimal, i.e., the best possible attack. We solve this problem by defining an objective function and then approach adaptive attacks as an optimization problem. The core idea of our adaptive attacks is to replicate secret watermarking keys locally by creating surrogate keys that are differentiable and can be used to optimize the attack's parameters. We demonstrate for Stable Diffusion models that such an attacker can break all five surveyed watermarking methods at no visible degradation in image quality. Optimizing our attacks is efficient and requires less than 1 GPU hour to reduce the detection accuracy to 6.3% or less. Our findings emphasize the need for more rigorous robustness testing against adaptive, learnable attackers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When There Is No Decoder: Removing Watermarks from Stable Diffusion Models in a No-box Setting

    cs.CR 2025-07 reject novelty 4.0 of 10

    Blur-plus-deblur and generator fine-tuning can push watermark bit accuracy toward chance, but only when the attacker can train a surrogate decoder that matches the target's architecture.

Pith tools