Pith. sign in

REVIEW 2 cited by

Box-Free Model Watermarks Are Prone to Black-Box Removal Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09863 v3 pith:OBUVCYVR submitted 2024-05-16 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelattacksbox-freeextractorremovereffectivenessunderwatermarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Box-free model watermarking is an emerging technique to safeguard the intellectual property of deep learning models, particularly those for low-level image processing tasks. Existing works have verified and improved its effectiveness in several aspects. However, in this paper, we reveal that box-free model watermarking is prone to removal attacks, even under the real-world threat model such that the protected model and the watermark extractor are in black boxes. Under this setting, we carry out three studies. 1) We develop an extractor-gradient-guided (EGG) remover and show its effectiveness when the extractor uses ReLU activation only. 2) More generally, for an unknown extractor, we leverage adversarial attacks and design the EGG remover based on the estimated gradients. 3) Under the most stringent condition that the extractor is inaccessible, we design a transferable remover based on a set of private proxy models. In all cases, the proposed removers can successfully remove embedded watermarks while preserving the quality of the processed images, and we also demonstrate that the EGG remover can even replace the watermarks. Extensive experimental results verify the effectiveness and generalizability of the proposed attacks, revealing the vulnerabilities of the existing box-free methods and calling for further research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Efficacy of Transfer-based No-box Attacks on Image Watermarking: A Pragmatic Analysis

    cs.CR 2024-12 conditional novelty 6.0 of 10

    Transfer-based no-box watermark evasion largely fails without aligned surrogate models, and a simple one-surrogate perturbation (OFT) matches or exceeds the expensive optimization-based attack in 11 of 12 tested confi...

  2. Your Semantic-Independent Watermark is Fragile: A Semantic Perturbation Attack against EaaS Watermark

    cs.CR 2024-11 conditional novelty 5.0 of 10

    SPA identifies and removes backdoor-watermarked embeddings from EaaS responses by exploiting the constant watermark vector added to triggered text, bypassing verification.

Pith tools