Pith. sign in

REVIEW 2 cited by

Large Language Model Watermark Stealing With Mixed Integer Programming

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19677 v1 pith:KUVHJGEQ submitted 2024-05-30 cs.CR cs.AI

classification cs.CRcs.AI
keywords watermarkattackgreenkeyslistmodelschemetext
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The Large Language Model (LLM) watermark is a newly emerging technique that shows promise in addressing concerns surrounding LLM copyright, monitoring AI-generated text, and preventing its misuse. The LLM watermark scheme commonly includes generating secret keys to partition the vocabulary into green and red lists, applying a perturbation to the logits of tokens in the green list to increase their sampling likelihood, thus facilitating watermark detection to identify AI-generated text if the proportion of green tokens exceeds a threshold. However, recent research indicates that watermarking methods using numerous keys are susceptible to removal attacks, such as token editing, synonym substitution, and paraphrasing, with robustness declining as the number of keys increases. Therefore, the state-of-the-art watermark schemes that employ fewer or single keys have been demonstrated to be more robust against text editing and paraphrasing. In this paper, we propose a novel green list stealing attack against the state-of-the-art LLM watermark scheme and systematically examine its vulnerability to this attack. We formalize the attack as a mixed integer programming problem with constraints. We evaluate our attack under a comprehensive threat model, including an extreme scenario where the attacker has no prior knowledge, lacks access to the watermark detector API, and possesses no information about the LLM's parameter settings or watermark injection/detection scheme. Extensive experiments on LLMs, such as OPT and LLaMA, demonstrate that our attack can successfully steal the green list and remove the watermark across all settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models

    cs.CL 2025-05 reject novelty 5.0 of 10

    An entropy-driven fusion of logits-based and sampling-based watermarks is claimed to outperform existing LLM watermarking methods on all four evaluation axes.

  2. SoK: On the Role and Future of AIGC Watermarking in the Era of Gen-AI

    cs.CR 2024-11 conditional novelty 5.0 of 10

    A systematization of AI-generated content watermarking that introduces a formal supply-chain definition and a property-based taxonomy.

Pith tools