Pith. sign in

REVIEW 4 cited by

Proving membership in LLM pretraining data via data watermarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.10892 v3 pith:OJHFYRGO submitted 2024-02-16 cs.CR cs.CLcs.LG

classification cs.CRcs.CLcs.LG
keywords datawatermarkswatermarkdetectionmodeldatasethasheshypothesis
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Detecting whether copyright holders' works were used in LLM pretraining is poised to be an important problem. This work proposes using data watermarks to enable principled detection with only black-box model access, provided that the rightholder contributed multiple training documents and watermarked them before public release. By applying a randomly sampled data watermark, detection can be framed as hypothesis testing, which provides guarantees on the false detection rate. We study two watermarks: one that inserts random sequences, and another that randomly substitutes characters with Unicode lookalikes. We first show how three aspects of watermark design -- watermark length, number of duplications, and interference -- affect the power of the hypothesis test. Next, we study how a watermark's detection strength changes under model and dataset scaling: while increasing the dataset size decreases the strength of the watermark, watermarks remain strong if the model size also increases. Finally, we view SHA hashes as natural watermarks and show that we can robustly detect hashes from BLOOM-176B's training data, as long as they occurred at least 90 times. Together, our results point towards a promising future for data watermarks in real world use.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings

    cs.LG 2025-04 conditional novelty 7.0 of 10

    STAMP detects dataset membership in LLMs by comparing model perplexity on a publicly released watermarked rephrasing against private watermarked rephrasings of the same documents.

  2. Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning

    cs.CR 2025-06 conditional novelty 6.0 of 10

    Indirect data poisoning (gradient-matching prompts) makes LLMs learn secret prompt-response pairs absent from training data, detectable with certified p-values and under 0.005% contaminated tokens.

  3. Data Watermarking for Sequential Recommender Systems

    cs.IR 2024-11 conditional novelty 6.0 of 10

    Inserting short consecutive item sequences into user interaction histories lets a data owner detect whether a sequential recommender was trained on the protected dataset.

  4. SoK: Watermarking for AI-Generated Content

    cs.CR 2024-11 conditional novelty 3.0 of 10

    A systematization of knowledge on watermarking for AI-generated content, unifying definitions, threat models, evaluation methods, and representative schemes across modalities.

Pith tools