Pith. sign in

REVIEW 1 cited by

Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.04036 v3 pith:BU733HFJ submitted 2025-03-06 cs.CR cs.CLcs.LG

classification cs.CRcs.CLcs.LG
keywords datawatermarkswatermarkingduringfictitioustrainingaccessapi-only
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data watermarking in language models injects traceable signals, such as specific token sequences or stylistic patterns, into copyrighted text, allowing copyright holders to track and verify training data ownership. Previous data watermarking techniques primarily focus on effective memorization during pretraining, while overlooking challenges that arise in other stages of the LLM lifecycle, such as the risk of watermark filtering during data preprocessing and verification difficulties due to API-only access. To address these challenges, we propose a novel data watermarking approach that injects plausible yet fictitious knowledge into training data using generated passages describing a fictitious entity and its associated attributes. Our watermarks are designed to be memorized by the LLM through seamlessly integrating in its training data, making them harder to detect lexically during preprocessing. We demonstrate that our watermarks can be effectively memorized by LLMs, and that increasing our watermarks' density, length, and diversity of attributes strengthens their memorization. We further show that our watermarks remain effective after continual pretraining and supervised finetuning. Finally, we show that our data watermarks can be evaluated even under API-only access via question answering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synchronization-Free Algebraic Fingerprints for Large Language Models: From Autoregressive to Diffusion Models

    cs.CR 2026-07 reject novelty 6.0 of 10

    A proposed synchronization-free LLM watermark embeds identity bits as parity values of a Reed-Solomon polynomial evaluated at token-pair hashes, but its probabilistic guarantees rely on an unsupported balanced-evaluat...

Pith tools