Pith. sign in

REVIEW 10 cited by

Double-I Watermark: Protecting Model Copyright for LLM Fine-tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.14883 v3 pith:C4GYXK5S submitted 2024-02-22 cs.CR cs.AIcs.LG

classification cs.CRcs.AIcs.LG
keywords fine-tuningapproachcustomizeddouble-imodelownerswatermarkbackdoor
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

To support various applications, a prevalent and efficient approach for business owners is leveraging their valuable datasets to fine-tune a pre-trained LLM through the API provided by LLM owners or cloud servers. However, this process carries a substantial risk of model misuse, potentially resulting in severe economic consequences for business owners. Thus, safeguarding the copyright of these customized models during LLM fine-tuning has become an urgent practical requirement, but there are limited existing solutions to provide such protection. To tackle this pressing issue, we propose a novel watermarking approach named ``Double-I watermark''. Specifically, based on the instruct-tuning data, two types of backdoor data paradigms are introduced with trigger in the instruction and the input, respectively. By leveraging LLM's learning capability to incorporate customized backdoor samples into the dataset, the proposed approach effectively injects specific watermarking information into the customized model during fine-tuning, which makes it easy to inject and verify watermarks in commercial scenarios. We evaluate the proposed "Double-I watermark" under various fine-tuning methods, demonstrating its harmlessness, robustness, uniqueness, imperceptibility, and validity through both quantitative and qualitative analyses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks

    cs.LG 2025-05 conditional novelty 7.0 of 10

    SIRA is a black-box paraphrase attack that masks high-self-information tokens and fills the gaps with an LLM, achieving near-100% watermark removal on seven schemes.

  2. CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation Backdoor

    cs.CL 2025-09 conditional novelty 6.0 of 10

    CTCC embeds LLM ownership fingerprints in cross-turn semantic contradictions: the model fires a secret response only when a user contradicts an earlier statement, with higher robustness and stealth than single-turn triggers.

  3. EverTracer: Hunting Stolen Large Language Models via Stealthy and Robust Probabilistic Fingerprint

    cs.CR 2025-09 conditional novelty 6.0 of 10

    A gray-box LLM fingerprinting method that detects memorized private text via calibrated probability variation, instead of using backdoor triggers.

  4. Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model Distribution

    cs.CR 2025-07 conditional novelty 6.0 of 10

    A branch-swapping mechanism over low-rank add-on modules stamps each distributed model copy with a unique binary user ID, achieving 100% reported verification accuracy with under 1% extra parameters.

  5. Towards Provable (In)Secure Model Weight Release Schemes

    cs.CR 2025-06 conditional novelty 6.0 of 10

    Defines game-based security for weight release schemes and breaks TaylorMLP with a near-complete, low-cost parameter extraction attack.

  6. AttenTrack: Mobile User Attention Awareness Based on Context and External Distractions

    cs.HC 2025-09 conditional novelty 5.0 of 10

    AttenTrack predicts a smartphone user's attention state from context and notification-response features, reaching cold-start F1 up to 80% in leave-one-user-out tests.

  7. MEraser: An Effective Fingerprint Erasure Approach for Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    By fine-tuning on mismatched pairs and then clean pairs, MEraser drops fingerprint success rate to zero on three backdoor-based fingerprinting schemes across multiple LLMs, with a reusable LoRA adapter for transfer.

  8. I'm Spartacus, No, I'm Spartacus: Measuring and Understanding LLM Identity Confusion

    cs.CR 2024-11 reject novelty 5.0 of 10

    Seven of 27 tested LLMs (25.93%) exhibited identity confusion, which the authors link to hallucination and show reduces user trust, especially in critical tasks.

  9. Unlocking the Effectiveness of LoRA-FP for Seamless Transfer Implantation of Fingerprints in Downstream Models

    cs.CR 2025-08 conditional novelty 3.0 of 10

    Backdoor fingerprints trained into LoRA adapters on a base LLM transfer to derivative models with 100% trigger success and, in several scenarios, greater robustness than directly injected fingerprints.

  10. CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems

    cs.CL 2025-05 reject novelty 3.0 of 10

    A trigger-based watermark for multi-agent reasoning traces detects only the injected phrase, not the reproduction of copyrighted content.

Pith tools