Book-level memorization varies sharply across LLMs: most books escape most models, yet Llama 3.1 70B can reproduce Harry Potter almost entirely from a six-token seed prompt.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2025 3roles
background 1polarities
background 1representative citing papers
LLM watermarking adoption is limited by misaligned stakeholder incentives; incentive-aligned approaches such as in-context watermarking can enable practical use in targeted domains like education and peer review.
Authors introduce MLM and CLM specialization methods that avoid memorizing identifiers in sensitive training data while aiming for a privacy-utility tradeoff on medical datasets.
citing papers explorer
-
Extracting memorized pieces of (copyrighted) books from open-weight language models
Book-level memorization varies sharply across LLMs: most books escape most models, yet Llama 3.1 70B can reproduce Harry Potter almost entirely from a six-token seed prompt.
-
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
LLM watermarking adoption is limited by misaligned stakeholder incentives; incentive-aligned approaches such as in-context watermarking can enable practical use in targeted domains like education and peer review.
-
Towards the Anonymization of the Language Modeling
Authors introduce MLM and CLM specialization methods that avoid memorizing identifiers in sensitive training data while aiming for a privacy-utility tradeoff on medical datasets.