Diffusion models on hierarchical synthetic data preferentially memorize prototypical samples composed of common substrings, with delayed memorization in fat-tailed distributions.
Gummadi, and Evimaria Terzi
2 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.
2
Pith papers citing it
2
external citations · external index
verdicts
UNVERDICTED 2representative citing papers
Authors introduce MLM and CLM specialization methods that avoid memorizing identifiers in sensitive training data while aiming for a privacy-utility tradeoff on medical datasets.
citing papers explorer
-
Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?
Diffusion models on hierarchical synthetic data preferentially memorize prototypical samples composed of common substrings, with delayed memorization in fat-tailed distributions.
-
Towards the Anonymization of the Language Modeling
Authors introduce MLM and CLM specialization methods that avoid memorizing identifiers in sensitive training data while aiming for a privacy-utility tradeoff on medical datasets.