Min-K% Prob detects pretraining data in LLMs by flagging outlier low-probability words in text, achieving 7.4% better performance than prior methods on the new WIKIMIA benchmark.
arXiv preprint arXiv:2302.07956 , year=
3 Pith papers cite this work, alongside 14 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
New canary crafting via greedy influence-based init and bilevel optimization for diversity in embedding space yields stronger one-run privacy leakage estimates at lower cost.
Under heterogeneous detectability, naive privacy-budget allocation in audits lets a strategic developer hide harm in low-detectability dimensions, inflating the welfare-weighted under-detection gap.
citing papers explorer
-
Detecting Pretraining Data from Large Language Models
Min-K% Prob detects pretraining data in LLMs by flagging outlier low-probability words in text, achieving 7.4% better performance than prior methods on the new WIKIMIA benchmark.
-
Detectability in Diversity: Improved Canary Crafting for Privacy Auditing in One Run
New canary crafting via greedy influence-based init and bilevel optimization for diversity in embedding space yields stronger one-run privacy leakage estimates at lower cost.
-
Differentially Private Auditing Under Strategic Response
Under heterogeneous detectability, naive privacy-budget allocation in audits lets a strategic developer hide harm in low-detectability dimensions, inflating the welfare-weighted under-detection gap.