Min-K% Prob detects pretraining data in LLMs by flagging outlier low-probability words in text, achieving 7.4% better performance than prior methods on the new WIKIMIA benchmark.
Proof of unlearning: Definitions and instantiation
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
representative citing papers
Verification of machine unlearning is fragile because model providers can use adversarial unlearning to pass checks while keeping data influence.
citing papers explorer
-
Detecting Pretraining Data from Large Language Models
Min-K% Prob detects pretraining data in LLMs by flagging outlier low-probability words in text, achieving 7.4% better performance than prior methods on the new WIKIMIA benchmark.
-
Verification of Machine Unlearning is Fragile
Verification of machine unlearning is fragile because model providers can use adversarial unlearning to pass checks while keeping data influence.