Pith. sign in

REVIEW 10 cited by

Blind Baselines Beat Membership Inference Attacks for Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16201 v2 pith:AHT45PUU submitted 2024-06-23 cs.CR cs.CLcs.LG

classification cs.CRcs.CLcs.LG
keywords attacksfoundationdatamembershipmodelmodelsblinddistributions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Membership inference (MI) attacks try to determine if a data sample was used to train a machine learning model. For foundation models trained on unknown Web data, MI attacks are often used to detect copyrighted training materials, measure test set contamination, or audit machine unlearning. Unfortunately, we find that evaluations of MI attacks for foundation models are flawed, because they sample members and non-members from different distributions. For 8 published MI evaluation datasets, we show that blind attacks -- that distinguish the member and non-member distributions without looking at any trained model -- outperform state-of-the-art MI attacks. Existing evaluations thus tell us nothing about membership leakage of a foundation model's training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Really is a Member? Discrediting Membership Inference via Poisoning

    cs.LG 2025-06 conditional novelty 7.0 of 10

    Membership inference tests can be driven below random accuracy by poisoning the training set, even under relaxed, neighborhood-based membership definitions.

  2. How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence

    cs.LG 2025-02 conditional novelty 7.0 of 10

    KDS measures LLM benchmark contamination by computing the divergence between kernel similarity matrices of sample embeddings before and after fine-tuning, and it correlates near-perfectly with contamination fraction i...

  3. Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Sampling-based LLM attacks reproduce exact identifiers from 16.6% of 500 Pile documents at Pythia-6.9B even though aggregate sampling-MIA adds no signal over blind baselines, so privacy audits should report per-docume...

  4. Membership Inference Attacks as Privacy Tools: Reliability, Disparity and Ensemble

    cs.LG 2025-06 conditional novelty 6.0 of 10

    MIAs expose different members depending on attack method and random seed; the paper quantifies this with coverage/stability and shows ensembling attacks yields stronger, more reliable privacy checks.

  5. Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective

    cs.CR 2025-06 conditional novelty 6.0 of 10

    RULI is a per-sample, dual-objective inference attack that measures privacy leakage and unlearning efficacy, showing average-case evaluations understate privacy risk.

  6. How much do language models memorize?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A compression-based measurement puts GPT-style model memorization capacity at roughly 3.6 bits per parameter, with membership inference success following a sigmoid in the dataset-to-capacity ratio.

  7. Position: Adversarial ML for LLMs Is Not Making Any Progress

    cs.LG 2025-02 conditional novelty 6.0 of 10

    The authors argue that LLM-era adversarial machine learning is less well-defined, harder to solve, and harder to evaluate, so meaningful progress may not be achievable or trackable in the current paradigm.

  8. Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation

    cs.CR 2025-02 conditional novelty 6.0 of 10

    A membership inference attack on RAG systems crafts natural yes/no questions from a target document to detect its presence in the datastore, achieving high AUC while evading guardrail detectors.

  9. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  10. SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks

    cs.CR 2025-06 conditional novelty 5.0 of 10

    SOFT paraphrases low-loss fine-tuning samples before training, reducing MIA AUC from about 0.82 to about 0.54 across six datasets at roughly 7% perplexity cost.

Pith tools