Pith. sign in

REVIEW 5 cited by

SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17975 v3 pith:WEHPMDJY submitted 2024-06-25 cs.CL cs.CRcs.LG

classification cs.CLcs.CRcs.LG
keywords llmsmiasdatasetspost-hocdistributionmethodsrandomizedshifts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Whether LLMs memorize their training data and what this means, from measuring privacy leakage to detecting copyright violations, has become a rapidly growing area of research. In the last few months, more than 10 new methods have been proposed to perform Membership Inference Attacks (MIAs) against LLMs. Contrary to traditional MIAs which rely on fixed-but randomized-records or models, these methods are mostly trained and tested on datasets collected post-hoc. Sets of members and non-members, used to evaluate the MIA, are constructed using informed guesses after the release of a model. This lack of randomization raises concerns of a distribution shift between members and non-members. In this work, we first extensively review the literature on MIAs against LLMs and show that, while most work focuses on sequence-level MIAs evaluated in post-hoc setups, a range of target models, motivations and units of interest are considered. We then quantify distribution shifts present in 6 datasets used in the literature using a model-less bag of word classifier and show that all datasets constructed post-hoc suffer from strong distribution shifts. These shifts invalidate the claims of LLMs memorizing strongly in real-world scenarios and, potentially, also the methodological contributions of the recent papers based on these datasets. Yet, all hope might not be lost. We introduce important considerations to properly evaluate MIAs against LLMs and discuss, in turn, potential ways forwards: randomized test splits, injections of randomized (unique) sequences, randomized fine-tuning, and several post-hoc control methods. While each option comes with its advantages and limitations, we believe they collectively provide solid grounds to guide MIA development and study LLM memorization. We conclude with an overview of recommended approaches to benchmark sequence-level and document-level MIAs against LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Contaminated Is Your Benchmark? Quantifying Dataset Leakage in Large Language Models with Kernel Divergence

    cs.LG 2025-02 conditional novelty 7.0 of 10

    KDS measures LLM benchmark contamination by computing the divergence between kernel similarity matrices of sample embeddings before and after fine-tuning, and it correlates near-perfectly with contamination fraction i...

  2. Causal Evaluation of Membership Inference Attacks

    cs.LG 2026-02 reject novelty 6.0 of 10

    The causal framing of MIA evaluation is new, but the headline zero-run IPW estimator is mathematically wrong for imbalanced data, and the claimed LLM validation is absent from the experiments.

  3. Membership Inference Attacks as Privacy Tools: Reliability, Disparity and Ensemble

    cs.LG 2025-06 conditional novelty 6.0 of 10

    MIAs expose different members depending on attack method and random seed; the paper quantifies this with coverage/stability and shows ensembling attacks yields stronger, more reliable privacy checks.

  4. Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation

    cs.CR 2025-02 conditional novelty 6.0 of 10

    A membership inference attack on RAG systems crafts natural yes/no questions from a target document to detect its presence in the datastore, achieving high AUC while evading guardrail detectors.

  5. SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks

    cs.CR 2025-06 conditional novelty 5.0 of 10

    SOFT paraphrases low-loss fine-tuning samples before training, reducing MIA AUC from about 0.82 to about 0.54 across six datasets at roughly 7% perplexity cost.

Pith tools