Pith. sign in

Membership Inference Attacks From First Principles

6 Pith papers cite this work, alongside 30 external citations. Polarity classification is still indexing.

6 Pith papers citing it
30 external citations · Pith
abstract

A membership inference attack allows an adversary to query a trained machine learning model to predict whether or not a particular example was contained in the model's training dataset. These attacks are currently evaluated using average-case "accuracy" metrics that fail to characterize whether the attack can confidently identify any members of the training set. We argue that attacks should instead be evaluated by computing their true-positive rate at low (e.g., <0.1%) false-positive rates, and find most prior attacks perform poorly when evaluated in this way. To address this we develop a Likelihood Ratio Attack (LiRA) that carefully combines multiple ideas from the literature. Our attack is 10x more powerful at low false-positive rates, and also strictly dominates prior attacks on existing metrics.

citation-role summary

background 1

citation-polarity summary

years

2026 5 2025 1

roles

background 1

polarities

background 1

representative citing papers

Auditing of Unlearning Algorithms

cs.LG · 2026-07-07 · accept · novelty 6.0

An auditor based on membership inference attacks computes valid lower bounds on the unlearning parameter ε, empirically separating certified unlearning methods (small bounds) from heuristic ones (large bounds).

citing papers explorer

Showing 6 of 6 citing papers.