Pith. sign in

REVIEW 5 cited by

Data Plagiarism Index: Characterizing the Privacy Risk of Data-Copying in Tabular Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13012 v1 pith:537RW2AJ submitted 2024-06-18 cs.LG cs.CRstat.ML

classification cs.LGcs.CRstat.ML
keywords datadata-copyingprivacytabulargenerativemodelshighindex
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The promise of tabular generative models is to produce realistic synthetic data that can be shared and safely used without dangerous leakage of information from the training set. In evaluating these models, a variety of methods have been proposed to measure the tendency to copy data from the training dataset when generating a sample. However, these methods suffer from either not considering data-copying from a privacy threat perspective, not being motivated by recent results in the data-copying literature or being difficult to make compatible with the high dimensional, mixed type nature of tabular data. This paper proposes a new similarity metric and Membership Inference Attack called Data Plagiarism Index (DPI) for tabular data. We show that DPI evaluates a new intuitive definition of data-copying and characterizes the corresponding privacy risk. We show that the data-copying identified by DPI poses both privacy and fairness threats to common, high performing architectures; underscoring the necessity for more sophisticated generative modeling techniques to mitigate this issue.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation

    cs.LG 2025-12 conditional novelty 7.0 of 10

    LLM tabular generators leak memorized numeric strings, allowing a no-box attack to achieve near-perfect membership inference on some state-of-the-art models.

  2. Ensembling Membership Inference Attacks Against Tabular Generative Models

    cs.CR 2025-09 conditional novelty 6.0 of 10

    No single membership inference attack dominates across tabular generative models, and unsupervised ensembles of attacks achieve better average rankings.

  3. Privacy Auditing Synthetic Data Release through Local Likelihood Attacks

    cs.LG 2025-08 unverdicted novelty 6.0 of 10

    Gen-LRA is a computationally efficient no-box MIA that exploits local overfitting in tabular generative models to produce a closed-form density-ratio statistic with a provable mean-score gap between members and non-members.

  4. Risk In Context: Benchmarking Privacy Leakage of Foundation Models in Synthetic Tabular Data Generation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    LLM-based tabular generators reproduce seed rows often enough that membership-inference attacks succeed more against them than against GAN, VAE, or diffusion baselines.

  5. Quantifying the Privacy of Counterfactuals by Leveraging Membership Inference Attacks Against Synthetic Data

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Membership inference attacks adapted from synthetic data succeed on counterfactuals using only the counterfactuals themselves, without model access.

Pith tools