Pith. sign in

REVIEW 2 cited by

Is My Text in Your AI Model? Gradient-based Membership Inference Test applied to LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.07384 v2 pith:LO6W6GQ2 submitted 2025-03-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords datamodelgradient-basedlearningmachinemodelstextclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work adapts and studies the gradient-based Membership Inference Test (gMINT) to the classification of text based on LLMs. MINT is a general approach intended to determine if given data was used for training machine learning models, and this work focuses on its application to the domain of Natural Language Processing. Using gradient-based analysis, the MINT model identifies whether particular data samples were included during the language model training phase, addressing growing concerns about data privacy in machine learning. The method was evaluated in seven Transformer-based models and six datasets comprising over 2.5 million sentences, focusing on text classification tasks. Experimental results demonstrate MINTs robustness, achieving AUC scores between 85% and 99%, depending on data size and model architecture. These findings highlight MINTs potential as a scalable and reliable tool for auditing machine learning models, ensuring transparency, safeguarding sensitive data, and fostering ethical compliance in the deployment of AI/NLP technologies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A black-box audit detects unauthorized fine-tuning by measuring a joint semantic-lexical distributional fingerprint in model outputs, robust to paraphrasing and distillation.

  2. PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER)

    cs.CL 2025-06 conditional novelty 4.0 of 10

    NER-based removal of person and location entities from resumes preserves occupancy-prediction accuracy on FairCVdb, and combined with a debiasing module yields gender-balanced shortlists.

Pith tools