Pith. sign in

REVIEW 2 cited by

Membership Inference Attacks against Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02902 v1 pith:5PCMPUIQ submitted 2024-11-05 cs.CV cs.AIcs.CLcs.CRcs.LG

classification cs.CVcs.AIcs.CLcs.CRcs.LG
keywords datavllmsdatasetsdetectionimageinferencelargemembership
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large vision-language models (VLLMs) exhibit promising capabilities for processing multi-modal tasks across various application scenarios. However, their emergence also raises significant data security concerns, given the potential inclusion of sensitive information, such as private photos and medical records, in their training datasets. Detecting inappropriately used data in VLLMs remains a critical and unresolved issue, mainly due to the lack of standardized datasets and suitable methodologies. In this study, we introduce the first membership inference attack (MIA) benchmark tailored for various VLLMs to facilitate training data detection. Then, we propose a novel MIA pipeline specifically designed for token-level image detection. Lastly, we present a new metric called MaxR\'enyi-K%, which is based on the confidence of the model output and applies to both text and image data. We believe that our work can deepen the understanding and methodology of MIAs in the context of VLLMs. Our code and datasets are available at https://github.com/LIONS-EPFL/VL-MIA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

    cs.CL 2026-07 conditional novelty 6.0 of 10

    FindMyText detects near-verbatim text containment via fingerprint chains and outperforms shared-fingerprint, BM25 and dense-retrieval baselines on Wikipedia, ArXiv and HPLT web data.

  2. LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states

    cs.CR 2024-11 conditional novelty 5.0 of 10

    LUMIA uses per-layer linear probes on LLM activations to detect training-data membership, outperforming output-based attacks and extending to multimodal models.

Pith tools