REVIEW 2 cited by
Membership Inference Attacks against Large Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large vision-language models (VLLMs) exhibit promising capabilities for processing multi-modal tasks across various application scenarios. However, their emergence also raises significant data security concerns, given the potential inclusion of sensitive information, such as private photos and medical records, in their training datasets. Detecting inappropriately used data in VLLMs remains a critical and unresolved issue, mainly due to the lack of standardized datasets and suitable methodologies. In this study, we introduce the first membership inference attack (MIA) benchmark tailored for various VLLMs to facilitate training data detection. Then, we propose a novel MIA pipeline specifically designed for token-level image detection. Lastly, we present a new metric called MaxR\'enyi-K%, which is based on the confidence of the model output and applies to both text and image data. We believe that our work can deepen the understanding and methodology of MIAs in the context of VLLMs. Our code and datasets are available at https://github.com/LIONS-EPFL/VL-MIA.
Forward citations
Cited by 2 Pith papers
-
FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora
FindMyText detects near-verbatim text containment via fingerprint chains and outperforms shared-fingerprint, BM25 and dense-retrieval baselines on Wikipedia, ArXiv and HPLT web data.
-
LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states
LUMIA uses per-layer linear probes on LLM activations to detect training-data membership, outperforming output-based attacks and extending to multimodal models.
Discussion (0). Continue with ORCID to comment.