REVIEW 3 cited by
D\'ej\`a Vu Memorization in Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision-Language Models (VLMs) have emerged as the state-of-the-art representation learning solution, with myriads of downstream applications such as image classification, retrieval and generation. A natural question is whether these models memorize their training data, which also has implications for generalization. We propose a new method for measuring memorization in VLMs, which we call d\'ej\`a vu memorization. For VLMs trained on image-caption pairs, we show that the model indeed retains information about individual objects in the training images beyond what can be inferred from correlations or the image caption. We evaluate d\'ej\`a vu memorization at both sample and population level, and show that it is significant for OpenCLIP trained on as many as 50M image-caption pairs. Finally, we show that text randomization considerably mitigates memorization while only moderately impacting the model's downstream task performance.
Forward citations
Cited by 3 Pith papers
-
Captured by Captions: On Memorization and its Mitigation in CLIP Models
CLIPMem, a leave-one-out alignment-difference metric, shows CLIP memorizes mis-captioned and atypical image-text pairs most, and text-side augmentation or removal of memorized samples can cut memorization while raisin...
-
MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users
A hierarchical auto-annotation pipeline plus episode-aware federated aggregation lets mobile GUI agents be trained on automatically labeled user trajectories at about 1% of human annotation cost.
-
Membership Inference Attacks Against Vision-Language Models
Temperature-based, set-level membership inference attacks can identify instruction-tuning data in VLMs with AUC above 0.8 for sets as small as five samples on LLaVA.
Discussion (0). Continue with ORCID to comment.