A defense that aligns image patches with LLM-generated knowledge elements and downweights poisoned samples reduces backdoor and poisoning attack success in contrastive vision-language models to near zero on tested benchmarks.
V ATT: trans- formers for multimodal self-supervised learning from raw video, audio and text
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment
A defense that aligns image patches with LLM-generated knowledge elements and downweights poisoned samples reduces backdoor and poisoning attack success in contrastive vision-language models to near zero on tested benchmarks.