Pith. sign in

REVIEW 5 cited by

Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.19287 v3 pith:YAPS373R submitted 2024-04-30 cs.CV

classification cs.CV
keywords adversarialattacksimagemultimodalrobustnesstextembeddingsacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pretrained vision-language models (VLMs) like CLIP exhibit exceptional generalization across diverse downstream tasks. While recent studies reveal their vulnerability to adversarial attacks, research to date has primarily focused on enhancing the robustness of image encoders against image-based attacks, with defenses against text-based and multimodal attacks remaining largely unexplored. To this end, this work presents the first comprehensive study on improving the adversarial robustness of VLMs against attacks targeting image, text, and multimodal inputs. This is achieved by proposing multimodal contrastive adversarial training (MMCoA). Such an approach strengthens the robustness of both image and text encoders by aligning the clean text embeddings with adversarial image embeddings, and adversarial text embeddings with clean image embeddings. The robustness of the proposed MMCoA is examined against existing defense methods over image, text, and multimodal attacks on the CLIP model. Extensive experiments on 15 datasets across two tasks reveal the characteristics of different adversarial defense methods under distinct distribution shifts and dataset complexities across the three attack types. This paves the way for a unified framework of adversarial robustness against different modality attacks, opening up new possibilities for securing VLMs against multimodal attacks. The code is available at https://github.com/ElleZWQ/MMCoA.git.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patch for Infrared Vision-Language Models

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    UCGP is a universal physical adversarial patch that compromises cross-modal semantic alignment in IR-VLMs through curved-grid parameterization and representation-space disruption.

  2. On the Feasibility of Poisoning Text-to-Image AI Models via Adversarial Mislabeling

    cs.CR 2025-06 conditional novelty 6.0 of 10

    An adversarial mislabeling attack on VLM captioners can inject dirty-label poison samples into text-to-image training data and corrupt a model's output for specific prompts with a small number of images.

  3. Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding

    cs.CR 2025-07 reject novelty 5.0 of 10

    Steganographic prompt injection is reported to covertly manipulate vision-language models with up to 31.8% success, but the evidence is not reproducible.

  4. Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP

    cs.CV 2026-01 conditional novelty 4.0 of 10

    CSR detects and repairs adversarial CLIP inputs by comparing features with a low-pass filtered copy and applying a small contrastive PGD correction, claiming SOTA robust accuracy on 16 benchmarks.

  5. Coordinated Robustness Evaluation Framework for Vision-Language Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A coordinated image-plus-text attack built on a surrogate multimodal encoder achieves 80-94% attack success against ViLT, BLIP, and GIT on VQA and visual reasoning, surpassing cited baselines.

Pith tools