Pith. sign in

REVIEW 3 cited by

Exploring Transferability of Multimodal Adversarial Samples for Vision-Language Pre-training Models with Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.12636 v5 pith:VXRHX2ZN submitted 2023-08-24 cs.MM

classification cs.MM
keywords adversarialmodelsmultimodalattackcontrastiveimage-textvision-languagelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The integration of visual and textual data in Vision-Language Pre-training (VLP) models is crucial for enhancing vision-language understanding. However, the adversarial robustness of these models, especially in the alignment of image-text features, has not yet been sufficiently explored. In this paper, we introduce a novel gradient-based multimodal adversarial attack method, underpinned by contrastive learning, to improve the transferability of multimodal adversarial samples in VLP models. This method concurrently generates adversarial texts and images within imperceptive perturbation, employing both image-text and intra-modal contrastive loss. We evaluate the effectiveness of our approach on image-text retrieval and visual entailment tasks, using publicly available datasets in a black-box setting. Extensive experiments indicate a significant advancement over existing single-modal transfer-based adversarial attack methods and current multimodal adversarial attack approaches.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SimVLA, a simplified three-step adversarial attack on vision-language models, improves transferable attack success by 8-15 points while using ~36% of the time and ~46% of the VRAM of the prior SOTA.

  2. Attacking Attention of Foundation Models Disrupts Downstream Tasks

    cs.CR 2025-06 conditional novelty 5.0 of 10

    A task-agnostic attack that perturbs attention and embeddings of CLIP/ViT backbones degrades classification, retrieval, captioning, segmentation, and depth estimation without using labels or text.

  3. Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    An alternating image-text optimization produces a universal adversarial suffix and image that transfer across open multimodal LLMs more effectively than single-modality jailbreaks.

Pith tools