Pith. sign in

REVIEW 6 cited by

OT-Attack: Enhancing Adversarial Transferability of Vision-Language Models via Optimal Transport Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04403 v1 pith:5HYQ7DC5 submitted 2023-12-07 cs.CV

classification cs.CV
keywords adversarialexamplesoptimalmodelstransferabilityimageimage-textot-attack
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language pre-training (VLP) models demonstrate impressive abilities in processing both images and text. However, they are vulnerable to multi-modal adversarial examples (AEs). Investigating the generation of high-transferability adversarial examples is crucial for uncovering VLP models' vulnerabilities in practical scenarios. Recent works have indicated that leveraging data augmentation and image-text modal interactions can enhance the transferability of adversarial examples for VLP models significantly. However, they do not consider the optimal alignment problem between dataaugmented image-text pairs. This oversight leads to adversarial examples that are overly tailored to the source model, thus limiting improvements in transferability. In our research, we first explore the interplay between image sets produced through data augmentation and their corresponding text sets. We find that augmented image samples can align optimally with certain texts while exhibiting less relevance to others. Motivated by this, we propose an Optimal Transport-based Adversarial Attack, dubbed OT-Attack. The proposed method formulates the features of image and text sets as two distinct distributions and employs optimal transport theory to determine the most efficient mapping between them. This optimal mapping informs our generation of adversarial examples to effectively counteract the overfitting issues. Extensive experiments across various network architectures and datasets in image-text matching tasks reveal that our OT-Attack outperforms existing state-of-the-art methods in terms of adversarial transferability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SimVLA, a simplified three-step adversarial attack on vision-language models, improves transferable attack success by 8-15 points while using ~36% of the time and ~46% of the VRAM of the prior SOTA.

  2. One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single adversarial image can make a unified vision-language model misclassify the same object across captioning, detection, region classification, and localization, and the new CrossVLAD benchmark and CRAFT attack m...

  3. Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LightD creates natural adversarial relighting images with GPT-selected lighting parameters and gradient optimization, outperforming prior non-suspicious attacks on vision-language models.

  4. GeoDetect: Geometric Adversarial Detection for VLPs

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Adversarial inputs in vision-language models sit farther from clean reference embeddings than normal inputs, so geometric scores can flag them — near-perfect on standard attacks, much weaker under adaptive ones.

  5. Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    The submission's abstract claims a quantum-inspired GNN with a CP-decomposition layer reaches 74.8% F2 on blockchain fraud detection, but the uploaded full text is an unrelated paper on VLM agent security.

  6. Understanding Knowledge Transferability for Transfer Learning: A Survey

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A survey that classifies transferability metrics by knowledge modality (dataset vs. model) and granularity (task vs. instance), with a theoretical primer and applications to eight learning paradigms.

Pith tools