Pith. sign in

REVIEW 12 cited by

BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.12075 v3 pith:O7J57WLO submitted 2023-11-20 cs.CV

classification cs.CV
keywords backdoorattacksdefenseslearningattackcontrastivemultimodalpatterns
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Studying backdoor attacks is valuable for model copyright protection and enhancing defenses. While existing backdoor attacks have successfully infected multimodal contrastive learning models such as CLIP, they can be easily countered by specialized backdoor defenses for MCL models. This paper reveals the threats in this practical scenario that backdoor attacks can remain effective even after defenses and introduces the \emph{\toolns} attack, which is resistant to backdoor detection and model fine-tuning defenses. To achieve this, we draw motivations from the perspective of the Bayesian rule and propose a dual-embedding guided framework for backdoor attacks. Specifically, we ensure that visual trigger patterns approximate the textual target semantics in the embedding space, making it challenging to detect the subtle parameter variations induced by backdoor learning on such natural trigger patterns. Additionally, we optimize the visual trigger patterns to align the poisoned samples with target vision features in order to hinder the backdoor unlearning through clean fine-tuning. Extensive experiments demonstrate that our attack significantly outperforms state-of-the-art baselines (+45.3% ASR) in the presence of SoTA backdoor defenses, rendering these mitigation and detection strategies virtually ineffective. Furthermore, our approach effectively attacks some more rigorous scenarios like downstream tasks. We believe that this paper raises awareness regarding the potential threats associated with the practical application of multimodal contrastive learning and encourages the development of more robust defense mechanisms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous Vehicles

    cs.CV 2025-08 conditional novelty 6.0 of 10

    ScenGE generates more collision-prone autonomous driving test scenarios by combining LLM-suggested adversarial events with optimized background traffic, beating prior generators on CARLA benchmarks.

  2. PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    PromptSafe uses LLM-rewritten safe prompts to train a gated soft prompt that suppresses NSFW content in text-to-image generation without image supervision or inference overhead.

  3. No Query, No Access

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new attack scenario, Victim Data-based Attack (VDBA), generates transferable adversarial examples for text classifiers using only unlabeled victim texts, achieving over 40% attack success on several models with zero...

  4. T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    An LLM-driven discrete optimization with prompt mutation can rewrite unsafe prompts to bypass text-to-video safety filters and produce harmful videos with higher success than existing methods.

  5. Manipulating Multimodal Agents via Cross-Modal Prompt Injection

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A coordinated attack that embeds malicious cues in both visual and textual inputs can hijack black-box multimodal agents, outperforming single-modality prompt injection attacks.

  6. Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving

    cs.CV 2025-01 conditional novelty 6.0 of 10

    CAD is a transfer-based black-box attack using CLIP embeddings and ChatGPT-generated deceptive reasoning text to make vision-language autonomous driving models take unsafe actions.

  7. Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Environmental illusions cause 5-7% accuracy drops in lane detection models and can trigger collisions in closed-loop simulation, with a proposed defense (MIDA) recovering ~4% robustness.

  8. Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MSEA+ARC, a multi-scale and ranking-based residualization method, claims consistent F1-IoU gains over TAM for token-level MLLM visual attribution.

  9. Multimodal Fine-grained Reasoning for Post Quality Evaluation

    cs.LG 2025-07 reject novelty 5.0 of 10

    MFTRR combines local-global cross-modal attention, gating, and graph-based evidence reasoning to rank forum post quality, reporting NDCG@3 gains of up to 9.5 points over text-only baselines on new private datasets.

  10. CopyrightShield: Enhancing Diffusion Model Security against Copyright Infringement Attacks

    cs.AI 2024-12 reject novelty 5.0 of 10

    A defense framework that uses masked image similarity and data attribution to detect and mitigate copyright-infringing backdoor samples in diffusion models.

  11. Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Training a driving VLM on a small set of images with faint reflection overlays and long prefixed answers makes the model generate verbose responses on triggered reflections, increasing response length while leaving cl...

  12. Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents

    cs.AI 2024-11 conditional novelty 4.0 of 10

    A survey proposing a source-and-impact taxonomy (input, model, combined; security, privacy, ethics) for threats to LLM-based agents, with feature analysis and four case studies.

Pith tools