Pith. sign in

REVIEW 8 cited by

Poisoning and Backdooring Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09667 v2 pith:QQXQN67S submitted 2021-06-17 cs.LG

classification cs.LG
keywords poisoningattacksdatasetimagesjustcontrastivedatasetseven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal contrastive learning methods like CLIP train on noisy and uncurated training datasets. This is cheaper than labeling datasets manually, and even improves out-of-distribution robustness. We show that this practice makes backdoor and poisoning attacks a significant threat. By poisoning just 0.01% of a dataset (e.g., just 300 images of the 3 million-example Conceptual Captions dataset), we can cause the model to misclassify test images by overlaying a small patch. Targeted poisoning attacks, whereby the model misclassifies a particular test input with an adversarially-desired label, are even easier requiring control of 0.0001% of the dataset (e.g., just three out of the 3 million images). Our attacks call into question whether training on noisy and uncurated Internet scrapes is desirable.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DisarmRAG: Stealthy Retriever-Centric Poisoning to Disable Self-Correction in Retrieval-Augmented Generation (Extended Version)

    cs.CR 2025-08 conditional novelty 7.0 of 10

    DisarmRAG compromises the retriever to inject anti-self-correction instructions, achieving over 90% attack success across six LLMs while evading basic detection.

  2. VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems

    cs.CR 2026-02 conditional novelty 6.0 of 10

    Synchronized text+image poisoning steers multimodal LLM recommenders to promote target items, reaching 0.73 mean exposure@20.

  3. Dataset Ownership Verification for Pre-trained Masked Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DOV4MM detects whether a masked pre-trained model was trained on a given dataset via relative embedding reconstruction difficulty, reporting p<0.05 in tests on ImageNet-1K and WikiText-103.

  4. InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning

    cs.CR 2025-06 conditional novelty 6.0 of 10

    InverTune removes backdoors from CLIP models by identifying the target label via adversarial perturbations, inverting the trigger, and selectively tuning backdoor-sensitive neurons, reducing attack success rates to ne...

  5. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  6. M2Restore: Mixture-of-Experts-based Mamba-CNN Fusion Framework for All-in-One Image Restoration

    cs.CV 2025-06 conditional novelty 5.0 of 10

    M2Restore is a CLIP-guided Mixture-of-Experts Mamba-CNN model that reports state-of-the-art results on the All-weather all-in-one image restoration benchmark.

  7. Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise

    cs.AI 2025-05 reject novelty 4.0 of 10

    A multi-agent router that selects culturally specialized LLM personas reports a jump in self-scored cultural alignment from 0.208 to 0.820, but the metric and the claimed method are not independently validated.

  8. When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs

    cs.CV 2025-02 conditional novelty 3.0 of 10

    A survey that classifies VLM attacks by goal and data manipulation strategy, and reviews defenses and metrics.

Pith tools