Pith. sign in

REVIEW 5 cited by

PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1710.10766 v3 pith:DMOTVDA3 submitted 2017-10-30 cs.LG

classification cs.LG
keywords modelsimageadversarialmethodpixeldefendattackattackingclassifier
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial perturbations of normal images are usually imperceptible to humans, but they can seriously confuse state-of-the-art machine learning models. What makes them so special in the eyes of image classifiers? In this paper, we show empirically that adversarial examples mainly lie in the low probability regions of the training distribution, regardless of attack types and targeted models. Using statistical hypothesis testing, we find that modern neural density models are surprisingly good at detecting imperceptible image perturbations. Based on this discovery, we devised PixelDefend, a new approach that purifies a maliciously perturbed image by moving it back towards the distribution seen in the training data. The purified image is then run through an unmodified classifier, making our method agnostic to both the classifier and the attacking method. As a result, PixelDefend can be used to protect already deployed models and be combined with other model-specific defenses. Experiments show that our method greatly improves resilience across a wide variety of state-of-the-art attacking methods, increasing accuracy on the strongest attack from 63% to 84% for Fashion MNIST and from 32% to 70% for CIFAR-10.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking

    cs.RO 2026-08 conditional novelty 6.0 of 10

    AGSD patches make OpenVLA fail on all LIBERO suites; SARF fine-tuning cuts that failure rate from 100% to 28.6% average while preserving clean performance and improving real-robot success.

  2. On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Aligning adversarial perturbations with the near-null singular directions of intermediate linear layers in transformer VLMs yields stronger attacks than existing feature- and output-space methods.

  3. Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Silencer adds a nearly invisible disturbance to portraits that makes LDM-based talking-head models keep the mouth silent, and it survives several image-purification countermeasures.

  4. Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

    cs.CL 2024-02 conditional novelty 6.0 of 10

    DPOP is a new loss function that prevents DPO from lowering preferred response likelihoods and outperforms standard DPO on diverse datasets, MT-Bench, and enables Smaug-72B to exceed 80% on the Open LLM Leaderboard.

  5. Adaptive Mask-guided K-space Diffusion for Accelerated MRI Reconstruction

    eess.IV 2025-06 reject novelty 4.0 of 10

    AMDM reconstructs undersampled MRI by masking k-space frequency components with adaptive masks inside a diffusion model, and reports large PSNR gains over baseline methods.

Pith tools