Pith. sign in

REVIEW 27 cited by

Adversarial Attacks and Defences: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.00069 v1 pith:POL5T4J7 submitted 2018-09-28 cs.LG cs.CRstat.ML

classification cs.LGcs.CRstat.ML
keywords learningdeepadversarialadversariesrecenttypesattackscountermeasures
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep learning has emerged as a strong and efficient framework that can be applied to a broad spectrum of complex learning problems which were difficult to solve using the traditional machine learning techniques in the past. In the last few years, deep learning has advanced radically in such a way that it can surpass human-level performance on a number of tasks. As a consequence, deep learning is being extensively used in most of the recent day-to-day applications. However, security of deep learning systems are vulnerable to crafted adversarial examples, which may be imperceptible to the human eye, but can lead the model to misclassify the output. In recent times, different types of adversaries based on their threat model leverage these vulnerabilities to compromise a deep learning system where adversaries have high incentives. Hence, it is extremely important to provide robustness to deep learning algorithms against these adversaries. However, there are only a few strong countermeasures which can be used in all types of attack scenarios to design a robust deep learning system. In this paper, we attempt to provide a detailed discussion on different types of adversarial attacks with various threat models and also elaborate the efficiency and challenges of recent countermeasures against them.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agnostic Learning under Targeted Poisoning: Optimal Rates and the Role of Randomness

    cs.LG 2025-06 conditional novelty 8.0 of 10

    The optimal excess error for agnostic learning under instance-targeted poisoning is eTheta(sqrt(d eta)), achieved by a randomized learner and unavoidable even against adversaries who see the learner's random seed.

  2. AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    AdvNav is a gradient-free attack that overlays Perlin noise on a VLN agent's camera and uses behavior feedback plus genetic search, breaking 49.70-87.30% of successful R2R navigations.

  3. Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A structured dual-target attack can force targeted misclassification of time series while keeping the explainer aligned with a reference rationale, showing explanation stability is not a reliable robustness proxy.

  4. Exploring Visual Prompting: Robustness Inheritance and Beyond

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Visual prompts built on robust source models inherit adversarial robustness but lose standard accuracy; a max-pooling over logit blocks (PBL) improves accuracy while keeping most robustness.

  5. FABLE: A Localized, Targeted Adversarial Attack on Weather Forecasting Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    FABLE applies 3D discrete wavelet decomposition to generate localized adversarial perturbations that steer deep learning weather forecasting models toward chosen forecast outcomes while keeping inputs close to the originals.

  6. Adversarial Coevolutionary Illumination with Generational Adversarial MAP-Elites

    cs.NE 2025-05 unverdicted novelty 6.0 of 10

    GAME is a new adversarial coevolutionary QD algorithm using generational alternation and vision embeddings that outperforms one-sided baselines across battle, wrestling, and deck-building tasks while revealing arms-ra...

  7. Fooling the Decoder: An Adversarial Attack on Quantum Error Correction

    quant-ph 2025-04 conditional novelty 6.0 of 10

    A white-box attack selects the syndrome volumes a trained RL surface-code decoder is most pessimistic about, shrinking its average logical qubit lifetime from 100,000 to 60 cycles.

  8. Comprehensive Evaluation of Cloaking Backdoor Attacks on Object Detector in Real-World

    cs.CR 2025-01 conditional novelty 6.0 of 10

    A blue bear-logo T-shirt used as a data-poisoning trigger can erase people from YOLO, CenterNet, and even Faster R-CNN detectors in real video with near-100% success and no drop in clean accuracy.

  9. Jailbroken: How Does LLM Safety Training Fail?

    cs.LG 2023-07 unverdicted novelty 6.0 of 10

    LLM safety training fails due to competing objectives and mismatched generalization, enabling new jailbreaks that succeed on all unsafe prompts from red-teaming sets in GPT-4 and Claude.

  10. Scaling Laws for Reward Model Overoptimization

    cs.LG 2022-10 unverdicted novelty 6.0 of 10

    Synthetic measurements show that gold-standard performance degrades according to distinct functional forms when optimizing proxy reward models via RL or best-of-n, with coefficients scaling smoothly by reward model pa...

  11. Key Protected Classification for Collaborative Learning

    cs.LG 2019-08 reject novelty 6.0 of 10

    Class-specific private random keys, combined with a fixed random projection layer, stop random-key GAN attacks in collaborative learning, but the security guarantee does not cover adversaries that can estimate keys fr...

  12. Signal-based Model Access Risk Analysis for AI System Operations Security

    cs.CR 2026-07 conditional novelty 5.0 of 10

    Introduces SMART, a six-level signal-based access taxonomy linking AI deployment interfaces to evasion risk and procurement guidance.

  13. Stabilizing Data-Free Model Extraction

    cs.LG 2025-09 conditional novelty 5.0 of 10

    MetaDFME stabilizes data-free model extraction by training the generator with Reptile-style meta-learning, achieving higher and less oscillating substitute accuracy.

  14. Diffusion-based Cumulative Adversarial Purification for Vision Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DiffCAP purifies adversarial images for vision-language models by injecting cumulative Gaussian noise until embeddings stabilize, then denoising, and outperforms prior defenses on captioning, VQA, and classification b...

  15. Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models

    cs.CV 2025-04 conditional novelty 5.0 of 10

    Hydra, an agentic reasoning framework that iteratively queries and critiques multiple vision models, improves object-level hallucination accuracy and maintains performance under adversarial image attacks on four LVLMs...

  16. Survival of the Cheapest: Cost-Aware Hardware Adaptation for Adversarial Robustness

    cs.CR 2024-09 unverdicted novelty 5.0 of 10

    A decision-support framework applies AFT models to show Nvidia L4 GPUs yield 20% longer adversarial survival time at 75% lower cost than V100, with inference latency as the strongest robustness predictor.

  17. Trojan Attacks on Neural Network Controllers for Robotic Systems

    eess.SY 2026-02 conditional novelty 4.0 of 10

    A parallel 'Trojan' neural network that multiply-gates wheel-speed commands can silently immobilize or dangerously accelerate a differential-drive robot inside a chosen trigger region, shown in simulation.

  18. Position: Certified Robustness Does Not (Yet) Imply Model Security

    cs.CR 2025-06 conditional novelty 4.0 of 10

    A certified robustness radius says nothing about whether a sample is clean or correctly predicted, so certification does not yet imply model security.

  19. Assessing the Resilience of Automotive Intrusion Detection Systems to Adversarial Manipulation

    cs.CR 2025-06 conditional novelty 4.0 of 10

    Gradient-based evasion attacks can lower the detection rate of CAN-bus intrusion detection systems, with effectiveness depending on attacker knowledge, dataset, and detector architecture.

  20. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.

  21. Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters

    cs.LG 2025-01 reject novelty 4.0 of 10

    An RL-based attack platform with custom distortion filters reports dramatically lower query counts, but its query accounting and victim-specific training undermine the comparison.

  22. Safety Monitoring of Machine Learning Perception Functions: a Survey

    cs.LG 2024-12 accept novelty 4.0 of 10

    A survey that organizes research on runtime safety monitors for ML perception into threat identification, requirements, detection, reaction, and evaluation, and lists open challenges.

  23. Fall Leaf Adversarial Attack on Traffic Sign Classification

    cs.CV 2024-11 reject novelty 4.0 of 10

    A leaf-shaped occlusion can flip some traffic sign classifications in the LISA-CNN model, but the evidence is limited to five signs and best-case placements.

  24. Crafting Adversarial Examples for Deep Learning Based Prognostics (Extended Version)

    cs.LG 2020-09 conditional novelty 4.0 of 10

    Adversarial examples crafted with FGSM and BIM substantially degrade remaining useful life predictions in LSTM, GRU, and CNN prognostics models on the NASA C-MAPSS FD001 dataset, and these attacks transfer across models.

  25. Protego: Detecting Adversarial Examples for Vision Transformers via Intrinsic Capabilities

    cs.CV 2025-01 reject novelty 3.0 of 10

    PROTEGO is a proposed plug-in detector that classifies adversarial versus normal images for Vision Transformers using the difference between adversarial and clean token features, reporting AUC over 0.95.

  26. C-LEAD: Contrastive Learning for Enhanced Adversarial Defense

    cs.CV 2025-10 reject novelty 2.0 of 10

    Contrastive learning with adversarial perturbations as positive pairs improves robustness of ResNet models on CIFAR-10, but evidence is weakened by missing baselines and inconsistent reporting.

  27. Emerging Security Challenges of Large Language Models

    cs.CR 2024-12 unverdicted

    A workshop report summarizes adversarial attack surfaces of LLMs, from training data poisoning to prompt injection, and calls for more research on defenses.

Pith tools