REVIEW 27 cited by
Adversarial Attacks and Defences: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Deep learning has emerged as a strong and efficient framework that can be applied to a broad spectrum of complex learning problems which were difficult to solve using the traditional machine learning techniques in the past. In the last few years, deep learning has advanced radically in such a way that it can surpass human-level performance on a number of tasks. As a consequence, deep learning is being extensively used in most of the recent day-to-day applications. However, security of deep learning systems are vulnerable to crafted adversarial examples, which may be imperceptible to the human eye, but can lead the model to misclassify the output. In recent times, different types of adversaries based on their threat model leverage these vulnerabilities to compromise a deep learning system where adversaries have high incentives. Hence, it is extremely important to provide robustness to deep learning algorithms against these adversaries. However, there are only a few strong countermeasures which can be used in all types of attack scenarios to design a robust deep learning system. In this paper, we attempt to provide a detailed discussion on different types of adversarial attacks with various threat models and also elaborate the efficiency and challenges of recent countermeasures against them.
Forward citations
Cited by 27 Pith papers
-
Agnostic Learning under Targeted Poisoning: Optimal Rates and the Role of Randomness
The optimal excess error for agnostic learning under instance-targeted poisoning is eTheta(sqrt(d eta)), achieved by a randomized learner and unavoidable even against adversaries who see the learner's random seed.
-
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
AdvNav is a gradient-free attack that overlays Perlin noise on a VLN agent's camera and uses behavior feedback plus genetic search, breaking 49.70-87.30% of successful R2R navigations.
-
Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks
A structured dual-target attack can force targeted misclassification of time series while keeping the explainer aligned with a reference rationale, showing explanation stability is not a reliable robustness proxy.
-
Exploring Visual Prompting: Robustness Inheritance and Beyond
Visual prompts built on robust source models inherit adversarial robustness but lose standard accuracy; a max-pooling over logit blocks (PBL) improves accuracy while keeping most robustness.
-
FABLE: A Localized, Targeted Adversarial Attack on Weather Forecasting Models
FABLE applies 3D discrete wavelet decomposition to generate localized adversarial perturbations that steer deep learning weather forecasting models toward chosen forecast outcomes while keeping inputs close to the originals.
-
Adversarial Coevolutionary Illumination with Generational Adversarial MAP-Elites
GAME is a new adversarial coevolutionary QD algorithm using generational alternation and vision embeddings that outperforms one-sided baselines across battle, wrestling, and deck-building tasks while revealing arms-ra...
-
Fooling the Decoder: An Adversarial Attack on Quantum Error Correction
A white-box attack selects the syndrome volumes a trained RL surface-code decoder is most pessimistic about, shrinking its average logical qubit lifetime from 100,000 to 60 cycles.
-
Comprehensive Evaluation of Cloaking Backdoor Attacks on Object Detector in Real-World
A blue bear-logo T-shirt used as a data-poisoning trigger can erase people from YOLO, CenterNet, and even Faster R-CNN detectors in real video with near-100% success and no drop in clean accuracy.
-
Jailbroken: How Does LLM Safety Training Fail?
LLM safety training fails due to competing objectives and mismatched generalization, enabling new jailbreaks that succeed on all unsafe prompts from red-teaming sets in GPT-4 and Claude.
-
Scaling Laws for Reward Model Overoptimization
Synthetic measurements show that gold-standard performance degrades according to distinct functional forms when optimizing proxy reward models via RL or best-of-n, with coefficients scaling smoothly by reward model pa...
-
Key Protected Classification for Collaborative Learning
Class-specific private random keys, combined with a fixed random projection layer, stop random-key GAN attacks in collaborative learning, but the security guarantee does not cover adversaries that can estimate keys fr...
-
Signal-based Model Access Risk Analysis for AI System Operations Security
Introduces SMART, a six-level signal-based access taxonomy linking AI deployment interfaces to evasion risk and procurement guidance.
-
Stabilizing Data-Free Model Extraction
MetaDFME stabilizes data-free model extraction by training the generator with Reptile-style meta-learning, achieving higher and less oscillating substitute accuracy.
-
Diffusion-based Cumulative Adversarial Purification for Vision Language Models
DiffCAP purifies adversarial images for vision-language models by injecting cumulative Gaussian noise until embeddings stabilize, then denoising, and outperforms prior defenses on captioning, VQA, and classification b...
-
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
Hydra, an agentic reasoning framework that iteratively queries and critiques multiple vision models, improves object-level hallucination accuracy and maintains performance under adversarial image attacks on four LVLMs...
-
Survival of the Cheapest: Cost-Aware Hardware Adaptation for Adversarial Robustness
A decision-support framework applies AFT models to show Nvidia L4 GPUs yield 20% longer adversarial survival time at 75% lower cost than V100, with inference latency as the strongest robustness predictor.
-
Trojan Attacks on Neural Network Controllers for Robotic Systems
A parallel 'Trojan' neural network that multiply-gates wheel-speed commands can silently immobilize or dangerously accelerate a differential-drive robot inside a chosen trigger region, shown in simulation.
-
Position: Certified Robustness Does Not (Yet) Imply Model Security
A certified robustness radius says nothing about whether a sample is clean or correctly predicted, so certification does not yet imply model security.
-
Assessing the Resilience of Automotive Intrusion Detection Systems to Adversarial Manipulation
Gradient-based evasion attacks can lower the detection rate of CAN-bus intrusion detection systems, with effectiveness depending on attacker knowledge, dataset, and detector architecture.
-
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.
-
Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters
An RL-based attack platform with custom distortion filters reports dramatically lower query counts, but its query accounting and victim-specific training undermine the comparison.
-
Safety Monitoring of Machine Learning Perception Functions: a Survey
A survey that organizes research on runtime safety monitors for ML perception into threat identification, requirements, detection, reaction, and evaluation, and lists open challenges.
-
Fall Leaf Adversarial Attack on Traffic Sign Classification
A leaf-shaped occlusion can flip some traffic sign classifications in the LISA-CNN model, but the evidence is limited to five signs and best-case placements.
-
Crafting Adversarial Examples for Deep Learning Based Prognostics (Extended Version)
Adversarial examples crafted with FGSM and BIM substantially degrade remaining useful life predictions in LSTM, GRU, and CNN prognostics models on the NASA C-MAPSS FD001 dataset, and these attacks transfer across models.
-
Protego: Detecting Adversarial Examples for Vision Transformers via Intrinsic Capabilities
PROTEGO is a proposed plug-in detector that classifies adversarial versus normal images for Vision Transformers using the difference between adversarial and clean token features, reporting AUC over 0.95.
-
C-LEAD: Contrastive Learning for Enhanced Adversarial Defense
Contrastive learning with adversarial perturbations as positive pairs improves robustness of ResNet models on CIFAR-10, but evidence is weakened by missing baselines and inconsistent reporting.
-
Emerging Security Challenges of Large Language Models
A workshop report summarizes adversarial attack surfaces of LLMs, from training data poisoning to prompt injection, and calls for more research on defenses.
Discussion (0). Continue with ORCID to comment.