Pith. sign in

Gradient Similarity: An Explainable Approach to Detect Adversarial Attacks against Deep Learning

1 Pith paper cite this work, alongside 7 external citations. Polarity classification is still indexing.

1 Pith paper citing it
7 external citations · Pith
abstract

Deep neural networks are susceptible to small-but-specific adversarial perturbations capable of deceiving the network. This vulnerability can lead to potentially harmful consequences in security-critical applications. To address this vulnerability, we propose a novel metric called \emph{Gradient Similarity} that allows us to capture the influence of training data on test inputs. We show that \emph{Gradient Similarity} behaves differently for normal and adversarial inputs, and enables us to detect a variety of adversarial attacks with a near perfect ROC-AUC of 95-100\%. Even white-box adversaries equipped with perfect knowledge of the system cannot bypass our detector easily. On the MNIST dataset, white-box attacks are either detected with a high ROC-AUC of 87-96\%, or require very high distortion to bypass our detector.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Influence Dynamics and Stagewise Data Attribution

cs.LG · 2025-10-14 · conditional · novelty 7.0

Using Bayesian influence functions and singular learning theory, the authors show that a sample's influence on a model varies non-monotonically over training, peaking and flipping sign at phase transitions.

citing papers explorer

Showing 1 of 1 citing paper.

  • Influence Dynamics and Stagewise Data Attribution cs.LG · 2025-10-14 · conditional · none · ref 58 · internal anchor

    Using Bayesian influence functions and singular learning theory, the authors show that a sample's influence on a model varies non-monotonically over training, peaking and flipping sign at phase transitions.