Pith. sign in

Adversarial Phenomenon in the Eyes of Bayesian Deep Learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Deep Learning models are vulnerable to adversarial examples, i.e.\ images obtained via deliberate imperceptible perturbations, such that the model misclassifies them with high confidence. However, class confidence by itself is an incomplete picture of uncertainty. We therefore use principled Bayesian methods to capture model uncertainty in prediction for observing adversarial misclassification. We provide an extensive study with different Bayesian neural networks attacked in both white-box and black-box setups. The behaviour of the networks for noise, attacks and clean test data is compared. We observe that Bayesian neural networks are uncertain in their predictions for adversarial perturbations, a behaviour similar to the one observed for random Gaussian perturbations. Thus, we conclude that Bayesian neural networks can be considered for detecting adversarial examples.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

OFAL: An Oracle-Free Active Learning Framework

cs.LG · 2025-08-11 · conditional · novelty 6.0

OFAL improves an MNIST classifier from 93.0% to 95.7% test accuracy without oracle labels, by generating uncertain synthetic samples from confident seeds using a VAE and dropout uncertainty.

citing papers explorer

Showing 1 of 1 citing paper.

  • OFAL: An Oracle-Free Active Learning Framework cs.LG · 2025-08-11 · conditional · none · ref 16 · internal anchor

    OFAL improves an MNIST classifier from 93.0% to 95.7% test accuracy without oracle labels, by generating uncertain synthetic samples from confident seeds using a VAE and dropout uncertainty.