A gradually shrinking teacher signal during knowledge distillation improves 2-bit quantized ResNet20 accuracy on CIFAR-10 and CIFAR-100 compared with fixed-coefficient distillation.
Knowledge distillation using unlabeled mismatched images
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Current approaches for Knowledge Distillation (KD) either directly use training data or sample from the training data distribution. In this paper, we demonstrate effectiveness of 'mismatched' unlabeled stimulus to perform KD for image classification networks. For illustration, we consider scenarios where this is a complete absence of training data, or mismatched stimulus has to be used for augmenting a small amount of training data. We demonstrate that stimulus complexity is a key factor for distillation's good performance. Our examples include use of various datasets for stimulating MNIST and CIFAR teachers.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Knowledge distillation for optimization of quantized deep neural networks
A gradually shrinking teacher signal during knowledge distillation improves 2-bit quantized ResNet20 accuracy on CIFAR-10 and CIFAR-100 compared with fixed-coefficient distillation.