A gradually shrinking teacher signal during knowledge distillation improves 2-bit quantized ResNet20 accuracy on CIFAR-10 and CIFAR-100 compared with fixed-coefficient distillation.
We found that the teacher needs not be a quantized neural network
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Knowledge distillation for optimization of quantized deep neural networks
A gradually shrinking teacher signal during knowledge distillation improves 2-bit quantized ResNet20 accuracy on CIFAR-10 and CIFAR-100 compared with fixed-coefficient distillation.