Stochastic self-distillation generates multiple teacher representations through dropout and uses student-guided attention to distill task-relevant knowledge, improving accuracy over single-model baselines at no extra inference cost.
IEEE/CAA Journal of Automatica Sinica (2019)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning from Stochastic Teacher Representations Using Student-Guided Knowledge Distillation
Stochastic self-distillation generates multiple teacher representations through dropout and uses student-guided attention to distill task-relevant knowledge, improving accuracy over single-model baselines at no extra inference cost.