Stochastic self-distillation generates multiple teacher representations through dropout and uses student-guided attention to distill task-relevant knowledge, improving accuracy over single-model baselines at no extra inference cost.
In: Neural IPS Datasets and Benchmarks Track (2023)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning from Stochastic Teacher Representations Using Student-Guided Knowledge Distillation
Stochastic self-distillation generates multiple teacher representations through dropout and uses student-guided attention to distill task-relevant knowledge, improving accuracy over single-model baselines at no extra inference cost.