On BabyLM 10M/100M data, teacher-less mutual learning with learned student weights improves over RoBERTa-base by 1-3% on BLiMP, yet simple self-distillation outperforms the proposed DWML.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
When Babies Teach Babies: Can student knowledge sharing outperform Teacher-Guided Distillation on small datasets?
On BabyLM 10M/100M data, teacher-less mutual learning with learned student weights improves over RoBERTa-base by 1-3% on BLiMP, yet simple self-distillation outperforms the proposed DWML.