Distilling multilingual BERT into a 3-layer, 256-unit student yields a CPU-fast sequence labeler that is within about one F1 point of the teacher and beats a strong LSTM baseline.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Small and Practical BERT Models for Sequence Labeling
Distilling multilingual BERT into a 3-layer, 256-unit student yields a CPU-fast sequence labeler that is within about one F1 point of the teacher and beats a strong LSTM baseline.