A 3-layer student transformer distilled from compressed BERT vectors retains roughly 90 percent of teacher performance on Math23K, according to the authors, though the teacher baseline is not shown.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
We Need Knowledge Distillation for Solving Math Word Problems
A 3-layer student transformer distilled from compressed BERT vectors retains roughly 90 percent of teacher performance on Math23K, according to the authors, though the teacher baseline is not shown.