A single NxM Transformer model trained with softmax losses from all layer combinations can decode with any smaller layer count, closely approximating 36 separately trained models.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-Layer Softmaxing during Training Neural Machine Translation for Flexible Decoding with Fewer Layers
A single NxM Transformer model trained with softmax losses from all layer combinations can decode with any smaller layer count, closely approximating 36 separately trained models.