Supervising two Transformer self-attention heads with dependency-tree parent/child adjacency matrices improves BLEU by 0.5-1.5 points across four translation pairs.
Neural machine translation of rare words with subword units
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Source Dependency-Aware Transformer with Supervised Self-Attention
Supervising two Transformer self-attention heads with dependency-tree parent/child adjacency matrices improves BLEU by 0.5-1.5 points across four translation pairs.