Sequence-level knowledge-distilled NMT students memorize more of the original corpus and hallucinate more than same-size baselines trained directly on that corpus, despite never seeing it.
Thank you for your visit at our website
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Memorization Inheritance in Sequence-Level Knowledge Distillation for Neural Machine Translation
Sequence-level knowledge-distilled NMT students memorize more of the original corpus and hallucinate more than same-size baselines trained directly on that corpus, despite never seeing it.