Pith. sign in

REVIEW

On the Sparsity of Neural Machine Translation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02646 v1 pith:CBM5ILKG submitted 2020-10-06 cs.CL

classification cs.CL
keywords parametersmachinemodelsneuralrejuvenatedtranslationabilityachieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern neural machine translation (NMT) models employ a large number of parameters, which leads to serious over-parameterization and typically causes the underutilization of computational resources. In response to this problem, we empirically investigate whether the redundant parameters can be reused to achieve better performance. Experiments and analyses are systematically conducted on different datasets and NMT architectures. We show that: 1) the pruned parameters can be rejuvenated to improve the baseline model by up to +0.8 BLEU points; 2) the rejuvenated parameters are reallocated to enhance the ability of modeling low-level lexical information.

Discussion (0). Continue with ORCID to comment.

Pith tools