Pith. sign in

REVIEW

Transformer on a Diet

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.06170 v1 pith:KXDFJ2B4 submitted 2020-02-14 cs.CL cs.LG

classification cs.CLcs.LG
keywords transformerarchitecturescompetitivelightresultsabilityavailablebeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer has been widely used thanks to its ability to capture sequence information in an efficient way. However, recent developments, such as BERT and GPT-2, deliver only heavy architectures with a focus on effectiveness. In this paper, we explore three carefully-designed light Transformer architectures to figure out whether the Transformer with less computations could produce competitive results. Experimental results on language model benchmark datasets hint that such trade-off is promising, and the light Transformer reduces 70% parameters at best, while obtains competitive perplexity compared to standard Transformer. The source code is publicly available.

Discussion (0). Sign in to comment.

Pith tools