Pith. sign in

REVIEW 1 cited by

Compression of Generative Pre-trained Language Models via Quantization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.10705 v2 pith:DFAUMTN7 submitted 2022-03-21 cs.CL cs.CV

classification cs.CLcs.CV
keywords generativecompressionplmscompressmethodsmodelsquantizationembeddings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing size of generative Pre-trained Language Models (PLMs) has greatly increased the demand for model compression. Despite various methods to compress BERT or its variants, there are few attempts to compress generative PLMs, and the underlying difficulty remains unclear. In this paper, we compress generative PLMs by quantization. We find that previous quantization methods fail on generative tasks due to the \textit{homogeneous word embeddings} caused by reduced capacity, and \textit{varied distribution of weights}. Correspondingly, we propose a token-level contrastive distillation to learn distinguishable word embeddings, and a module-wise dynamic scaling to make quantizers adaptive to different modules. Empirical results on various tasks show that our proposed method outperforms the state-of-the-art compression methods on generative PLMs by a clear margin. With comparable performance with the full-precision models, we achieve 14.4x and 13.4x compression rates on GPT-2 and BART, respectively.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Compression for Word and Sentence Embeddings using Discrete Wavelet Transform

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Keeping only the low-frequency DWT coefficients of word and sentence embeddings preserves most of their semantic quality at 50 to 93 percent fewer dimensions.

Pith tools