Pith. sign in

REVIEW 1 cited by

A Survey on Model Compression and Acceleration for Pretrained Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.07105 v2 pith:X52R4IOA submitted 2022-02-15 cs.CL cs.AIcs.LG

A Survey on Model Compression and Acceleration for Pretrained Language Models

classification cs.CL cs.AIcs.LG
keywords includinginferencelanguagemodelmodelspretrainedaccelerationcompression
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Despite achieving state-of-the-art performance on many NLP tasks, the high energy cost and long inference delay prevent Transformer-based pretrained language models (PLMs) from seeing broader adoption including for edge and mobile computing. Efficient NLP research aims to comprehensively consider computation, time and carbon emission for the entire life-cycle of NLP, including data preparation, model training and inference. In this survey, we focus on the inference stage and review the current state of model compression and acceleration for pretrained language models, including benchmarks, metrics and methodology.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GenAI for Energy-Efficient and Interference-Aware Compressed Sensing of GNSS Signals on a Google Edge TPU

    cs.LG 2026-05 unverdicted novelty 3.0

    VAE models quantized for Edge TPUs deliver over 42x compression of GNSS signals with F2-score 0.915 for classifying 72 interference types, nearly matching uncompressed performance of 0.923.