Pith. sign in

REVIEW 1 cited by

Quantization-Aware and Tensor-Compressed Training of Transformers for Natural Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01076 v2 pith:NBIWZ3MI submitted 2023-06-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords trainingmodelmodelstensor-compressedlanguagenaturalquantization-awaretransformer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Fine-tuned transformer models have shown superior performances in many natural language tasks. However, the large model size prohibits deploying high-performance transformer models on resource-constrained devices. This paper proposes a quantization-aware tensor-compressed training approach to reduce the model size, arithmetic operations, and ultimately runtime latency of transformer-based models. We compress the embedding and linear layers of transformers into small low-rank tensor cores, which significantly reduces model parameters. A quantization-aware training with learnable scale factors is used to further obtain low-precision representations of the tensor-compressed models. The developed approach can be used for both end-to-end training and distillation-based training. To improve the convergence, a layer-by-layer distillation is applied to distill a quantized and tensor-compressed student model from a pre-trained transformer. The performance is demonstrated in two natural language understanding tasks, showing up to $63\times$ compression ratio, little accuracy loss and remarkable inference and training speedup.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A tensor-compressed transformer training accelerator on FPGA that stores all parameters on-chip, claiming 20-51x memory reduction and up to 4x energy savings per epoch versus an RTX 3090.

Pith tools