Pith. sign in

REVIEW

USM RNN-T model weights binarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02887 v2 pith:5Q2A5SBM submitted 2024-06-05 eess.AS cs.SD

classification eess.AScs.SD
keywords modelsizecostweightsbinarizationgrowsmodelsonly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale universal speech models (USM) are already used in production. However, as the model size grows, the serving cost grows too. Serving cost of large models is dominated by model size that is why model size reduction is an important research topic. In this work we are focused on model size reduction using weights only quantization. We present the weights binarization of USM Recurrent Neural Network Transducer (RNN-T) and show that its model size can be reduced by 15.9x times at cost of word error rate (WER) increase by only 1.9% in comparison to the float32 model. It makes it attractive for practical applications.

Discussion (0). Continue with ORCID to comment.

Pith tools