Pith. sign in

REVIEW 1 cited by

Resource-Efficient Speech Quality Prediction through Quantization Aware Training and Binary Activation Maps

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04578 v1 pith:WLNYGVGQ submitted 2024-07-05 cs.SD cs.NEeess.AS

classification cs.SDcs.NEeess.AS
keywords speechbinaryqualityactivationquantizationawaredevicesmaps
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As speech processing systems in mobile and edge devices become more commonplace, the demand for unintrusive speech quality monitoring increases. Deep learning methods provide high-quality estimates of objective and subjective speech quality metrics. However, their significant computational requirements are often prohibitive on resource-constrained devices. To address this issue, we investigated binary activation maps (BAMs) for speech quality prediction on a convolutional architecture based on DNSMOS. We show that the binary activation model with quantization aware training matches the predictive performance of the baseline model. It further allows using other compression techniques. Combined with 8-bit weight quantization, our approach results in a 25-fold memory reduction during inference, while replacing almost all dot products with summations. Our findings show a path toward substantial resource savings by supporting mixed-precision binary multiplication in hard- and software.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scalable Speech Enhancement with Dynamic Channel Pruning

    eess.AS 2024-12 conditional novelty 5.0 of 10

    A custom convolutional speech enhancement network with a learned gating module skips individual channels at runtime, saving up to 29.6% of MACs on VoiceBank+DEMAND with a negligible PESQ drop.

Pith tools