Pith. sign in

REVIEW 2 cited by

Efficient and Effective Methods for Mixed Precision Neural Network Quantization for Faster, Energy-efficient Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.13330 v2 pith:TRKB3S4E submitted 2023-01-30 cs.LG cs.CV

classification cs.LGcs.CV
keywords precisionlayernetworksperformanceaccuracymethodsnetworkquantization
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

For efficient neural network inference, it is desirable to achieve state-of-the-art accuracy with the simplest networks requiring the least computation, memory, and power. Quantizing networks to lower precision is a powerful technique for simplifying networks. As each layer of a network may have different sensitivity to quantization, mixed precision quantization methods selectively tune the precision of individual layers to achieve a minimum drop in task performance (e.g., accuracy). To estimate the impact of layer precision choice on task performance, two methods are introduced: i) Entropy Approximation Guided Layer selection (EAGL) is fast and uses the entropy of the weight distribution, and ii) Accuracy-aware Layer Precision Selection (ALPS) is straightforward and relies on single epoch fine-tuning after layer precision reduction. Using EAGL and ALPS for layer precision selection, full-precision accuracy is recovered with a mix of 4-bit and 2-bit layers for ResNet-50, ResNet-101 and BERT-base transformer networks, demonstrating enhanced performance across the entire accuracy-throughput frontier. The techniques demonstrate better performance than existing techniques in several commensurate comparisons. Notably, this is accomplished with significantly lesser computational time required to reach a solution.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A probabilistic framework for dynamic quantization

    cs.LG 2025-05 conditional novelty 6.0 of 10

    An input-adaptive 8-bit quantization scheme estimates activation ranges via a lightweight probabilistic surrogate, achieving near-dynamic accuracy with static-like memory overhead.

  2. ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs

    cs.CL 2025-04 conditional novelty 5.0 of 10

    ImPart sparsifies delta weights in SVD space by assigning lower drop rates to high-importance singular vectors, reporting better accuracy than DARE and LowRank at high compression ratios.

Pith tools