Pith. sign in

REVIEW 2 cited by

Mixed Low-precision Deep Learning Inference using Dynamic Fixed Point

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1701.08978 v2 pith:OL6ATQWI submitted 2017-01-31 cs.LG cs.NE

classification cs.LGcs.NE
keywords accuracyfullprecisionweightsternaryclustermethodtop-1
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a cluster-based quantization method to convert pre-trained full precision weights into ternary weights with minimal impact on the accuracy. In addition, we also constrain the activations to 8-bits thus enabling sub 8-bit full integer inference pipeline. Our method uses smaller clusters of N filters with a common scaling factor to minimize the quantization loss, while also maximizing the number of ternary operations. We show that with a cluster size of N=4 on Resnet-101, can achieve 71.8% TOP-1 accuracy, within 6% of the best full precision results while replacing ~85% of all multiplications with 8-bit accumulations. Using the same method with 4-bit weights achieves 76.3% TOP-1 accuracy which within 2% of the full precision result. We also study the impact of the size of the cluster on both performance and accuracy, larger cluster sizes N=64 can replace ~98% of the multiplications with ternary operations but introduces significant drop in accuracy which necessitates fine tuning the parameters with retraining the network at lower precision. To address this we have also trained low-precision Resnet-50 with 8-bit activations and ternary weights by pre-initializing the network with full precision weights and achieve 68.9% TOP-1 accuracy within 4 additional epochs. Our final quantized model can run on a full 8-bit compute pipeline, with a potential 16x improvement in performance compared to baseline full-precision models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. tuGEMM: Area-Power-Efficient Temporal Unary GEMM Architecture for Low-Precision Edge AI

    cs.AR 2024-12 conditional novelty 6.0 of 10

    A temporal unary GEMM architecture performs exact low-precision matrix multiply with post-synthesis area and power reductions of up to 15x and 11x versus stochastic uGEMM, at the cost of data-dependent latency.

  2. Pushing the Limits of BFP on Narrow Precision LLM Inference

    cs.AR 2025-01 conditional novelty 5.0 of 10

    A dynamic block floating-point format with pivot-focus and adaptive grouping, plus a hierarchical lookup table, lets attention Softmax run in integer-only hardware with negligible accuracy loss.

Pith tools