Pith. sign in

REVIEW 3 cited by

And the Bit Goes Down: Revisiting the Quantization of Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.05686 v5 pith:7QYIN62N submitted 2019-07-12 cs.CV

classification cs.CV
keywords quantizationapproachfactormemorymethodnetworkpreservingreconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we address the problem of reducing the memory footprint of convolutional network architectures. We introduce a vector quantization method that aims at preserving the quality of the reconstruction of the network outputs rather than its weights. The principle of our approach is that it minimizes the loss reconstruction error for in-domain inputs. Our method only requires a set of unlabelled data at quantization time and allows for efficient inference on CPU by using byte-aligned codebooks to store the compressed weights. We validate our approach by quantizing a high performing ResNet-50 model to a memory size of 5MB (20x compression factor) while preserving a top-1 accuracy of 76.1% on ImageNet object classification and by compressing a Mask R-CNN with a 26x factor.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVQ:Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Masked vector quantization (MVQ) prunes unimportant weights before clustering and uses masked k-means to build codebooks, improving accuracy over conventional VQ while cutting FLOPs and enabling a smaller, more effici...

  2. VQ4ALL: Efficient Neural Network Representation via a Universal Codebook

    cs.LG 2024-12 conditional novelty 6.0 of 10

    VQ4ALL builds a single universal codebook from the weight distributions of several networks, then learns per-network assignments to reconstruct low-bit weights while keeping accuracy close to the original models.

  3. Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects

    eess.SY 2025-09 conditional novelty 2.0 of 10

    A review paper synthesizes evidence that AI data center electricity demand is large, bursty, and power-electronics-dominated, creating multi-timescale grid challenges.

Pith tools