Pith. sign in

REVIEW 99 cited by

Quantizing deep convolutional networks for efficient inference: A whitepaper

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.08342 v1 pith:D44WXUKW submitted 2018-06-21 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords weightsnetworksquantizationactivationsbitspointquantizingtraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present an overview of techniques for quantizing convolutional neural networks for inference with integer weights and activations. Per-channel quantization of weights and per-layer quantization of activations to 8-bits of precision post-training produces classification accuracies within 2% of floating point networks for a wide variety of CNN architectures. Model sizes can be reduced by a factor of 4 by quantizing weights to 8-bits, even when 8-bit arithmetic is not supported. This can be achieved with simple, post training quantization of weights.We benchmark latencies of quantized networks on CPUs and DSPs and observe a speedup of 2x-3x for quantized implementations compared to floating point on CPUs. Speedups of up to 10x are observed on specialized processors with fixed point SIMD capabilities, like the Qualcomm QDSPs with HVX. Quantization-aware training can provide further improvements, reducing the gap to floating point to 1% at 8-bit precision. Quantization-aware training also allows for reducing the precision of weights to four bits with accuracy losses ranging from 2% to 10%, with higher accuracy drop for smaller networks.We introduce tools in TensorFlow and TensorFlowLite for quantizing convolutional networks and review best practices for quantization-aware training to obtain high accuracy with quantized weights and activations. We recommend that per-channel quantization of weights and per-layer quantization of activations be the preferred quantization scheme for hardware acceleration and kernel optimization. We also propose that future processors and hardware accelerators for optimized inference support precisions of 4, 8 and 16 bits.

Discussion (0). Continue with ORCID to comment.

Forward citations

Showing 60 of 99 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 765 citations worldwide. See all 99 Pith citations

  1. Neural Network Quantization by Learning Low-Loss Subspaces

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Learning quantization-aware linear paths in weight space yields a midpoint whose direct quantization matches quantization-aware training performance without using straight-through estimators.

  2. Signed Symmetric Quantization for Few-Bit Integers

    cs.LG 2026-06 accept novelty 7.0 of 10

    Signed absmax quantization chooses the scale sign to protect the dominant outlier, is conditionally bound-optimal for L2 error on 88-99% of LLM weight groups, and improves few-bit accuracy at zero extra inference cost.

  3. Understanding Quantization-Aware Training: Gradients at Quantized Weights Bias to the Low-Loss Basin

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    A geometric local landscape model explains PTQ basin-crossing failure at aggressive bitwidths and proves finite-time QAT recovery via straight-through estimator gradient bias under quantizer-compatibility assumptions.

  4. FTerViT: Fully Ternary Vision Transformer

    cs.CV 2026-05 conditional novelty 7.0 of 10

    FTerViT introduces fully ternary Vision Transformers with TernaryBitConv2d and TernaryLayerNorm operators, achieving 82.43% ImageNet top-1 at 6.09 MB with 15x compression.

  5. Q-ARVD: Quantizing Autoregressive Video Diffusion Models

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Q-ARVD introduces final-quality-aware frame weighting and outlier-aware adaptive dual-scale quantization to enable accurate low-bit inference for autoregressive video diffusion models.

  6. When Bits Break Recourse: Counterfactual-Faithful Quantization

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Quantization can preserve accuracy while invalidating algorithmic recourse; CFQ trains the quantized model at teacher recourse points and preserves counterfactual validity and recourse cost.

  7. QSLM: A Performance- and Memory-aware Quantization Framework with Tiered Search Strategy for Spike-driven Language Models

    cs.NE 2026-01 unverdicted novelty 7.0 of 10

    QSLM automates tiered quantization of spike-driven language models via sensitivity analysis and multi-objective search, delivering up to 86.5% memory reduction and 20% power savings while keeping accuracy close to the...

  8. DPQuant: Efficient and Differentially-Private Model Training via Dynamic Quantization Scheduling

    cs.LG 2025-09 unverdicted novelty 7.0 of 10

    DPQuant uses epoch-wise probabilistic layer rotation and DP loss sensitivity to quantize only a changing subset of layers, reducing accuracy degradation from quantization noise in DP-SGD and delivering up to 2.21x thr...

  9. FPTQuant: Function-Preserving Transforms for LLM Quantization

    cs.LG 2025-06 conditional novelty 7.0 of 10

    FPTQuant introduces function-preserving transforms that make transformer activations amenable to static 4-bit quantization with minimal inference overhead.

  10. SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Embedding a lightweight draft model into the idle GPU cycles of weight-offloaded LLM inference yields a 2.54x throughput gain over the best baseline.

  11. Behavior Backdoor for Deep Learning Models

    cs.LG 2024-12 conditional novelty 7.0 of 10

    A new backdoor attack keeps a model normal until it is quantized, after which it flips to a chosen target prediction for all inputs.

  12. Reclaiming Residual Knowledge: A Novel Paradigm to Low-Bit Quantization

    cs.CV 2024-08 unverdicted novelty 7.0 of 10

    CoRa reclaims quantization residuals in pre-trained ConvNets by searching low-rank adapter architectures instead of weights, matching SOTA accuracy on ImageNet in 3-4 bit settings with under 250 iterations on 1600 images.

  13. SpinQuant: LLM quantization with learned rotations

    cs.LG 2024-05 conditional novelty 7.0 of 10

    SpinQuant learns optimal rotations to enable accurate 4-bit quantization of LLM weights, activations, and KV cache, reducing the zero-shot gap to full precision to 2.9 points on LLaMA-2 7B.

  14. A Theoretical Framework for Stochastic Activity Prediction in Tensor Accelerator Wallace-Tree Multipliers

    cs.AR 2026-07 reject novelty 6.5 of 10

    SAP predicts Wallace-tree switching from a one-bit Bernoulli proxy of operand Hamming weight and freezes inputs under a deterministic safety controller, with Lipschitz, information-retention, and optimality theorems.

  15. LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs

    cs.DC 2026-08 conditional novelty 6.0 of 10

    Moving neighbor sampling and FP16 quantization to a BlueField-2 SmartNIC reduces transferred data and delivers measured GNN training speedups on three datasets.

  16. INT8 Quantization Makes ARM Edge Inference Dispatch-Invariant

    cs.ET 2026-07 conditional novelty 6.0 of 10

    INT8 QDQ CNNs are byte-exact across ARM Cortex-A53/A72/A76 under XNNPACK even with different SIMD kernels, unlike FP32 or x86 INT8.

  17. DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction

    cs.AR 2026-07 conditional novelty 6.0 of 10

    DSTAR reports 7.33x latency speedup and 41.89x energy savings over an A100 GPU on seven diffusion transformers by quantizing differential activations to as few as 2 bits and reusing block-wise sparse attention scores.

  18. Moving Like a Human: Ego-Motion-Normalized Temporal Signatures for Real-Time Aerial Person Tracking on Milliwatt-Class Hardware

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A 22k-parameter, 7.6-MFLOP stateless detector fed ego-motion-normalized residual-motion channels tracks small aerial persons at 31.85 FPS on a Pi Zero 2W, beating YOLOv8n by 2.7x AP25 with 16x lower latency.

  19. Minimizing Quantized Semantic Age of Information (QSAoI) in Foundation Model-Based Semantic Communications

    eess.SP 2026-06 unverdicted novelty 6.0 of 10

    Defines QSAoI metric and develops foundation model-based optimization to minimize its expected value via mixed-precision quantization and blocklength adaptation over fading channels.

  20. Privacy-Preserving Federated Autoencoder for ECG Anomaly Detection on Edge Devices

    cs.CR 2026-06 conditional novelty 6.0 of 10

    A federated system combining autoencoders, FedAvg, Renyi DP-SGD, and INT8 quantization matches centralized AUROC performance (0.782 for ConvAE) on PTB-XL while halving model size and cutting edge latency by up to 44% ...

  21. You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

    cs.CL 2025-11 reject novelty 6.0 of 10

    TAQ estimates per-layer importance from hidden representations and output sensitivity on task calibration data to allocate mixed precision in a training-free PTQ setting, outperforming task-agnostic baselines on accur...

  22. Enabling Vibration-Based Gesture Recognition on Everyday Furniture via Energy-Efficient FPGA Implementation of 1D Convolutional Networks

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Raw waveform input with lightweight 1D-CNN and 1D-SepCNN models, integer quantization, and hardware-aware search enable 0.95+ accuracy swipe recognition on Spartan-7 FPGAs at under 10 ms latency and 1.2 mJ energy.

  23. Full Integer Arithmetic Online Training for Spiking Neural Networks

    cs.NE 2025-09 conditional novelty 6.0 of 10

    An integer-only, online training algorithm for spiking neural networks uses mixed-precision shadow weights and bit-shift operations to match full-precision accuracy with over 60% lower memory usage.

  24. NeurStore: Efficient In-database Deep Learning Model Management System

    cs.DB 2025-09 conditional novelty 6.0 of 10

    A database model-management system that deduplicates tensors across models and uses delta quantization to cut storage and speed loading.

  25. Task-Specific Zero-shot Quantization-Aware Training for Object Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A zero-shot quantization-aware training method for object detectors that synthesizes task-specific images with bounding-box labels via adaptive label sampling, then distills task-specific knowledge into the quantized network.

  26. Spatial Lifting for Dense Prediction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Replicating 2D images into a 3D volume and processing with a shallow 3D U-Net gives competitive dense prediction accuracy at a fraction of the parameter count, with slice-consistency used as a free quality score.

  27. MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A joint low-rank and mixed-precision quantization method that assigns rank and bit-width per transformer layer under a memory constraint, showing state-of-the-art compression accuracy.

  28. MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MPQ-DMv2 adds binary residual quantization, temporal relation distillation, and SVD-initialized LoRA to mixed-precision quantization, improving low-bit diffusion model generation quality.

  29. How Weight Resampling and Optimizers Shape the Dynamics of Continual Learning and Forgetting in Neural Networks

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Zapping the last layer during pretraining speeds a model's recovery after transfer, and Adam produces different learning and forgetting patterns than SGD in continual learning.

  30. FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new Fisher-information-based reconstruction loss, DPLR-FIM, improves low-bit post-training quantization accuracy for Vision Transformers without specialized quantizers.

  31. Starting Positions Matter: A Study on Better Weight Initialization for Neural Network Quantization

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Weight initialization measurably affects quantized CNN accuracy, and a graph hypernetwork finetuned on quantized networks (GHN-QAT) can predict parameters that survive 4-bit and even 2-bit quantization better than ran...

  32. Chameleon: A Multiplier-Free Temporal Convolutional Network Accelerator for End-to-End Few-Shot and Continual Learning from Sequential Data

    cs.AR 2025-05 conditional novelty 6.0 of 10

    Chameleon is a 40-nm CMOS accelerator that performs end-to-end few-shot and continual learning on-chip using prototypical networks and TCN embeddings, and runs keyword spotting at 3.1 uW.

  33. Refining Datapath for Microscaling ViTs

    cs.AR 2025-05 conditional novelty 6.0 of 10

    MXInt-based datapath designs put all ViT nonlinear operators on an FPGA at 2-5 bit mantissas, with under 1% ImageNet accuracy loss on DeiT models.

  34. Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A mix of signed and unsigned 4-bit floating-point formats, timestep-aware LoRA experts, and a denoising-weighted loss keeps diffusion-model image quality close to full precision.

  35. NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    NeUQI improves low-bit uniform quantization of LLMs by relaxing the integer zero-point constraint and efficiently searching a near-optimal scale, beating existing PTQ baselines at 2-4 bits.

  36. Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A saliency-weighted quantization-aware training method lets 4-bit quantized imitation-learning policies match full-precision success rates across robot manipulation, driving, and control benchmarks.

  37. A probabilistic framework for dynamic quantization

    cs.LG 2025-05 conditional novelty 6.0 of 10

    An input-adaptive 8-bit quantization scheme estimates activation ranges via a lightweight probabilistic surrogate, achieving near-dynamic accuracy with static-like memory overhead.

  38. Quantized Approximate Signal Processing (QASP): Towards Homomorphic Encryption for audio

    eess.AS 2025-05 conditional novelty 6.0 of 10

    The authors demonstrate the first fully homomorphic encryption pipeline that computes STFT, Mel, MFCC, and gammatone features directly from raw audio, with approximate variants that improve accuracy in some private au...

  39. AHCQ-SAM: Toward Accurate and Hardware-Compatible Post-Training Segment Anything Model Quantization

    cs.CV 2025-03 unverdicted novelty 6.0 of 10

    AHCQ-SAM introduces ACNR, HLUQ, CAG, and LNQ quantization techniques that deliver 15.2% mAP gain on 4-bit SAM-B and 14.01% J&F gain on 4-bit SAM2-Tiny versus prior PTQ methods.

  40. Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A discrete hill-climbing search over permutation, scaling, and rotation invariances improves 2-bit quantized OPT models when applied on top of GPTQ, AWQ, and OmniQuant.

  41. Quantized Spike-driven Transformer

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A 4-bit quantized spike-driven transformer with multi-bit training and binary inference achieves 80.3% ImageNet accuracy with 6.8M parameters.

  42. MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A quantization framework using KLT-enhanced and smooth-fused rotations lets Mamba models run at 8-bit precision with near-full accuracy and at 4-bit weights with moderate loss.

  43. DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DGQ quantizes text-to-image diffusion models to 4-8 bits without fine-tuning by preserving activation outliers and applying prompt-specific log quantization to cross-attention scores.

  44. Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart

    cs.LG 2024-12 conditional novelty 6.0 of 10

    BWRF improves quantization-aware training by grafting full-precision blocks onto the low-precision model during training, producing mixed-precision guides that raise ImageNet and CIFAR-10 accuracy at 2 to 4 bits.

  45. AutoRank: MCDA Based Rank Personalization for LoRA-Enabled Distributed Learning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    AutoRank assigns per-participant LoRA ranks from data complexity metrics (loss entropy, distribution entropy, Gini-Simpson) via TOPSIS and CRITIC, improving federated learning accuracy and convergence.

  46. MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MPQ-DM combines kurtosis-based intra-layer mixed-precision weight quantization with time-smoothed relation distillation to keep diffusion models accurate at 2 to 4 bit widths.

  47. PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models

    cs.LG 2024-12 conditional novelty 6.0 of 10

    PTSBench benchmarks post-training sparsity techniques and model families, finding learning-based allocation and block-wise reconstruction most effective, and attention-based models most sparsity-friendly.

  48. Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens

    cs.LG 2024-11 conditional novelty 6.0 of 10

    The paper derives a scaling law for quantization-induced loss increase as a function of model size, training tokens, and bit width, and uses it to argue that low-bit quantization will hurt future fully trained LLMs.

  49. Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A trainable soft-tanh quantizer with learned steepness and clipping improves 1-4 bit network accuracy and yields fast ARM kernels.

  50. Cheetah: Mixed Low-Precision Hardware & Software Co-Design Framework for DNNs on the Edge

    cs.LG 2019-08 conditional novelty 6.0 of 10

    16-bit posits train MNIST and Fashion-MNIST feedforward networks with less accuracy loss than 16-bit floating point, and 5 to 8 bit posits yield better inference accuracy and energy-delay tradeoffs than float or fixed point.

  51. Efficient Reasoning on the Edge

    cs.LG 2026-03 accept novelty 5.5 of 10

    LoRA adapters, budget-forced GRPO, dynamic switching, parallel verification and FPTQuant enable practical chain-of-thought reasoning on quantized Qwen2.5-7B for edge devices.

  52. Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Across four NAS/DNN-predictor regression benchmarks, GEN (deep graph convolution) achieves the best average rank over 11 GNN message-passing layers, though attention GATv2 wins on the largest graphs.

  53. Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n

    cs.CV 2026-07 accept novelty 5.0 of 10

    Adding 695 synthetic or derived training images did not improve a YOLOv8n campus-waste detector over the real-only baseline; image source mattered more than image count.

  54. Quantization in Federated Learning: Methods, Challenges and Future Directions

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    This survey introduces a taxonomy for quantization in federated learning organized around client heterogeneity, aggregation consistency, non-IID robustness, privacy integration, and hardware co-optimization, while ana...

  55. OffQ: Taming Structured Outliers in LLM Quantization by Offsetting

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    OffQ mitigates structured activation outliers in LLMs via PCA-based rotation and shared offset absorption to support effective W4A4KV4 uniform quantization.

  56. The Thermodynamic Costs of Simple Linear Regression

    cond-mat.stat-mech 2026-05 unverdicted novelty 5.0 of 10

    Thermodynamic lower bounds are approximated for exact and SGD linear regression, producing energy-aware scaling laws for optimal training dataset size given a target generalization error.

  57. MicroBi-ConvLSTM: An Ultra-Lightweight Efficient Model for Human Activity Recognition on Resource Constrained Devices

    cs.CV 2026-02 conditional novelty 5.0 of 10

    MicroBi-ConvLSTM is a convolutional-recurrent model with 11.4K parameters that delivers competitive accuracy on eight HAR benchmarks and full INT8 deployment coverage on Raspberry Pi Pico 2 and ESP32.

  58. Performance and Complexity Trade-off Optimization of Speech Models During Training

    cs.SD 2026-01 conditional novelty 5.0 of 10

    By turning each layer's width into a continuous, noise-smoothed parameter, the authors train speech models whose sizes shrink during training, reducing FLOPs and size by roughly 80–90% in their case studies.

  59. COMET: Co-Optimization of a CNN Model using Efficient-Hardware OBC Techniques

    eess.SP 2025-10 unverdicted novelty 5.0 of 10

    COMET co-optimizes CNN inference via OBC Schemes A/B on inputs/weights, four LUT techniques, and an im2col-based GEMM core to deliver efficient FPGA deployment with negligible accuracy loss on LeNet-5 and All-CNN-C.

  60. FedBiF: Communication-Efficient Federated Learning via Bits Freezing

    cs.LG 2025-09 conditional novelty 5.0 of 10

    FedBiF trains federated models by updating a single bit of each quantized weight per round, achieving 1 bit-per-parameter uplink and 3-4 bits downlink with accuracy close to FedAvg.

See all 99 Pith citations

Pith tools