REVIEW 3 cited by
Neural Network Quantization for Efficient Inference: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
As neural networks have become more powerful, there has been a rising desire to deploy them in the real world; however, the power and accuracy of neural networks is largely due to their depth and complexity, making them difficult to deploy, especially in resource-constrained devices. Neural network quantization has recently arisen to meet this demand of reducing the size and complexity of neural networks by reducing the precision of a network. With smaller and simpler networks, it becomes possible to run neural networks within the constraints of their target hardware. This paper surveys the many neural network quantization techniques that have been developed in the last decade. Based on this survey and comparison of neural network quantization techniques, we propose future directions of research in the area.
Forward citations
Cited by 3 Pith papers
-
Lyapunov-Guided Training for Hardware-Safe Neural Networks Under Fixed-Point Arithmetic
Monotone Lyapunov projection of layerwise hidden-state energy suppresses two's-complement overflow under wrapping fixed-point QAT/PTQ, recovering 86.55% MNIST accuracy where unconstrained models collapse to chance.
-
Automatic mixed precision for optimizing gained time with constrained loss mean-squared-error based on model partition to sequential sub-graphs
The paper derives an additive loss-MSE sensitivity metric and a hardware-aware time-gain model, then uses integer programming to assign per-layer FP8 or BF16 formats for LLM inference.
-
Efficient Split Learning LSTM Models for FPGA-based Edge IoT Devices
A case study showing how knowledge distillation, pruning, and quantization let a small LSTM run on a low-end FPGA, with three split configurations trading off latency, power, and resource usage.
Discussion (0). Continue with ORCID to comment.