REVIEW 3 cited by
Learning Sparse Low-Precision Neural Networks With Learnable Regularization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We consider learning deep neural networks (DNNs) that consist of low-precision weights and activations for efficient inference of fixed-point operations. In training low-precision networks, gradient descent in the backward pass is performed with high-precision weights while quantized low-precision weights and activations are used in the forward pass to calculate the loss function for training. Thus, the gradient descent becomes suboptimal, and accuracy loss follows. In order to reduce the mismatch in the forward and backward passes, we utilize mean squared quantization error (MSQE) regularization. In particular, we propose using a learnable regularization coefficient with the MSQE regularizer to reinforce the convergence of high-precision weights to their quantized values. We also investigate how partial L2 regularization can be employed for weight pruning in a similar manner. Finally, combining weight pruning, quantization, and entropy coding, we establish a low-precision DNN compression pipeline. In our experiments, the proposed method yields low-precision MobileNet and ShuffleNet models on ImageNet classification with the state-of-the-art compression ratios of 7.13 and 6.79, respectively. Moreover, we examine our method for image super resolution networks to produce 8-bit low-precision models at negligible performance loss.
Forward citations
Cited by 3 Pith papers
-
Training High-Performance and Large-Scale Deep Neural Networks with Full 8-bit Integers
WAGEUBN trains ResNet models on ImageNet using 8-bit integers for weights, activations, gradients, errors, batch normalization, and the Momentum optimizer, with moderate accuracy loss.
-
Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
Two-stage and gradually decreasing quantization, stochastic precision sampling, and joint teacher-student distillation each improve low-bit CNN accuracy on ImageNet and CIFAR-100, with the largest gains when combined.
-
Contrast & Compress: Learning Lightweight Embeddings for Short Trajectories
A small Transformer trained with a cosine-based triplet loss learns 16-dimensional embeddings that retrieve similar short driving trajectories from Argoverse 2 substantially better than FFT-based triplet training.
Discussion (0). Continue with ORCID to comment.