REVIEW 4 cited by
On Uniform Scalar Quantization for Learned Image Compression
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learned image compression possesses a unique challenge when incorporating non-differentiable quantization into the gradient-based training of the networks. Several quantization surrogates have been proposed to fulfill the training, but they were not systematically justified from a theoretical perspective. We fill this gap by contrasting uniform scalar quantization, the most widely used category with rounding being its simplest case, and its training surrogates. In principle, we find two factors crucial: one is the discrepancy between the surrogate and rounding, leading to train-test mismatch; the other is gradient estimation risk due to the surrogate, which consists of bias and variance of the gradient estimation. Our analyses and simulations imply that there is a tradeoff between the train-test mismatch and the gradient estimation risk, and the tradeoff varies across different network structures. Motivated by these analyses, we present a method based on stochastic uniform annealing, which has an adjustable temperature coefficient to control the tradeoff. Moreover, our analyses enlighten us as to two subtle tricks: one is to set an appropriate lower bound for the variance parameter of the estimated quantized latent distribution, which effectively reduces the train-test mismatch; the other is to use zero-center quantization with partial stop-gradient, which reduces the gradient estimation variance and thus stabilize the training. Our method with the tricks is verified to outperform the existing practices of quantization surrogates on a variety of representative image compression networks.
Forward citations
Cited by 4 Pith papers
-
The Gap Between Principle and Practice of Lossy Image Coding
The paper attributes the gap between ideal and practical lossy image coding to five effects, and reports an estimated rate-distortion upper bound that beats VTM by up to 35 percent on Kodak.
-
Sparse Point Clouds Assisted Learned Image Compression
Projecting sparse LiDAR depth into predicted structural features and injecting them into learned image codecs consistently improves rate-distortion performance on KITTI and Waymo.
-
Generalized Gaussian Model for Learned Image Compression
A generalized Gaussian entropy model with a learned shape parameter and two training fixes improves rate-distortion performance of learned image codecs compared to Gaussian and mixture models.
-
An Information-Theoretic Regularizer for Lossy Neural Image Compression
A regularizer that maximizes conditional source entropy gives BD-rate gains of 0.9% to 3.0% across five neural compression models, but may be equivalent to reweighting the rate term.
Discussion (0). Continue with ORCID to comment.