Pith. sign in

REVIEW 4 cited by

On Uniform Scalar Quantization for Learned Image Compression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.17051 v1 pith:ETL5PMIJ submitted 2023-09-29 cs.CV

classification cs.CV
keywords quantizationestimationgradienttraininganalysescompressionimagemismatch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learned image compression possesses a unique challenge when incorporating non-differentiable quantization into the gradient-based training of the networks. Several quantization surrogates have been proposed to fulfill the training, but they were not systematically justified from a theoretical perspective. We fill this gap by contrasting uniform scalar quantization, the most widely used category with rounding being its simplest case, and its training surrogates. In principle, we find two factors crucial: one is the discrepancy between the surrogate and rounding, leading to train-test mismatch; the other is gradient estimation risk due to the surrogate, which consists of bias and variance of the gradient estimation. Our analyses and simulations imply that there is a tradeoff between the train-test mismatch and the gradient estimation risk, and the tradeoff varies across different network structures. Motivated by these analyses, we present a method based on stochastic uniform annealing, which has an adjustable temperature coefficient to control the tradeoff. Moreover, our analyses enlighten us as to two subtle tricks: one is to set an appropriate lower bound for the variance parameter of the estimated quantized latent distribution, which effectively reduces the train-test mismatch; the other is to use zero-center quantization with partial stop-gradient, which reduces the gradient estimation variance and thus stabilize the training. Our method with the tricks is verified to outperform the existing practices of quantization surrogates on a variety of representative image compression networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Gap Between Principle and Practice of Lossy Image Coding

    cs.IT 2025-01 conditional novelty 6.0 of 10

    The paper attributes the gap between ideal and practical lossy image coding to five effects, and reports an estimated rate-distortion upper bound that beats VTM by up to 35 percent on Kodak.

  2. Sparse Point Clouds Assisted Learned Image Compression

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Projecting sparse LiDAR depth into predicted structural features and injecting them into learned image codecs consistently improves rate-distortion performance on KITTI and Waymo.

  3. Generalized Gaussian Model for Learned Image Compression

    eess.IV 2024-11 conditional novelty 6.0 of 10

    A generalized Gaussian entropy model with a learned shape parameter and two training fixes improves rate-distortion performance of learned image codecs compared to Gaussian and mixture models.

  4. An Information-Theoretic Regularizer for Lossy Neural Image Compression

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A regularizer that maximizes conditional source entropy gives BD-rate gains of 0.9% to 3.0% across five neural compression models, but may be equivalent to reweighting the rate term.

Pith tools