REVIEW 4 cited by
Entroformer: A Transformer-based Entropy Model for Learned Image Compression
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
One critical component in lossy deep image compression is the entropy model, which predicts the probability distribution of the quantized latent representation in the encoding and decoding modules. Previous works build entropy models upon convolutional neural networks which are inefficient in capturing global dependencies. In this work, we propose a novel transformer-based entropy model, termed Entroformer, to capture long-range dependencies in probability distribution estimation effectively and efficiently. Different from vision transformers in image classification, the Entroformer is highly optimized for image compression, including a top-k self-attention and a diamond relative position encoding. Meanwhile, we further expand this architecture with a parallel bidirectional context model to speed up the decoding process. The experiments show that the Entroformer achieves state-of-the-art performance on image compression while being time-efficient.
Forward citations
Cited by 4 Pith papers
-
LANCE: Locally Adaptive Neural Context Estimation for Overfitted Image Compression
LANCE extends OIC frameworks with a spatial hyperprior and predictive coding scheme, reporting BD-rate gains of 1.4-3% over Cool-Chic 4.0 on Kodak and CLIC.
-
StableCodec: Taming One-Step Diffusion for Extreme Image Compression
A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.
-
Generative Image Compression by Estimating Gradients of the Rate-variable Feature Distribution
A single rate-variable generative compression model treats quantization as a forward corruption and reverses it with a two-step denoiser, outperforming prior generative codecs on perceptual quality benchmarks.
-
Spectral and Spatial Graph Learning for Multispectral Solar Image Compression
A graph-based learned codec modeling wavelength-to-wavelength relationships plus windowed spatial attention reports modest PSNR/MS-SSIM gains and 20.15% lower MSID on six-channel solar images than two self-defined baselines.
Discussion (0). Sign in to comment.