Pith. sign in

REVIEW 1 cited by

Correcting Diffusion-Based Perceptual Image Compression with Privileged End-to-End Decoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.04916 v2 pith:XLAZHPOQ submitted 2024-04-07 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords diffusionmodelscompressiondecoderend-to-endperceptualdistortionencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The images produced by diffusion models can attain excellent perceptual quality. However, it is challenging for diffusion models to guarantee distortion, hence the integration of diffusion models and image compression models still needs more comprehensive explorations. This paper presents a diffusion-based image compression method that employs a privileged end-to-end decoder model as correction, which achieves better perceptual quality while guaranteeing the distortion to an extent. We build a diffusion model and design a novel paradigm that combines the diffusion model and an end-to-end decoder, and the latter is responsible for transmitting the privileged information extracted at the encoder side. Specifically, we theoretically analyze the reconstruction process of the diffusion models at the encoder side with the original images being visible. Based on the analysis, we introduce an end-to-end convolutional decoder to provide a better approximation of the score function $\nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t)$ at the encoder side and effectively transmit the combination. Experiments demonstrate the superiority of our method in both distortion and perception compared with previous perceptual compression methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

    eess.IV 2026-07 conditional novelty 7.0 of 10

    A generative video codec that transmits selected latent anchors and a text prompt, then synthesizes the rest with a diffusion transformer, reaches <0.005 bpp with strong perceptual quality.

Pith tools