Pith. sign in

REVIEW 2 cited by

Extreme Video Compression with Pre-trained Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08934 v1 pith:PXEPEGBB submitted 2024-02-14 eess.IV cs.CV

classification eess.IVcs.CV
keywords videomodelsqualitycompressiondiffusionframesimageperceptual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have achieved remarkable success in generating high quality image and video data. More recently, they have also been used for image compression with high perceptual quality. In this paper, we present a novel approach to extreme video compression leveraging the predictive power of diffusion-based generative models at the decoder. The conditional diffusion model takes several neural compressed frames and generates subsequent frames. When the reconstruction quality drops below the desired level, new frames are encoded to restart prediction. The entire video is sequentially encoded to achieve a visually pleasing reconstruction, considering perceptual quality metrics such as the learned perceptual image patch similarity (LPIPS) and the Frechet video distance (FVD), at bit rates as low as 0.02 bits per pixel (bpp). Experimental results demonstrate the effectiveness of the proposed scheme compared to standard codecs such as H.264 and H.265 in the low bpp regime. The results showcase the potential of exploiting the temporal relations in video data using generative models. Code is available at: https://github.com/ElesionKyrie/Extreme-Video-Compression-With-Prediction-Using-Pre-trainded-Diffusion-Models-

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffusion-based Perceptual Neural Video Compression with Temporal Diffusion Information Reuse

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DiffVC integrates Stable Diffusion into a conditional neural video codec, with temporal reuse of diffusion predictions for speed and quantization-parameter prompting for variable bitrate, achieving state-of-the-art pe...

  2. Multi-Modal Learning meets Genetic Programming: Analyzing Alignment in Latent Space Optimization

    cs.NE 2026-04 unverdicted novelty 5.0 of 10

    SNIP's symbolic-numeric alignment stays coarse and does not improve during optimization, so multi-modal LSO does not yet deliver effective bi-modal search for symbolic regression.

Pith tools