A screen-and-natural-image codec that losslessly encodes OCR text and uses a diffusion renderer with three levels of conditioning achieves state-of-the-art perceptual quality and high text accuracy at 0.005 to 0.05 bpp.
Rethinking lossy com- pression: The rate-distortion-perception tradeoff
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PICD: Versatile Perceptual Image Compression with Diffusion Rendering
A screen-and-natural-image codec that losslessly encodes OCR text and uses a diffusion renderer with three levels of conditioning achieves state-of-the-art perceptual quality and high text accuracy at 0.005 to 0.05 bpp.