Pith. sign in

REVIEW 11 cited by

Text + Sketch: Image Compression at Ultra Low Rates

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.01944 v1 pith:DOSTW36W submitted 2023-07-04 cs.LG cs.CVcs.ITmath.IT

classification cs.LGcs.CVcs.ITmath.IT
keywords modelscompressiontextdescriptionsgenerateimagepre-trainedtraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advances in text-to-image generative models provide the ability to generate high-quality images from short text descriptions. These foundation models, when pre-trained on billion-scale datasets, are effective for various downstream tasks with little or no further training. A natural question to ask is how such models may be adapted for image compression. We investigate several techniques in which the pre-trained models can be directly used to implement compression schemes targeting novel low rate regimes. We show how text descriptions can be used in conjunction with side information to generate high-fidelity reconstructions that preserve both semantics and spatial structure of the original. We demonstrate that at very low bit-rates, our method can significantly improve upon learned compressors in terms of perceptual and semantic fidelity, despite no end-to-end training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors

    cs.CV 2026-03 conditional novelty 6.5 of 10

    Ultra-low-bitrate image decoding is cast as one-step next-frame prediction from a compact anchor using adapted video diffusion priors, yielding large perceptual bitrate savings versus DiffC.

  2. StableCodec: Taming One-Step Diffusion for Extreme Image Compression

    eess.IV 2025-06 conditional novelty 6.0 of 10

    A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.

  3. Generative Image Compression by Estimating Gradients of the Rate-variable Feature Distribution

    eess.IV 2025-05 conditional novelty 6.0 of 10

    A single rate-variable generative compression model treats quantization as a forward corruption and reverses it with a two-step denoiser, outperforming prior generative codecs on perceptual quality benchmarks.

  4. Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ResULIC combines residual text captions, optimized prompts, and bitrate-aware diffusion steps to reach state-of-the-art perceptual quality at ultra-low bitrates.

  5. PICD: Versatile Perceptual Image Compression with Diffusion Rendering

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A screen-and-natural-image codec that losslessly encodes OCR text and uses a diffusion renderer with three levels of conditioning achieves state-of-the-art perceptual quality and high text accuracy at 0.005 to 0.05 bpp.

  6. Diffusion-based Perceptual Neural Video Compression with Temporal Diffusion Information Reuse

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DiffVC integrates Stable Diffusion into a conditional neural video codec, with temporal reuse of diffusion predictions for speed and quantization-parameter prompting for variable bitrate, achieving state-of-the-art pe...

  7. Lossy Compression with Pretrained Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A complete, zero-shot implementation of the DiffC algorithm lets pretrained Stable Diffusion models act as lossy image compressors at ultra-low bitrates.

  8. Semantics-Guided Diffusion for Deep Joint Source-Channel Coding in Wireless Image Transmission

    cs.IT 2025-01 conditional novelty 6.0 of 10

    SGD-JSCC guides a diffusion denoiser with transmitted text or edge maps to clean the received latent features of a DeepJSCC image codec, improving perceptual quality at low SNR without pilot-based CSI.

  9. SMIC: Semantic Multi-Item Compression based on CLIP dictionary

    eess.IV 2024-12 conditional novelty 6.0 of 10

    A dictionary-based multi-item codec sparsely projects CLIP image embeddings onto learned semantic atoms and regenerates images with unCLIP, reaching about 1e-4 BPP per image on a 5000-image collection.

  10. SDGIC: A Semantic Disambiguation-Guided Generative Image Compression Method for Ultra-Low Bitrates

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A diffusion-based image codec guided by text, a highly compressed image, and CLIP-derived semantic pseudo-words improves semantic consistency at bitrates below 0.05 bpp.

  11. Latent Feature-Guided Conditional Diffusion for Generative Image Semantic Communication

    cs.MM 2025-04 conditional novelty 5.0 of 10

    An ROI-weighted latent feature space with conditional diffusion reconstruction is claimed to reduce perceptual error by 43.3% versus DeepJSCC in semantic image communication.

Pith tools