Pith. sign in

REVIEW 1 cited by

Perceptual Image Compression with Cooperative Cross-Modal Side Information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.13847 v2 pith:VNL2PG35 submitted 2023-11-23 cs.CV cs.ITeess.IVmath.IT

classification cs.CVcs.ITeess.IVmath.IT
keywords imagecompressioninformationsidetextimagesperceptualsemantic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The explosion of data has resulted in more and more associated text being transmitted along with images. Inspired by from distributed source coding, many works utilize image side information to enhance image compression. However, existing methods generally do not consider using text as side information to enhance perceptual compression of images, even though the benefits of multimodal synergy have been widely demonstrated in research. This begs the following question: How can we effectively transfer text-level semantic dependencies to help image compression, which is only available to the decoder? In this work, we propose a novel deep image compression method with text-guided side information to achieve a better rate-perception-distortion tradeoff. Specifically, we employ the CLIP text encoder and an effective Semantic-Spatial Aware block to fuse the text and image features. This is done by predicting a semantic mask to guide the learned text-adaptive affine transformation at the pixel level. Furthermore, we design a text-conditional generative adversarial networks to improve the perceptual quality of reconstructed images. Extensive experiments involving four datasets and ten image quality assessment metrics demonstrate that the proposed approach achieves superior results in terms of rate-perception trade-off and semantic distortion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniMIC: Towards Universal Multi-modality Perceptual Image Compression

    eess.IV 2024-12 conditional novelty 6.0 of 10

    A single text-conditioned diffusion refiner improves perceptual quality (FID, LPIPS) of images from eight different base codecs and extends to unseen codecs.

Pith tools