REVIEW 2 cited by
MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. In recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC.
Forward citations
Cited by 2 Pith papers
-
UniMIC: Towards Universal Multi-modality Perceptual Image Compression
A single text-conditioned diffusion refiner improves perceptual quality (FID, LPIPS) of images from eight different base codecs and extends to unseen codecs.
-
LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression
A single LMM generates and compresses image captions for ultra-low-bitrate learned image compression, improving LPIPS BD-rate by 41.58% over MISC.
Discussion (0). Continue with ORCID to comment.