Pith. sign in

REVIEW 5 cited by

MambaVC: Learned Visual Compression with Selective State Spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15413 v3 pith:GUD5CXIS submitted 2024-05-24 eess.IV cs.CVcs.ITmath.IT

classification eess.IVcs.CVcs.ITmath.IT
keywords mambavccompressionvisualdesignsefficiencylearnedmemorynetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learned visual compression is an important and active task in multimedia. Existing approaches have explored various CNN- and Transformer-based designs to model content distribution and eliminate redundancy, where balancing efficacy (i.e., rate-distortion trade-off) and efficiency remains a challenge. Recently, state-space models (SSMs) have shown promise due to their long-range modeling capacity and efficiency. Inspired by this, we take the first step to explore SSMs for visual compression. We introduce MambaVC, a simple, strong and efficient compression network based on SSM. MambaVC develops a visual state space (VSS) block with a 2D selective scanning (2DSS) module as the nonlinear activation function after each downsampling, which helps to capture informative global contexts and enhances compression. On compression benchmark datasets, MambaVC achieves superior rate-distortion performance with lower computational and memory overheads. Specifically, it outperforms CNN and Transformer variants by 9.3% and 15.6% on Kodak, respectively, while reducing computation by 42% and 24%, and saving 12% and 71% of memory. MambaVC shows even greater improvements with high-resolution images, highlighting its potential and scalability in real-world applications. We also provide a comprehensive comparison of different network designs, underscoring MambaVC's advantages. Code is available at https://github.com/QinSY123/2024-MambaVC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors

    cs.CV 2026-03 conditional novelty 6.5 of 10

    Ultra-low-bitrate image decoding is cast as one-step next-frame prediction from a compact anchor using adapted video diffusion priors, yielding large perceptual bitrate savings versus DiffC.

  2. DCVC-MB: Neural B-Frame Video Compression using State Space Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    DCVC-MB, a neural B-frame video codec using Mamba state-space fusion, reports BD-rate savings up to 8.98% over prior neural codecs and up to 30.45% over VTM-19.0-LDP.

  3. Understanding Rate-Distortion Performance in Distributed Transformer Inference

    cs.LG 2026-01 unverdicted novelty 6.0 of 10

    Deeper transformer layers produce intermediate representations that are harder to lossy-compress, and the paper links this to growing covariance and Rademacher complexity.

  4. Linear Attention Modeling for Learned Image Compression

    cs.CV 2025-02 conditional novelty 5.0 of 10

    LALIC replaces transformer and Mamba blocks in learned image compression with bidirectional RWKV linear-attention blocks, reporting BD-rate gains over VTM-9.1 while keeping decoder latency moderate.

  5. CMamba: Learned Image Compression with State Space Models

    eess.IV 2025-02 conditional novelty 5.0 of 10

    A hybrid CNN and Mamba (state space model) image compression codec reports BD-Rate savings of 14.95% to 18.83% over VVC with fewer parameters, FLOPs, and lower decoding time than the prior best learned method.

Pith tools