REVIEW 9 cited by
VmambaIR: Visual State Space Model for Image Restoration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Image restoration is a critical task in low-level computer vision, aiming to restore high-quality images from degraded inputs. Various models, such as convolutional neural networks (CNNs), generative adversarial networks (GANs), transformers, and diffusion models (DMs), have been employed to address this problem with significant impact. However, CNNs have limitations in capturing long-range dependencies. DMs require large prior models and computationally intensive denoising steps. Transformers have powerful modeling capabilities but face challenges due to quadratic complexity with input image size. To address these challenges, we propose VmambaIR, which introduces State Space Models (SSMs) with linear complexity into comprehensive image restoration tasks. We utilize a Unet architecture to stack our proposed Omni Selective Scan (OSS) blocks, consisting of an OSS module and an Efficient Feed-Forward Network (EFFN). Our proposed omni selective scan mechanism overcomes the unidirectional modeling limitation of SSMs by efficiently modeling image information flows in all six directions. Furthermore, we conducted a comprehensive evaluation of our VmambaIR across multiple image restoration tasks, including image deraining, single image super-resolution, and real-world image super-resolution. Extensive experimental results demonstrate that our proposed VmambaIR achieves state-of-the-art (SOTA) performance with much fewer computational resources and parameters. Our research highlights the potential of state space models as promising alternatives to the transformer and CNN architectures in serving as foundational frameworks for next-generation low-level visual tasks.
Forward citations
Cited by 9 Pith papers
-
EAMamba: Efficient All-Around Vision State Space Model for Image Restoration
EAMamba introduces channel-grouped multi-head selective scanning whose cost does not grow with the number of scan directions, and reports 31-89% FLOPs reductions with comparable PSNR across denoising, super-resolution...
-
DH-Mamba: Exploring Dual-domain Hierarchical State Space Models for MRI Reconstruction
DH-Mamba is a dual-domain hierarchical Mamba network that uses circular k-space scanning and local diversity enhancement to outperform prior MRI reconstruction methods on three public datasets.
-
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
AVS-Mamba applies Mamba with temporal and cross-modal scanning to audio-visual segmentation, reporting top scores on AVSBench-object but not on AVSBench-semantic with the stronger backbone.
-
MambaIRv2: Attentive State Space Restoration
MambaIRv2 modifies Mamba's state-space output with semantic prompts and reorders tokens by semantic group, achieving single-scan non-causal restoration that beats several transformer baselines.
-
RAWMamba: Unified sRGB-to-RAW De-rendering With State Space Model
RAWMamba unifies image and video sRGB-to-RAW de-rendering with a metadata embedding module and a local tone-aware Mamba backbone, reporting 2.26 dB to 3.37 dB PSNR gains over prior task-specific models.
-
Atmos-Bench: 3D Atmospheric Structures for Climate Insight
Atmos-Bench introduces a synthetic 3D benchmark for satellite LiDAR backscatter recovery, and FourCastX, a frequency-MoE inpainting model, reports substantially higher PSNR/SSIM than six baselines on it.
-
CMamba: Learned Image Compression with State Space Models
A hybrid CNN and Mamba (state space model) image compression codec reports BD-Rate savings of 14.95% to 18.83% over VVC with fewer parameters, FLOPs, and lower decoding time than the prior best learned method.
-
Teacher-Guided Causal Interventions for Image Denoising: Orthogonal Content-Noise Disentanglement in Vision Transformers
TCD-Net couples a ViT denoiser with "causal" regularizers (de-centering, orthogonality, Nano Banana Pro distillation) and reports marginal PSNR shifts of ≤0.08 dB on some benchmarks, with no error bars or code.
-
CWNet: Causal Wavelet Network for Low-Light Image Enhancement
CWNet mixes wavelet-based frequency enhancement, Mamba-style high-frequency scanning, and two semantic consistency losses to produce competitive low-light image enhancement with 1.23 million parameters.
Discussion (0). Continue with ORCID to comment.