REVIEW 2 cited by
A Hybrid Transformer-Mamba Network for Single Image Deraining
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Existing deraining Transformers employ self-attention mechanisms with fixed-range windows or along channel dimensions, limiting the exploitation of non-local receptive fields. In response to this issue, we introduce a novel dual-branch hybrid Transformer-Mamba network, denoted as TransMamba, aimed at effectively capturing long-range rain-related dependencies. Based on the prior of distinct spectral-domain features of rain degradation and background, we design a spectral-banded Transformer blocks on the first branch. Self-attention is executed within the combination of the spectral-domain channel dimension to improve the ability of modeling long-range dependencies. To enhance frequency-specific information, we present a spectral enhanced feed-forward module that aggregates features in the spectral domain. In the second branch, Mamba layers are equipped with cascaded bidirectional state space model modules to additionally capture the modeling of both local and global information. At each stage of both the encoder and decoder, we perform channel-wise concatenation of dual-branch features and achieve feature fusion through channel reduction, enabling more effective integration of the multi-scale information from the Transformer and Mamba branches. To better reconstruct innate signal-level relations within clean images, we also develop a spectral coherence loss. Extensive experiments on diverse datasets and real-world images demonstrate the superiority of our method compared against the state-of-the-art approaches.
Forward citations
Cited by 2 Pith papers
-
Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining
The paper introduces a dual-branch state-space video deraining model with a dynamic stacking filter and semi-supervised median stacking loss, showing top PSNR across three benchmarks and new downstream task gains on a...
-
MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection
MGDFIS combines three attention and mixing modules and reports accuracy gains on UAV detection benchmarks, but the paper's internal inconsistencies and missing details prevent verification.
Discussion (0). Continue with ORCID to comment.