Pith. sign in

REVIEW 2 cited by

A Hybrid Transformer-Mamba Network for Single Image Deraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00410 v1 pith:H7IPW25A submitted 2024-08-31 cs.CV

classification cs.CV
keywords channelfeaturesinformationspectralbranchdependenciesderainingdual-branch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing deraining Transformers employ self-attention mechanisms with fixed-range windows or along channel dimensions, limiting the exploitation of non-local receptive fields. In response to this issue, we introduce a novel dual-branch hybrid Transformer-Mamba network, denoted as TransMamba, aimed at effectively capturing long-range rain-related dependencies. Based on the prior of distinct spectral-domain features of rain degradation and background, we design a spectral-banded Transformer blocks on the first branch. Self-attention is executed within the combination of the spectral-domain channel dimension to improve the ability of modeling long-range dependencies. To enhance frequency-specific information, we present a spectral enhanced feed-forward module that aggregates features in the spectral domain. In the second branch, Mamba layers are equipped with cascaded bidirectional state space model modules to additionally capture the modeling of both local and global information. At each stage of both the encoder and decoder, we perform channel-wise concatenation of dual-branch features and achieve feature fusion through channel reduction, enabling more effective integration of the multi-scale information from the Transformer and Mamba branches. To better reconstruct innate signal-level relations within clean images, we also develop a spectral coherence loss. Extensive experiments on diverse datasets and real-world images demonstrate the superiority of our method compared against the state-of-the-art approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semi-Supervised State-Space Model with Dynamic Stacking Filter for Real-World Video Deraining

    cs.CV 2025-05 conditional novelty 5.0 of 10

    The paper introduces a dual-branch state-space video deraining model with a dynamic stacking filter and semi-supervised median stacking loss, showing top PSNR across three benchmarks and new downstream task gains on a...

  2. MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection

    cs.CV 2025-06 reject novelty 4.0 of 10

    MGDFIS combines three attention and mixing modules and reports accuracy gains on UAV detection benchmarks, but the paper's internal inconsistencies and missing details prevent verification.

Pith tools