Pith. sign in

REVIEW 3 cited by

MambaDFuse: A Mamba-based Dual-phase Model for Multi-modality Image Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08406 v1 pith:AZKHBH3E submitted 2024-04-12 cs.CV

classification cs.CV
keywords fusionimagefeaturesdual-phasefeaturefusedmambadfusetasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modality image fusion (MMIF) aims to integrate complementary information from different modalities into a single fused image to represent the imaging scene and facilitate downstream visual tasks comprehensively. In recent years, significant progress has been made in MMIF tasks due to advances in deep neural networks. However, existing methods cannot effectively and efficiently extract modality-specific and modality-fused features constrained by the inherent local reductive bias (CNN) or quadratic computational complexity (Transformers). To overcome this issue, we propose a Mamba-based Dual-phase Fusion (MambaDFuse) model. Firstly, a dual-level feature extractor is designed to capture long-range features from single-modality images by extracting low and high-level features from CNN and Mamba blocks. Then, a dual-phase feature fusion module is proposed to obtain fusion features that combine complementary information from different modalities. It uses the channel exchange method for shallow fusion and the enhanced Multi-modal Mamba (M3) blocks for deep fusion. Finally, the fused image reconstruction module utilizes the inverse transformation of the feature extraction to generate the fused result. Through extensive experiments, our approach achieves promising fusion results in infrared-visible image fusion and medical image fusion. Additionally, in a unified benchmark, MambaDFuse has also demonstrated improved performance in downstream tasks such as object detection. Code with checkpoints will be available after the peer-review process.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    No verifiable result can be extracted because the body is a different paper on medical image fusion, not the NSFW detection study described in the abstract.

  2. Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

    cs.CV 2025-08 conditional novelty 4.0 of 10

    AdaSFFuse combines a learnable wavelet transform and a spatial-frequency Mamba block to report state-of-the-art fusion scores on infrared-visible, multi-exposure, multi-focus, and medical image pairs.

  3. MM2CT: MR-to-CT translation for multi-modal image fusion with mamba

    eess.IV 2025-08 conditional novelty 4.0 of 10

    MM2CT fuses T1- and T2-weighted MRI with Mamba blocks to synthesize CT images, reporting PSNR 25.72 and SSIM 89.54 on a pelvis dataset.

Pith tools