Pith. sign in

REVIEW 3 cited by

ReMamber: Referring Image Segmentation with Mamba Twister

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17839 v2 pith:67NYX3K6 submitted 2024-03-26 cs.CV cs.AI

classification cs.CVcs.AI
keywords mambaremambermulti-modaltwisterarchitecturechannelefficientfusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Referring Image Segmentation~(RIS) leveraging transformers has achieved great success on the interpretation of complex visual-language tasks. However, the quadratic computation cost makes it resource-consuming in capturing long-range visual-language dependencies. Fortunately, Mamba addresses this with efficient linear complexity in processing. However, directly applying Mamba to multi-modal interactions presents challenges, primarily due to inadequate channel interactions for the effective fusion of multi-modal data. In this paper, we propose ReMamber, a novel RIS architecture that integrates the power of Mamba with a multi-modal Mamba Twister block. The Mamba Twister explicitly models image-text interaction, and fuses textual and visual features through its unique channel and spatial twisting mechanism. We achieve competitive results on three challenging benchmarks with a simple and efficient architecture. Moreover, we conduct thorough analyses of ReMamber and discuss other fusion designs using Mamba. These provide valuable perspectives for future research. The code has been released at: https://github.com/yyh-rain-song/ReMamber.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    DETRIS uses dense mixtures of convolutions and cross-attention adapters to tune a frozen DINOv2/CLIP pair, achieving top reported IoU on three referring image segmentation benchmarks while updating only a small fracti...

  2. AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation

    cs.CV 2025-01 reject novelty 6.0 of 10

    AVS-Mamba applies Mamba with temporal and cross-modal scanning to audio-visual segmentation, reporting top scores on AVSBench-object but not on AVSBench-semantic with the stronger backbone.

  3. EDMB: Edge Detector with Mamba

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A Mamba-based edge detector achieves SOTA results on BSDS500 and produces multi-granularity edges on single-label datasets using an ELBO-supervised Gaussian decoder.

Pith tools