Pith. sign in

REVIEW 13 cited by

Visual Mamba: A Survey and New Outlooks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.18861 v3 pith:JZFMW4KZ submitted 2024-04-29 cs.CV

classification cs.CV
keywords mambavisualchallengesmodelapplicationscomputationaldatafoundation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Mamba, a recent selective structured state space model, excels in long sequence modeling, which is vital in the large model era. Long sequence modeling poses significant challenges, including capturing long-range dependencies within the data and handling the computational demands caused by their extensive length. Mamba addresses these challenges by overcoming the local perception limitations of convolutional neural networks and the quadratic computational complexity of Transformers. Given its advantages over these mainstream foundation architectures, Mamba exhibits great potential to be a visual foundation architecture. Since January 2024, Mamba has been actively applied to diverse computer vision tasks, yielding numerous contributions. To help keep pace with the rapid advancements, this paper reviews visual Mamba approaches, analyzing over 200 papers. This paper begins by delineating the formulation of the original Mamba model. Subsequently, it delves into representative backbone networks, and applications categorized using different modalities, including image, video, point cloud, and multi-modal data. Particularly, we identify scanning techniques as critical for adapting Mamba to vision tasks, and decouple these scanning techniques to clarify their functionality and enhance their flexibility across various applications. Finally, we discuss the challenges and future directions, providing insights into new outlooks in this fast evolving area. A comprehensive list of visual Mamba models reviewed in this work is available at https://github.com/Ruixxxx/Awesome-Vision-Mamba-Models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond BEV: Optimizing Point-Level Tokens for Collaborative Perception

    cs.CV 2025-08 conditional novelty 7.0 of 10

    CoPLOT replaces BEV features with semantically ordered, frequency-enhanced point-level tokens for collaborative perception, improving 3D detection while cutting overhead.

  2. DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone

    cs.LG 2025-11 conditional novelty 6.0 of 10

    A masked diffusion language model built on a bidirectional Mamba backbone matches Transformer-based denoisers on quality while decoding with near-linear time scaling.

  3. Focus Through Motion: RGB-Event Collaborative Token Sparsification for Efficient Object Detection

    cs.CV 2025-09 conditional novelty 6.0 of 10

    FocusMamba uses event-camera activity to adaptively prune uninformative tokens in both RGB and event streams, improving detection accuracy and cutting FLOPs.

  4. ECP-Mamba: An Efficient Multi-scale Self-supervised Contrastive Learning Method with State Space Model for PolSAR Image Classification

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ECP-Mamba, a Mamba-based network with a spiral scan and multi-scale self-distillation, reports state-of-the-art PolSAR image classification accuracy at label rates as low as 0.2%.

  5. SMamba: Sparse Mamba for Event-based Object Detection

    cs.CV 2025-01 conditional novelty 6.0 of 10

    SMamba prunes uninformative event tokens using a spatiotemporal continuity score, then scans the survivors with Mamba blocks to match or beat prior event detectors at lower compute.

  6. DH-Mamba: Exploring Dual-domain Hierarchical State Space Models for MRI Reconstruction

    eess.IV 2025-01 conditional novelty 6.0 of 10

    DH-Mamba is a dual-domain hierarchical Mamba network that uses circular k-space scanning and local diversity enhancement to outperform prior MRI reconstruction methods on three public datasets.

  7. DIME-Net: A Dual-Illumination Adaptive Enhancement Network Based on Retinex and Mixture-of-Experts

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A single Retinex-based network with sparse mixture-of-experts tone curves, trained on a mixed low-light/backlit dataset, improves PSNR/SSIM/LPIPS on LOLv1 and BAID without dataset-specific retraining.

  8. Few-Shot Object Detection via Spatial-Channel State Space Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A Mamba-based channel sequence model combined with spatial attention improves few-shot object detection on VOC and COCO.

  9. MambaHSI: Spatial-Spectral Mamba for Hyperspectral Image Classification

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A pixel-level Mamba model with spatial and spectral branches reports state-of-the-art hyperspectral image classification on four datasets.

  10. A Survey on Large Language Model Acceleration based on KV Cache Management

    cs.AI 2024-12 conditional novelty 4.0 of 10

    A survey that classifies KV cache management techniques for faster LLM inference into token-level, model-level, and system-level categories, with benchmark resources.

  11. Selective State Space Memory for Large Vision-Language Models

    cs.CV 2024-12 reject novelty 4.0 of 10

    SSMI inserts Mamba-based state space modules into LVLMs, fine-tunes only 0.5% of parameters, and reports higher captioning, VQA, and retrieval scores than its baselines.

  12. BadScan: An Architectural Backdoor Attack on Visual State Space Models

    cs.CV 2024-11 reject novelty 4.0 of 10

    BadScan is a trigger-activated architectural backdoor for VMamba that replaces the standard 2D selective scan with malformed scans at inference time.

  13. A Survey on Mamba Architecture for Vision Applications

    cs.CV 2025-02 conditional novelty 1.0 of 10

    A survey of Mamba-based vision models that summarizes scanning mechanisms, key architectures, and benchmark results, contributing no new experimental findings.

Pith tools