Pith. sign in

REVIEW 2 cited by

Dynamic Vision Mamba

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.04787 v1 pith:2QBT7HBM submitted 2025-04-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords visionmambamodelsredundancytokenblockinferenceblocks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mamba-based vision models have gained extensive attention as a result of being computationally more efficient than attention-based models. However, spatial redundancy still exists in these models, represented by token and block redundancy. For token redundancy, we analytically find that early token pruning methods will result in inconsistency between training and inference or introduce extra computation for inference. Therefore, we customize token pruning to fit the Mamba structure by rearranging the pruned sequence before feeding it into the next Mamba block. For block redundancy, we allow each image to select SSM blocks dynamically based on an empirical observation that the inference speed of Mamba-based vision models is largely affected by the number of SSM blocks. Our proposed method, Dynamic Vision Mamba (DyVM), effectively reduces FLOPs with minor performance drops. We achieve a reduction of 35.2\% FLOPs with only a loss of accuracy of 1.7\% on Vim-S. It also generalizes well across different Mamba vision model architectures and different vision tasks. Our code will be made public.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RAUM-Net: Regional Attention and Uncertainty-aware Mamba Network

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RAUM-Net combines Mamba features, region attention, and MC-dropout uncertainty filtering to improve semi-supervised fine-grained classification under occlusion and label scarcity.

  2. Mammo-Mamba: A Hybrid State-Space and Transformer Architecture with Sequential Mixture of Experts for Multi-View Mammography

    eess.IV 2025-07 conditional novelty 4.0 of 10

    Mammo-Mamba, a gated MambaVision model with sequential mixture-of-expert-style depth routing, reports 0.8696 accuracy and 0.9089 AUC on the CBIS-DDSM mass classification benchmark.

Pith tools