REVIEW 2 cited by
Dynamic Vision Mamba
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Mamba-based vision models have gained extensive attention as a result of being computationally more efficient than attention-based models. However, spatial redundancy still exists in these models, represented by token and block redundancy. For token redundancy, we analytically find that early token pruning methods will result in inconsistency between training and inference or introduce extra computation for inference. Therefore, we customize token pruning to fit the Mamba structure by rearranging the pruned sequence before feeding it into the next Mamba block. For block redundancy, we allow each image to select SSM blocks dynamically based on an empirical observation that the inference speed of Mamba-based vision models is largely affected by the number of SSM blocks. Our proposed method, Dynamic Vision Mamba (DyVM), effectively reduces FLOPs with minor performance drops. We achieve a reduction of 35.2\% FLOPs with only a loss of accuracy of 1.7\% on Vim-S. It also generalizes well across different Mamba vision model architectures and different vision tasks. Our code will be made public.
Forward citations
Cited by 2 Pith papers
-
RAUM-Net: Regional Attention and Uncertainty-aware Mamba Network
RAUM-Net combines Mamba features, region attention, and MC-dropout uncertainty filtering to improve semi-supervised fine-grained classification under occlusion and label scarcity.
-
Mammo-Mamba: A Hybrid State-Space and Transformer Architecture with Sequential Mixture of Experts for Multi-View Mammography
Mammo-Mamba, a gated MambaVision model with sequential mixture-of-expert-style depth routing, reports 0.8696 accuracy and 0.9089 AUC on the CBIS-DDSM mass classification benchmark.
Discussion (0). Continue with ORCID to comment.