Pith. sign in

REVIEW 2 cited by

Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.09808 v3 pith:GZG7EUQX submitted 2024-09-15 cs.CV cs.AI

classification cs.CVcs.AI
keywords famba-vcross-layermambamodelstrainingefficiencyfusiontoken
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mamba and Vision Mamba (Vim) models have shown their potential as an alternative to methods based on Transformer architecture. This work introduces Fast Mamba for Vision (Famba-V), a cross-layer token fusion technique to enhance the training efficiency of Vim models. The key idea of Famba-V is to identify and fuse similar tokens across different Vim layers based on a suit of cross-layer strategies instead of simply applying token fusion uniformly across all the layers that existing works propose. We evaluate the performance of Famba-V on CIFAR-100. Our results show that Famba-V is able to enhance the training efficiency of Vim models by reducing both training time and peak memory usage during training. Moreover, the proposed cross-layer strategies allow Famba-V to deliver superior accuracy-efficiency trade-offs. These results all together demonstrate Famba-V as a promising efficiency enhancement technique for Vim models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UAVD-Mamba: Deformable Token Fusion Vision Mamba for Multimodal UAV Detection

    cs.CV 2025-07 conditional novelty 6.0 of 10

    UAVD-Mamba, a Mamba-based multimodal detector with deformable token blocks, reaches 83.0% mAP on DroneVehicle, outperforming OAFA by 3.6%.

  2. A Survey on Mamba Architecture for Vision Applications

    cs.CV 2025-02 conditional novelty 1.0 of 10

    A survey of Mamba-based vision models that summarizes scanning mechanisms, key architectures, and benchmark results, contributing no new experimental findings.

Pith tools