REVIEW 3 cited by
ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multimodal Large Language Models (MLLMs) have attracted much attention for their multifunctionality. However, traditional Transformer architectures incur significant overhead due to their secondary computational complexity. To address this issue, we introduce ML-Mamba, a multimodal language model, which utilizes the latest and efficient Mamba-2 model for inference. Mamba-2 is known for its linear scalability and fast processing of long sequences. We replace the Transformer-based backbone with a pre-trained Mamba-2 model and explore methods for integrating 2D visual selective scanning mechanisms into multimodal learning while also trying various visual encoders and Mamba-2 model variants. Our extensive experiments in various multimodal benchmark tests demonstrate the competitive performance of ML-Mamba and highlight the potential of state space models in multimodal tasks. The experimental results show that: (1) we empirically explore how to effectively apply the 2D vision selective scan mechanism for multimodal learning. We propose a novel multimodal connector called the Mamba-2 Scan Connector (MSC), which enhances representational capabilities. (2) ML-Mamba achieves performance comparable to state-of-the-art methods such as TinyLaVA and MobileVLM v2 through its linear sequential modeling while faster inference speed; (3) Compared to multimodal models utilizing Mamba-1, the Mamba-2-based ML-Mamba exhibits superior inference performance and effectiveness.
Forward citations
Cited by 3 Pith papers
-
scMamba: A Scalable Foundation Model for Single-Cell Multi-Omics Integration Beyond Highly Variable Feature Selection
A Mamba-based model integrates paired single-cell RNA and ATAC data using all genomic features, with patch tokenization and contrastive learning, and reports better integration and downstream performance than seven ex...
-
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
The authors distill a diffusion transformer into a mostly-Mamba hybrid model, reaching teacher-level GenEval scores while generating up to 4K images with linear-complexity speed.
-
SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
SSD-Poser reconstructs full-body SMPL poses from three sparse HMD signals with a hybrid state-space/attention encoder and frequency-aware decoder, reporting state-of-the-art accuracy at real-time speed on AMASS.
Discussion (0). Continue with ORCID to comment.