Pith. sign in

REVIEW 13 cited by

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14520 v4 pith:MFG5AZXP submitted 2024-03-21 cs.CV

classification cs.CV
keywords cobraefficientcomplexitylanguagemambamllmmodelachieves
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In recent years, the application of multimodal large language models (MLLM) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, current MLLMs are composed of the well-known Transformer network, which has a less efficient quadratic computation complexity. To improve the efficiency of such basic models, we propose Cobra, a linear computational complexity MLLM. Specifically, Cobra integrates the efficient Mamba language model into the visual modality. Moreover, we explore and study various modal fusion schemes to create an effective multi-modal Mamba. Extensive experiments demonstrate that (1) Cobra achieves extremely competitive performance with current computationally efficient state-of-the-art methods, e.g., LLaVA-Phi, TinyLLaVA, and MobileVLM v2, and has faster speed due to Cobra's linear sequential modeling. (2) Interestingly, the results of closed-set challenging prediction benchmarks show that Cobra performs well in overcoming visual illusions and spatial relationship judgments. (3) Notably, Cobra even achieves comparable performance to LLaVA with about 43% of the number of parameters. We will make all codes of Cobra open-source and hope that the proposed method can facilitate future research on complexity problems in MLLM. Our project page is available at: https://sites.google.com/view/cobravlm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Auras, a perception-generation disaggregation framework with a public context buffer and asynchronous pipeline executor, raises embodied-agent throughput by 2.54x on average without losing accuracy (102.7%).

  2. scMamba: A Scalable Foundation Model for Single-Cell Multi-Omics Integration Beyond Highly Variable Feature Selection

    q-bio.CB 2025-06 conditional novelty 6.0 of 10

    A Mamba-based model integrates paired single-cell RNA and ATAC data using all genomic features, with patch tokenization and contrastive learning, and reports better integration and downstream performance than seven ex...

  3. Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors distill a diffusion transformer into a mostly-Mamba hybrid model, reaching teacher-level GenEval scores while generating up to 4K images with linear-complexity speed.

  4. MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A quantization framework using KLT-enhanced and smooth-fused rotations lets Mamba models run at 8-bit precision with near-full accuracy and at 4-bit weights with moderate loss.

  5. MBQ: Modality-Balanced Quantization for Large Vision-Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A modality-weighted quantization method improves accuracy of 3-bit and 4-bit vision-language models by protecting sensitive language tokens during calibration.

  6. Efficient Self-Supervised Video Hashing with Selective State Spaces

    cs.CV 2024-12 conditional novelty 6.0 of 10

    S5VH uses bidirectional Mamba layers and a hash-center alignment loss to improve self-supervised video hashing accuracy and efficiency.

  7. CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A coarse-to-fine autoregressive policy with multi-scale action tokenization matches or beats diffusion policies on robot manipulation benchmarks at roughly 10x lower inference cost.

  8. Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

    cs.CV 2026-07 accept novelty 5.5 of 10

    A training-free subtraction-then-attention cascade matches or beats full fusion recall on LEVIR-CC at 10–15× lower query cost; Mamba is no faster than attention at L=196; TBF cuts parameters 2.3× for a 0.007 BLEU-1 cost.

  9. HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    HAMF feeds learnable future motion tokens into the scene encoder alongside road and agent tokens, then uses a Mamba decoder to output six diverse trajectories, achieving competitive Argoverse 2 results with 3.0M parameters.

  10. PerPO: Perceptual Preference Optimization via Discriminative Rewarding

    cs.AI 2025-02 conditional novelty 5.0 of 10

    PerPO trains multimodal LLMs by ranking their candidate answers with deterministic visual rewards (IoU, edit distance) and using the reward differences as margins in listwise preference optimization.

  11. Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity

    cs.LG 2025-01 conditional novelty 5.0 of 10

    Modality-specific projection weights let a Mamba model match dense multimodal baselines at the same loss using 25% to 65% of the training compute.

  12. QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning

    cs.RO 2024-12 conditional novelty 5.0 of 10

    Compressing 10-step action chunks into discrete latent codes lets an 8B multimodal model drive a quadruped at controller frequency and raises average task success by about 65%.

  13. AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment

    cs.CV 2024-12 conditional novelty 5.0 of 10

    AlignMamba fuses audio, video, and language by matching tokens to a language anchor and enforcing distribution similarity, reporting small accuracy gains with large efficiency gains on MOSI and MOSEI.

Pith tools