Pith. sign in

REVIEW 4 cited by

Scaling Vision-Language Models with Sparse Mixture of Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07226 v1 pith:UEFROMZ5 submitted 2023-03-13 cs.CV cs.CL

classification cs.CVcs.CL
keywords modelsscalingvision-languagebecomelarge-scalemodelperformanceresearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The field of natural language processing (NLP) has made significant strides in recent years, particularly in the development of large-scale vision-language models (VLMs). These models aim to bridge the gap between text and visual information, enabling a more comprehensive understanding of multimedia data. However, as these models become larger and more complex, they also become more challenging to train and deploy. One approach to addressing this challenge is the use of sparsely-gated mixture-of-experts (MoE) techniques, which divide the model into smaller, specialized sub-models that can jointly solve a task. In this paper, we explore the effectiveness of MoE in scaling vision-language models, demonstrating its potential to achieve state-of-the-art performance on a range of benchmarks over dense models of equivalent computational cost. Our research offers valuable insights into stabilizing the training of MoE models, understanding the impact of MoE on model interpretability, and balancing the trade-offs between compute performance when scaling VLMs. We hope our work will inspire further research into the use of MoE for scaling large-scale vision-language models and other multimodal machine learning applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection

    cs.CV 2025-07 conditional novelty 7.0 of 10

    MoE fine-tuning with decomposed pre-trained FFN experts lets a real-time open-vocabulary detector beat a much larger-data baseline with similar active parameter count.

  2. Mixture of Cognitive Experts in Large Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Routing CV experts into atomic evidence then Bloom-staged verbalization improves LVLM benchmarks and yields measurable query-conditioned reasoning traces.

  3. Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.5 of 10

    A monolithic multimodal LLM that cuts pre-training data by 58% and first-token latency by up to 69% while matching or beating its predecessor on 15 benchmarks.

  4. SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities

    cs.CL 2025-06 conditional novelty 5.0 of 10

    SMAR, a KL-based regularizer on per-modality expert routing, retains 86.6% of a Mixtral 8x7B's language score during visual instruction tuning with only 2.5% pure-text data.

Pith tools