Pith. sign in

REVIEW 1 cited by

MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04081 v2 pith:MRLI72XK submitted 2024-01-08 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords mambamodelsmoe-mambaperformanceexpertsmixturemodelrecent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

State Space Models (SSMs) have become serious contenders in the field of sequential modeling, challenging the dominance of Transformers. At the same time, Mixture of Experts (MoE) has significantly improved Transformer-based Large Language Models, including recent state-of-the-art open models. We propose that to unlock the potential of SSMs for scaling, they should be combined with MoE. We showcase this on Mamba, a recent SSM-based model that achieves remarkable performance. Our model, MoE-Mamba, outperforms both Mamba and baseline Transformer-MoE. In particular, MoE-Mamba reaches the same performance as Mamba in $2.35\times$ fewer training steps while preserving the inference performance gains of Mamba against Transformer.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Routing Mamba applies mixture-of-experts to Mamba projection layers with one shared router, reporting perplexity parity with dense Mamba at roughly half the active parameters on 20B-token pretraining.

Pith tools