Pith. sign in

REVIEW 6 cited by

CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02526 v1 pith:ZOTKHO63 submitted 2024-02-04 cs.LG

classification cs.LG
keywords competitioncompetesmoeeffectiveexpertsrepresentationsmoecollapseenjoys
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Sparse mixture of experts (SMoE) offers an appealing solution to scale up the model complexity beyond the mean of increasing the network's depth or width. However, effective training of SMoE has proven to be challenging due to the representation collapse issue, which causes parameter redundancy and limited representation potentials. In this work, we propose a competition mechanism to address this fundamental challenge of representation collapse. By routing inputs only to experts with the highest neural response, we show that, under mild assumptions, competition enjoys the same convergence rate as the optimal estimator. We further propose CompeteSMoE, an effective and efficient algorithm to train large language models by deploying a simple router that predicts the competition outcomes. Consequently, CompeteSMoE enjoys strong performance gains from the competition routing policy while having low computation overheads. Our extensive empirical evaluations on two transformer architectures and a wide range of tasks demonstrate the efficacy, robustness, and scalability of CompeteSMoE compared to state-of-the-art SMoE strategies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models

    math.ST 2026-08 conditional novelty 7.0 of 10

    Kernels annihilated by a non-trivial parameter differential or difference-differential operator make the mixing measure in infinite mixtures non-identifiable, with a minimax lower bound ruling out consistent estimation.

  2. Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps

    stat.ML 2025-10 conditional novelty 6.0 of 10

    For softmax-gated Gaussian mixtures of experts, merging duplicate fitted atoms along a dendrogram and choosing the level by a height-likelihood score consistently recovers the true number of experts at parametric rate...

  3. Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures

    stat.ML 2025-05 conditional novelty 6.0 of 10

    A single overfitted GGMoE fit is converted into a dendrogram of merged experts, and a criterion based on merge heights and likelihood consistently recovers the true number of experts.

  4. Expert Merging in Sparse Mixture of Experts with Nash Bargaining

    cs.LG 2025-10 conditional novelty 4.0 of 10

    Expert merging in MoE models can be improved by setting per-expert weights with the Nash bargaining solution instead of simple averaging.

  5. On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation

    cs.LG 2025-02 conditional novelty 4.0 of 10

    Zero-initialized attention is cast as a mixture-of-experts model, with polynomial sample-complexity rates proven for estimating linear and non-linear prompts and the gating factor.

  6. Mixture of Experts (MoE): A Big Data Perspective

    cs.LG 2025-01 conditional novelty 2.0 of 10

    A survey of MoE methods for big data that catalogs architectures, use cases, and open challenges without adding new results.

Pith tools