REVIEW 6 cited by
CompeteSMoE -- Effective Training of Sparse Mixture of Experts via Competition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Sparse mixture of experts (SMoE) offers an appealing solution to scale up the model complexity beyond the mean of increasing the network's depth or width. However, effective training of SMoE has proven to be challenging due to the representation collapse issue, which causes parameter redundancy and limited representation potentials. In this work, we propose a competition mechanism to address this fundamental challenge of representation collapse. By routing inputs only to experts with the highest neural response, we show that, under mild assumptions, competition enjoys the same convergence rate as the optimal estimator. We further propose CompeteSMoE, an effective and efficient algorithm to train large language models by deploying a simple router that predicts the competition outcomes. Consequently, CompeteSMoE enjoys strong performance gains from the competition routing policy while having low computation overheads. Our extensive empirical evaluations on two transformer architectures and a wide range of tasks demonstrate the efficacy, robustness, and scalability of CompeteSMoE compared to state-of-the-art SMoE strategies.
Forward citations
Cited by 6 Pith papers
-
Partial Differential Equation Barriers to Identifiability in Infinite Mixture Models
Kernels annihilated by a non-trivial parameter differential or difference-differential operator make the mixing measure in infinite mixtures non-identifiable, with a minimax lower bound ruling out consistent estimation.
-
Dendrograms of Mixing Measures for Softmax-Gated Gaussian Mixture of Experts: Consistency Without Model Sweeps
For softmax-gated Gaussian mixtures of experts, merging duplicate fitted atoms along a dendrogram and choosing the level by a height-likelihood score consistently recovers the true number of experts at parametric rate...
-
Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures
A single overfitted GGMoE fit is converted into a dendrogram of merged experts, and a criterion based on merge heights and likelihood consistently recovers the true number of experts.
-
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
Expert merging in MoE models can be improved by setting per-expert weights with the Nash bargaining solution instead of simple averaging.
-
On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation
Zero-initialized attention is cast as a mixture-of-experts model, with polynomial sample-complexity rates proven for estimating linear and non-linear prompts and the gating factor.
-
Mixture of Experts (MoE): A Big Data Perspective
A survey of MoE methods for big data that catalogs architectures, use cases, and open challenges without adding new results.
Discussion (0). Continue with ORCID to comment.