Pith. sign in

REVIEW 2 cited by

On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02935 v2 pith:KU6CHKQI submitted 2024-10-03 stat.ML cs.LG

classification stat.MLcs.LG
keywords gatinghmoeexpertexpertsmixturemodelssoftmaxtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the growing prominence of the Mixture of Experts (MoE) architecture in developing large-scale foundation models, we investigate the Hierarchical Mixture of Experts (HMoE), a specialized variant of MoE that excels in handling complex inputs and improving performance on targeted tasks. Our analysis highlights the advantages of using the Laplace gating function over the traditional Softmax gating within the HMoE frameworks. We theoretically demonstrate that applying the Laplace gating function at both levels of the HMoE model helps eliminate undesirable parameter interactions caused by the Softmax gating and, therefore, accelerates the expert convergence as well as enhances the expert specialization. Empirical validation across diverse scenarios supports these theoretical claims. This includes large-scale multimodal tasks, image classification, and latent domain discovery and prediction tasks, where our modified HMoE models show great performance improvements compared to the conventional HMoE models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective

    cs.LG 2025-02 reject novelty 5.0 of 10

    The paper derives convergence rates for sigmoid gating mixture-of-experts with quadratic scores and uses them to argue sigmoid self-attention is more sample-efficient than softmax, but the link to attention is an unpr...

  2. On Zero-Initialized Attention: Optimal Prompt and Gating Factor Estimation

    cs.LG 2025-02 conditional novelty 4.0 of 10

    Zero-initialized attention is cast as a mixture-of-experts model, with polynomial sample-complexity rates proven for estimating linear and non-linear prompts and the gating factor.

Pith tools