Pith. sign in

REVIEW 2 cited by

DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.03760 v3 pith:UX4SGRUF submitted 2021-06-07 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords dselect-kexpertsgatesparsedifferentiablegateslearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The Mixture-of-Experts (MoE) architecture is showing promising results in improving parameter sharing in multi-task learning (MTL) and in scaling high-capacity neural networks. State-of-the-art MoE models use a trainable sparse gate to select a subset of the experts for each input example. While conceptually appealing, existing sparse gates, such as Top-k, are not smooth. The lack of smoothness can lead to convergence and statistical performance issues when training with gradient-based methods. In this paper, we develop DSelect-k: a continuously differentiable and sparse gate for MoE, based on a novel binary encoding formulation. The gate can be trained using first-order methods, such as stochastic gradient descent, and offers explicit control over the number of experts to select. We demonstrate the effectiveness of DSelect-k on both synthetic and real MTL datasets with up to $128$ tasks. Our experiments indicate that DSelect-k can achieve statistically significant improvements in prediction and expert selection over popular MoE gates. Notably, on a real-world, large-scale recommender system, DSelect-k achieves over $22\%$ improvement in predictive performance compared to Top-k. We provide an open-source implementation of DSelect-k.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    MoDE, a mixture-of-experts diffusion transformer with noise-conditioned routing, reports state-of-the-art results on CALVIN and LIBERO with lower inference FLOPs than dense baselines.

  2. No More Tuning: Prioritized Multi-Task Learning with Lagrangian Differential Multiplier Methods

    cs.LG 2024-12 reject novelty 4.0 of 10

    NMT optimizes lower-priority tasks under a Lagrangian penalty that keeps the primary task loss near its pre-trained optimum, with no manual balancing weights in the loss combination.

Pith tools