REVIEW 2 cited by
DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
The Mixture-of-Experts (MoE) architecture is showing promising results in improving parameter sharing in multi-task learning (MTL) and in scaling high-capacity neural networks. State-of-the-art MoE models use a trainable sparse gate to select a subset of the experts for each input example. While conceptually appealing, existing sparse gates, such as Top-k, are not smooth. The lack of smoothness can lead to convergence and statistical performance issues when training with gradient-based methods. In this paper, we develop DSelect-k: a continuously differentiable and sparse gate for MoE, based on a novel binary encoding formulation. The gate can be trained using first-order methods, such as stochastic gradient descent, and offers explicit control over the number of experts to select. We demonstrate the effectiveness of DSelect-k on both synthetic and real MTL datasets with up to $128$ tasks. Our experiments indicate that DSelect-k can achieve statistically significant improvements in prediction and expert selection over popular MoE gates. Notably, on a real-world, large-scale recommender system, DSelect-k achieves over $22\%$ improvement in predictive performance compared to Top-k. We provide an open-source implementation of DSelect-k.
Forward citations
Cited by 2 Pith papers
-
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
MoDE, a mixture-of-experts diffusion transformer with noise-conditioned routing, reports state-of-the-art results on CALVIN and LIBERO with lower inference FLOPs than dense baselines.
-
No More Tuning: Prioritized Multi-Task Learning with Lagrangian Differential Multiplier Methods
NMT optimizes lower-priority tasks under a Lagrangian penalty that keeps the primary task loss near its pre-trained optimum, with no manual balancing weights in the loss combination.
Discussion (0). Continue with ORCID to comment.