Pith. sign in

REVIEW 3 cited by

Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.07689 v1 pith:34EBOULN submitted 2022-04-16 cs.LG cs.CL

classification cs.LGcs.CL
keywords taskslearningmulti-taskactivateddifferentmodelsparselyweights
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traditional multi-task learning (MTL) methods use dense networks that use the same set of shared weights across several different tasks. This often creates interference where two or more tasks compete to pull model parameters in different directions. In this work, we study whether sparsely activated Mixture-of-Experts (MoE) improve multi-task learning by specializing some weights for learning shared representations and using the others for learning task-specific information. To this end, we devise task-aware gating functions to route examples from different tasks to specialized experts which share subsets of network weights conditioned on the task. This results in a sparsely activated multi-task model with a large number of parameters, but with the same computational cost as that of a dense model. We demonstrate such sparse networks to improve multi-task learning along three key dimensions: (i) transfer to low-resource tasks from related tasks in the training mixture; (ii) sample-efficient generalization to tasks not seen during training by making use of task-aware routing from seen related tasks; (iii) robustness to the addition of unrelated tasks by avoiding catastrophic forgetting of existing tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SciGPT: A Large Language Model for Scientific Literature Understanding and Knowledge Discovery

    cs.CL 2025-09 reject novelty 4.0 of 10

    SciGPT, a fine-tuned Qwen3 model for scientific literature, is reported to outperform GPT-4 on a new ScienceBench benchmark, but the evaluation is unreliable due to missing artifacts and contradictory numbers.

  2. Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A supervised mixture of experts, routed by fixed bandwidth and task labels, improves multi-task ASR/ST over hard parameter sharing while keeping active parameters constant.

  3. Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure

    cs.DC 2025-07 reject novelty 4.0 of 10

    A CXL-based disaggregated memory architecture with hybrid XLink interconnects is proposed and prototyped, claiming large speedups for memory-bound AI and HPC workloads.

Pith tools