REVIEW 6 cited by
Task-Specific Expert Pruning for Sparse Mixture-of-Experts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The sparse Mixture-of-Experts (MoE) model is powerful for large-scale pre-training and has achieved promising results due to its model capacity. However, with trillions of parameters, MoE is hard to be deployed on cloud or mobile environment. The inference of MoE requires expert parallelism, which is not hardware-friendly and communication expensive. Especially for resource-limited downstream tasks, such sparse structure has to sacrifice a lot of computing efficiency for limited performance gains. In this work, we observe most experts contribute scarcely little to the MoE fine-tuning and inference. We further propose a general method to progressively drop the non-professional experts for the target downstream task, which preserves the benefits of MoE while reducing the MoE model into one single-expert dense model. Our experiments reveal that the fine-tuned single-expert model could preserve 99.3% benefits from MoE across six different types of tasks while enjoying 2x inference speed with free communication cost.
Forward citations
Cited by 6 Pith papers
-
AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding
AcceptMoE is a verifier-side expert selector that uses commitment-weighted router demand and an entropy-based set size, reducing expert traffic while keeping mean accuracy within 0.27 percentage points of EAGLE-3 with...
-
Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference
Communication-aware expert placement plus device-level pruning yields 1.23–1.86× MoE inference throughput and better accuracy at equal speedup than load-balance or sequential baselines.
-
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
DERN prunes SMoE LLMs by decomposing removed experts into neuron segments, reassigning the best-matching ones to kept experts, and clustering them into compact replacements, beating prior pruning baselines without retraining.
-
Unified Start, Personalized End: Progressive Pruning for Efficient 3D Medical Image Segmentation
PSP-Seg prunes redundant modules during training to make 3D segmentation networks much smaller and faster without losing accuracy.
-
MoENAS: Mixture-of-Expert based Neural Architecture Search for jointly Accurate, Fair, and Robust Edge Deep Neural Networks
MoENAS, a mixture-of-experts neural architecture search, produces MobileViTv2 variants with reported accuracy, fairness, robustness, and generalization gains over state-of-the-art edge DNNs on person classification.
-
Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia
Deleting syntax- or semantics-specialized experts in a Mixture-of-Experts language model reproduces Broca's- and Wernicke's-like aphasia, and retraining the remaining experts models functional recovery.
Discussion (0). Continue with ORCID to comment.