FoMoE partitions expert layers across workers in MoE LLMs, skips non-resident experts, and reports up to 1.42x lower communication than baselines plus 1.4x throughput gains while maintaining stable routing.
Federated mixture of experts
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 3verdicts
UNVERDICTED 3roles
method 1polarities
use method 1representative citing papers
FedSQ stabilizes federated weight averaging under heterogeneous data by fixing binary gating masks derived from a pretrained model's structure while optimizing only quantitative parameters.
FLEX-MoE proposes client-expert fitness scores and an optimization algorithm to jointly maximize specialization and enforce balanced expert utilization in federated MoE for edge computing under non-IID data and capacity constraints.
citing papers explorer
-
FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs
FoMoE partitions expert layers across workers in MoE LLMs, skips non-resident experts, and reports up to 1.42x lower communication than baselines plus 1.4x throughput gains while maintaining stable routing.
-
FedSQ: Optimized Weight Averaging via Fixed Gating
FedSQ stabilizes federated weight averaging under heterogeneous data by fixing binary gating masks derived from a pretrained model's structure while optimizing only quantitative parameters.
-
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
FLEX-MoE proposes client-expert fitness scores and an optimization algorithm to jointly maximize specialization and enforce balanced expert utilization in federated MoE for edge computing under non-IID data and capacity constraints.