FlexMoE produces nested pruned subnetworks for MoE LLMs across budgets via channel importance ranking and discrete action learning, plus one mid-budget recovery fine-tune, retaining 99.8% performance at 50% expert parameter pruning.
Towards resiliency in large language model serving with KevlarFlow
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
LUMEN coordinates failure recovery in distributed LLM serving via load-aware decisions on checkpoints, request routing, and model reload, showing improved serving and recovery times in prototypes and simulations.
citing papers explorer
-
FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models
FlexMoE produces nested pruned subnetworks for MoE LLMs across budgets via channel importance ranking and discrete action learning, plus one mid-budget recovery fine-tune, retaining 99.8% performance at 50% expert parameter pruning.
-
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving
LUMEN coordinates failure recovery in distributed LLM serving via load-aware decisions on checkpoints, request routing, and model reload, showing improved serving and recovery times in prototypes and simulations.