A single-GPU serving system for multiple fine-tuned MoE LLMs achieves near-single-model throughput by sharing similar experts and reconfiguring non-expert layers at runtime.
K., El-Araby, E., and El-Ghazawi, T
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration
A single-GPU serving system for multiple fine-tuned MoE LLMs achieves near-single-model throughput by sharing similar experts and reconfiguring non-expert layers at runtime.