A proposal for a MoE-based edge LLM deployment framework called CoEL, with a proof-of-concept showing that distributed inference across two edge servers is about 1.7 times slower than a single dual-GPU server.
The model compression and token compression are proposed in CoEL to address these issues
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.NI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
A proposal for a MoE-based edge LLM deployment framework called CoEL, with a proof-of-concept showing that distributed inference across two edge servers is about 1.7 times slower than a single dual-GPU server.