EPD Disaggregation separates encoding from prefill and decode in LMM serving, enabling parallel encoding, cached token transfer, and dynamic resource shifts that improve TTFT, memory, batch size, and SLO attainment.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Efficiently Serving Large Multimodal Models Using EPD Disaggregation
EPD Disaggregation separates encoding from prefill and decode in LMM serving, enabling parallel encoding, cached token transfer, and dynamic resource shifts that improve TTFT, memory, batch size, and SLO attainment.