BALANCE jointly schedules users between autoregressive and speculative decoding on a single edge GPU, with a 1/2-approximation algorithm, and reports simulated throughput gains of 27 to 39 percent over either mode alone.
QLLMS: Quantization-adaptive LLM scheduling for partially informed edge serving systems,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.NI 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
BALANCE jointly schedules users between autoregressive and speculative decoding on a single edge GPU, with a 1/2-approximation algorithm, and reports simulated throughput gains of 27 to 39 percent over either mode alone.