BALANCE jointly schedules users between autoregressive and speculative decoding on a single edge GPU, with a 1/2-approximation algorithm, and reports simulated throughput gains of 27 to 39 percent over either mode alone.
Network edge inference for large language models: Principles, tech- niques, and opportunities,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.NI 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
BALANCE jointly schedules users between autoregressive and speculative decoding on a single edge GPU, with a 1/2-approximation algorithm, and reports simulated throughput gains of 27 to 39 percent over either mode alone.