LAPS-SD schedules speculative decoding requests by combining LAS-style preemption during early unstable acceptance rates with SJF ordering once acceptance rates stabilize, cutting average inference latency by about 39% in experiments.
Luan, Zhou Su, and Jing Deng
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
LAPS-SD schedules speculative decoding requests by combining LAS-style preemption during early unstable acceptance rates with SJF ordering once acceptance rates stabilize, cutting average inference latency by about 39% in experiments.