LEDE reframes speculative decoding as an MDP and applies offline RL to learn dynamic policies for exit layer and speculation length selection, delivering 2.0-2.7x speedups over autoregressive decoding on Llama-2/3 models.
Accelerating inference in large language models with a unified layer skipping strategy,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Experience-Driven Dynamic Exits for LLMs with Reinforcement Learning
LEDE reframes speculative decoding as an MDP and applies offline RL to learn dynamic policies for exit layer and speculation length selection, delivering 2.0-2.7x speedups over autoregressive decoding on Llama-2/3 models.