Embedding a lightweight draft model into the idle GPU cycles of weight-offloaded LLM inference yields a 2.54x throughput gain over the best baseline.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices
Embedding a lightweight draft model into the idle GPU cycles of weight-offloaded LLM inference yields a 2.54x throughput gain over the best baseline.