Pith. sign in

S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Speculative decoding (SD) has attracted a significant amount of research attention due to the substantial speedup it can achieve for LLM inference. However, despite the high speedups they offer, speculative decoding methods often achieve optimal performance on high-end devices or with a substantial GPU memory overhead. Given limited memory and the necessity of quantization, a high-performing model on a high-end GPU can slow down by up to 7 times. To this end, we propose Skippy Simultaneous Speculative Decoding (or S3D), a cost-effective self-speculative SD method based on simultaneous multi-token decoding and mid-layer skipping. When compared against recent effective open-source SD systems, our method has achieved one of the top performance-memory ratios while requiring minimal architecture changes and training data. Leveraging our memory efficiency, we created a smaller yet more effective SD model based on Phi-3. It is 1.4 to 2 times faster than the quantized EAGLE model and operates in half-precision while using less VRAM.

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Accelerating Large Language Model Reasoning via Speculative Search

cs.CL · 2025-05-03 · conditional · novelty 6.0

SpecSearch speeds up tree-search LLM reasoning by drafting thoughts with a small model, rejecting low-quality thoughts with a PRM-based threshold, and correcting them with a large model, achieving up to 2.12x speedup over token-level speculative decoding.

citing papers explorer

Showing 1 of 1 citing paper.

  • Accelerating Large Language Model Reasoning via Speculative Search cs.CL · 2025-05-03 · conditional · none · ref 42 · internal anchor

    SpecSearch speeds up tree-search LLM reasoning by drafting thoughts with a small model, rejecting low-quality thoughts with a PRM-based threshold, and correcting them with a large model, achieving up to 2.12x speedup over token-level speculative decoding.