Pith. sign in

Optimizing speculative decoding for serving large language models using goodput

10 Pith papers cite this work. Polarity classification is still indexing.

10 Pith papers citing it

citation-role summary

background 2

citation-polarity summary

years

2026 8 2025 2

verdicts

UNVERDICTED 10

roles

background 2

polarities

background 2

representative citing papers

Regulating Branch Parallelism in LLM Serving

cs.DC · 2026-05-07 · unverdicted · novelty 7.0

TAPER regulates LLM branch parallelism by admitting extra branches opportunistically when predicted externality fits slack, delivering 1.48-1.77x higher goodput than eager or fixed-cap baselines on Qwen3-32B while keeping over 95% SLO attainment.

An Interpretable Latency Model for Speculative Decoding in LLM Serving

cs.LG · 2026-05-14 · unverdicted · novelty 6.0

The paper presents an interpretable latency model for speculative decoding that infers effective batch size via Little's Law and decomposes demand to predict and explain performance across serving loads, validated on vLLM measurements.

citing papers explorer

Showing 10 of 10 citing papers.