DFlare replaces DFlash's shared fused representation with per-draft-layer attention to distinct target-layer combinations, enabling deeper drafts and 2.4M training samples for 5-11% higher speedups than DFlash on Qwen3 and GPT-OSS models.
arXiv preprint arXiv:2509.22134 (2025)
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
SMART expands speculative decoding trees only when a node's marginal benefit-cost ratio exceeds current tree-level speedup, claiming ~15–20% extra wall-clock speedup without quality loss.
citing papers explorer
-
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
DFlare replaces DFlash's shared fused representation with per-draft-layer attention to distinct target-layer combinations, enabling deeper drafts and 2.4M training samples for 5-11% higher speedups than DFlash on Qwen3 and GPT-OSS models.
-
SMART: When is it Actually Worth Expanding a Speculative Tree?
SMART expands speculative decoding trees only when a node's marginal benefit-cost ratio exceeds current tree-level speedup, claiming ~15–20% extra wall-clock speedup without quality loss.