REVIEW 4 major objections 5 minor 84 references
D-cut, a training-free pruning layer for batched speculative decoding, restores and extends speedups under high concurrency by verifying only the draft tokens most likely to be accepted.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 01:28 UTC pith:NHJFATVX
load-bearing objection D-cut is a solid, incremental systems paper: the negative result it targets (DFlash below AR at high concurrency) is real, and the confidence-based cross-request pruning mostly recovers speedup, but the evaluation lacks error bars and the cost-table profiling is under-specified. the 4 major comments →
D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that verification compute—not draft quality—is the scarce resource in batched speculative decoding, and that pruning draft depth across the batch by confidence, guided by a profiled hardware cost model, yields larger throughput gains than any fixed draft depth. The paper demonstrates that acceptance lengths vary widely across concurrent requests and that verification cost curves depend sharply on GPU and parallelism, so the optimal per-request keep-depth varies both across requests and across deployments. D-cut computes prefix-product confidence scores for every draft position, sums the top-K scores to estimate the batch's expected token advance for each budget ratio, an
What carries the argument
The prefix-product confidence score si,k = product of ci,t over t=1..k—the drafter's confidence that the block survives to depth k—is the ranking signal; global top-K selection over these scores yields contiguous prefixes per request, and the argmax over a discrete ratio set of (sum of top-K scores divided by profiled cost C(B,rho)) is the budget-selection rule. This identity decouples the algorithmic benefit (confidence-estimated mean accepted tokens) from the hardware cost (profiled step latency), letting the method adapt pruning depth to both the current batch's confidence and the deployment's cost curve.
Load-bearing premise
The drafter's prefix-product confidence reliably ranks which draft tokens the target model would accept, and the startup-profiled dummy-step latencies faithfully predict real serving latency at every batch shape—if either fails, the selected verification budget can be suboptimal and the measured speedups may not transfer.
What would settle it
Run D-cut on a deployment where the drafter's confidence is decorrelated from the target's acceptance (e.g., a distribution shift or a drafter known to be miscalibrated), measure whether the selected budget ratio still tracks the throughput-optimal ratio; or compare the profiled cost table against real step latencies under mixed-length requests and varying memory pressure to see whether C(B,rho) remains accurate.
If this is right
- D-cut makes block-parallel speculative decoding safe at high concurrency, preventing throughput from falling below autoregressive decoding on dense models.
- The method is training-free and preserves the target distribution exactly, so it can be dropped into existing serving stacks that already support variable-length verification batches.
- The runtime cost model removes manual per-deployment tuning: the same code picks aggressive pruning on compute-bound GPUs and conservative pruning on memory-bound ones.
- Because it selects only a global budget and ranks per-position, D-cut extends to any block-parallel drafter that exposes token confidences.
- Pruned verification shortens step time enough to offset selector overhead (measured at 2-3% of the step).
Where Pith is reading between the lines
- The confidence prefix-product is used as a ranking signal, not a calibrated probability; if a drafter's confidence is poorly correlated with target acceptance (drift, distribution shift, adversarial prompts), the top-K selection could prune accepted tokens, and the inferred speedups may not transfer. The paper's own monotonicity validation on three datasets is the only support.
- The startup cost table is profiled from dummy steps; at serve time, variable input lengths, preemption, or memory pressure could change the real latency curve, making the profiled argmax suboptimal—an online re-profiling or adaptive cost estimate would be a natural extension.
- The discrete ratio set {0.25,0.50,0.75,1.00} limits granularity; a continuous budget or a learned mapping from batch features to budget could extract more of the available speedup.
- The dynamic verification shape conflicts with static CUDA-graph capture in current engines, as the paper notes; co-designing the engine to accept variable packed shapes is a concrete engineering path to realizing the full benefit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes D-cut, a training-free scheduling layer for batched block-parallel speculative decoding. For each request, D-cut computes prefix-product confidence scores from the drafter, globally ranks all candidate draft positions across the batch, and selects one of a small set of verification budgets by maximizing an estimated speedup ratio: a confidence-based estimate of expected accepted tokens divided by a startup-profiled end-to-end step-latency table. The method is applied to DFlash-style drafters and evaluated on dense and MoE models across batch sizes and two GPU platforms. The paper claims that this restores acceleration where fixed-depth long drafts fall below autoregressive decoding at high concurrency, improves average speedup from 1.26x to 1.65x, and reaches up to 3.0x speedup over autoregressive decoding on MoE models, while preserving the target output distribution.
Significance. The verification-cost explosion under high concurrency is a real and increasingly important problem for speculative decoding, and the proposed cross-request, runtime-cost-aware budget allocation is a sensible and potentially impactful idea. The paper is strongest where it is concrete: Algorithm 1 is clearly specified, per-step selector complexity is given, the fixed-ratio ablation in Table 2 directly tests the budget-selection mechanism, the confidence signal is validated with Spearman/AUROC, and the authors provide an implementation link and a description of CUDA-graph-compatible packing. If the empirical claims hold, this is a practical contribution to LLM serving. The main risks are that the startup cost table is not shown to be representative of real serving-time sequence lengths, that the losslessness guarantee is stated more broadly than the proven setting, and that throughput results are reported without uncertainty estimates.
major comments (4)
- [Section C (Cost-table profiling protocol) and Eq. (5)] The cost table C(B,rho) is profiled on 'dummy speculative steps' at startup, but the sequence length and KV-cache state of those dummy tensors are never specified. In the served setup, context length reaches 8,192 tokens and outputs reach 2,048 tokens; verification latency in FlashAttention grows with sequence length. If the dummy profile uses near-empty sequences, C(B,rho) systematically underestimates verification cost during later decoding steps. Since Eq. (5) divides the confidence-derived benefit by C(B,rho), an underestimated denominator biases rho* toward larger verification budgets, which is exactly the regime where the paper claims D-cut restores acceleration. As written, the method is adaptive to static hardware and batch size, not to the runtime sequence state, contrary to the 'runtime-adaptive' claim. Please report the context length used in profiling and either include seque
- [Section 3.2 and Section D (Figure 7)] Eq. (3) treats the prefix-product si,k as a quantitative estimate of Pr(L_i >= k), and Eq. (5) sums these scores as an estimate of expected tokens advanced. The validation in Figure 7 is ordinal: Spearman correlation and AUROC establish monotonic ranking, not calibration. If the drafter confidences are systematically overconfident at depth, the sum in Eq. (5) overweights deep positions and selects too large a ratio. To support the argmax in Eq. (5), please report a reliability/calibration analysis (e.g., binned empirical acceptance vs. si,k, or the ratio of the sum estimate to true MAT) and show that the chosen ratio is robust to monotone rescaling of the scores.
- [Tables 1-2 and Figures 4-5] All throughput numbers are reported as single points with no repetitions, seeds, or confidence intervals. Several of the central comparisons are small: for example, Table 1 shows D-cut(8) vs. DFlash(8) on Qwen3.5-27B averaging 1.24x vs. 1.21x, and D-cut(8) vs. DFlash(8) on Hy3-295B-A21B GSM8K gives 2.51x vs. 2.50x; some individual entries are worse (e.g., Qwen3.5-35B-A3B GSM8K D-cut(8) 2.45x vs. DFlash 2.65x). Without repeated runs or error bars, the aggregate claim 'from 1.26x to 1.65x' is not statistically supported. Please add multiple seeds/runs and report mean and variance, or at least mark differences that are within measurement noise.
- [Section A, Section B, Abstract, and Section 6] The losslessness guarantee in Section A is proven for the evaluated setting: deterministic draft proposals and target-only verification. Section B then shows that applying the unshifted score si,k from Eq. (3) can change the output distribution under standard rejection sampling with stochastic proposals, and it prescribes shifted confidence and causal early stopping to fix this. Algorithm 1 and the main text use the unshifted score and do not incorporate that fix. Yet the abstract states 'without compromising output quality' and Section 6 states 'strictly preserving the target model's original output distribution' without these scope conditions. Please qualify the losslessness claim to the deterministic-draft/target-only setting, and state clearly whether the §4.5 temperature-1 results use the causal variant or rely on greedy draft proposals.
minor comments (5)
- [Section 4.1 / Section 4.5] 'Target-only verification' is used as a key procedural term but is never precisely defined in the main text. Please specify the acceptance/residual-sampling protocol, especially for the temperature-1 experiment in Figure 5.
- [Figure 2b] The abbreviation 'P50' is used but not defined; please state that it is the median over steps/repetitions and how many steps/repetitions were used.
- [Algorithm 1] Line 5 says 'ties toward smaller k', but the tie-breaking rule across requests is not fully specified. Section B mentions deterministic tie-breaking by depth and request index; please make Algorithm 1 consistent with the implemented rule.
- [Figure 4] The 5x6 panel layout is very dense; axis labels and legend entries are hard to read at print size. Consider enlarging panels or splitting into subfigures.
- [Section C] The statement 'batch sizes that fall between captured points reuse the nearest captured entry' should clarify whether the nearest is by batch size only, and how this interacts with the sequence-length dependence raised in the major comments.
Circularity Check
No significant circularity: the speedup results are measured end-to-end; confidence scores and the profiled cost table are external inputs rather than fitted outputs.
full rationale
D-cut's selection rule (Eq. 5) is an online argmax over two inputs: prefix-product drafter confidences si,k (Eq. 3) and a startup-profiled latency table C(B,rho). Neither input is fitted to the reported throughput numbers. The confidence signal is validated in Section D against acceptance labels from full verification on three benchmarks (Spearman 0.579-0.770, AUROC 0.957-0.972), so the ranking premise has in-paper empirical support rather than resting on the method's own selections. The cost table is an end-to-end dummy-step measurement described in Section C; any mismatch with real serving (e.g., sequence-length-dependent verification cost) would be a measurement-validity limitation of the adaptivity claim, not a circular reduction. Reported speedups in Tables 1-2 and Figures 4-5 are measured output-token throughput over AR, not values implied by Eq. (5). Losslessness in Section A follows from prefix truncation leaving the target's conditional distributions unchanged, and Section B supplies an explicit causal construction for the rejection-sampling extension; neither proof assumes what it derives. Self-citations (PEARL, TALON, ECHO, DOUBLE) appear in motivation and related work, but the load-bearing premises (acceptance-length variance, runtime cost variation, confidence monotonicity) are re-established in Figures 1-2 and Section D, so the self-citations are contextual rather than load-bearing.
Axiom & Free-Parameter Ledger
free parameters (1)
- Keep-ratio buckets R =
{0.25, 0.50, 0.75, 1.00}
axioms (3)
- ad hoc to paper Prefix-product confidence si,k = prod_{t=1}^k ci,t estimates Pr(L_i >= k), the probability the target verifier accepts draft positions 1..k.
- domain assumption Startup-profiled cost table C(B,rho) measured with dummy steps equals real serving step latency at the same batch shape.
- domain assumption Target-only verification with truncated prefixes preserves the target output distribution.
read the original abstract
Speculative decoding accelerates large language model (LLM) inference without compromising output quality. Recent parallel drafting methods further improve single-request performance by decoupling draft length from drafting latency, enabling longer drafts and higher mean accepted tokens (MAT). However, under high request concurrency, long drafts waste substantial computation on rejected tokens, increasing verification cost and potentially making speculative decoding slower than autoregressive decoding. We present D-Cut, an adaptive pruning method that selects draft tokens jointly across the batch and concentrates the verification budget on tokens most likely to be accepted. D-Cut is motivated by two observations. First, acceptance lengths vary considerably across concurrent requests; D-Cut therefore performs cross-request pruning, allocating the verification budget adaptively according to draft confidence. Second, verification cost depends strongly on the deployment environment, including GPU architecture and parallelism strategy; D-Cut incorporates a runtime cost model to adapt its pruning depth to the target environment. Experiments on dense and mixture-of-experts (MoE) models show that, under high concurrency, D-Cut improves the average speedup from \(1.26\times\) to \(1.65\times\), restores acceleration in dense-model configurations where long-draft baselines are slower than autoregressive decoding, and achieves up to \(3.0\times\) speedup over autoregressive decoding on MoE models.
Figures
Reference graph
Works this paper leans on
-
[1]
2026 , eprint=
DFlash: Block Diffusion for Flash Speculative Decoding , author=. 2026 , eprint=
2026
-
[2]
2026 , howpublished =
GPT-5.5 System Card , author =. 2026 , howpublished =
2026
-
[3]
2025 , month = nov, howpublished =
Gemini 3 Pro Model Card , author =. 2025 , month = nov, howpublished =
2025
-
[4]
2025 , month = nov, howpublished =
System Card: Claude Opus 4.5 , author =. 2025 , month = nov, howpublished =
2025
-
[5]
Qwen3.5: Accelerating Productivity with Native Multimodal Agents , url =
-
[6]
2026 , url=
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence , author=. 2026 , url=
2026
-
[7]
2025 , eprint=
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models , author=. 2025 , eprint=
2025
-
[8]
2024 , eprint=
Better & Faster Large Language Models via Multi-token Prediction , author=. 2024 , eprint=
2024
-
[9]
2022 , eprint=
Efficiently Scaling Transformer Inference , author=. 2022 , eprint=
2022
-
[10]
Proceedings of the 40th International Conference on Machine Learning , pages =
Fast Inference from Transformers via Speculative Decoding , author =. Proceedings of the 40th International Conference on Machine Learning , pages =. 2023 , editor =
2023
-
[11]
2023 , eprint=
Accelerating Large Language Model Decoding with Speculative Sampling , author=. 2023 , eprint=
2023
-
[12]
2025 , url=
Tianyu Liu and Yun Li and Qitan Lv and Kai Liu and Jianchen Zhu and Winston Hu and Xiao Sun , booktitle=. 2025 , url=
2025
-
[13]
Miao, Xupeng and Oliaro, Gabriele and Zhang, Zhihao and Cheng, Xinhao and Wang, Zeyu and Zhang, Zhengxin and Wong, Rae Ying Yee and Zhu, Alan and Yang, Lijie and Shi, Xiaoxiang and Shi, Chunan and Chen, Zhuoming and Arfeen, Daiyaan and Abhyankar, Reyna and Jia, Zhihao , title =. Proceedings of the 29th ACM International Conference on Architectural Support...
arXiv 2024
-
[14]
and Chen, Deming and Dao, Tri , title =
Cai, Tianle and Li, Yuhong and Geng, Zhengyang and Peng, Hongwu and Lee, Jason D. and Chen, Deming and Dao, Tri , title =. Proceedings of the 41st International Conference on Machine Learning , articleno =. 2024 , publisher =
2024
-
[15]
2025 , eprint=
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty , author=. 2025 , eprint=
2025
-
[16]
EAGLE -2: Faster Inference of Language Models with Dynamic Draft Trees
Li, Yuhui and Wei, Fangyun and Zhang, Chao and Zhang, Hongyang. EAGLE -2: Faster Inference of Language Models with Dynamic Draft Trees. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.422
-
[17]
2025 , eprint=
EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test , author=. 2025 , eprint=
2025
-
[18]
2024 , eprint=
OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure , author=. 2024 , eprint=
2024
-
[19]
Proceedings of the 41st International Conference on Machine Learning , articleno =
Fu, Yichao and Bailis, Peter and Stoica, Ion and Zhang, Hao , title =. Proceedings of the 41st International Conference on Machine Learning , articleno =. 2024 , publisher =
2024
-
[20]
2025 , eprint=
DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference , author=. 2025 , eprint=
2025
-
[21]
Gonzalez and Ion Stoica , booktitle=
Lianmin Zheng and Wei-Lin Chiang and Ying Sheng and Siyuan Zhuang and Zhanghao Wu and Yonghao Zhuang and Zi Lin and Zhuohan Li and Dacheng Li and Eric Xing and Hao Zhang and Joseph E. Gonzalez and Ion Stoica , booktitle=. Judging. 2023 , url=
2023
-
[22]
Proceedings of the 41st International Conference on Machine Learning , articleno =
Du, Cunxiao and Jiang, Jing and Yuanchen, Xu and Wu, Jiawei and Yu, Sicheng and Li, Yongqi and Li, Shenggui and Xu, Kai and Nie, Liqiang and Tu, Zhaopeng and You, Yang , title =. Proceedings of the 41st International Conference on Machine Learning , articleno =. 2024 , publisher =
2024
-
[23]
Turning Up the Heat: Min-p Sampling for Creative and Coherent
Nguyen Nhat Minh and Andrew Baker and Clement Neo and Allen G Roush and Andreas Kirsch and Ravid Shwartz-Ziv , booktitle=. Turning Up the Heat: Min-p Sampling for Creative and Coherent. 2025 , url=
2025
-
[24]
2024 , eprint=
The Llama 3 Herd of Models , author=. 2024 , eprint=
2024
-
[25]
2023 , eprint=
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations , author=. 2023 , eprint=
2023
-
[26]
2021 , eprint=
Evaluating Large Language Models Trained on Code , author=. 2021 , eprint=
2021
-
[27]
2021 , eprint=
Training Verifiers to Solve Math Word Problems , author=. 2021 , eprint=
2021
-
[28]
Abstractive Text Summarization using Sequence-to-sequence RNN s and Beyond
Nallapati, Ramesh and Zhou, Bowen and dos Santos, Cicero and Gu l c ehre, C a g lar and Xiang, Bing. Abstractive Text Summarization using Sequence-to-sequence RNN s and Beyond. Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning. 2016. doi:10.18653/v1/K16-1028
-
[29]
and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav , title =
Kwiatkowski, Tom and Palomaki, Jennimaria and Redfield, Olivia and Collins, Michael and Parikh, Ankur and Alberti, Chris and Epstein, Danielle and Polosukhin, Illia and Devlin, Jacob and Lee, Kenton and Toutanova, Kristina and Jones, Llion and Kelcey, Matthew and Chang, Ming-Wei and Dai, Andrew M. and Uszkoreit, Jakob and Le, Quoc and Petrov, Slav , title...
-
[31]
REST : Retrieval-Based Speculative Decoding
He, Zhenyu and Zhong, Zexuan and Cai, Tianle and Lee, Jason and He, Di. REST : Retrieval-Based Speculative Decoding. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. doi:10.18653/v1/2024.naacl-long.88
-
[32]
2023 , month =
Prompt Lookup Decoding , author =. 2023 , month =
2023
-
[33]
The Thirteenth International Conference on Learning Representations , year=
Block Verification Accelerates Speculative Decoding , author=. The Thirteenth International Conference on Learning Representations , year=
-
[34]
2025 , eprint=
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths , author=. 2025 , eprint=
2025
-
[35]
2024 , eprint=
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures , author=. 2024 , eprint=
2024
-
[36]
2024 , eprint=
DySpec: Faster Speculative Decoding with Dynamic Token Tree Structure , author=. 2024 , eprint=
2024
-
[37]
2025 , eprint=
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding , author=. 2025 , eprint=
2025
-
[38]
Gao, Xiangxiang and Xie, Weisheng and Xiang, Yiwei and Ji, Feng , title =. Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Symposium on Educational Advances in Artificial Intelligence , articleno =. 2025 , isbn =. doi:10.1609/aaai.v...
-
[39]
2025 , eprint=
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding , author=. 2025 , eprint=
2025
-
[40]
2024 , eprint=
Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding , author=. 2024 , eprint=
2024
-
[41]
2024 , eprint=
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge , author=. 2024 , eprint=
2024
-
[42]
2024 , eprint=
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism , author=. 2024 , eprint=
2024
-
[43]
2024 , eprint=
Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling , author=. 2024 , eprint=
2024
-
[44]
2024 , eprint=
SAM Decoding: Speculative Decoding via Suffix Automaton , author=. 2024 , eprint=
2024
-
[45]
2025 , eprint=
Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment , author=. 2025 , eprint=
2025
-
[46]
2025 , eprint=
CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter , author=. 2025 , eprint=
2025
-
[47]
The Thirteenth International Conference on Learning Representations , year=
Learning Harmonized Representations for Speculative Sampling , author=. The Thirteenth International Conference on Learning Representations , year=
-
[48]
2024 , eprint=
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models , author=. 2024 , eprint=
2024
-
[49]
2024 , eprint=
Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens , author=. 2024 , eprint=
2024
-
[50]
L ogit S pec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation
Liu, Tianyu and Lv, Qitan and Li, Hao and Gao, Xing and Sun, Xiao and Sun, Xiaoyan. L ogit S pec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation. Findings of the A ssociation for C omputational L inguistics: ACL 2026. 2026. doi:10.18653/v1/2026.findings-acl.1655
-
[51]
Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and Rémi Louf and Morgan Funtowicz and Joe Davison and Sam Shleifer and Patrick von Platen and Clara Ma and Yacine Jernite and Julien Plu and Canwen Xu and Teven Le Scao and Sylvain Gugger and Mariama Drame and Quentin L...
2020
-
[52]
2019 , eprint=
PyTorch: An Imperative Style, High-Performance Deep Learning Library , author=. 2019 , eprint=
2019
-
[53]
, title=
NVIDIA and Vingelmann, Péter and Fitzek, Frank H.P. , title=. 2020 , url=
2020
-
[54]
International Conference on Learning Representations , year=
The Curious Case of Neural Text Degeneration , author=. International Conference on Learning Representations , year=
-
[55]
Constrained Decoding with Speculative Lookaheads
Nakshatri, Nishanth Sridhar and Roy, Shamik and Das, Rajarshi and Chaidaroon, Suthee and Boytsov, Leonid and Gangadharaiah, Rashmi. Constrained Decoding with Speculative Lookaheads. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)...
-
[56]
and Barrett, Clark and Sheng, Ying
Zheng, Lianmin and Yin, Liangsheng and Xie, Zhiqiang and Sun, Chuyue and Huang, Jeff and Yu, Cody Hao and Cao, Shiyi and Kozyrakis, Christos and Stoica, Ion and Gonzalez, Joseph E. and Barrett, Clark and Sheng, Ying. SGLang : Efficient Execution of Structured Language Model Programs. Advances in Neural Information Processing Systems. 2024. doi:10.52202/07...
-
[57]
FlashAttention-3 : Fast and Accurate Attention with Asynchrony and Low-precision
Shah, Jay and Bikshandi, Ganesh and Zhang, Ying and Thakkar, Vijay and Ramani, Pradeep and Dao, Tri. FlashAttention-3 : Fast and Accurate Attention with Asynchrony and Low-precision. 2024. arXiv:2407.08608
Pith/arXiv arXiv 2024
-
[58]
and Kaiser, Lukasz and Polosukhin, Illia , title =
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser, Lukasz and Polosukhin, Illia , title =. Advances in Neural Information Processing Systems , volume =. 2017 , url =
2017
-
[59]
Brown, Tom B. and Mann, Benjamin and Ryder, Nick and Subbiah, Melanie and Kaplan, Jared and Dhariwal, Prafulla and Neelakantan, Arvind and Shyam, Pranav and Sastry, Girish and Askell, Amanda and Agarwal, Sandhini and Herbert-Voss, Ariel and Krueger, Gretchen and Henighan, Tom and Child, Rewon and Ramesh, Aditya and Ziegler, Daniel M. and Wu, Jeffrey and W...
2020
-
[60]
arXiv preprint arXiv:2401.07851 , year =
Xia, Heming and Yang, Zhe and Dong, Qingxiu and Wang, Peiyi and Li, Yongqi and Ge, Tao and Liu, Tianyu and Li, Wenjie and Sui, Zhifang , title =. arXiv preprint arXiv:2401.07851 , year =
-
[61]
2024 , url =
DeepSeek-V3 Technical Report , journal =. 2024 , url =
2024
-
[62]
arXiv preprint arXiv:2602.06036 , year =
Chen, Jian and Liang, Yesheng and Liu, Zhijian , title =. arXiv preprint arXiv:2602.06036 , year =
-
[63]
2026 , eprint=
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification , author=. 2026 , eprint=
2026
-
[64]
Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding , url =
Zhang, Jun and Wang, Jue and Li, Huan and Shou, Lidan and Chen, Ke and Chen, Gang and Mehrotra, Sharad , year =. Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding , url =. doi:10.18653/v1/2024.acl-long.607 , booktitle =
-
[65]
2026 , eprint=
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding , author=. 2026 , eprint=
2026
-
[66]
2025 , eprint=
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration , author=. 2025 , eprint=
2025
-
[67]
KNN - SSD : Enabling Dynamic Self-Speculative Decoding via Nearest Neighbor Layer Set Optimization
Song, Mingbo and Xia, Heming and Zhang, Jun and Leong, Chak Tou and Xu, Qiancheng and Li, Wenjie and Li, Sujian. KNN - SSD : Enabling Dynamic Self-Speculative Decoding via Nearest Neighbor Layer Set Optimization. Findings of the Association for Computational Linguistics : EACL 2026. 2026. doi:10.18653/v1/2026.findings-eacl.31
-
[68]
2026 , eprint=
TALON: Confidence-Aware Speculative Decoding with Adaptive Token Trees , author=. 2026 , eprint=
2026
-
[69]
2026 , eprint=
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism , author=. 2026 , eprint=
2026
-
[70]
2026 , eprint=
Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism , author=. 2026 , eprint=
2026
-
[71]
2025 , publisher=
SpecForge: Train speculative decoding models effortlessly , author=. 2025 , publisher=
2025
-
[72]
2023 , howpublished =
2023
-
[73]
2026 , eprint=
P-EAGLE: Parallel-Drafting EAGLE with Scalable Training , author=. 2026 , eprint=
2026
-
[74]
2026 , eprint=
Speculative Decoding: Performance or Illusion? , author=. 2026 , eprint=
2026
-
[75]
2026 , eprint=
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios , author=. 2026 , eprint=
2026
-
[76]
2026 , eprint=
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding , author=. 2026 , eprint=
2026
-
[77]
2021 , eprint=
Program Synthesis with Large Language Models , author=. 2021 , eprint=
2021
-
[78]
Let s Verify Step by Step , url =
Lightman, Hunter and Kosaraju, Vineet and Burda, Yuri and Edwards, Harrison and Baker, Bowen and Lee, Teddy and Leike, Jan and Schulman, John and Sutskever, Ilya and Cobbe, Karl , booktitle =. Let s Verify Step by Step , url =
-
[79]
2025 , eprint=
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding , author=. 2025 , eprint=
2025
-
[80]
arXiv preprint arXiv:2606.02091 , year=
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding , author=. arXiv preprint arXiv:2606.02091 , year=
-
[81]
arXiv preprint arXiv:2605.29707 , year=
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding , author=. arXiv preprint arXiv:2605.29707 , year=
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.