REVIEW 4 major objections 5 minor 35 references
Speculative decoding remains lossless but no single drafter wins across workloads; this paper shows that specializing the drafter's training data, architecture, and serving-time verification budget jointly yields 1.98–2.40× throughput over
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 01:17 UTC pith:VDCYKN25
load-bearing objection Workload-specialized drafting plus batch-level verification budgeting is a credible engineering win; the one real uncertainty is whether the D-cut confidence proxy travels to new workloads. the 4 major comments →
AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that no single drafting structure is optimal across real-world workloads, so the drafter, its training distribution, and the verification depth must be co-specialized. Training-level specialization pairs an MTP drafter with high-entropy conversational data and a block-diffusion drafter with code and mathematics data; architecture-level specialization introduces DFly, whose hybrid target-conditioning backbone combines a shared nonlinear context with layer-specific target views, and whose hidden-correction head conditions each parallel draft prediction on the previously sampled draft token; inference-level specialization introduces D-cut, which estimates prefix survival as a p
What carries the argument
Three mechanisms carry the argument. Training-Time Test (TTT) for MTP: the shared MTP block is autoregressively unrolled during training on its own predictions, with depth-shifted teacher targets, so later draft positions learn to recover from realistic upstream errors; an acceptance-aligned objective (cold-start adaptive KL–TV blend, then end-to-end total-variation loss) ties training directly to speculative acceptance. DFly: a parallel block-diffusion drafter whose hybrid target-conditioning backbone adds layer-specific weighted target views to a shared transformed context, and whose hidden-correction head applies a lightweight sequential correction to each draft position conditioned on th
Load-bearing premise
The load-bearing premise is that a request's probability of accepting the k-th draft token is well approximated by the product of the drafter's own token-level confidence scores; if that confidence is miscalibrated on a new workload, D-cut will prune tokens the target would accept and keep tokens it would reject.
What would settle it
Train a drafter on one domain (e.g., chat) and deploy D-cut on a code-heavy traffic mix where the drafter's softmax confidence is systematically overconfident; if measured per-position acceptance drops much more than the roughly 1.5% observed in the paper, or throughput gains invert at high concurrency, the confidence-as-survival assumption is violated. A direct offline check: compare D-cut's estimated prefix-survival s_{i,k} with empirically measured acceptance rates at each draft depth across domains; divergence would predict where pruning fails.
If this is right
- MTP post-training with TTT improves acceptance mainly at draft positions 2 and 3, converting previously low-value suffixes into accepted progress while preserving first-position acceptance.
- DFly outperforms both fully parallel and semi-autoregressive block-drafting baselines on math and code benchmarks, and on Hy3-A21B raises average accepted length by roughly 30% over the block-parallel baseline.
- Domain-specific data expansion strengthens code and math acceptance with only a slight chat degradation, confirming that workload specialization, not a universal gain, drives the improvement.
- D-cut preserves acceptance within about 1.5% of the unpruned baseline while raising throughput at high concurrency (up to +15.7% at concurrency 64), extending the useful concurrency range before verification-bound saturation.
- Training and evaluation must be matched in thinking mode: a no-thinking drafter loses substantial acceptance when evaluated under high-thinking prompts, so separate drafters are needed per mode.
- The default DFly configuration, 5 draft layers with block size 8, matches the acceptance of a deeper 7-layer drafter at lower proposer-side latency, indicating a favorable efficiency–quality trade-off.
Where Pith is reading between the lines
- If the confidence-surrogate assumption holds across workloads, D-cut's top-K pruning is a general principle: any batched drafter exposing per-token confidence could use the same budget rule, which suggests extending the mechanism to tree-based or retrieval-based drafters, not only block-parallel ones. (editorial inference)
- The workload-specialization thesis implies that a single serving system should maintain multiple drafters and route requests by predicted output structure (e.g., code vs. chat); the paper evaluates separate drafters but does not itself design such a router, making an online request classifier a natural testable next step. (editorial inference)
- The two-stage loss schedule (cold-start LK then end-to-end TV) suggests that directly optimizing acceptance is unstable from poor initializations; the schedule may transfer to other draft–target pairs, and testing it on smaller models would clarify whether the instability is intrinsic or specific to block diffusion. (editorial inference)
- Because D-cut is evaluated on live production traffic, its gains may be sensitive to traffic mix; a testable extension is measuring whether the throughput margin persists on workloads with unusually short or unusually long continuations. (editorial inference)
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. AngelSpec proposes a unified training and serving framework for speculative decoding that combines two complementary drafting paradigms: an autoregressive MTP drafter specialized for high-entropy conversation and a block-parallel diffusion drafter (DFly) specialized for code and mathematics. The training component introduces TTT unrolling, acceptance-aligned losses, and domain-specialized data. The inference component, D-cut, treats target verification as a shared batch resource and adaptively prunes draft depth across concurrent requests using a confidence-based survival estimate and a profiled runtime cost model. Experiments are reported on Qwen3-8B and the proprietary Hy3-A21B model, with throughput sweeps on benchmark datasets and on claimed Hunyuan production traffic. The headline results are that DFly increases average accepted length by roughly 30% over DFlash on Hy3-A21B and delivers 1.98–2.40× speedup over autoregressive decoding, and that D-cut further improves throughput at high concurrency while reducing mean acceptance length by only about 1.5%.
Significance. If the reported results hold, this is a substantial practical contribution. The paper identifies workload heterogeneity as a first-order design axis, provides a coherent architecture (DFly) with a clear ablation story, and offers an open-source training framework. The internal arithmetic of the acceptance metrics, the e2e-TV objective, the break-even analysis, and the reported speedup ranges is consistent, and the ablation tables are carefully controlled. The main value is in showing that drafter structure and training data can be co-specialized by domain, and that verification depth can be adapted at serving time. The claims are, however, partially built on proprietary models and private production traffic, and the D-cut component relies on a confidence-to-survival proxy that is not validated in this manuscript. Those issues limit the transferability and verifiability of the strongest deployment claims.
major comments (4)
- [§4.2, Eq. (20)] The central D-cut mechanism substitutes the drafter's token-level confidence product s_{i,k} = prod c_{i,t} for the prefix-survival probability Pr(L_i >= k) in Eq. (19). This substitution is not justified by the training objectives in §2.3/§3.3: aligning q with p in TV/KL does not make q(x) a calibrated estimator of the target acceptance probability a(x)=min(1,p(x)/q(x)) (or of argmax-match probability under deterministic verification). A well-calibrated drafter can be arbitrarily miscalibrated as an acceptance predictor. Since Eq. (24) selects the verification ratio based on rankings of s_{i,k}, any systematic miscalibration can either retain tokens the target rejects or prune tokens it would accept. The live-traffic result that acceptance drops only ~1.5% is a single empirical point; no calibration plot, no comparison of estimated vs. empirical survival probabilities, and no sensitivit
- [§4.2 and §7] The manuscript states that 'Further algorithmic and correctness details are provided in the companion paper (Liu et al., 2026d)' and the related-work section cites D-cut as if it were an external method (Liu et al., 2026e). D-cut is presented in this paper as one of its three main contributions, yet the present text does not contain enough detail to verify correctness claims such as the preservation of the target output distribution under pruning, the handling of unequal prefixes in the packed verification batch, or the tie-breaking behavior of top-K selection. The correctness of the method is a load-bearing part of the paper, so the relevant proofs or full algorithmic descriptions should be included here rather than deferred to a companion paper by overlapping authors.
- [§4.4, Table 7, Figure 4] The throughput comparisons report only point estimates: each cell in Table 7 uses 3×120s windows, yet no standard deviations, confidence intervals, or per-window values are given. The D-cut gains in Figure 4 are as small as +3.0% (c48) and -0.9% at one point, so without variance information it is not possible to determine whether the 'highest throughput at every tested concurrency' claim is statistically meaningful. The authors should report variability across windows or repetitions, and ideally a paired comparison for the D-cut vs. DFly difference, since the same request mix and concurrency are being compared.
- [§4.4, Figure 4] The live-traffic benchmark is described only as 'Hunyuan production traffic' against a proprietary Hy3-295B-A21B model. No details are given about the request mix, prompt-length distribution, generation-length distribution, arrival pattern, or the number of requests replayed. This makes the headline D-cut results non-reproducible and complicates comparison with the benchmark sweep. At minimum, the paper should describe the traffic characteristics and, if possible, release an anonymized or synthetic version of the workload. This is not a request for a different experimental design, but the current level of detail is insufficient for a reader to assess whether the reported high-concurrency benefit transfers to other serving conditions.
minor comments (5)
- [§8 / Abstract] The conclusion refers to 'Hy3-A20B' while the abstract and body use 'Hy3-A21B'. The same paragraph also appears twice verbatim; one copy should be removed.
- [§5.1] The text says 'Hy3-DFlash-L3-B5 with four draft tokens' and later explains that block size 5 corresponds to four draft tokens. This is correct but confusingly worded; consider consistently saying 'B5 gives four draft tokens' in the first mention.
- [§7] The sentence 'DSpark (Cheng et al., 2026) and D-cut (Liu et al., 2026e) instead treats verification tokens...' has a subject-verb agreement error; it should be 'treat'.
- [§4.3] The statement that the comparison 'understates D-cut' because D-cut uses piecewise CUDA graph capture while DFly uses full-and-piecewise capture is a useful caveat, but the direction should be made more explicit: D-cut is evaluated with a known implementation overhead, so an improved implementation would likely widen the observed gap.
- [§2.5] The metric definition says 'Avg = (p1 + p2 + p3)/3', which is an average of cumulative acceptance rates. This is fine, but it is easy to misread as an average of per-position marginal acceptance rates; a short clarification that p_i denotes Pr(L >= i) would help.
Circularity Check
No significant circularity; central claims rest on measured end-to-end comparisons, not on fitted inputs or self-citation chains.
full rationale
The headline results — DFly raising accepted length by ~30% and achieving 1.98–2.40× over AR and 10.5–11.8% over DFlash — are measured end-to-end in Table 7 and Figure 4 against external baselines (AR, MTP-3, DFlash) on Hy3-295B-A21B. No parameter is fitted to the final throughput number and then renamed a prediction. D-cut's Eq. (20) substitutes the drafter's token-confidence product s_{i,k} for the unknown prefix-survival probability Pr(L_i>=k); the text explicitly says the true probabilities are unknown and that D-cut 'estimates them,' so this is an acknowledged modeling assumption rather than a hidden reduction. The subsequent throughput gain is measured empirically, and the risk that the confidence proxy is miscalibrated on new workloads is a validity/correctness concern, not circularity. Self-citations exist — D-cut is credited to Liu et al. 2026d/e and the DFlare backbone to Zhang et al. 2026a, both with overlapping authors — but they are not load-bearing: the selection equations, loss equations, ablations, and deployment measurements are all present in this paper, and the companion citation provides 'further algorithmic and correctness details' rather than the only support. Section 5's break-even model is explicitly labeled 'a device-side prediction and must ultimately be validated using end-to-end engine measurements,' so it is not used as circular proof. No uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in solely via self-citation. The derivation chain is self-contained with respect to the empirical claims made.
Axiom & Free-Parameter Ledger
free parameters (5)
- MTP draft depth D=3 and DFly block size B=8 =
D=3 / B=8
- D-cut ratio set R =
{0.25,0.50,0.75,1.00}
- D-PACE smoothing coefficient rho =
not reported
- LK cold-start to e2e TV switch =
not specified
- Top-K support K for distribution matching =
10000
axioms (5)
- standard math Rejection sampling preserves the target distribution exactly
- domain assumption Target hidden states are available to the drafter at inference
- domain assumption Code/math continuations are lower entropy than chat, making block-diffusion preferable for them
- ad hoc to paper Drafter confidence products estimate target-acceptance survival probabilities
- domain assumption Open-PerfectBlend benchmarks and Hunyuan traffic replay represent real deployment workloads
read the original abstract
Speculative decoding accelerates large language model inference without changing the target distribution, but no single drafting structure performs best across real-world workloads. Autoregressive multi-token prediction (MTP) is a lightweight, stable proposal mechanism, whereas block-parallel diffusion amortizes drafting latency over much longer candidate sequences; the better choice depends strongly on the output distribution. We present AngelSpec, a unified training framework for MTP and block-parallel speculative decoding that addresses this heterogeneity at three levels. At the training level, rather than fitting one universal drafter to a uniform data mixture, we co-specialize structure and data: the MTP drafter is trained on diverse conversational data for high-entropy open-ended chat, and the block-diffusion drafter on code and mathematics data for longer predictable continuations. At the architecture level, we propose DFly, a block-diffusion framework combining a hybrid target-conditioning backbone with a predecessor-conditioned autoregressive head, improving target-feature utilization and intra-block dependency modeling while keeping generation parallel. At the inference level, both acceptance length and verification cost vary with domain, request, online load, and hardware, so DFly treats verification as a shared batch-level resource: it reallocates compute toward high-confidence prefixes across requests and combines expected utility with a profiled cost model to adapt verification depth online. Across the Hy3 series, DFly raises the average accepted length on Hy3-A21B by roughly 30% and attains the highest average throughput at every tested concurrency from 4 to 64, a 1.98-2.40x speedup over autoregressive decoding and 10.5-11.8% higher throughput than DFlash. We release AngelSpec to support training and extending these methods.
Figures
Reference graph
Works this paper leans on
-
[1]
Opencodeinstruct: A large-scale instruction tuning dataset for code llms
Wasi Uddin Ahmad, Aleksander Ficek, Mehrzad Samadi, Jocelyn Huang, Vahid Noroozi, Somshubra Majum- dar, and Boris Ginsburg. Opencodeinstruct: A large-scale instruction tuning dataset for code llms. 2025a. URLhttps://arxiv.org/abs/2504.04030. Wasi Uddin Ahmad, Sean Narenthiran, Somshubra Majumdar, Aleksander Ficek, Siddhartha Jain, Jocelyn Huang, Vahid Nor...
-
[3]
URL https: //arxiv.org/abs/2504.18583. Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models,
-
[7]
URLhttps://arxiv.org/abs/2107.03374. Zhuoming Chen, Avner May, Ruslan Svirschevski, Yuhsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen. Sequoia: Scalable, robust, and hardware-aware speculative decoding,
-
[8]
URL https://arxiv.org/abs/ 2402.12374. Xin Cheng, Xingkai Yu, Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, Qinyu Chen, Chengqi Deng, Jiping Yu, Damai Dai, Zhengyan Zhang, Yixuan Wei, Yixuan Tan, Wenkai Yang, Runxin Xu, Yu Wu, Zhean Xu, Xuanyu Wang, Muyang Chen, Rui Tian, Xiao Bi, Zhewen Hao, Shaoyuan C...
-
[9]
URLhttps://arxiv.org/abs/2607.05147. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems,
-
[10]
URLhttps://arxiv.org/abs/2110.14168. DeepSeek-AI. Deepseek-v4: Towards highly efficient million-token context intelligence,
-
[12]
URL https: //arxiv.org/abs/2103.03874. Xinyi Hu, Yuhao Shen, Baolin Zhang, Hengxin Zhang, Jun Dai, Shuang Ge, Lei Chen, Yue Li, and Mingcheng Wan. Echo: Elastic speculative decoding with sparse gating for high-concurrency scenarios,
-
[13]
Jianuo Huang, Yaojie Zhang, Qituan Zhang, Hao Lin, Hanlin Xu, and Linfeng Zhang
URL https://arxiv.org/abs/2604.09603. Jianuo Huang, Yaojie Zhang, Qituan Zhang, Hao Lin, Hanlin Xu, and Linfeng Zhang. Domino: Decoupling causal modeling from autoregressive drafting in speculative decoding.arXiv preprint arXiv:2605.29707,
-
[14]
URLhttps://arxiv.org/abs/2405.19715. Mude Hui, Xin Huang, Jaime Campos Salas, Yue Sun, Nathan Pemberton, Xiang Song, Ashish Khetan, and George Karypis. P-eagle: Parallel-drafting eagle with scalable training,
-
[15]
URL https://arxiv.org/ abs/2602.01469. Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar- Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code,
-
[16]
Yaniv Leviathan, Matan Kalman, and Yossi Matias
URLhttps://arxiv.org/abs/2403.07974. Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (eds.),Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of ...
-
[17]
URL https: //proceedings.mlr.press/v202/leviathan23a.html. Yucheng Li, Huiqiang Jiang, Yang Xu, Jianxin Yang, Yi Zhang, Yizhong Cao, Yuhao Shen, Fan Zhou, Rui Men, Jianwei Zhang, et al. Breaking entropy bounds: Accelerating rl training via mtp with rejection sampling.arXiv preprint arXiv:2606.12370,
-
[18]
EAGLE-2: Faster inference of language models with dynamic draft trees
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE-2: Faster inference of language models with dynamic draft trees. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.),Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 7421–7432, Miami, Florida, USA, November
2024
-
[19]
doi: 10.18653/v1/2024.emnlp-main.422
Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.422. URLhttps://aclanthology.org/2024.emnlp-main.422/. Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. Eagle: Speculative sampling requires rethinking feature uncertainty, 2025a. URLhttps://arxiv.org/abs/2401.15077. Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. Eag...
Pith/arXiv arXiv 2024
-
[20]
Blog: https://pytorch.org/blog/ torchspec-speculative-decoding-training-at-scale/. Fuliang Liu, Xue Li, Ketai Zhao, Yinxi Gao, Ziyan Zhou, Zhonghui Zhang, Zhibin Wang, Wanchun Dou, Sheng Zhong, and Chen Tian. Dart: Diffusion-inspired speculative decoding for fast llm inference.arXiv preprint arXiv:2601.19278, 2026a. URLhttps://arxiv.org/abs/2601.19278. Ti...
-
[21]
Tianyu Liu, Qitan Lv, Hao Li, Xing Gao, Xiao Sun, and Xiaoyan Sun
URLhttps://openreview.net/forum?id=QOXrVMiHGK. Tianyu Liu, Qitan Lv, Hao Li, Xing Gao, Xiao Sun, and Xiaoyan Sun. LogitSpec: Accelerating retrieval-based speculative decoding via next next token speculation. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David Jurgens (eds.),Findings of the Association for Computational Linguistics: ACL 2026, pp....
arXiv 2026
-
[22]
Xiaoxuan Liu, Jiaxiang Yu, Jongseok Park, Ion Stoica, and Alvin Cheung
URLhttps://arxiv.org/abs/2310.07177. Xiaoxuan Liu, Jiaxiang Yu, Jongseok Park, Ion Stoica, and Alvin Cheung. Speculative decoding: Performance or illusion?, 2026f. URLhttps://arxiv.org/abs/2601.11580. Jonathan Mamou, Oren Pereg, Daniel Korat, Moshe Berchansky, Nadav Timor, Moshe Wasserblat, and Roy Schwartz. Dynamic speculation lookahead accelerates specu...
-
[23]
URLhttps://arxiv.org/abs/2405.04304. Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. Specinfer: Accelerating large language model serving with tree-based speculative inference and verifica...
-
[24]
Association for Computing Machinery. ISBN 9798400703867. doi: 10.1145/3620666.3651335. URLhttps://doi.org/10.1145/3620666.3651335. Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Heyi Tang, Feng Ren, Teng Ma, Shangming Cai, Yineng Zhang, Mingxing Zhang, et al. Mooncake: A kvcache-centric disaggregated architecture for llm serving. ACM Transactions on Storage,
-
[25]
URLhttps://arxiv.org/abs/2606.03819. Alexander Samarin, Sergei Krutikov, Anton Shevtsov, Sergei Skvortsov, Filipp Fisin, and Alexander Golubev. Lk losses: Direct acceptance rate optimization for speculative decoding,
-
[26]
URL https://arxiv.org/ abs/2602.23881. Yuhao Shen, Tianyu Liu, Xinyi Hu, Quan Kong, Baolin Zhang, Jun Dai, Jun Zhang, Shuang Ge, Lei Chen, Yue Li, et al. Draft less, retrieve more: Hybrid tree construction for speculative decoding.arXiv preprint arXiv:2605.20104, 2026a. Yuhao Shen, Junyi Shen, Quan Kong, Tianyu Liu, Yao Lu, and Cong Wang. Specbranch: Spec...
Pith/arXiv arXiv 2025
-
[27]
Association for Computational Linguistics. ISBN 979-8-89176-335-7. doi: 10.18653/v1/2025.findings-emnlp.1326. URL https://aclanthology.org/2025. findings-emnlp.1326/. Qwen Team. Qwen3 technical report,
-
[28]
Jikai Wang, Yi Su, Juntao Li, Qingrong Xia, Zi Ye, Xinyu Duan, Zhefeng Wang, and Min Zhang
URLhttps://arxiv.org/abs/2505.09388. Jikai Wang, Yi Su, Juntao Li, Qingrong Xia, Zi Ye, Xinyu Duan, Zhefeng Wang, and Min Zhang. Opt-tree: Speculative decoding with adaptive draft tree structure,
-
[29]
URL https://arxiv.org/abs/2406.17276. – 25 – Tencent Hy Tianyu Wu, Yu Yao, Zhenting Qi, Han Zheng, Zhuohan Wang, Haoran Ma, Lawrence Liao, Himabindu Lakkaraju, Ju Li, and Yilun Du. D-pace: Dynamic position-aware cross-entropy for parallel speculative drafting.arXiv preprint arXiv:2605.18810,
-
[31]
Yunfan Xiong, Ruoyu Zhang, Yanzeng Li, Tianhao Wu, and Lei Zou
URLhttps://arxiv.org/abs/2401.07851. Yunfan Xiong, Ruoyu Zhang, Yanzeng Li, Tianhao Wu, and Lei Zou. Dyspec: Faster speculative decoding with dynamic token tree structure,
-
[32]
URLhttps://arxiv.org/abs/2410.11744. Tengyu Xu, Eryk Helenowski, Karthik Abinav Sankararaman, Di Jin, Kaiyan Peng, Eric Han, Shaoliang Nie, Chen Zhu, Hejia Zhang, Wenxuan Zhou, et al. The perfect blend: Redefining rlhf with mixture of judges. arXiv preprint arXiv:2409.20370,
-
[33]
Glm-5: from vibe coding to agentic engineering.arXiv preprint arXiv:2602.15763,
Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, et al. Glm-5: from vibe coding to agentic engineering.arXiv preprint arXiv:2602.15763,
-
[34]
Dflare: Scaling up draft capacity for block diffusion speculative decoding
Jiebin Zhang, Zhenghan Yu, Song Liu, Eugene J Yu, Zheng Li, Dawei Zhu, Jiangshan Duo, Weimin Xiong, Yifan Song, Guanghua Yu, et al. Dflare: Scaling up draft capacity for block diffusion speculative decoding. arXiv preprint arXiv:2606.02091, 2026a. Jiebin Zhang, Zhenghan Yu, Liang Wang, Nan Yang, Eugene J Yu, Zheng Li, Yifan Song, Dawei Zhu, Xingxing Zhang...
-
[35]
URLhttps://arxiv.org/abs/2306.05685. Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, and Rishabh Agarwal. Distillspec: Improving speculative decoding via knowledge distillation,
-
[36]
URLhttps://arxiv.org/abs/2310.08461. – 26 –
-
[2021]
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D
URLhttps://arxiv.org/abs/2108.07732. Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao. Medusa: Simple llm inference acceleration framework with multiple decoding heads. InProceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org,
-
[2023]
Jian Chen, Yesheng Liang, and Zhijian Liu
URL https://arxiv.org/ abs/2302.01318. Jian Chen, Yesheng Liang, and Zhijian Liu. Dflash: Block diffusion for flash speculative decoding,
-
[2024]
URL https://arxiv.org/abs/2404.19737. Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset,
-
[2025]
Zihao An, Huajun Bai, Ziqiong Liu, Dong Li, and Emad Barsoum
URLhttps://arxiv.org/abs/2502.17387. Zihao An, Huajun Bai, Ziqiong Liu, Dong Li, and Emad Barsoum. Pard: Accelerating llm inference with low-cost parallel draft model adaptation.arXiv preprint arXiv:2504.18583,
-
[2026]
URLhttps://arxiv.org/abs/2602.06036. Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Po...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.