REVIEW 5 minor 32 references
Two fixes to Speculative Jacobi Decoding raise lossless text-to-image speedup from ~2× to 3.8×.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 22:29 UTC pith:WFZGETRG
load-bearing objection Clean 3.8 imes lossless wall-clock speedup for AR T2I by fixing SJD's short-acceptance bottleneck with two simple, proved-correct tricks; solid systems paper.
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
SJD-PAC demonstrates that the skewed acceptance-length distribution of Speculative Jacobi Decoding can be reshaped by two complementary, distribution-preserving mechanisms—Proactive Drafting of a shallow K-ary tree after each rejection and Adaptive Continuation that keeps verifying tokens past the first failure—raising step compression to 4.51× and wall-clock speedup to 3.8× on Lumina-mGPT (and 3.25× on Emu3) while FID and CLIP-Score remain indistinguishable from the unaccelerated baseline.
What carries the argument
SJD-PAC: the joint application of Proactive Drafting (local K-ary tree of depth D after a rejection, then a long chain) and Adaptive Continuation (rejection sampling continues over the entire window with stale distributions, recording only the first rejection index).
Load-bearing premise
Image tokens are locally insensitive to distant context changes, so continuing verification with stale probabilities after a rejection still keeps most later tokens without needing a new forward pass.
What would settle it
Measure Total Variation distance between image-token distributions before and after a context perturbation at increasing offsets on a new tokenizer or architecture; if the distance stays large even at modest offsets, Adaptive Continuation’s retention rate collapses and the claimed speedup disappears while remaining correct only in distribution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SJD-PAC, a training-free, lossless enhancement of Speculative Jacobi Decoding for autoregressive text-to-image models. It addresses the skewed acceptance-length distribution of vanilla SJD (frequent single-token steps) via two mechanisms: Proactive Drafting, which builds a shallow K-ary tree of diverse local proposals after a rejection to raise subsequent acceptance probability, and Adaptive Continuation, which continues verification past the first rejection using stale distributions (justified by measured image-token locality) while resampling only rejected tokens. The combination is shown to raise average acceptance length, yielding 4.51 imes step compression / 3.80 imes wall-clock speedup on Lumina-mGPT and 4.31 imes / 3.25 imes on Emu3 while preserving FID and CLIP-Score of the unaccelerated baseline on MS-COCO and PartiPrompts. Correctness is established by reduction to the classical rejection-sampling identity (Supplementary A), independent of draft quality.
Significance. If the reported speedups and losslessness hold, the work is a clear advance for practical AR T2I inference: it removes the need for a separate draft model, remains strictly distribution-preserving, and delivers wall-clock gains competitive with or better than recent lossy alternatives (LANTERN++, GSD, SJD2) without quality degradation. The explicit ablations isolating PD and AC, the locality measurement (Fig. 4), the hyper-parameter study, and the public code strengthen reproducibility. The supplementary proofs that both components recover the exact target marginal are a welcome formal contribution that elevates the paper above purely empirical acceleration claims.
minor comments (5)
- In §4.1 the tree-construction description (sample K candidates without replacement from p(·|X^{t-1}_{<j})) is clear, but a short pseudocode fragment or reference to the PD call in Algorithm 1 would make the hybrid tree-plus-chain construction easier to re-implement.
- Fig. 4 reports TV distance for Lumina-mGPT only; a one-sentence note confirming the same rapid decay was observed on Emu3 (or a small inset) would strengthen the claim that AC’s efficiency rests on a general image-token property.
- Table 1 lists both step-compression and wall-clock numbers; adding a brief remark on why SJD2’s higher step ratio does not translate into higher wall-clock speed (window length 128 vs. 64) would help readers interpret the comparison without consulting the text.
- Minor notation: the residual distribution is written both as max(p-q,0) and as the normalized p'_res; a single consistent symbol after the first definition would improve readability.
- In the qualitative caption of Fig. 5 the step counts (~517 vs. ~2390) are useful; stating the corresponding wall-clock times (or noting that they scale similarly) would complete the visual evidence.
Circularity Check
No significant circularity; losslessness follows from classical rejection sampling and empirical claims are measured against independent public baselines.
full rationale
The paper's central correctness claim (that SJD-PAC samples exactly from the target distribution p) is proved in Supplementary A by direct application of the classical Rejection Sampling Lemma (von Neumann et al., external and standard). Theorems 1 (AC) and 2 (PD) show that even with stale draft distributions q the acceptance/rejection/resampling steps recover p exactly, by algebraic identity of the residual distribution; this is a verification of the algorithm, not a reduction of the claim to its own inputs. Speedups (3.8 imes wall-clock, 4.51 imes step compression) and quality preservation (FID/CLIP identical to AR baseline) are measured on public MS-COCO and PartiPrompts against independently implemented baselines (EAGLE-2, SJD, LANTERN++, GSD, SJD2). Hyperparameters (K, D, L) are selected by ablation and reported as such; no fitted constant is later re-presented as a prediction. Citations to prior SJD work are to distinct author groups and serve only as the starting point being improved; they are not load-bearing uniqueness theorems. The locality observation (Fig. 4) governs only the magnitude of the speedup, not correctness, and is directly measured. The derivation chain is therefore self-contained and non-circular.
Axiom & Free-Parameter Ledger
free parameters (3)
- K (branching factor of proactive tree) =
4
- D (depth of proactive tree) =
3
- L (Jacobi window length) =
64
axioms (3)
- standard math Rejection sampling recovers the exact target marginal regardless of the quality of the proposal distribution q.
- domain assumption Image-token distributions exhibit strong spatial locality (TV distance decays rapidly with context offset).
- domain assumption The target model’s forward pass on a window of length L produces the correct conditional distributions p_i for every position inside the window.
read the original abstract
Speculative Jacobi Decoding (SJD) offers a draft-model-free approach to accelerate autoregressive text-to-image synthesis. However, the high-entropy nature of visual generation yields low draft-token acceptance rates in complex regions, creating a bottleneck that severely limits overall throughput. To overcome this, we introduce SJD-PAC, an enhanced SJD framework. First, SJD-PAC employs a proactive drafting strategy to improve local acceptance rates in these challenging high-entropy regions. Second, we introduce an adaptive continuation mechanism that sustains sequence validation after an initial rejection, bypassing the need for full resampling. Working in tandem, these optimizations significantly increase the average acceptance length per step, boosting inference speed while strictly preserving the target distribution. Experiments on standard text-to-image benchmarks demonstrate that SJD-PAC achieves a $3.8\times$ speedup with lossless image quality. Code is available at https://github.com/KangJialiang/SJD-PAC.
Reference graph
Works this paper leans on
-
[1]
Ethan Chern, Jiadi Su, Yan Ma, and Pengfei Liu. Anole: An open, autoregressive, native large multimodal mod- els for interleaved image-text generation.arXiv preprint arXiv:2407.06135, 2024. 1, 2
Pith/arXiv arXiv 2024
-
[2]
Emerging properties in unified multimodal pretraining.arXiv preprint arXiv:2505.14683, 2025
Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al. Emerging properties in unified multimodal pretraining.arXiv preprint arXiv:2505.14683, 2025. 1, 2
Pith/arXiv arXiv 2025
-
[3]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12882, 2021. 1
2021
-
[4]
Sur la distance de deux lois de probabilit´e
Maurice Fr ´echet. Sur la distance de deux lois de probabilit´e. InAnnales de l’ISUP, pages 183–198, 1957. 5, 6
1957
-
[5]
Partiprompts benchmark.https:// parti.research.google/, 2022
Google Research. Partiprompts benchmark.https:// parti.research.google/, 2022. 5, 6
2022
-
[6]
Clipscore: A reference-free evaluation met- ric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. InProceedings of the 2021 confer- ence on empirical methods in natural language processing, pages 7514–7528, 2021. 6
2021
-
[7]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017. 6
2017
-
[8]
Openclip.Zenodo, 2021
Gabriel Ilharco, Mitchell Wortsman, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, et al. Openclip.Zenodo, 2021. 6
2021
-
[9]
Doohyuk Jang, Sihwan Park, June Yong Yang, Yeonsung Jung, Jihun Yun, Souvik Kundu, Sung-Yub Kim, and Eunho Yang. Lantern: Accelerating visual autoregressive mod- els with relaxed speculative decoding.arXiv preprint arXiv:2410.03355, 2024. 1, 3, 6
Pith/arXiv arXiv 2024
-
[10]
Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai, and Xinghao Chen. Vispec: Accelerating vision-language mod- els with vision-aware speculative decoding.arXiv preprint arXiv:2509.15235, 2025. 3
arXiv 2025
-
[11]
Fast inference from transformers via speculative decoding
Yaron Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. InIn- ternational Conference on Machine Learning, pages 21184– 21195. PMLR, 2023. 1, 2
2023
-
[12]
Eagle-2: Faster inference of language models with dynamic draft trees
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. Eagle-2: Faster inference of language models with dynamic draft trees. InProceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing, pages 7421– 7432, 2024. 1, 5, 6, 7
2024
-
[13]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, Ja Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 4, 5, 6, 8
2014
-
[14]
Dongyang Liu, Shitian Zhao, Le Zhuo, Weifeng Lin, Yi Xin, Xinyue Li, Qi Qin, Yu Qiao, Hongsheng Li, and Peng Gao. Lumina-mgpt: Illuminate flexible photorealistic text- to-image generation with multimodal generative pretraining. arXiv preprint arXiv:2408.02657, 2024. 1, 2, 4, 5, 6, 7, 8
Pith/arXiv arXiv 2024
-
[15]
Specinfer: Accel- erating large language model serving with tree-based specu- lative inference and verification
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al. Specinfer: Accel- erating large language model serving with tree-based specu- lative inference and verification. InProceedings of the 29th ACM International Conference on Architectural Support for Programming ...
2024
-
[16]
Temperature- centric investigation of speculative decoding with knowledge distillation
Siru Ouyang, Shuohang Wang, Minhao Jiang, Ming Zhong, Donghan Yu, Jiawei Han, and Yelong Shen. Temperature- centric investigation of speculative decoding with knowledge distillation. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 13125–13137, 2024. 1
2024
-
[17]
Sihwan Park, Doohyuk Jang, Sungyub Kim, Souvik Kundu, and Eunho Yang. Lantern++: Enhancing relaxed speculative decoding with static tree drafting for visual auto-regressive models.arXiv preprint arXiv:2502.06352, 2025. 1, 3, 5, 6, 7, 8
Pith/arXiv arXiv 2025
-
[18]
Zero-shot text-to-image generation
Aditya Rah, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. InInternational confer- ence on machine learning, pages 8821–8831. Pmlr, 2021. 1, 2
2021
-
[19]
Grouped speculative decoding for autoregressive im- age generation
Junhyuk So, Juncheol Shin, Hyunho Kook, and Eunhyeok Park. Grouped speculative decoding for autoregressive im- age generation. InProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 15375–15384,
-
[20]
Spectr: Fast spec- ulative decoding via optimal transport.Advances in Neural Information Processing Systems, 36:30222–30242, 2023
Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami, Himanshu Jain, and Felix Yu. Spectr: Fast spec- ulative decoding via optimal transport.Advances in Neural Information Processing Systems, 36:30222–30242, 2023. 5
2023
-
[21]
Chameleon: Mixed-modal early-fusion foundation models.arXiv preprint arXiv:2405.09818, 2024
Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models.arXiv preprint arXiv:2405.09818, 2024. 1, 2
Pith/arXiv arXiv 2024
-
[22]
Yao Teng, Han Shi, Xian Liu, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu. Accelerating auto- regressive text-to-image generation with training-free spec- ulative jacobi decoding.arXiv preprint arXiv:2410.01699,
-
[23]
1, 2, 3, 4, 5, 6, 7, 8
-
[24]
Yao Teng, Fuyun Wang, Xian Liu, Zhekai Chen, Han Shi, Yu Wang, Zhenguo Li, Weiyang Liu, Difan Zou, and Xi- hui Liu. Speculative jacobi-denoising decoding for acceler- ating autoregressive text-to-image generation.arXiv preprint arXiv:2510.08994, 2025. 2, 3, 5, 6, 7, 8
arXiv 2025
-
[25]
Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017
Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017. 2
2017
-
[26]
Attention is all you need.Advances in neural information processing systems, 30, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in neural information processing systems, 30, 2017. 2
2017
-
[27]
Various techniques used in con- nection with random digits.John von Neumann, Collected Works, 5:768–770, 1963
John V on Neumann et al. Various techniques used in con- nection with random digits.John von Neumann, Collected Works, 5:768–770, 1963. 3 9
1963
-
[28]
Emu3: Next-token prediction is all you need.arXiv preprint arXiv:2409.18869, 2024
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need.arXiv preprint arXiv:2409.18869, 2024. 1, 2, 5, 6, 7, 8
Pith/arXiv arXiv 2024
-
[29]
Janus: Decoupling visual encod- ing for unified multimodal understanding and generation
Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, et al. Janus: Decoupling visual encod- ing for unified multimodal understanding and generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 12966–12977, 2025
2025
-
[30]
Lumina-mgpt 2.0: Stand- alone autoregressive image modeling.arXiv preprint arXiv:2507.17801, 2025
Yi Xin, Juncheng Yan, Qi Qin, Zhen Li, Dongyang Liu, Shicheng Li, Victor Shea-Jay Huang, Yupeng Zhou, Ren- rui Zhang, Le Zhuo, et al. Lumina-mgpt 2.0: Stand- alone autoregressive image modeling.arXiv preprint arXiv:2507.17801, 2025. 2
Pith/arXiv arXiv 2025
-
[31]
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin- fei Yang, Burcu Karagol Ayan, et al. Scaling autoregres- sive models for content-rich text-to-image generation.arXiv preprint arXiv:2206.10789, 2(3):5, 2022. 1, 2
Pith/arXiv arXiv 2022
-
[32]
Guohui Zhang, Hu Yu, Xiaoxiao Ma, JingHao Zhang, Yaning Pan, Mingde Yao, Jie Xiao, Linjiang Huang, and Feng Zhao. Group critical-token policy optimiza- tion for autoregressive image generation.arXiv preprint arXiv:2509.22485, 2025. 1 10 SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation Supplementary Materia...
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.