Pith. sign in

REVIEW 5 minor 32 references

Two fixes to Speculative Jacobi Decoding raise lossless text-to-image speedup from ~2× to 3.8×.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 22:29 UTC pith:WFZGETRG

load-bearing objection Clean 3.8 imes lossless wall-clock speedup for AR T2I by fixing SJD's short-acceptance bottleneck with two simple, proved-correct tricks; solid systems paper.

arxiv 2603.18599 v2 pith:WFZGETRG submitted 2026-03-19 cs.CV

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation

classification cs.CV
keywords speculative decodingJacobi decodingtext-to-imageautoregressive generationlossless accelerationproactive draftingadaptive continuation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Autoregressive text-to-image models must emit thousands of tokens one by one, so they remain slow even when generation quality is high. Speculative Jacobi Decoding tries to break that sequential bottleneck without an extra draft model by treating generation as a fixed-point iteration and verifying many tokens in parallel; in practice it still stalls on single-token steps in high-entropy image regions, capping real speedup near 2×. SJD-PAC adds two training-free changes that work together: after a rejection it immediately drafts a small tree of alternative continuations (Proactive Drafting), and it never aborts the verification loop early so that later tokens that remain valid under the new context can still be kept (Adaptive Continuation). The combination lifts average accepted length per model call, producing a measured 3.8× wall-clock speedup on a 7 B model while leaving FID and CLIP scores statistically identical to ordinary sampling. The result shows that lossless acceleration of visual autoregressive models can match or beat several lossy methods without sacrificing fidelity.

Core claim

SJD-PAC demonstrates that the skewed acceptance-length distribution of Speculative Jacobi Decoding can be reshaped by two complementary, distribution-preserving mechanisms—Proactive Drafting of a shallow K-ary tree after each rejection and Adaptive Continuation that keeps verifying tokens past the first failure—raising step compression to 4.51× and wall-clock speedup to 3.8× on Lumina-mGPT (and 3.25× on Emu3) while FID and CLIP-Score remain indistinguishable from the unaccelerated baseline.

What carries the argument

SJD-PAC: the joint application of Proactive Drafting (local K-ary tree of depth D after a rejection, then a long chain) and Adaptive Continuation (rejection sampling continues over the entire window with stale distributions, recording only the first rejection index).

Load-bearing premise

Image tokens are locally insensitive to distant context changes, so continuing verification with stale probabilities after a rejection still keeps most later tokens without needing a new forward pass.

What would settle it

Measure Total Variation distance between image-token distributions before and after a context perturbation at increasing offsets on a new tokenizer or architecture; if the distance stays large even at modest offsets, Adaptive Continuation’s retention rate collapses and the claimed speedup disappears while remaining correct only in distribution.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper introduces SJD-PAC, a training-free, lossless enhancement of Speculative Jacobi Decoding for autoregressive text-to-image models. It addresses the skewed acceptance-length distribution of vanilla SJD (frequent single-token steps) via two mechanisms: Proactive Drafting, which builds a shallow K-ary tree of diverse local proposals after a rejection to raise subsequent acceptance probability, and Adaptive Continuation, which continues verification past the first rejection using stale distributions (justified by measured image-token locality) while resampling only rejected tokens. The combination is shown to raise average acceptance length, yielding 4.51 imes step compression / 3.80 imes wall-clock speedup on Lumina-mGPT and 4.31 imes / 3.25 imes on Emu3 while preserving FID and CLIP-Score of the unaccelerated baseline on MS-COCO and PartiPrompts. Correctness is established by reduction to the classical rejection-sampling identity (Supplementary A), independent of draft quality.

Significance. If the reported speedups and losslessness hold, the work is a clear advance for practical AR T2I inference: it removes the need for a separate draft model, remains strictly distribution-preserving, and delivers wall-clock gains competitive with or better than recent lossy alternatives (LANTERN++, GSD, SJD2) without quality degradation. The explicit ablations isolating PD and AC, the locality measurement (Fig. 4), the hyper-parameter study, and the public code strengthen reproducibility. The supplementary proofs that both components recover the exact target marginal are a welcome formal contribution that elevates the paper above purely empirical acceleration claims.

minor comments (5)
  1. In §4.1 the tree-construction description (sample K candidates without replacement from p(·|X^{t-1}_{<j})) is clear, but a short pseudocode fragment or reference to the PD call in Algorithm 1 would make the hybrid tree-plus-chain construction easier to re-implement.
  2. Fig. 4 reports TV distance for Lumina-mGPT only; a one-sentence note confirming the same rapid decay was observed on Emu3 (or a small inset) would strengthen the claim that AC’s efficiency rests on a general image-token property.
  3. Table 1 lists both step-compression and wall-clock numbers; adding a brief remark on why SJD2’s higher step ratio does not translate into higher wall-clock speed (window length 128 vs. 64) would help readers interpret the comparison without consulting the text.
  4. Minor notation: the residual distribution is written both as max(p-q,0) and as the normalized p'_res; a single consistent symbol after the first definition would improve readability.
  5. In the qualitative caption of Fig. 5 the step counts (~517 vs. ~2390) are useful; stating the corresponding wall-clock times (or noting that they scale similarly) would complete the visual evidence.

Circularity Check

0 steps flagged

No significant circularity; losslessness follows from classical rejection sampling and empirical claims are measured against independent public baselines.

full rationale

The paper's central correctness claim (that SJD-PAC samples exactly from the target distribution p) is proved in Supplementary A by direct application of the classical Rejection Sampling Lemma (von Neumann et al., external and standard). Theorems 1 (AC) and 2 (PD) show that even with stale draft distributions q the acceptance/rejection/resampling steps recover p exactly, by algebraic identity of the residual distribution; this is a verification of the algorithm, not a reduction of the claim to its own inputs. Speedups (3.8 imes wall-clock, 4.51 imes step compression) and quality preservation (FID/CLIP identical to AR baseline) are measured on public MS-COCO and PartiPrompts against independently implemented baselines (EAGLE-2, SJD, LANTERN++, GSD, SJD2). Hyperparameters (K, D, L) are selected by ablation and reported as such; no fitted constant is later re-presented as a prediction. Citations to prior SJD work are to distinct author groups and serve only as the starting point being improved; they are not load-bearing uniqueness theorems. The locality observation (Fig. 4) governs only the magnitude of the speedup, not correctness, and is directly measured. The derivation chain is therefore self-contained and non-circular.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 0 invented entities

The paper rests on the classical rejection-sampling identity, the empirical locality of image tokens, and three hand-chosen hyper-parameters that control the draft tree and window. No new physical entities or free constants are fitted to produce the quality numbers; the free parameters affect only speed.

free parameters (3)
  • K (branching factor of proactive tree) = 4
    Chosen by ablation (optimal K=4); controls diversity vs. chain length under fixed compute budget.
  • D (depth of proactive tree) = 3
    Chosen by ablation (optimal D=3); same trade-off as K.
  • L (Jacobi window length) = 64
    Set to 64 after observing that L=32 becomes a bottleneck once AC is added; larger L hurts wall-clock on the authors’ hardware.
axioms (3)
  • standard math Rejection sampling recovers the exact target marginal regardless of the quality of the proposal distribution q.
    Used throughout §3.2 and the supplementary proofs; classical Monte-Carlo fact.
  • domain assumption Image-token distributions exhibit strong spatial locality (TV distance decays rapidly with context offset).
    Empirically measured in Fig. 4 on Lumina-mGPT; underpins the efficiency claim of Adaptive Continuation.
  • domain assumption The target model’s forward pass on a window of length L produces the correct conditional distributions p_i for every position inside the window.
    Standard AR modeling assumption required for any speculative method.

pith-pipeline@v1.1.0-grok45 · 20124 in / 2289 out tokens · 19266 ms · 2026-07-13T22:29:39.191637+00:00 · methodology

0 comments
read the original abstract

Speculative Jacobi Decoding (SJD) offers a draft-model-free approach to accelerate autoregressive text-to-image synthesis. However, the high-entropy nature of visual generation yields low draft-token acceptance rates in complex regions, creating a bottleneck that severely limits overall throughput. To overcome this, we introduce SJD-PAC, an enhanced SJD framework. First, SJD-PAC employs a proactive drafting strategy to improve local acceptance rates in these challenging high-entropy regions. Second, we introduce an adaptive continuation mechanism that sustains sequence validation after an initial rejection, bypassing the need for full resampling. Working in tandem, these optimizations significantly increase the average acceptance length per step, boosting inference speed while strictly preserving the target distribution. Experiments on standard text-to-image benchmarks demonstrate that SJD-PAC achieves a $3.8\times$ speedup with lossless image quality. Code is available at https://github.com/KangJialiang/SJD-PAC.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 10 linked inside Pith

  1. [1]

    Anole: An open, autoregressive, native large multimodal mod- els for interleaved image-text generation.arXiv preprint arXiv:2407.06135, 2024

    Ethan Chern, Jiadi Su, Yan Ma, and Pengfei Liu. Anole: An open, autoregressive, native large multimodal mod- els for interleaved image-text generation.arXiv preprint arXiv:2407.06135, 2024. 1, 2

  2. [2]

    Emerging properties in unified multimodal pretraining.arXiv preprint arXiv:2505.14683, 2025

    Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al. Emerging properties in unified multimodal pretraining.arXiv preprint arXiv:2505.14683, 2025. 1, 2

  3. [3]

    Taming transformers for high-resolution image synthesis

    Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12882, 2021. 1

  4. [4]

    Sur la distance de deux lois de probabilit´e

    Maurice Fr ´echet. Sur la distance de deux lois de probabilit´e. InAnnales de l’ISUP, pages 183–198, 1957. 5, 6

  5. [5]

    Partiprompts benchmark.https:// parti.research.google/, 2022

    Google Research. Partiprompts benchmark.https:// parti.research.google/, 2022. 5, 6

  6. [6]

    Clipscore: A reference-free evaluation met- ric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. InProceedings of the 2021 confer- ence on empirical methods in natural language processing, pages 7514–7528, 2021. 6

  7. [7]

    Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017. 6

  8. [8]

    Openclip.Zenodo, 2021

    Gabriel Ilharco, Mitchell Wortsman, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, et al. Openclip.Zenodo, 2021. 6

  9. [9]

    Lantern: Accelerating visual autoregressive mod- els with relaxed speculative decoding.arXiv preprint arXiv:2410.03355, 2024

    Doohyuk Jang, Sihwan Park, June Yong Yang, Yeonsung Jung, Jihun Yun, Souvik Kundu, Sung-Yub Kim, and Eunho Yang. Lantern: Accelerating visual autoregressive mod- els with relaxed speculative decoding.arXiv preprint arXiv:2410.03355, 2024. 1, 3, 6

  10. [10]

    Vispec: Accelerating vision-language mod- els with vision-aware speculative decoding.arXiv preprint arXiv:2509.15235, 2025

    Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai, and Xinghao Chen. Vispec: Accelerating vision-language mod- els with vision-aware speculative decoding.arXiv preprint arXiv:2509.15235, 2025. 3

  11. [11]

    Fast inference from transformers via speculative decoding

    Yaron Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. InIn- ternational Conference on Machine Learning, pages 21184– 21195. PMLR, 2023. 1, 2

  12. [12]

    Eagle-2: Faster inference of language models with dynamic draft trees

    Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. Eagle-2: Faster inference of language models with dynamic draft trees. InProceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing, pages 7421– 7432, 2024. 1, 5, 6, 7

  13. [13]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, Ja Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 4, 5, 6, 8

  14. [14]

    Lumina-mgpt: Illuminate flexible photorealistic text- to-image generation with multimodal generative pretraining

    Dongyang Liu, Shitian Zhao, Le Zhuo, Weifeng Lin, Yi Xin, Xinyue Li, Qi Qin, Yu Qiao, Hongsheng Li, and Peng Gao. Lumina-mgpt: Illuminate flexible photorealistic text- to-image generation with multimodal generative pretraining. arXiv preprint arXiv:2408.02657, 2024. 1, 2, 4, 5, 6, 7, 8

  15. [15]

    Specinfer: Accel- erating large language model serving with tree-based specu- lative inference and verification

    Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al. Specinfer: Accel- erating large language model serving with tree-based specu- lative inference and verification. InProceedings of the 29th ACM International Conference on Architectural Support for Programming ...

  16. [16]

    Temperature- centric investigation of speculative decoding with knowledge distillation

    Siru Ouyang, Shuohang Wang, Minhao Jiang, Ming Zhong, Donghan Yu, Jiawei Han, and Yelong Shen. Temperature- centric investigation of speculative decoding with knowledge distillation. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 13125–13137, 2024. 1

  17. [17]

    Lantern++: Enhancing relaxed speculative decoding with static tree drafting for visual auto-regressive models.arXiv preprint arXiv:2502.06352, 2025

    Sihwan Park, Doohyuk Jang, Sungyub Kim, Souvik Kundu, and Eunho Yang. Lantern++: Enhancing relaxed speculative decoding with static tree drafting for visual auto-regressive models.arXiv preprint arXiv:2502.06352, 2025. 1, 3, 5, 6, 7, 8

  18. [18]

    Zero-shot text-to-image generation

    Aditya Rah, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. InInternational confer- ence on machine learning, pages 8821–8831. Pmlr, 2021. 1, 2

  19. [19]

    Grouped speculative decoding for autoregressive im- age generation

    Junhyuk So, Juncheol Shin, Hyunho Kook, and Eunhyeok Park. Grouped speculative decoding for autoregressive im- age generation. InProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 15375–15384,

  20. [20]

    Spectr: Fast spec- ulative decoding via optimal transport.Advances in Neural Information Processing Systems, 36:30222–30242, 2023

    Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami, Himanshu Jain, and Felix Yu. Spectr: Fast spec- ulative decoding via optimal transport.Advances in Neural Information Processing Systems, 36:30222–30242, 2023. 5

  21. [21]

    Chameleon: Mixed-modal early-fusion foundation models.arXiv preprint arXiv:2405.09818, 2024

    Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models.arXiv preprint arXiv:2405.09818, 2024. 1, 2

  22. [22]

    Accelerating auto- regressive text-to-image generation with training-free spec- ulative jacobi decoding.arXiv preprint arXiv:2410.01699,

    Yao Teng, Han Shi, Xian Liu, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu. Accelerating auto- regressive text-to-image generation with training-free spec- ulative jacobi decoding.arXiv preprint arXiv:2410.01699,

  23. [23]

    1, 2, 3, 4, 5, 6, 7, 8

  24. [24]

    Speculative jacobi-denoising decoding for acceler- ating autoregressive text-to-image generation.arXiv preprint arXiv:2510.08994, 2025

    Yao Teng, Fuyun Wang, Xian Liu, Zhekai Chen, Han Shi, Yu Wang, Zhenguo Li, Weiyang Liu, Difan Zou, and Xi- hui Liu. Speculative jacobi-denoising decoding for acceler- ating autoregressive text-to-image generation.arXiv preprint arXiv:2510.08994, 2025. 2, 3, 5, 6, 7, 8

  25. [25]

    Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017

    Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning.Advances in neural information pro- cessing systems, 30, 2017. 2

  26. [26]

    Attention is all you need.Advances in neural information processing systems, 30, 2017

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in neural information processing systems, 30, 2017. 2

  27. [27]

    Various techniques used in con- nection with random digits.John von Neumann, Collected Works, 5:768–770, 1963

    John V on Neumann et al. Various techniques used in con- nection with random digits.John von Neumann, Collected Works, 5:768–770, 1963. 3 9

  28. [28]

    Emu3: Next-token prediction is all you need.arXiv preprint arXiv:2409.18869, 2024

    Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need.arXiv preprint arXiv:2409.18869, 2024. 1, 2, 5, 6, 7, 8

  29. [29]

    Janus: Decoupling visual encod- ing for unified multimodal understanding and generation

    Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, et al. Janus: Decoupling visual encod- ing for unified multimodal understanding and generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 12966–12977, 2025

  30. [30]

    Lumina-mgpt 2.0: Stand- alone autoregressive image modeling.arXiv preprint arXiv:2507.17801, 2025

    Yi Xin, Juncheng Yan, Qi Qin, Zhen Li, Dongyang Liu, Shicheng Li, Victor Shea-Jay Huang, Yupeng Zhou, Ren- rui Zhang, Le Zhuo, et al. Lumina-mgpt 2.0: Stand- alone autoregressive image modeling.arXiv preprint arXiv:2507.17801, 2025. 2

  31. [31]

    Scaling autoregres- sive models for content-rich text-to-image generation.arXiv preprint arXiv:2206.10789, 2(3):5, 2022

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin- fei Yang, Burcu Karagol Ayan, et al. Scaling autoregres- sive models for content-rich text-to-image generation.arXiv preprint arXiv:2206.10789, 2(3):5, 2022. 1, 2

  32. [32]

    Group critical-token policy optimiza- tion for autoregressive image generation.arXiv preprint arXiv:2509.22485, 2025

    Guohui Zhang, Hu Yu, Xiaoxiao Ma, JingHao Zhang, Yaning Pan, Mingde Yao, Jie Xiao, Linjiang Huang, and Feng Zhao. Group critical-token policy optimiza- tion for autoregressive image generation.arXiv preprint arXiv:2509.22485, 2025. 1 10 SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation Supplementary Materia...