REVIEW 5 cited by
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Nearly every recent image synthesis approach, including diffusion, masked-token prediction, and next-token prediction, uses a Transformer network architecture. Despite this common backbone, there has been no direct, compute controlled comparison of how these approaches affect performance and efficiency. We analyze the scalability of each approach through the lens of compute budget measured in FLOPs. We find that token prediction methods, led by next-token prediction, significantly outperform diffusion on prompt following. On image quality, while next-token prediction initially performs better, scaling trends suggest it is eventually matched by diffusion. We compare the inference compute efficiency of each approach and find that next token prediction is by far the most efficient. Based on our findings we recommend diffusion for applications targeting image quality and low latency; and next-token prediction when prompt following or throughput is more important.
Forward citations
Cited by 5 Pith papers
-
Spanning Tree Autoregressive Visual Generation
Breadth-first traversal of random spanning trees as token order keeps autoregressive image quality and enables connected-mask inpainting without changing the transformer architecture.
-
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
Aligning a VAE's latent space with DINOv2 features resolves the reconstruction-generation trade-off in latent diffusion, enabling faster DiT training and a state-of-the-art ImageNet FID of 1.35.
-
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
A coarse-to-fine autoregressive policy with multi-scale action tokenization matches or beats diffusion policies on robot manipulation benchmarks at roughly 10x lower inference cost.
-
[MASK] is All You Need
Discrete Interpolants frames image generation, segmentation, and video generation as unmasking discrete [MASK] tokens, connecting masked generative models and discrete diffusion models.
-
High-Resolution Image Synthesis via Next-Token Prediction
An autoregressive model with continuous tokens, a new positional embedding (VoPE), and a data-feedback training strategy achieves strong text-to-image benchmarks at resolutions up to 4K.
Discussion (0). Continue with ORCID to comment.