REVIEW 14 cited by
DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion models have become the dominant approach for visual generation. They are trained by denoising a Markovian process which gradually adds noise to the input. We argue that the Markovian property limits the model's ability to fully utilize the generation trajectory, leading to inefficiencies during training and inference. In this paper, we propose DART, a transformer-based model that unifies autoregressive (AR) and diffusion within a non-Markovian framework. DART iteratively denoises image patches spatially and spectrally using an AR model that has the same architecture as standard language models. DART does not rely on image quantization, which enables more effective image modeling while maintaining flexibility. Furthermore, DART seamlessly trains with both text and image data in a unified model. Our approach demonstrates competitive performance on class-conditioned and text-to-image generation tasks, offering a scalable, efficient alternative to traditional diffusion models. Through this unified framework, DART sets a new benchmark for scalable, high-quality image synthesis.
Forward citations
Cited by 14 Pith papers
-
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
Distilling frozen Motion-JEPA features into a compact 32-D latent whose geometry is coupled to the decoder lets a standard non-autoregressive flow-matching DiT reach state-of-the-art text-to-motion quality on HumanML3...
-
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
A masked-autoregressive diffusion model trained on a compact essential-feature latent space claims state-of-the-art text-to-motion generation under a new essential-dimension evaluation protocol.
-
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
Iterative latent thought refinement plus terminal text grounding lets diffusion models solve multi-step visual reasoning tasks at 92.1% average accuracy, beating DiffThinker by 8.3 points.
-
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
A latent-space transformer autoregressive flow with one deep block plus shallow refiners, tuned noise injection, and score-based guidance reaches competitive FID in high-resolution image synthesis, the first at this s...
-
SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation
SRDiffusion accelerates video diffusion by switching from a large model to a smaller sibling model after early high-noise steps, using an adaptive threshold for the switch.
-
Next Patch Prediction for Autoregressive Visual Generation
Averaging neighboring image tokens into patches during training lets autoregressive image models train faster and generate higher-quality images, with inference unchanged.
-
CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
A coarse-to-fine autoregressive policy with multi-scale action tokenization matches or beats diffusion policies on robot manipulation benchmarks at roughly 10x lower inference cost.
-
Normalizing Flows are Capable Generative Models
TarFlow, a stack of alternating-direction causal Transformer autoregressive flows with Gaussian noise training, score-based denoising, and guidance, sets a new likelihood record and near-diffusion sample quality for n...
-
Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
CoDe speeds up Visual Auto-Regressive image generation by using a 2B model for early coarse scales and a 0.3B model for later fine scales, with 1.7x-2.9x speedup and only a small FID increase.
-
Continuous Speculative Decoding for Autoregressive Image Generation
Continuous speculative decoding accelerates continuous autoregressive image generation by over 2x while approximately maintaining output quality.
-
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
ScaleKV cuts KV cache memory for Visual Autoregressive text-to-image generation to 10% by classifying layers as drafters or refiners per scale and pruning low-attention tokens while keeping benchmark scores nearly unchanged.
-
Visual Autoregressive Modeling for Image Super-Resolution
VARSR shows that next-scale visual autoregressive prediction, augmented with diffusion-based quantization residual refinement, can produce competitive perceptual-quality super-resolution at roughly ten times lower inf...
-
Causal Diffusion Transformers for Generative Modeling
A decoder-only transformer that factors generation over both token order and noise level, coupling autoregressive and diffusion training, achieves competitive ImageNet generation and in-context editing.
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.