REVIEW 74 cited by
Autoregressive Image Generation without Vector Quantization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Conventional wisdom holds that autoregressive models for image generation are typically accompanied by vector-quantized tokens. We observe that while a discrete-valued space can facilitate representing a categorical distribution, it is not a necessity for autoregressive modeling. In this work, we propose to model the per-token probability distribution using a diffusion procedure, which allows us to apply autoregressive models in a continuous-valued space. Rather than using categorical cross-entropy loss, we define a Diffusion Loss function to model the per-token probability. This approach eliminates the need for discrete-valued tokenizers. We evaluate its effectiveness across a wide range of cases, including standard autoregressive models and generalized masked autoregressive (MAR) variants. By removing vector quantization, our image generator achieves strong results while enjoying the speed advantage of sequence modeling. We hope this work will motivate the use of autoregressive generation in other continuous-valued domains and applications. Code is available at: https://github.com/LTH14/mar.
Forward citations
Showing 60 of 74 Pith papers that cite this
-
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
Distilling frozen Motion-JEPA features into a compact 32-D latent whose geometry is coupled to the decoder lets a standard non-autoregressive flow-matching DiT reach state-of-the-art text-to-motion quality on HumanML3...
-
World Modeling with Probabilistic Structure Integration
A single probabilistic video model extracts optical flow, depth, and segments via counterfactual prompts, then integrates those structures as new token types to improve its own video predictions.
-
Hita: Holistic Tokenizer for Autoregressive Image Generation
Hita's holistic-to-local tokenization lets vanilla autoregressive image models generate global tokens first, improving FID, convergence, and enabling zero-shot style transfer and inpainting.
-
Language-Guided Image Tokenization for Generation
TexTok conditions image tokenization and detokenization on text captions, improving reconstruction and enabling state-of-the-art class-conditional generation FID with far fewer latent tokens.
-
TinyFusion: Diffusion Transformers Learned Shallow
A learnable depth-pruning method that optimizes post-fine-tuning recoverability produces a 14-layer DiT-XL with FID 2.86 and a 2x speedup at 7% of the original training cost.
-
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
A masked-autoregressive diffusion model trained on a compact essential-feature latent space claims state-of-the-art text-to-motion generation under a new essential-dimension evaluation protocol.
-
Token Radius Attention for Efficient Video Generation
Video diffusion transformers can run ~1.5-2x faster with competitive quality by converting each query's attention entropy into a spatially decayed retention radius instead of dense attention.
-
Beyond Token-Level Cross-Entropy: Fr\'echet Distributional Post-Training for Autoregressive Image Generation
FD-loss post-training with detached rollout replay and a probability-level straight-through estimator improves FID and FD_r6 across eight ImageNet configurations.
-
SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation
SRC-Flow compresses RAE features via a Semantic Representation Compressor into a low-dimensional space, enabling normalizing flows to reach gFID 1.65 on ImageNet 256x256 and 2.07 on 512x512 while retaining exact likelihoods.
-
ELT: Elastic Looped Transformers for Visual Generation
Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.
-
InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
An autoregressive transformer with inertial-frame tokenization and geometric rotary positional encoding reports state-of-the-art validity and stability on QM9, GEOM-Drugs, and B3LYP, plus strong functional-group-condi...
-
Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization
A training-only ViT-based projector, VQBridge, combined with learning annealing, achieves full codebook utilization in vector-quantized networks at large codebook sizes, improving reconstruction and autoregressive ima...
-
Decoupling High and Low Frequencies for Faithful Image Generation with Fine Details
A wavelet-based VAE that trains low- and high-frequency branches separately improves image reconstruction and diffusion generation.
-
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
VARIN uses a Location-aware Argmax Inversion pseudo-inverse of Gumbel-max sampling to extract editable discrete noises, enabling training-free prompt-guided editing for visual autoregressive models.
-
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
ROVI's pre-detection VLM-LLM re-captioning yields richer open-vocabulary box labels on 1M curated images, and a GLIGEN model trained on ROVI improves instance grounding, prompt fidelity, and aesthetic quality in the p...
-
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
A causal audio language model with continuous-valued tokens and masked next-token prediction matches diffusion-based text-to-audio quality with smaller, streamable models.
-
CaO$_2$: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation
CaO2 selects confident diffusion-generated samples and optimizes their latents against the denoising objective, achieving state-of-the-art distilled-dataset accuracy on ImageNet subsets.
-
Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
HybridSep combines CLAP embeddings, SSL features, and adversarial consistency training to improve language-queried audio separation, achieving higher SDR and semantic scores than AudioSep and FlowSep in their reported setup.
-
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
StreamMel interleaves text tokens with continuous mel frames in one autoregressive model, reaching state-of-the-art streaming latency with quality comparable to offline zero-shot TTS on LibriSpeech.
-
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
A latent-space transformer autoregressive flow with one deep block plus shallow refiners, tuned noise injection, and score-based guidance reaches competitive FID in high-resolution image synthesis, the first at this s...
-
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
Context-as-Memory conditions video generation on selected historical frames chosen by camera FOV overlap, improving scene consistency in long generated videos.
-
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.
-
Scalable Autoregressive 3D Molecule Generation
Quetzal is an autoregressive 3D molecule generator that matches diffusion-model sample quality on QM9 and GEOM while sampling much faster and enabling exact likelihood computation.
-
Context-Aware Autoregressive Models for Multi-Conditional Image Generation
ContextAR enables multi-conditional autoregressive image generation by embedding canny, depth, HED, pose, and subject conditions into a single token sequence with condition-specific attention masking and positional encoding.
-
Continuous Visual Autoregressive Generation via Score Maximization
Energy-based AutoRegression (EAR) uses a strictly proper energy score to train a masked autoregressive Transformer on continuous image tokens, reaching ImageNet 256x256 FID 1.97 at 937M parameters while generating ima...
-
Phenotype-Guided Generative Model for High-Fidelity Cardiac MRI Synthesis: Advancing Pretraining and Clinical Applications
A phenotype-conditional masked autoregressive diffusion model generates high-fidelity cardiac MRI cines that boost downstream disease classification and phenotype regression when used as pretraining data.
-
Capturing Conditional Dependence via Auto-regressive Diffusion Models
Auto-regressive diffusion models provably control conditional-distribution sampling error with only a factor-K increase in inference cost, unlike vanilla diffusion where conditional error can blow up despite small joi...
-
EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling
Enforcing scale and rotation equivariance in pretrained image autoencoders via a reconstruction loss on transformed latents speeds up and improves latent generative models.
-
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
HMA is a masked autoregressive transformer that predicts future video and actions across many robot embodiments, running up to 15x faster than prior diffusion-based video simulators while matching or improving visual ...
-
Diffusion Autoencoders are Scalable Image Tokenizers
A single diffusion L2 loss can train scalable image tokenizers that match or outperform GAN-LPIPS tokenizers for reconstruction and downstream generation.
-
Visual Generation Without Guidance
GFT trains a single β-conditioned network that reproduces Classifier-Free Guidance's sampling distribution, matching CFG FID scores across five model families with half the inference cost.
-
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Applying test-time verifiers, DPO preference alignment, and a new adaptive reward model (PARM) to autoregressive image generators improves GenEval score from 53% to 77%.
-
Recurrent Diffusion for Large-Scale Parameter Generation
RPG generates full weights for models up to 200M parameters, including ConvNeXt-L and LLaMA LoRA adapters, at accuracy comparable to trained checkpoints, using recurrent token prototypes to condition a 1D diffusion model.
-
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
Aligning a VAE's latent space with DINOv2 features resolves the reconstruction-generation trade-off in latent diffusion, enabling faster DiT training and a state-of-the-art ImageNet FID of 1.35.
-
Dual Diffusion for Unified Image Generation and Understanding
A single diffusion transformer trained with a joint image-flow and masked-text-diffusion loss performs text-to-image generation, image captioning, and visual question answering without any autoregressive text decoder.
-
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
A recurrent token-prediction model that adds noise during quantization and generates images by predicting discrete codes over ten steps, reaching FID 2.56 on ImageNet 256x256.
-
Next Patch Prediction for Autoregressive Visual Generation
Averaging neighboring image tokens into patches during training lets autoregressive image models train faster and generate higher-quality images, with inference unchanged.
-
Parallelized Autoregressive Visual Generation
Grouping spatially distant visual tokens into parallel prediction steps reduces autoregressive generation steps by 3.9x to 11.3x with modest FID/FVD loss.
-
E-CAR: Efficient Continuous Autoregressive Image Generation via Multistage Modeling
A stage-wise continuous autoregressive model with multistage flow matching gets large speedups on 256x256 ImageNet generation, but with a clear FID cost versus DiT and MAR.
-
Normalizing Flows are Capable Generative Models
TarFlow, a stack of alternating-direction causal Transformer autoregressive flows with Gaussian noise training, score-based denoising, and guidance, sets a new likelihood record and near-diffusion sample quality for n...
-
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.
-
ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality
ZipAR is a training-free decoding method that uses spatial locality to generate multiple visual tokens per forward pass, cutting autoregressive image generation steps by up to 91%.
-
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
A robot policy generates its own language reasoning before acting and injects it into a diffusion action decoder, outperforming several VLA baselines on real-robot manipulation.
-
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Orthus combines an LM head and a diffusion head in one transformer to generate interleaved text and images, claiming better GenEval and MME-P scores than Chameleon and Show-o.
-
Hotspot-Driven Peptide Design via Multi-Fragment Autoregressive Extension
PepHAR generates peptide binders by first sampling hot-spot residues from a learned energy model, then autoregressively extending fragments via dihedral angles, then refining the full structure.
-
Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient
CoDe speeds up Visual Auto-Regressive image generation by using a 2B model for early coarse scales and a 0.3B model for later fine scales, with 1.7x-2.9x speedup and only a small FID increase.
-
Image Generation Diversity Issues and How to Tame Them
The paper proposes a retrieval-based diversity metric (IRS), finds that state-of-the-art diffusion models retrieve at most 77% of training images, and introduces feature-conditioned DiADM to improve unconditional diversity.
-
Continuous Speculative Decoding for Autoregressive Image Generation
Continuous speculative decoding accelerates continuous autoregressive image generation by over 2x while approximately maintaining output quality.
-
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
CLEAR is a zero-shot TTS model that autoregressively predicts compact continuous audio latents with a per-token rectified flow head, reaching 1.88% WER on LibriSpeech Subset-B with an RTF of 0.29 and a 96 ms streaming delay.
-
Towards Consistent Long-Term Pose Generation
A one-stage Transformer with placeholder tokens generates continuous 2D pose sequences from a single image and text, avoiding autoregressive error accumulation.
-
Parameterized Diffusion Optimization enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
AOR-DR decomposes DR severity grading into conditional binary steps modeled by a diffusion decoder and reports higher accuracy and macro-F1 than six ordinal regression methods on four datasets.
-
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
Align Your Flow distills flow maps with new continuous-time objectives and autoguidance, achieving state-of-the-art few-step FID on ImageNet and strong text-to-image results.
-
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
On AudioCaps, IMPACT reports the best Fréchet Distance and Fréchet Audio Distance among the compared systems while generating audio faster than diffusion baselines.
-
LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework
A conditional 3D generation framework that combines masked autoencoding and diffusion in token space, with prefix learning and reconstruction-guided sampling, reports state-of-the-art results on ShapeNet and Objaverse.
-
CoC: Chain-of-Cancer based on Cross-Modal Autoregressive Traction for Survival Prediction
Chain-of-Cancer, a cross-modal autoregressive model using text prompts and four clinical modalities, achieves the highest reported survival prediction C-index on five public cancer datasets.
-
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
ScaleKV cuts KV cache memory for Visual Autoregressive text-to-image generation to 10% by classifying layers as drafters or refiners per scale and pruning low-attention tokens while keeping benchmark scores nearly unchanged.
-
Unified Spatial-Temporal Edge-Enhanced Graph Networks for Pedestrian Trajectory Prediction
UniEdge combines a unified spatial-temporal graph, edge-to-edge graph convolution, and a transformer encoder predictor to achieve state-of-the-art ADE/FDE on ETH, UCY, and SDD.
-
Exploring Representation-Aligned Latent Space for Better Generation
Aligning VAE latents with DINOv2 semantic features improves latent diffusion model image generation on ImageNet by about 15% FID.
-
Visual Autoregressive Modeling for Image Super-Resolution
VARSR shows that next-scale visual autoregressive prediction, augmented with diffusion-based quantization residual refinement, can produce competitive perceptual-quality super-resolution at roughly ten times lower inf...
-
LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers
LazyDiT learns small gates that decide when to reuse cached layer outputs, cutting diffusion transformer compute by up to half while matching or beating DDIM quality.
Discussion (0). Continue with ORCID to comment.