{"work":{"id":"41efe203-9377-4c63-b1d6-e499cd6e46f6","openalex_id":null,"doi":null,"arxiv_id":"2406.06525","raw_key":null,"title":"Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation","authors":null,"authors_text":"Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo","year":2024,"venue":"cs.CV","abstract":"We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an affirmative answer to whether vanilla autoregressive models, e.g., Llama, without inductive biases on visual signals can achieve state-of-the-art image generation performance if scaling properly. We reexamine design spaces of image tokenizers, scalability properties of image generation models, and their training data quality. The outcome of this exploration consists of: (1) An image tokenizer with downsample ratio of 16, reconstruction quality of 0.94 rFID and codebook usage of 97% on ImageNet benchmark. (2) A series of class-conditional image generation models ranging from 111M to 3.1B parameters, achieving 2.18 FID on ImageNet 256x256 benchmarks, outperforming the popular diffusion models such as LDM, DiT. (3) A text-conditional image generation model with 775M parameters, from two-stage training on LAION-COCO and high aesthetics quality images, demonstrating competitive performance of visual quality and text alignment. (4) We verify the effectiveness of LLM serving frameworks in optimizing the inference speed of image generation models and achieve 326% - 414% speedup. We release all models and codes to facilitate open-source community of visual generation and multimodal foundation models.","external_url":"https://arxiv.org/abs/2406.06525","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-05-25T04:20:19.385133+00:00","pith_arxiv_id":"2406.06525","created_at":"2026-05-09T22:44:15.232984+00:00","updated_at":"2026-05-25T04:20:19.385133+00:00","title_quality_ok":true,"display_title":"Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation","render_title":"Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation"},"hub":{"state":{"work_id":"41efe203-9377-4c63-b1d6-e499cd6e46f6","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":76,"external_cited_by_count":null,"distinct_field_count":8,"first_pith_cited_at":"2024-08-22T16:32:32+00:00","last_pith_cited_at":"2026-05-22T17:59:42+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-05-26T18:06:43.839447+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":10},{"context_role":"baseline","n":6},{"context_role":"method","n":3}],"polarity_counts":[{"context_polarity":"background","n":9},{"context_polarity":"baseline","n":6},{"context_polarity":"use_method","n":3},{"context_polarity":"unclear","n":1}],"runs":{"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:00:14.729674+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling","work_id":"67d9e391-26d1-459e-ab56-07e60511c886","shared_citers":11},{"title":"Chameleon: Mixed-Modal Early-Fusion Foundation Models","work_id":"2661b9a6-25cc-41a1-8100-612d2b801289","shared_citers":10},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":10},{"title":"Emu3: Next-Token Prediction is All You Need","work_id":"720d288e-fac0-464c-9929-19efd9a52afc","shared_citers":10},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":8},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":8},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":7},{"title":"Show-o: One Single Transformer to Unify Multimodal Understanding and Generation","work_id":"1393dc24-a6b2-44e1-b5d7-7009d1fa4811","shared_citers":7},{"title":"ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment","work_id":"94248955-4bc5-4517-98a0-66224a36d865","shared_citers":6},{"title":"Finite scalar quantization: Vq-vae made simple","work_id":"34dd22bc-0de9-4e11-9a1b-1358e10fbfe1","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation","work_id":"c81c5be2-0655-4234-a3e9-6c32753f136b","shared_citers":6},{"title":"Scaling Autoregressive Models for Content-Rich Text-to-Image Generation","work_id":"0a105815-ff2e-43ce-8566-966cdcae1af4","shared_citers":6},{"title":"Seed-x: Multimodal models with unified multi-granularity comprehension and generation","work_id":"15953092-dd9e-49ae-9f72-e28fc93a6068","shared_citers":6},{"title":"Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model","work_id":"c2bb4d2d-29de-4bf2-9150-3d6373ff358f","shared_citers":6},{"title":"BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset","work_id":"86d896d2-592f-4d9b-938e-dfeb11f9388f","shared_citers":5},{"title":"Diffusion Transformers with Representation Autoencoders","work_id":"c1a2d4de-4439-4005-8c56-e0124e4ed5fa","shared_citers":5},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":5},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":5},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":5},{"title":"Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think","work_id":"1aff8ef8-079b-4afe-9e6a-148e6fd08e6a","shared_citers":5},{"title":"SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features","work_id":"50eec732-2d41-432f-9dcf-ac7fff235ea5","shared_citers":5}],"time_series":[{"n":2,"year":2024},{"n":4,"year":2025},{"n":29,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:00:02.428328+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T17:59:49.424835+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation","claims":[{"claim_text":"We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an affirmative answer to whether vanilla autoregressive models, e.g., Llama, without inductive biases on visual signals can achieve state-of-the-art image generation performance if scaling properly. We reexamine design spaces of image tokenizers, scalability properties of image generation models, and their training data quality. The outcome of this exploration consists of: (1) An image tokenizer with downsample ratio o","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T17:59:44.966904+00:00"}},"summary":{"title":"Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation","claims":[{"claim_text":"We introduce LlamaGen, a new family of image generation models that apply original ``next-token prediction'' paradigm of large language models to visual generation domain. It is an affirmative answer to whether vanilla autoregressive models, e.g., Llama, without inductive biases on visual signals can achieve state-of-the-art image generation performance if scaling properly. We reexamine design spaces of image tokenizers, scalability properties of image generation models, and their training data quality. The outcome of this exploration consists of: (1) An image tokenizer with downsample ratio o","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling","work_id":"67d9e391-26d1-459e-ab56-07e60511c886","shared_citers":11},{"title":"Chameleon: Mixed-Modal Early-Fusion Foundation Models","work_id":"2661b9a6-25cc-41a1-8100-612d2b801289","shared_citers":10},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":10},{"title":"Emu3: Next-Token Prediction is All You Need","work_id":"720d288e-fac0-464c-9929-19efd9a52afc","shared_citers":10},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":8},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":8},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":7},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":7},{"title":"Show-o: One Single Transformer to Unify Multimodal Understanding and Generation","work_id":"1393dc24-a6b2-44e1-b5d7-7009d1fa4811","shared_citers":7},{"title":"ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment","work_id":"94248955-4bc5-4517-98a0-66224a36d865","shared_citers":6},{"title":"Finite scalar quantization: Vq-vae made simple","work_id":"34dd22bc-0de9-4e11-9a1b-1358e10fbfe1","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation","work_id":"c81c5be2-0655-4234-a3e9-6c32753f136b","shared_citers":6},{"title":"Scaling Autoregressive Models for Content-Rich Text-to-Image Generation","work_id":"0a105815-ff2e-43ce-8566-966cdcae1af4","shared_citers":6},{"title":"Seed-x: Multimodal models with unified multi-granularity comprehension and generation","work_id":"15953092-dd9e-49ae-9f72-e28fc93a6068","shared_citers":6},{"title":"Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model","work_id":"c2bb4d2d-29de-4bf2-9150-3d6373ff358f","shared_citers":6},{"title":"BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset","work_id":"86d896d2-592f-4d9b-938e-dfeb11f9388f","shared_citers":5},{"title":"Diffusion Transformers with Representation Autoencoders","work_id":"c1a2d4de-4439-4005-8c56-e0124e4ed5fa","shared_citers":5},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":5},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":5},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":5},{"title":"Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think","work_id":"1aff8ef8-079b-4afe-9e6a-148e6fd08e6a","shared_citers":5},{"title":"SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features","work_id":"50eec732-2d41-432f-9dcf-ac7fff235ea5","shared_citers":5}],"time_series":[{"n":2,"year":2024},{"n":4,"year":2025},{"n":29,"year":2026}],"dependency_candidates":[]},"authors":[]}}