{"work":{"id":"4dc55d76-271e-42dd-878f-c20546599c69","openalex_id":"https://openalex.org/W4392538976","doi":"10.48550/arxiv.2403.03206","arxiv_id":"2403.03206","raw_key":null,"title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis","authors":null,"authors_text":"Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M\\\"uller, Harry Saini","year":2024,"venue":"cs.CV","abstract":"Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and noise in a straight line. Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales. Through a large-scale study, we demonstrate the superior performance of this approach compared to established diffusion formulations for high-resolution text-to-image synthesis. Additionally, we present a novel transformer-based architecture for text-to-image generation that uses separate weights for the two modalities and enables a bidirectional flow of information between image and text tokens, improving text comprehension, typography, and human preference ratings. We demonstrate that this architecture follows predictable scaling trends and correlates lower validation loss to improved text-to-image synthesis as measured by various metrics and human evaluations. Our largest models outperform state-of-the-art models, and we will make our experimental data, code, and model weights publicly available.","external_url":"https://arxiv.org/abs/2403.03206","cited_by_count":86,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2403.03206","created_at":"2026-05-09T06:25:47.769843+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis","render_title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis"},"hub":{"state":{"work_id":"4dc55d76-271e-42dd-878f-c20546599c69","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":90,"external_cited_by_count":86,"distinct_field_count":14,"first_pith_cited_at":"2024-05-14T16:33:25+00:00","last_pith_cited_at":"2026-07-08T14:27:31+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T06:49:29.597988+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":13},{"context_role":"method","n":5},{"context_role":"baseline","n":3},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":13},{"context_polarity":"use_method","n":4},{"context_polarity":"baseline","n":3},{"context_polarity":"unclear","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:09:30.684677+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":13},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":11},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":9},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":8},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":8},{"title":"Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis","work_id":"40702548-f094-4c67-a5db-a62f426f852e","shared_citers":7},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":6},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":5},{"title":"DanceGRPO: Unleashing GRPO on Visual Generation","work_id":"7404dd36-8f9c-478f-b089-ef9f8189c711","shared_citers":5},{"title":"ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment","work_id":"94248955-4bc5-4517-98a0-66224a36d865","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":5},{"title":"High-Resolution Image Synthesis with Latent Diffusion Models","work_id":"f0270d36-2952-47fb-84c1-95e3ec341126","shared_citers":5},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":5},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":5},{"title":"Training Diffusion Models with Reinforcement Learning","work_id":"67684dda-3930-452a-b91a-36cbb8e2e219","shared_citers":5},{"title":"Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer","work_id":"f1080a62-48e1-4255-b023-7556be57370d","shared_citers":5},{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","work_id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","shared_citers":4},{"title":"Mean Flows for One-step Generative Modeling","work_id":"07a52ad5-0f82-4095-9a66-559b09fea1ae","shared_citers":4},{"title":"One step diffusion via shortcut models","work_id":"4017f821-b436-40dd-a38d-0b3bb52f7d99","shared_citers":4},{"title":"PixArt-$\\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis","work_id":"77157568-e4be-4041-bb20-388177fc59d0","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":4},{"title":"Scalable Diffusion Models with Transformers","work_id":"a3a05169-18b1-42bb-8775-eada50163437","shared_citers":4}],"time_series":[{"n":2,"year":2024},{"n":2,"year":2025},{"n":30,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:10:09.102019+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:10:17.391938+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis","claims":[{"claim_text":"Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and noise in a straight line. Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales.","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Scaling Rectified Flow Transformers for High-Resolution Image Synthesis because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:10:09.107287+00:00"}},"summary":{"title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis","claims":[{"claim_text":"Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and noise in a straight line. Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales.","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Scaling Rectified Flow Transformers for High-Resolution Image Synthesis because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":13},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":11},{"title":"Classifier-Free Diffusion Guidance","work_id":"acf2c588-c088-4a6c-938e-150ad7c666d7","shared_citers":9},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":8},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":8},{"title":"Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis","work_id":"40702548-f094-4c67-a5db-a62f426f852e","shared_citers":7},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":6},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":5},{"title":"DanceGRPO: Unleashing GRPO on Visual Generation","work_id":"7404dd36-8f9c-478f-b089-ef9f8189c711","shared_citers":5},{"title":"ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment","work_id":"94248955-4bc5-4517-98a0-66224a36d865","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":5},{"title":"High-Resolution Image Synthesis with Latent Diffusion Models","work_id":"f0270d36-2952-47fb-84c1-95e3ec341126","shared_citers":5},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":5},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":5},{"title":"Training Diffusion Models with Reinforcement Learning","work_id":"67684dda-3930-452a-b91a-36cbb8e2e219","shared_citers":5},{"title":"Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer","work_id":"f1080a62-48e1-4255-b023-7556be57370d","shared_citers":5},{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","work_id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","shared_citers":4},{"title":"Mean Flows for One-step Generative Modeling","work_id":"07a52ad5-0f82-4095-9a66-559b09fea1ae","shared_citers":4},{"title":"One step diffusion via shortcut models","work_id":"4017f821-b436-40dd-a38d-0b3bb52f7d99","shared_citers":4},{"title":"PixArt-$\\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis","work_id":"77157568-e4be-4041-bb20-388177fc59d0","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":4},{"title":"Scalable Diffusion Models with Transformers","work_id":"a3a05169-18b1-42bb-8775-eada50163437","shared_citers":4}],"time_series":[{"n":2,"year":2024},{"n":2,"year":2025},{"n":30,"year":2026}],"dependency_candidates":[]},"authors":[]}}