REVIEW 23 cited by
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters challenges including training instability, slow inference, and extrapolation artifacts. In this paper, we present Lumina-Next, an improved version of Lumina-T2X, showcasing stronger generation performance with increased training and inference efficiency. We begin with a comprehensive analysis of the Flag-DiT architecture and identify several suboptimal components, which we address by introducing the Next-DiT architecture with 3D RoPE and sandwich normalizations. To enable better resolution extrapolation, we thoroughly compare different context extrapolation methods applied to text-to-image generation with 3D RoPE, and propose Frequency- and Time-Aware Scaled RoPE tailored for diffusion transformers. Additionally, we introduced a sigmoid time discretization schedule to reduce sampling steps in solving the Flow ODE and the Context Drop method to merge redundant visual tokens for faster network evaluation, effectively boosting the overall sampling speed. Thanks to these improvements, Lumina-Next not only improves the quality and efficiency of basic text-to-image generation but also demonstrates superior resolution extrapolation capabilities and multilingual generation using decoder-based LLMs as the text encoder, all in a zero-shot manner. To further validate Lumina-Next as a versatile generative framework, we instantiate it on diverse tasks including visual recognition, multi-view, audio, music, and point cloud generation, showcasing strong performance across these domains. By releasing all codes and model weights, we aim to advance the development of next-generation generative AI capable of universal modeling.
Forward citations
Cited by 23 Pith papers
-
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
Layer-normalized averaging of all decoder-only LLM hidden states, rather than last-layer embeddings, improves text-to-image compositional alignment and beats T5 on GenAI-Bench.
-
RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
A large captioned retinal dataset and a three-step flow-matching text-to-image model enable fine-grained, caption-controlled synthesis of realistic color fundus photographs.
-
Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds
Scene2Sound generates object-anchored, spatially consistent soundscapes for 3D Gaussian Splatting worlds without retraining, by merging multi-view detections whose rendering Gaussian sets overlap.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
The authors distill a diffusion transformer into a mostly-Mamba hybrid model, reaching teacher-level GenEval scores while generating up to 4K images with linear-complexity speed.
-
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.
-
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.
-
Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
An efficient text-to-video system produces 2K, 24 fps, 5-second videos with a 4B-parameter model by distilling a 13B teacher and guiding high-resolution generation with low-resolution features.
-
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.
-
Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models
Feeding large multimodal models a prompt in multiple languages, not just English, improves text-to-image alignment and human-preference scores across three benchmarks.
-
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
Controlled scaling shows a 2.3B self-attention U-ViT matches or slightly outperforms SDXL U-Net and larger cross-attention DiT variants, while long captions and larger datasets improve text-image alignment.
-
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.
-
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
SnapGen is a 379M-parameter UNet with cross-architecture distillation and a 1.38M-parameter decoder that generates 1024x1024 images on a phone in about 1.4 seconds, with GenEval 0.66 and ImageNet FID 2.06.
-
Learning Visual Generative Priors without Text
An image-to-image diffusion model pretrained on 190 million unlabeled images serves as a transferable visual generative prior for text-to-image, novel-view synthesis, and image-to-video tasks.
-
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
Switti shows that removing causality from a scale-wise transformer and disabling classifier-free guidance at the last scales yields faster, competitive text-to-image generation.
-
One Diffusion to Generate Them All
OneDiffusion shows that a single 2.8B-parameter diffusion model, trained by treating all tasks as frame sequences with varying noise scales, can handle image generation and image understanding tasks bidirectionally.
-
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
A decoupled-attention adapter transfers image-pair edits to new photos in diffusion transformers, trained with a new 218-task visual editing dataset.
-
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
Deep fusion of a frozen LLM with a DiT improves text-image alignment over shallow fusion baselines, and a scaled recipe (FuseDiT) achieves competitive results despite limited data and compute.
-
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
SANA-1.5 combines layer growth, depth pruning, and VLM-judged best-of-N sampling to push GenEval text-to-image alignment from 0.81 to 0.96.
-
LaVin-DiT: Large Vision Diffusion Transformer
A single diffusion transformer with a spatial-temporal VAE and in-context example pairs unifies over 20 vision tasks, with strong scores on several image benchmarks and uneven evidence on video tasks.
-
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
A lightweight open-source connector between a frozen multimodal LLM and a diffusion model yields a unified model that matches larger systems on image generation and understanding benchmarks, with the caveat that headl...
-
LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization
LiteVAR compresses VAR image generation via multi-diagonal windowed attention, CFG output sharing, and mixed-precision quantization, reporting up to 85% attention savings and 50% memory reduction with minimal FID change.
-
High-Resolution Image Synthesis via Next-Token Prediction
An autoregressive model with continuous tokens, a new positional embedding (VoPE), and a data-feedback training strategy achieves strong text-to-image benchmarks at resolutions up to 4K.
Discussion (0). Continue with ORCID to comment.