REVIEW 28 cited by
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sora unveils the potential of scaling Diffusion Transformer for generating photorealistic images and videos at arbitrary resolutions, aspect ratios, and durations, yet it still lacks sufficient implementation details. In this technical report, we introduce the Lumina-T2X family - a series of Flow-based Large Diffusion Transformers (Flag-DiT) equipped with zero-initialized attention, as a unified framework designed to transform noise into images, videos, multi-view 3D objects, and audio clips conditioned on text instructions. By tokenizing the latent spatial-temporal space and incorporating learnable placeholders such as [nextline] and [nextframe] tokens, Lumina-T2X seamlessly unifies the representations of different modalities across various spatial-temporal resolutions. This unified approach enables training within a single framework for different modalities and allows for flexible generation of multimodal data at any resolution, aspect ratio, and length during inference. Advanced techniques like RoPE, RMSNorm, and flow matching enhance the stability, flexibility, and scalability of Flag-DiT, enabling models of Lumina-T2X to scale up to 7 billion parameters and extend the context window to 128K tokens. This is particularly beneficial for creating ultra-high-definition images with our Lumina-T2I model and long 720p videos with our Lumina-T2V model. Remarkably, Lumina-T2I, powered by a 5-billion-parameter Flag-DiT, requires only 35% of the training computational costs of a 600-million-parameter naive DiT. Our further comprehensive analysis underscores Lumina-T2X's preliminary capability in resolution extrapolation, high-resolution editing, generating consistent 3D views, and synthesizing videos with seamless transitions. We expect that the open-sourcing of Lumina-T2X will further foster creativity, transparency, and diversity in the generative AI community.
Forward citations
Cited by 28 Pith papers
-
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
WeEdit trains a glyph-guided, RL-optimized image editor on a 330K-pair synthetic multilingual dataset and reports open-source SOTA on its own bilingual and multilingual text-editing benchmarks.
-
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
UniMC uses tokenized instance conditions (class, box, keypoints) and a timestep-aware modulator in a DiT backbone to control multi-class human and animal image generation, trained and evaluated on the new HAIG-2.9M dataset.
-
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
Layer-normalized averaging of all decoder-only LLM hidden states, rather than last-layer embeddings, improves text-to-image compositional alignment and beats T5 on GenAI-Bench.
-
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
A 3,068-prompt benchmark with per-instance Q&A scoring shows that current text-to-image models, including reasoning-enhanced ones, handle reasoning-driven prompts poorly, with mathematical reasoning near zero.
-
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
DiffSplat repurposes image diffusion models to generate multi-view Gaussian splat grids, using a rendering loss for 3D consistency and achieving state-of-the-art text- and image-conditioned 3D generation.
-
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
PhyGDPO uses groupwise direct preference optimization with real videos as winners to make text-to-video models generate more physically plausible videos.
-
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Z-Image is an efficient 6B-parameter foundation model for image generation that rivals larger commercial systems in photorealism and bilingual text rendering through a new single-stream diffusion transformer and strea...
-
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
iMontage repurposes a pretrained video diffusion model to generate coherent yet highly dynamic image sets from arbitrary numbers of input images.
-
Transition Models: Rethinking the Generative Learning Objective
TiM trains a single diffusion-type model on arbitrary time-interval transitions, achieving strong one-step and multi-step text-to-image generation with 865M parameters.
-
SADA: Stability-guided Adaptive Diffusion Acceleration
SADA accelerates ODE-based generative model sampling by adaptively combining step skipping and token pruning through a stability criterion, giving about 1.8 times speedup with minor fidelity loss.
-
FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
FreeCus is a training-free method that combines pivotal attention sharing, reversed noise shifting, and MLLM captions to personalize Flux.1 text-to-image generation from a single reference image.
-
Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models
A video-diffusion pipeline conditioned on geometry maps, followed by component-wise UV inpainting, produces more coherent and seam-free textures for 3D meshes than Text2Tex, Paint3D, and Meshy in the reported tests.
-
Enhance-A-Video: Better Generated Video for Free
Enhance-A-Video computes the mean off-diagonal temporal attention weight and uses it, scaled by a per-prompt temperature, to boost the attention residual in DiT models during inference.
-
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.
-
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
A new evaluation suite finds that CLIPScore, HPSv2, and Aesthetic Score misjudge challenging text-to-image outputs, while GPT-4o and human ratings favor FLUX.1 and Ideogram2.0.
-
Ingredients: Blending Custom Photos with Video Diffusion Transformers
Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.
-
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
Aligning a VAE's latent space with DINOv2 features resolves the reconstruction-generation trade-off in latent diffusion, enabling faster DiT training and a state-of-the-art ImageNet FID of 1.35.
-
Dual Diffusion for Unified Image Generation and Understanding
A single diffusion transformer trained with a joint image-flow and masked-text-diffusion loss performs text-to-image generation, image captioning, and visual question answering without any autoregressive text decoder.
-
CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
CLEAR replaces full attention in pre-trained diffusion transformers with local circular-window attention and distills the teacher into a student that keeps quality at a fraction of the compute.
-
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.
-
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
SnapGen is a 379M-parameter UNet with cross-architecture distillation and a 1.38M-parameter decoder that generates 1024x1024 images on a phone in about 1.4 seconds, with GenEval 0.66 and ImageNet FID 2.06.
-
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
A decoupled object/camera motion intensity estimator trained by contrastive ranking plus a diffusion transformer that injects the two scores to enable user-controllable video motion.
-
Mind the Time: Temporally-Controlled Multi-Event Video Generation
MinT generates multi-event videos where each event's timing is controlled by the user, using a fine-tuned video diffusion transformer with a time-aware rotary position embedding.
-
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
Metric-scale depth alignment plus scene-constrained noise shaping improves camera controllability and video quality for image-to-video generation on RealEstate10K.
-
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
Continual pre-training of a text-to-video model with block duplication plus an LLM-conditioned cross-attention improves benchmark scores, but aggregate gains hide several per-dimension regressions.
-
Causal Diffusion Transformers for Generative Modeling
A decoder-only transformer that factors generation over both token order and noise level, coupling autoregressive and diffusion training, achieves competitive ImageNet generation and in-context editing.
-
Video Diffusion Transformers are In-Context Learners
Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.
- The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation
Discussion (0). Continue with ORCID to comment.