Pith. sign in

REVIEW 23 cited by

Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18583 v1 pith:KQ44SNTV submitted 2024-06-05 cs.CV cs.LG

classification cs.CVcs.LG
keywords generationextrapolationlumina-nextlumina-t2xropearchitecturecapabilitiescontext
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters challenges including training instability, slow inference, and extrapolation artifacts. In this paper, we present Lumina-Next, an improved version of Lumina-T2X, showcasing stronger generation performance with increased training and inference efficiency. We begin with a comprehensive analysis of the Flag-DiT architecture and identify several suboptimal components, which we address by introducing the Next-DiT architecture with 3D RoPE and sandwich normalizations. To enable better resolution extrapolation, we thoroughly compare different context extrapolation methods applied to text-to-image generation with 3D RoPE, and propose Frequency- and Time-Aware Scaled RoPE tailored for diffusion transformers. Additionally, we introduced a sigmoid time discretization schedule to reduce sampling steps in solving the Flow ODE and the Context Drop method to merge redundant visual tokens for faster network evaluation, effectively boosting the overall sampling speed. Thanks to these improvements, Lumina-Next not only improves the quality and efficiency of basic text-to-image generation but also demonstrates superior resolution extrapolation capabilities and multilingual generation using decoder-based LLMs as the text encoder, all in a zero-shot manner. To further validate Lumina-Next as a versatile generative framework, we instantiate it on diverse tasks including visual recognition, multi-view, audio, music, and point cloud generation, showcasing strong performance across these domains. By releasing all codes and model weights, we aim to advance the development of next-generation generative AI capable of universal modeling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Layer-normalized averaging of all decoder-only LLM hidden states, rather than last-layer embeddings, improves text-to-image compositional alignment and beats T5 on GenAI-Bench.

  2. RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions

    eess.IV 2025-05 conditional novelty 7.0 of 10

    A large captioned retinal dataset and a three-step flow-matching text-to-image model enable fine-grained, caption-controlled synthesis of realistic color fundus photographs.

  3. Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Scene2Sound generates object-anchored, spatially consistent soundscapes for 3D Gaussian Splatting worlds without retraining, by merging multi-view detections whose rendering Gaussian sets overlap.

  4. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  5. Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors distill a diffusion transformer into a mostly-Mamba hybrid model, reaching teacher-level GenEval scores while generating up to 4K images with linear-complexity speed.

  6. ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.

  7. FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.

  8. Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis

    cs.CV 2025-04 conditional novelty 6.0 of 10

    An efficient text-to-video system produces 2K, 24 fps, 5-second videos with a 4B-parameter model by distilling a 13B teacher and guiding high-resolution generation with low-resolution features.

  9. Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.

  10. Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Feeding large multimodal models a prompt in multiple languages, not just English, improves text-to-image alignment and human-preference scores across three benchmarks.

  11. Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Controlled scaling shows a 2.3B self-attention U-ViT matches or slightly outperforms SDXL U-Net and larger cross-attention DiT variants, while long captions and larger datasets improve text-image alignment.

  12. VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.

  13. SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SnapGen is a 379M-parameter UNet with cross-architecture distillation and a 1.38M-parameter decoder that generates 1024x1024 images on a phone in about 1.4 seconds, with GenEval 0.66 and ImageNet FID 2.06.

  14. Learning Visual Generative Priors without Text

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An image-to-image diffusion model pretrained on 190 million unlabeled images serves as a transferable visual generative prior for text-to-image, novel-view synthesis, and image-to-video tasks.

  15. Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Switti shows that removing causality from a scale-wise transformer and disabling classifier-free guidance at the last scales yields faster, competitive text-to-image generation.

  16. One Diffusion to Generate Them All

    cs.CV 2024-11 conditional novelty 6.0 of 10

    OneDiffusion shows that a single 2.8B-parameter diffusion model, trained by treating all tasks as frame sequences with varying noise scales, can handle image generation and image understanding tasks bidirectionally.

  17. RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A decoupled-attention adapter transfers image-pair edits to new photos in diffusion transformers, trained with a new 218-task visual editing dataset.

  18. Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Deep fusion of a frozen LLM with a DiT improves text-image alignment over shallow fusion baselines, and a scaled recipe (FuseDiT) achieves competitive results despite limited data and compute.

  19. SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

    cs.CV 2025-01 conditional novelty 5.0 of 10

    SANA-1.5 combines layer growth, depth pruning, and VLM-judged best-of-N sampling to push GenEval text-to-image alignment from 0.81 to 0.96.

  20. LaVin-DiT: Large Vision Diffusion Transformer

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A single diffusion transformer with a spatial-temporal VAE and in-context example pairs unifies over 20 vision tasks, with strong scores on several image benchmarks and uneven evidence on video tasks.

  21. OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A lightweight open-source connector between a frozen multimodal LLM and a diffusion model yields a unified model that matches larger systems on image generation and understanding benchmarks, with the caveat that headl...

  22. LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization

    cs.CV 2024-11 conditional novelty 4.0 of 10

    LiteVAR compresses VAR image generation via multi-diagonal windowed attention, CFG output sharing, and mixed-precision quantization, reporting up to 85% attention savings and 50% memory reduction with minimal FID change.

  23. High-Resolution Image Synthesis via Next-Token Prediction

    cs.CV 2024-11 conditional novelty 4.0 of 10

    An autoregressive model with continuous tokens, a new positional embedding (VoPE), and a data-feedback training strategy achieves strong text-to-image benchmarks at resolutions up to 4K.

Pith tools