Pith. sign in

REVIEW 22 cited by

Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.02657 v3 pith:FWUGH6CH submitted 2024-08-05 cs.CV

classification cs.CV
keywords generationlumina-mgptmultimodaltasksflexibleimagelikevisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Lumina-mGPT, a family of multimodal autoregressive models capable of various vision and language tasks, particularly excelling in generating flexible photorealistic images from text descriptions. By initializing from multimodal Generative PreTraining (mGPT), we demonstrate that decoder-only Autoregressive (AR) model can achieve image generation performance comparable to modern diffusion models with high efficiency through Flexible Progressive Supervised Fine-tuning (FP-SFT). Equipped with our proposed Unambiguous image Representation (UniRep), Lumina-mGPT can flexibly generate high-quality images of varying aspect ratios. Building on the strong image generation capabilities, we further explore Ominiponent Supervised Fine-tuning (Omni-SFT), an initial attempt to elevate Lumina-mGPT into a unified multi-modal generalist. The resulting model demonstrates versatile multimodal capabilities, including visual generation tasks like text-to-image/multiview generation and controllable generation, visual recognition tasks like segmentation and depth estimation, and vision-language tasks like multi-turn visual question answering, showing the rosy potential of the technical direction. Codes and checkpoints are available at https://github.com/Alpha-VLLM/Lumina-mGPT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A training-free closed-loop PID controller iteratively corrects latent control signals so diffusion models stay consistent with ID, pose, or depth references better than matched open-loop sampling.

  2. SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation

    cs.CV 2026-03 accept novelty 6.0 of 10

    SJD-PAC combines proactive multi-path drafting and adaptive continuation to raise average acceptance length in Speculative Jacobi Decoding, delivering 3.8 imes lossless wall-clock speedup on Lumina-mGPT and Emu3.

  3. Exploring Autoregressive Vision Foundation Models for Image Compression

    eess.IV 2025-09 conditional novelty 6.0 of 10

    Pretrained autoregressive vision foundation models can be repurposed directly as image entropy coders, delivering competitive perceptual quality at very low bitrates with no fine-tuning.

  4. UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    UNCAGE is a training-free contrastive attention guidance method that improves compositional text-image alignment in masked generative transformers by prioritizing the unmasking of tokens that clearly represent individ...

  5. ChemMLLM: Chemical Multimodal Large Language Model

    cs.LG 2025-05 reject novelty 6.0 of 10

    A chemical multimodal LLM is trained to understand and generate molecule images alongside SMILES and text, with claims of state-of-the-art results on five new tasks.

  6. IA-T2I: Internet-Augmented Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    IA-T2I uses active retrieval, hierarchical image selection, and self-reflection to supply internet reference images to T2I models, improving generation accuracy on uncertain-knowledge prompts.

  7. Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.

  8. Parallelized Autoregressive Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Grouping spatially distant visual tokens into parallel prediction steps reduces autoregressive generation steps by 3.9x to 11.3x with modest FID/FVD loss.

  9. E-CAR: Efficient Continuous Autoregressive Image Generation via Multistage Modeling

    cs.CV 2024-12 reject novelty 6.0 of 10

    A stage-wise continuous autoregressive model with multistage flow matching gets large speedups on 256x256 ImageNet generation, but with a clear FID cost versus DiT and MAR.

  10. Doe-1: Closed-Loop Autonomous Driving with Large World Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.

  11. Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ATLAS adds a Think–Plan–Paint loop with shared positional tokens to unified MLLMs, plus RL-based layout alignment, achieving large reported gains over prior layout-based unified models on compositional image generatio...

  12. Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MaskGIL, a bidirectional-LLaMA masked autoregressive model, generates ImageNet 256x256 images with FID 3.71 in 8 steps, and also supports text-driven and speech-driven generation.

  13. DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DC-AR generates 512x512 images in 12 masked autoregressive steps plus 20 diffusion refinement steps, using a 32x compressed 2D tokenizer, and reports gFID 5.49 on MJHQ-30K.

  14. CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    CycleVAR adapts a pretrained visual autoregressive model to unpaired image translation using softmax-relaxed quantization and source-token prefixes, achieving FID scores competitive with CycleGAN-Turbo.

  15. Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A new benchmark and agent framework for complex text-to-image generation, with an unvalidated AI-judge evaluation and claims that the agent outperforms GPT-4o on the authors' own benchmark.

  16. UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A self-improving post-training method that uses a model's own generated images as training data, with SFT and GRPO, improves generation and understanding and reduces task imbalance.

  17. Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

    cs.LG 2025-05 conditional novelty 5.0 of 10

    ScaleKV cuts KV cache memory for Visual Autoregressive text-to-image generation to 10% by classifying layers as drafters or refiners per scale and pruning low-attention tokens while keeping benchmark scores nearly unchanged.

  18. HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

    cs.CV 2025-02 reject novelty 5.0 of 10

    HealthGPT unifies medical image comprehension and generation in a single autoregressive model using heterogeneous low-rank adaptation, reporting strong benchmark results.

  19. ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

    cs.CV 2024-12 conditional novelty 5.0 of 10

    ILLUME unifies visual understanding and generation in one LLM with a semantic vision tokenizer and a self-enhancing alignment scheme, reaching competitive benchmarks with only 15M pretraining pairs.

  20. SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping

    cs.CV 2025-06 conditional novelty 4.0 of 10

    SkipVAR selects, per sample, between step skipping and unconditional branch replacement using handcrafted frequency features and a trained logistic regression, to accelerate visual autoregressive generation.

  21. StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

    cs.CV 2025-05 reject novelty 4.0 of 10

    StyleAR enables autoregressive image generation models to do style-aligned text-to-image generation using only binary text-image data, via self-reconstruction training and style-enhanced tokens.

  22. LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models

    cs.CV 2025-02 conditional novelty 4.0 of 10

    LANTERN++ replaces dynamic tree drafting with static tree drafting plus a multiplicative relaxation bound, reporting up to 2.56x latency reduction for visual autoregressive image models at some image quality cost.

Pith tools