REVIEW 22 cited by
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Lumina-mGPT, a family of multimodal autoregressive models capable of various vision and language tasks, particularly excelling in generating flexible photorealistic images from text descriptions. By initializing from multimodal Generative PreTraining (mGPT), we demonstrate that decoder-only Autoregressive (AR) model can achieve image generation performance comparable to modern diffusion models with high efficiency through Flexible Progressive Supervised Fine-tuning (FP-SFT). Equipped with our proposed Unambiguous image Representation (UniRep), Lumina-mGPT can flexibly generate high-quality images of varying aspect ratios. Building on the strong image generation capabilities, we further explore Ominiponent Supervised Fine-tuning (Omni-SFT), an initial attempt to elevate Lumina-mGPT into a unified multi-modal generalist. The resulting model demonstrates versatile multimodal capabilities, including visual generation tasks like text-to-image/multiview generation and controllable generation, visual recognition tasks like segmentation and depth estimation, and vision-language tasks like multi-turn visual question answering, showing the rosy potential of the technical direction. Codes and checkpoints are available at https://github.com/Alpha-VLLM/Lumina-mGPT.
Forward citations
Cited by 22 Pith papers
-
From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation
A training-free closed-loop PID controller iteratively corrects latent control signals so diffusion models stay consistent with ID, pose, or depth references better than matched open-loop sampling.
-
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
SJD-PAC combines proactive multi-path drafting and adaptive continuation to raise average acceptance length in Speculative Jacobi Decoding, delivering 3.8 imes lossless wall-clock speedup on Lumina-mGPT and Emu3.
-
Exploring Autoregressive Vision Foundation Models for Image Compression
Pretrained autoregressive vision foundation models can be repurposed directly as image entropy coders, delivering competitive perceptual quality at very low bitrates with no fine-tuning.
-
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
UNCAGE is a training-free contrastive attention guidance method that improves compositional text-image alignment in masked generative transformers by prioritizing the unmasking of tokens that clearly represent individ...
-
ChemMLLM: Chemical Multimodal Large Language Model
A chemical multimodal LLM is trained to understand and generate molecule images alongside SMILES and text, with claims of state-of-the-art results on five new tasks.
-
IA-T2I: Internet-Augmented Text-to-Image Generation
IA-T2I uses active retrieval, hierarchical image selection, and self-reflection to supply internet reference images to T2I models, improving generation accuracy on uncertain-knowledge prompts.
-
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.
-
Parallelized Autoregressive Visual Generation
Grouping spatially distant visual tokens into parallel prediction steps reduces autoregressive generation steps by 3.9x to 11.3x with modest FID/FVD loss.
-
E-CAR: Efficient Continuous Autoregressive Image Generation via Multistage Modeling
A stage-wise continuous autoregressive model with multistage flow matching gets large speedups on 256x256 ImageNet generation, but with a clear FID cost versus DiT and MAR.
-
Doe-1: Closed-Loop Autonomous Driving with Large World Model
Doe-1 unifies perception, prediction, and planning in autonomous driving into a single autoregressive next-token generation model over image, text, and action tokens.
-
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
ATLAS adds a Think–Plan–Paint loop with shared positional tokens to unified MLLMs, plus RL-based layout alignment, achieving large reported gains over prior layout-based unified models on compositional image generatio...
-
Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
MaskGIL, a bidirectional-LLaMA masked autoregressive model, generates ImageNet 256x256 images with FID 3.71 in 8 steps, and also supports text-driven and speech-driven generation.
-
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
DC-AR generates 512x512 images in 12 masked autoregressive steps plus 20 diffusion refinement steps, using a 32x compressed 2D tokenizer, and reports gFID 5.49 on MJHQ-30K.
-
CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image Translation
CycleVAR adapts a pretrained visual autoregressive model to unpaired image translation using softmax-relaxed quantization and source-token prefixes, achieving FID scores competitive with CycleGAN-Turbo.
-
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation
A new benchmark and agent framework for complex text-to-image generation, with an unvalidated AI-judge evaluation and claims that the agent outperforms GPT-4o on the authors' own benchmark.
-
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
A self-improving post-training method that uses a model's own generated images as training data, with SFT and GRPO, improves generation and understanding and reduces task imbalance.
-
Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression
ScaleKV cuts KV cache memory for Visual Autoregressive text-to-image generation to 10% by classifying layers as drafters or refiners per scale and pruning low-attention tokens while keeping benchmark scores nearly unchanged.
-
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
HealthGPT unifies medical image comprehension and generation in a single autoregressive model using heterogeneous low-rank adaptation, reporting strong benchmark results.
-
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
ILLUME unifies visual understanding and generation in one LLM with a semantic vision tokenizer and a self-enhancing alignment scheme, reaching competitive benchmarks with only 15M pretraining pairs.
-
SkipVAR: Accelerating Visual Autoregressive Modeling via Adaptive Frequency-Aware Skipping
SkipVAR selects, per sample, between step skipping and unconditional branch replacement using handcrafted frequency features and a trained logistic regression, to accelerate visual autoregressive generation.
-
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
StyleAR enables autoregressive image generation models to do style-aligned text-to-image generation using only binary text-image data, via self-reconstruction training and style-enhanced tokens.
-
LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models
LANTERN++ replaces dynamic tree drafting with static tree drafting plus a multiplicative relaxation bound, reporting up to 2.56x latency reduction for visual autoregressive image models at some image quality cost.
Discussion (0). Continue with ORCID to comment.