REVIEW 6 cited by
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for visual content generation. However, existing works have insufficiently addressed the varying granularity demands of different image generation tasks within a unified MLLM paradigm - from the diversity required in text-to-image generation to the precise controllability needed in image manipulation. In this work, we propose PUMA, emPowering Unified MLLM with Multi-grAnular visual generation. PUMA unifies multi-granular visual features as both inputs and outputs of MLLMs, elegantly addressing the different granularity requirements of various image generation tasks within a unified MLLM framework. Following multimodal pretraining and task-specific instruction tuning, PUMA demonstrates proficiency in a wide range of multimodal tasks. This work represents a significant step towards a truly unified MLLM capable of adapting to the granularity demands of various visual tasks. The code and model will be released in https://github.com/rongyaofang/PUMA.
Forward citations
Cited by 6 Pith papers
-
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
The authors build a 6M-image, 20M-caption reasoning dataset with generation chain-of-thought and a 7-track VLM-judged benchmark, then rank 19 text-to-image models.
-
One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
Single-patch visual-token grounding with RL-based token selection and full-image geometry decoding reports large gains over multi-patch and coordinate-text grounding for MLLM scene text spotting.
-
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Reinforcement learning that jointly optimizes a text planning step and patch-by-patch image generation in one autoregressive model improves compositional text-to-image benchmarks by double-digit absolute percentage po...
-
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
An autoregressive model with group self-attention that separates learning from applying achieves state-of-the-art few-shot image manipulation on unseen instructions.
-
Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models
Uncertainty-o estimates uncertainty in large multimodal models by perturbing prompts and computing entropy over semantically clustered answers, improving hallucination detection across five modalities.
-
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
ILLUME unifies visual understanding and generation in one LLM with a semantic vision tokenizer and a self-enhancing alignment scheme, reaching competitive benchmarks with only 15M pretraining pairs.
Discussion (0). Continue with ORCID to comment.