REVIEW 12 cited by
GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Despite the success achieved by existing image generation and editing methods, current models still struggle with complex problems including intricate text prompts, and the absence of verification and self-correction mechanisms makes the generated images unreliable. Meanwhile, a single model tends to specialize in particular tasks and possess the corresponding capabilities, making it inadequate for fulfilling all user requirements. We propose GenArtist, a unified image generation and editing system, coordinated by a multimodal large language model (MLLM) agent. We integrate a comprehensive range of existing models into the tool library and utilize the agent for tool selection and execution. For a complex problem, the MLLM agent decomposes it into simpler sub-problems and constructs a tree structure to systematically plan the procedure of generation, editing, and self-correction with step-by-step verification. By automatically generating missing position-related inputs and incorporating position information, the appropriate tool can be effectively employed to address each sub-problem. Experiments demonstrate that GenArtist can perform various generation and editing tasks, achieving state-of-the-art performance and surpassing existing models such as SDXL and DALL-E 3, as can be seen in Fig. 1. Project page is https://zhenyuw16.github.io/GenArtist_page.
Forward citations
Cited by 12 Pith papers
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
X-Planner, an MLLM-based planner, decomposes complex image-editing instructions into localized sub-edits with masks and boxes, improving editing quality on standard and new complex benchmarks.
-
Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models
A multi-agent system that audits story images with a vision-language model and repairs inconsistencies with targeted diffusion edits improves multi-panel consistency over existing story visualization methods.
-
BrushEdit: All-In-One Image Inpainting and Editing
BrushEdit couples a multimodal language model and an object detector with a single arbitrary-mask inpainting model to turn free-form text instructions into interactive, multi-turn image edits.
-
SPAgent: Adaptive Task Decomposition and Model Selection for General Video Generation and Editing
SPAgent is an MLLM-based coordinator that decomposes user instructions, plans execution routes, and selects among open-source video generation and editing models, outperforming single models in MOS.
-
ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting
ChatGen-Evo, a three-stage training strategy, outperforms direct supervised fine-tuning on the new ChatGenBench benchmark for automatic text-to-image generation.
-
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
ATLAS adds a Think–Plan–Paint loop with shared positional tokens to unified MLLMs, plus RL-based layout alignment, achieving large reported gains over prior layout-based unified models on compositional image generatio...
-
Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation
ASR replaces the vision encoder of a multimodal LLM with graph-derived structural features to generate UI layouts, reporting better overlap and relation metrics than four prior methods.
-
GraPE: A Generate-Plan-Edit Framework for Compositional T2I Synthesis
A generate-plan-edit wrapper, using GPT-4o to plan atomic edits and a diffusion editor to execute them, improves compositional text-to-image fidelity across many T2I models.
-
MagicQuill: An Intelligent Interactive Image Editing System
MagicQuill combines brush-based edge and color control with an MLLM that guesses user intent, enabling fast interactive image edits without typing prompts.
-
Two-flow Feedback Multi-scale Progressive Generative Adversarial Network
A GAN paper that proposes several new modules but reports no actual experimental results, with placeholder dataset names and percentages.
-
Dynamic Double Space Tower
The paper claims a four-layer Gestalt-based tower can replace attention in VQA and lift a 3B model to state-of-the-art spatial reasoning, but provides no reproducible method or consistent results.
Discussion (0). Continue with ORCID to comment.