REVIEW 8 cited by
Controllable Text-to-Image Generation with GPT-4
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Current text-to-image generation models often struggle to follow textual instructions, especially the ones requiring spatial reasoning. On the other hand, Large Language Models (LLMs), such as GPT-4, have shown remarkable precision in generating code snippets for sketching out text inputs graphically, e.g., via TikZ. In this work, we introduce Control-GPT to guide the diffusion-based text-to-image pipelines with programmatic sketches generated by GPT-4, enhancing their abilities for instruction following. Control-GPT works by querying GPT-4 to write TikZ code, and the generated sketches are used as references alongside the text instructions for diffusion models (e.g., ControlNet) to generate photo-realistic images. One major challenge to training our pipeline is the lack of a dataset containing aligned text, images, and sketches. We address the issue by converting instance masks in existing datasets into polygons to mimic the sketches used at test time. As a result, Control-GPT greatly boosts the controllability of image generation. It establishes a new state-of-art on the spatial arrangement and object positioning generation and enhances users' control of object positions, sizes, etc., nearly doubling the accuracy of prior models. Our work, as a first attempt, shows the potential for employing LLMs to enhance the performance in computer vision tasks.
Forward citations
Cited by 8 Pith papers
-
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
LiveEdit distills a bidirectional video foundation model into a unidirectional streaming editor via three-stage training plus mask caching to reach 12.66 FPS with stable edits.
-
Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
Canvas3D lets users arrange objects in a 3D canvas generated from a text prompt, then feeds depth, skeleton, and lighting constraints to diffusion models to produce images that match the layout.
-
SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
SketchAgent automates sketch-to-diagram conversion with a three-agent pipeline, but its benchmark replaces hand-drawn sketches with simplified renderings of the very diagrams the system must produce.
-
Detail++: Training-Free Detail Enhancer for T2I Diffusion Models
Detail++ uses progressive multi-branch prompt injection and test-time attention optimization to improve attribute binding in text-to-image generation.
-
FlexControl: Computation-Aware ControlNet with Differentiable Router for Text-to-Image Generation
FlexControl learns per-timestep, per-block gating for ControlNet control branches, using a FLOPs budget loss to cut compute while preserving or improving image fidelity.
-
PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction
PersonaVlog auto-generates personalized vlogs from a theme and reference image using multimodal agents with a feedback-rollback loop, and introduces the ThemeVlogEval benchmark.
-
Detailed radial scale height profile of dust grains as probed by dust self-scattering in HL Tau
From the near-far side asymmetry and azimuthal contrast in HL Tau's polarized intensity, the authors infer a radial dust scale height profile and a turbulence parameter alpha increasing from 1e-5 at 100 au to 1e-2.5 at 20 au.
-
Test-time Prompt Refinement for Text-to-Image Models
A training-free closed loop, in which a multimodal LLM rewrites a text prompt after inspecting the generated image, improves overall text-to-image alignment but degrades some attribute and spatial categories.
Discussion (0). Continue with ORCID to comment.