Pith. sign in

REVIEW 8 cited by

Controllable Text-to-Image Generation with GPT-4

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18583 v1 pith:LI5H22A4 submitted 2023-05-29 cs.CV cs.AI

classification cs.CVcs.AI
keywords generationgpt-4modelssketchescontrol-gpttexttext-to-imagecode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current text-to-image generation models often struggle to follow textual instructions, especially the ones requiring spatial reasoning. On the other hand, Large Language Models (LLMs), such as GPT-4, have shown remarkable precision in generating code snippets for sketching out text inputs graphically, e.g., via TikZ. In this work, we introduce Control-GPT to guide the diffusion-based text-to-image pipelines with programmatic sketches generated by GPT-4, enhancing their abilities for instruction following. Control-GPT works by querying GPT-4 to write TikZ code, and the generated sketches are used as references alongside the text instructions for diffusion models (e.g., ControlNet) to generate photo-realistic images. One major challenge to training our pipeline is the lack of a dataset containing aligned text, images, and sketches. We address the issue by converting instance masks in existing datasets into polygons to mimic the sketches used at test time. As a result, Control-GPT greatly boosts the controllability of image generation. It establishes a new state-of-art on the spatial arrangement and object positioning generation and enhances users' control of object positions, sizes, etc., nearly doubling the accuracy of prior models. Our work, as a first attempt, shows the potential for employing LLMs to enhance the performance in computer vision tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

    cs.CV 2026-06 unverdicted novelty 6.5 of 10

    LiveEdit distills a bidirectional video foundation model into a unidirectional streaming editor via three-stage training plus mask caching to reach 12.66 FPS with stable edits.

  2. Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas

    cs.HC 2025-08 conditional novelty 6.0 of 10

    Canvas3D lets users arrange objects in a 3D canvas generated from a text prompt, then feeds depth, skeleton, and lighting constraints to diffusion models to produce images that match the layout.

  3. SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches

    cs.AI 2025-08 reject novelty 6.0 of 10

    SketchAgent automates sketch-to-diagram conversion with a three-agent pipeline, but its benchmark replaces hand-drawn sketches with simplified renderings of the very diagrams the system must produce.

  4. Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Detail++ uses progressive multi-branch prompt injection and test-time attention optimization to improve attribute binding in text-to-image generation.

  5. FlexControl: Computation-Aware ControlNet with Differentiable Router for Text-to-Image Generation

    cs.LG 2025-02 conditional novelty 6.0 of 10

    FlexControl learns per-timestep, per-block gating for ControlNet control branches, using a FLOPs budget loss to cut compute while preserving or improving image fidelity.

  6. PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction

    cs.CV 2025-08 conditional novelty 5.0 of 10

    PersonaVlog auto-generates personalized vlogs from a theme and reference image using multimodal agents with a feedback-rollback loop, and introduces the ThemeVlogEval benchmark.

  7. Detailed radial scale height profile of dust grains as probed by dust self-scattering in HL Tau

    astro-ph.EP 2025-08 unverdicted novelty 5.0 of 10

    From the near-far side asymmetry and azimuthal contrast in HL Tau's polarized intensity, the authors infer a radial dust scale height profile and a turbulence parameter alpha increasing from 1e-5 at 100 au to 1e-2.5 at 20 au.

  8. Test-time Prompt Refinement for Text-to-Image Models

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A training-free closed loop, in which a multimodal LLM rewrites a text prompt after inspecting the generated image, improves overall text-to-image alignment but degrades some attribute and spatial categories.

Pith tools