Pith. sign in

REVIEW 2 cited by

Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08541 v2 pith:JWUSFMO7 submitted 2023-10-12 cs.CV

classification cs.CV
keywords iterativeself-refinementdesigngenerationidea2imgimagemodelsmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce ``Idea to Image,'' a system that enables multimodal iterative self-refinement with GPT-4V(ision) for automatic image design and generation. Humans can quickly identify the characteristics of different text-to-image (T2I) models via iterative explorations. This enables them to efficiently convert their high-level generation ideas into effective T2I prompts that can produce good images. We investigate if systems based on large multimodal models (LMMs) can develop analogous multimodal self-refinement abilities that enable exploring unknown models or environments via self-refining tries. Idea2Img cyclically generates revised T2I prompts to synthesize draft images, and provides directional feedback for prompt revision, both conditioned on its memory of the probed T2I model's characteristics. The iterative self-refinement brings Idea2Img various advantages over vanilla T2I models. Notably, Idea2Img can process input ideas with interleaved image-text sequences, follow ideas with design instructions, and generate images of better semantic and visual qualities. The user preference study validates the efficacy of multimodal iterative self-refinement on automatic image design and generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    ToolArtist trains a unified multimodal model to reason, search the web, and generate images as one policy, improving scores on WISE and WorldGenBench-Humanities.

  2. Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A multi-agent GPT-4o pipeline that iteratively refines ad banners through self-critique and style voting outperforms single-shot generative baselines in the paper's evaluations, though the evaluation is partly self-re...

Pith tools