Pith. sign in

REVIEW 4 cited by

LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10640 v2 pith:KCF66LRS submitted 2023-10-16 cs.CV

classification cs.CV
keywords objectspromptsgenerationmodelstextualcomplexdescriptionsdetailed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in generating images from short, single-object descriptions, these models often struggle to faithfully capture all the nuanced details within longer and more elaborate textual inputs. In response, we present a novel approach leveraging Large Language Models (LLMs) to extract critical components from text prompts, including bounding box coordinates for foreground objects, detailed textual descriptions for individual objects, and a succinct background context. These components form the foundation of our layout-to-image generation model, which operates in two phases. The initial Global Scene Generation utilizes object layouts and background context to create an initial scene but often falls short in faithfully representing object characteristics as specified in the prompts. To address this limitation, we introduce an Iterative Refinement Scheme that iteratively evaluates and refines box-level content to align them with their textual descriptions, recomposing objects as needed to ensure consistency. Our evaluation on complex prompts featuring multiple objects demonstrates a substantial improvement in recall compared to baseline diffusion models. This is further validated by a user study, underscoring the efficacy of our approach in generating coherent and detailed scenes from intricate textual inputs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GOBench measures how well multimodal AI models generate and understand geometric optics, finding that even top models make frequent physical errors.

  2. DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Introduces DetailMaster, a 4,116-prompt benchmark with fine-grained evaluation of long-prompt text-to-image generation, finding that state-of-the-art models achieve only about 50% accuracy on attribute binding and spa...

  3. Test-time Prompt Refinement for Text-to-Image Models

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A training-free closed loop, in which a multimodal LLM rewrites a text prompt after inspecting the generated image, improves overall text-to-image alignment but degrades some attribute and spatial categories.

  4. Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration

    eess.SP 2025-06 conditional novelty 4.0 of 10

    The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...

Pith tools