Pith. sign in

REVIEW 2 cited by

Compositional Text-to-Image Generation with Dense Blob Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.08246 v1 pith:53DW6HKK submitted 2024-05-14 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords blobrepresentationsgenerationcompositionaltext-to-imagebetterblobgencontrollability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability. In this work, we propose to decompose a scene into visual primitives - denoted as dense blob representations - that contain fine-grained details of the scene while being modular, human-interpretable, and easy-to-construct. Based on blob representations, we develop a blob-grounded text-to-image diffusion model, termed BlobGEN, for compositional generation. Particularly, we introduce a new masked cross-attention module to disentangle the fusion between blob representations and visual features. To leverage the compositionality of large language models (LLMs), we introduce a new in-context learning approach to generate blob representations from text prompts. Our extensive experiments show that BlobGEN achieves superior zero-shot generation quality and better layout-guided controllability on MS-COCO. When augmented by LLMs, our method exhibits superior numerical and spatial correctness on compositional image generation benchmarks. Project page: https://blobgen-2d.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-Planner, an MLLM-based planner, decomposes complex image-editing instructions into localized sub-edits with masks and boxes, improving editing quality on standard and new complex benchmarks.

  2. LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A multimodal-LLM planner, polygon-based layout masks, and SVD-based structure injection combine to improve spatial and textual control of pre-trained text-to-image diffusion models.

Pith tools