Pith. sign in

REVIEW 3 cited by

Localized Text-to-Image Generation for Free via Cross Attention Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.14636 v1 pith:WH5PNRIL submitted 2023-06-26 cs.CV

classification cs.CV
keywords generationlocalizedtext-to-imagemodelsinferenceattentioncrosstime
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the tremendous success in text-to-image generative models, localized text-to-image generation (that is, generating objects or features at specific locations in an image while maintaining a consistent overall generation) still requires either explicit training or substantial additional inference time. In this work, we show that localized generation can be achieved by simply controlling cross attention maps during inference. With no additional training, model architecture modification or inference time, our proposed cross attention control (CAC) provides new open-vocabulary localization abilities to standard text-to-image models. CAC also enhances models that are already trained for localized generation when deployed at inference time. Furthermore, to assess localized text-to-image generation performance automatically, we develop a standardized suite of evaluations using large pretrained recognition models. Our experiments show that CAC improves localized generation performance with various types of location information ranging from bounding boxes to semantic segmentation maps, and enhances the compositional capability of state-of-the-art text-to-image generative models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Appearance pointers are compact tokens that let a diffusion transformer apply text, image, or combined prompts to specific image regions in a single pass.

  2. Compositional Video Generation via Inference-Time Guidance

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CVG improves compositional faithfulness in frozen text-to-video diffusion models by steering early denoising steps with gradients from a classifier trained on the model's own cross-attention features.

  3. Scaling Group Inference for Diverse and High-Quality Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Groups of generated images become more diverse while staying high-quality when K outputs are chosen from M candidates via a quadratic integer program with progressive pruning.

Pith tools