Pith. sign in

REVIEW 8 cited by

Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.16990 v1 pith:TXTJP6RJ submitted 2024-03-25 cs.CV cs.AIcs.GRcs.LG

classification cs.CVcs.AIcs.GRcs.LG
keywords subjectsattentionboundedgenerationleakagemultiplecomplexdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image diffusion models have an unprecedented ability to generate diverse and high-quality images. However, they often struggle to faithfully capture the intended semantics of complex input prompts that include multiple subjects. Recently, numerous layout-to-image extensions have been introduced to improve user control, aiming to localize subjects represented by specific tokens. Yet, these methods often produce semantically inaccurate images, especially when dealing with multiple semantically or visually similar subjects. In this work, we study and analyze the causes of these limitations. Our exploration reveals that the primary issue stems from inadvertent semantic leakage between subjects in the denoising process. This leakage is attributed to the diffusion model's attention layers, which tend to blend the visual features of different subjects. To address these issues, we introduce Bounded Attention, a training-free method for bounding the information flow in the sampling process. Bounded Attention prevents detrimental leakage among subjects and enables guiding the generation to promote each subject's individuality, even with complex multi-subject conditioning. Through extensive experimentation, we demonstrate that our method empowers the generation of multiple subjects that better align with given prompts and layouts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TACA scales cross-modal attention logits by a timestep-dependent temperature to rebalance text and visual tokens, improving T2I-CompBench alignment on FLUX and SD3.5.

  2. Multitwine: Multi-Object Compositing with Text and Layout Control

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single diffusion model simultaneously composites multiple objects into a scene with text and layout control, outperforming sequential insertion on interacting cases.

  3. PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An inversion-free, training-free diffusion editing method that anchors output latents to a pixel-manipulated copy of the image achieves consistent object repositioning, resizing, and pasting in 16 steps.

  4. All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Certain random seeds yield consistently more accurate compositional text-to-image outputs, and mining these seeds plus fine-tuning on the resulting self-generated images improves numerical and spatial composition accuracy.

  5. Stable Flow: Vital Layers for Training-Free Image Editing

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.

  6. Scaling Down Semantic Leakage: Investigating Associative Bias in Smaller Language Models

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Tests four Qwen2.5 models and finds the 0.5B model shows the least semantic leakage, but the size trend is non-monotonic.

  7. LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation

    cs.CV 2024-11 conditional novelty 4.0 of 10

    LocRef-Diffusion inserts two lightweight cross-attention modules into Stable Diffusion to control both layout and appearance of multiple instances without fine-tuning per object.

  8. Text-to-Image Synthesis: A Decade Survey

    cs.CV 2024-11 conditional novelty 1.0 of 10

    A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.

Pith tools