REVIEW 8 cited by
Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-image diffusion models have an unprecedented ability to generate diverse and high-quality images. However, they often struggle to faithfully capture the intended semantics of complex input prompts that include multiple subjects. Recently, numerous layout-to-image extensions have been introduced to improve user control, aiming to localize subjects represented by specific tokens. Yet, these methods often produce semantically inaccurate images, especially when dealing with multiple semantically or visually similar subjects. In this work, we study and analyze the causes of these limitations. Our exploration reveals that the primary issue stems from inadvertent semantic leakage between subjects in the denoising process. This leakage is attributed to the diffusion model's attention layers, which tend to blend the visual features of different subjects. To address these issues, we introduce Bounded Attention, a training-free method for bounding the information flow in the sampling process. Bounded Attention prevents detrimental leakage among subjects and enables guiding the generation to promote each subject's individuality, even with complex multi-subject conditioning. Through extensive experimentation, we demonstrate that our method empowers the generation of multiple subjects that better align with given prompts and layouts.
Forward citations
Cited by 8 Pith papers
-
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
TACA scales cross-modal attention logits by a timestep-dependent temperature to rebalance text and visual tokens, improving T2I-CompBench alignment on FLUX and SD3.5.
-
Multitwine: Multi-Object Compositing with Text and Layout Control
A single diffusion model simultaneously composites multiple objects into a scene with text and layout control, outperforming sequential insertion on interacting cases.
-
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
An inversion-free, training-free diffusion editing method that anchors output latents to a pixel-manipulated copy of the image achieves consistent object repositioning, resizing, and pasting in 16 steps.
-
All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds
Certain random seeds yield consistently more accurate compositional text-to-image outputs, and mining these seeds plus fine-tuning on the resulting self-generated images improves numerical and spatial composition accuracy.
-
Stable Flow: Vital Layers for Training-Free Image Editing
An automatic vital-layer selection for FLUX enables training-free, stable text-driven image editing via selective attention injection.
-
Scaling Down Semantic Leakage: Investigating Associative Bias in Smaller Language Models
Tests four Qwen2.5 models and finds the 0.5B model shows the least semantic leakage, but the size trend is non-monotonic.
-
LocRef-Diffusion:Tuning-Free Layout and Appearance-Guided Generation
LocRef-Diffusion inserts two lightweight cross-attention modules into Stable Diffusion to control both layout and appearance of multiple instances without fine-tuning per object.
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.