MaskAttn-SDXL adds learned binary token-location gates before softmax in SDXL cross-attention, improving compositional consistency on multi-object prompts.
Improving composi- tional attribute binding in text-to-image generative mod- els via enhanced text embeddings
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation
MaskAttn-SDXL adds learned binary token-location gates before softmax in SDXL cross-attention, improving compositional consistency on multi-object prompts.