Pith. sign in

REVIEW 5 cited by

Divide & Bind Your Attention for Improved Generative Semantic Nursing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.10864 v3 pith:WO5JW54V submitted 2023-07-20 cs.CV cs.AIcs.CLcs.LG

Divide & Bind Your Attention for Improved Generative Semantic Nursing

classification cs.CV cs.AIcs.CLcs.LG
keywords promptsattributebindingcomplexgenerativeimprovedlossaddress
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Emerging large-scale text-to-image generative models, e.g., Stable Diffusion (SD), have exhibited overwhelming results with high fidelity. Despite the magnificent progress, current state-of-the-art models still struggle to generate images fully adhering to the input prompt. Prior work, Attend & Excite, has introduced the concept of Generative Semantic Nursing (GSN), aiming to optimize cross-attention during inference time to better incorporate the semantics. It demonstrates promising results in generating simple prompts, e.g., "a cat and a dog". However, its efficacy declines when dealing with more complex prompts, and it does not explicitly address the problem of improper attribute binding. To address the challenges posed by complex prompts or scenarios involving multiple entities and to achieve improved attribute binding, we propose Divide & Bind. We introduce two novel loss objectives for GSN: a novel attendance loss and a binding loss. Our approach stands out in its ability to faithfully synthesize desired objects with improved attribute alignment from complex prompts and exhibits superior performance across multiple evaluation benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

    cs.CV 2024-03 unverdicted novelty 7.0

    ELLA introduces a timestep-aware semantic connector to link LLMs with diffusion models for improved dense prompt following, validated on a new 1K-prompt benchmark.

  2. AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis

    cs.CV 2026-05 unverdicted novelty 6.0

    AI-T2I improves text-to-image alignment in diffusion models by using aggregation and isolation losses on cross-attention maps to fix scattering and overlap issues.

  3. BindEdit: Taming Attention Leakage for Precise Multi-Object Image Editing

    cs.CV 2026-06 unverdicted novelty 5.0

    BindEdit suppresses two forms of attention leakage in diffusion-based editing by binding target tokens to regions, rebalancing cross-attention, and adding a region fidelity term, plus a new multi-object benchmark.

  4. STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model

    cs.CV 2026-06 unverdicted novelty 5.0

    STEDiff improves semantic alignment in text-to-image diffusion models via training-free embedding strengthening with the [EOT] token and a spatial semantic loss, showing gains on T2I-CompBench.

  5. DebFilter: Eradicating Biases Stashed in Value

    cs.CV 2026-05 unverdicted novelty 4.0

    DebFilter mitigates biases in text-to-image diffusion models by applying a fixed offset to the guidance embedding slice in cross-attention during inference.