Pith. sign in

REVIEW 4 major objections 7 minor 1 cited by

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Detail++ claims that training-free progressive detail injection fixes attribute binding in text-to-image diffusion models.

desk verdict A plausible training-free multi-branch attention method that improves attribute binding, but the reported evidence lacks error bars and a mask-quality check, so the 'significant outperformance' claim is not yet grounded. read the letter →

arxiv 2507.17853 v3 pith:4NO7LAZJ submitted 2025-07-23 cs.CV cs.AI

classification cs.CVcs.AI
keywords text-to-imagediffusionattributebindingtraining-freeenhancementprogressivedetailinjectioncross-attentionmasksself-attentionsharingcentroidalignmentlossstylecomposition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Detail++ claims that the attribute-binding failures of text-to-image diffusion models can be fixed without any training by generating a complex prompt in stages. A plain, attribute-free version of the prompt fixes the layout first, and then each additional attribute is injected only into the region belonging to its subject, using a binary mask read off the subject token's cross-attention map. A test-time centroid-alignment loss sharpens those masks so attributes do not spill across subjects. The paper reports that this training-free plugin improves color, texture, and style binding over strong baselines on SDXL, SD3, and Flux, and that the gains grow as the number of attributes increases.

What carries the argument

The load-bearing mechanism is the pairing of shared self-attention maps with Accumulative Latent Modification (ALM). Self-attention maps are treated as a layout blueprint: the maps computed in the full-prompt branch are reused by all sub-prompt branches for the first $S$ denoising steps, keeping the spatial composition identical across branches. For each subject $q_i$, the averaged cross-attention map $M_i$ is normalized and thresholded to a binary mask $B_i$, and the update $z^{t-1}_{i+1} = z^{t-1}_i + B_i \odot(\tilde{z}^{t-1}_{i+1} - z^{t-1}_i)$ copies the new attribute's latent only where the mask is one. The Centroid Alignment Loss $L_{\text{align}} = \sum_i \| p_{\text{centroid}}(q_i) - p_{\max}(q_i)\|_2$, combined with an entropy term, is minimized over the latent at test time to concentrate each subject's attention and thereby improve the masks.

What would settle it

Generate a large set of prompts whose subjects overlap heavily or contain hollow objects, and compare, per seed, the binary mask from the subject token against human- or detector-labeled subject regions. If the mask's overlap with the true region is low, or if per-region attribute scores show the attribute landing on the neighbor or inside the hole in a substantial fraction of attempts, the central claim that attention-derived masks localize injected attributes is refuted.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that progressive, region-restricted attribute injection solves detail binding in complex prompts. Starting from the same noise latent, Detail++ runs several branches in parallel: one branch uses the full prompt, one uses the prompt with all modifiers removed, and each remaining branch re-adds a single attribute. Self-attention maps from the full-prompt branch are shared with the other branches during the early denoising steps so every branch commits to the same layout. At each of those steps, the latent of branch $i+1$ is combined with the latent of branch $i$ only inside the binary mask of the subject that owns the new attribute, so the attribute is written into the right region and nowhere else. The paper further claims that a centroid-alignment loss, which pulls the brightest point of each subject's cross-attention map toward its centroid, makes the masks accurate enough that this training-free procedure outperforms existing methods on multiple-attribute and multi-style prompts.

Load-bearing premise

The whole pipeline stands on the assumption that a subject token's cross-attention map, once averaged and thresholded, is a reliable stencil of where that subject is in the image; when the stencil is blurred, overlapping, or misplaced, the injected attribute leaks onto the wrong object.

Editorial extensions

If this is right

  • Because the method is training-free and operates on attention maps, it can be dropped onto already-deployed U-Net and DiT diffusion models and improve binding on prompts they already handle.
  • Attribute-count robustness: the reported scores stay nearly flat as prompts go from two to four attributes, whereas baseline methods degrade noticeably.
  • Style composition becomes separable: a Lego-style foreground and an oil-painting background can be generated without the styles blending, as demonstrated qualitatively and on the proposed style-composition benchmark.
  • The branch scheduler and cached self-attention keep the multi-branch overhead small, so the benefit does not require a large compute budget.
  • The same machinery extends to DiT backbones by partitioning multimodal attention into image-to-image and text-to-image blocks, so the method is not tied to U-Net architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • This suggests a general design principle: parallel decoding of all attributes at once is what causes binding errors, so a sequential, factorized generation schedule may help even without attention-mask injection.
  • A straightforward extension would apply the same mask-and-copy update to text-guided image editing, where one attribute must be changed on one object while everything else stays fixed.
  • The style-composition scoring recipe introduced here, detect components, crop, and CLIP-score each against its style descriptor, could serve as a reusable automatic metric for style disentanglement beyond this benchmark.
  • A testable variant would replace single-token attention masks with multi-token or external-segmenter masks for hollow and heavily overlapping subjects, the failure cases the paper itself reports.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. Detail++ proposes a training-free, multi-branch method for improving attribute binding in text-to-image diffusion models. It decomposes a complex prompt into a set of sub-prompts, shares self-attention maps from a full-prompt branch to maintain layout, and progressively injects attributes using binary masks extracted from cross-attention maps (Eqs. 3-4). A test-time Centroid Alignment Loss combined with an entropy loss is used to sharpen the masks. The method is implemented on SDXL, SD3, and Flux.1-schnell and evaluated on T2I-CompBench++ and a newly introduced Style Composition Benchmark (SCB). The paper reports consistent improvements over strong baselines, especially on color and texture binding, along with ablations for the propagation step S, the mask threshold tau, and the loss terms.

Significance. The result, if robust, is significant: a training-free plugin that improves attribute binding across U-Net- and DiT-based backbones would be practically valuable, and the proposed efficient self-attention propagation makes multi-branch inference more affordable. The manuscript's strengths include evaluation on three backbones, ablations for the main hyperparameters, a user study, and unusually candid documentation of failure modes (Sec. 7.1, Fig. A10). The method does not require training or predefined layouts, and the core mechanism is clearly explained. However, the headline quantitative claims are not yet supported at the level needed for publication: the main comparison lacks uncertainty quantification, a key component's reliability is not directly measured, and one efficiency comparison is confounded by checkpoint and step-count differences.

major comments (4)
  1. [Sec. 5.2, Table 1] The headline claim that Detail++ "significantly outperforms existing methods" is not statistically supported. Table 1 reports point estimates only; the ablation section says five random seeds are used for ablation experiments, but no standard deviations, confidence intervals, or paired significance tests are reported for the main comparison. Moreover, the key hyperparameters S (Fig. 13 and Table 5), tau (Fig. 12), and lambda (Sec. 5.3) are selected on the same T2I-CompBench++ test subsets used to produce Table 1. Please report per-seed means and variances for all methods, run paired tests against the strongest baselines (at least R-Bind, T2I-R1, and TACA), and select hyperparameters on a separate validation split or otherwise show that the conclusions are stable across the hyperparameter plateau.
  2. [Sec. 4.2, Eqs. (3)-(4), Algorithm 1] The binary cross-attention masks B_i are the load-bearing component that decides where each attribute is written. The manuscript itself documents failure modes: Sec. 7.1 admits that early layout errors cannot be corrected by later injection, and Fig. A10 shows that the centroid alignment loss deforms a hollow wreath and that background attention maps can contain salt-and-pepper noise. Yet no quantitative evaluation of mask quality (e.g., IoU against reference segmentations, overlap between subject masks) is provided, and there is no oracle-mask ablation that would bound how much of the reported gain depends on mask fidelity. I request two additions: (i) mask-quality statistics on a sample of T2I-CompBench++ and SCB prompts, and (ii) an oracle variant in which masks are replaced by off-the-shelf detector boxes or segmentations. This directly tests the concern that diffuse or misplaced masks could leak attributes into neighboring subjects.
  3. [Sec. 5.2, Table 2 and Fig. 8] The efficiency comparison for the Flux variant is confounded by checkpoint and step-count differences. The text states that Detail++(Flux) runs on the 8-step distilled Flux.1-schnell, while the Flux baseline and other methods use a larger number of steps; the manuscript does not state the Flux baseline step count. The reported large negative time overhead for Detail++(Flux) therefore largely reflects distillation rather than the method's efficiency. Please compare Detail++(Flux) against Flux.1-schnell with the same step count, and report overhead relative to that matched baseline; similarly ensure the SDXL and SD3 comparisons use identical step counts, resolution, and precision settings for the method and its baseline.
  4. [Sec. 7.4 and Table 1 (SCB)] The SCB benchmark is introduced in this paper and is used to support a central claim about style composition. Its metric is a pipeline of Grounding DINO cropping plus CLIP scoring, whose failure modes (e.g., missed detections, imperfect crops) are not analyzed, and no evidence is given that the metric correlates with human judgment on the full 1,000-prompt set. The user study covers only 12 style prompts. Please (i) release the SCB prompts and evaluation code, (ii) report human correlation on a random subset of SCB, and (iii) include negative-control prompts or compare against an external style-composition evaluation to validate the metric. Without this, the style-composition superiority claim rests on a self-designed, unvalidated measure.
minor comments (7)
  1. [Algorithm 1] Line 7 contains the typo "bianry" for "binary"; the variable name "d z^{t-1}_{i+1}" is also confusing because it is first assigned the value of z^{t-1}_{i+1} and then overwritten.
  2. [Sec. 4.3, Eq. (6)-(9)] The coordinate convention in Eq. (6) should be clarified (w and h are used as spatial coordinates but not defined against the latent grid), and the implementation of Eq. (9) needs the number of gradient steps and the schedule of alpha_t for reproducibility.
  3. [Sec. 4.4] There is a typo "prohressive" that should be "progressive".
  4. [Sec. 5.4] In the user study, TACA is cited as [54], but [54] is the DiT editing paper by Shin et al.; the TACA method appears to be reference [39]. Please correct the citation.
  5. [Sec. 5.3, Table 4] The branch scheduler's quality drop (1.42% for SDXL and 1.33% for Flux) is reported without variance; given the small effect, report the per-seed spread to show it is within noise.
  6. [Table 1] The color coding for first/second/third highest scores is not accessible in grayscale or for color-blind readers; also, the empty SCB entries for Ranni should be explained in the table caption.
  7. [Sec. 7.2, Quantitative Results Analysis] The paper honestly notes that Detail++ has limited capacity to modify shape because self-attention maps are shared across branches; this limitation should also appear in the main paper's conclusion, since the abstract's "significantly outperforms" is not supported on the shape subset.

Circularity Check

1 steps flagged · score 3.0 of 10

Mild evaluation circularity: S and tau are selected on the T2I-CompBench++ test split whose scores are then reported as superiority evidence; the injection equations themselves are not circular.

  1. fitted input called prediction [Sec. 5.3; Supp. Sec. 7.2; Table 1]
    "Generally, the ablation studies in Sec. V-C adopt the evaluation metrics and test splits defined in T2I-CompBench++ [23]. We evaluate on its three standard attribute-binding subsets separately, namely color, texture, and shape. For each subset, we use 300 prompts taken from the test portion of T2I-CompBench++... We therefore set S at the saturation point for each backbone to achieve the best quality–efficiency trade-off, adopting S=0.25T for DiT-based models and S=0.5T for SDXL... we adopt τ=0.5 as a robust default across all backbones."

    The same test split used for ablations (Supp. Sec. 7.2, 300 prompts per color/texture/shape subset) is the basis of the headline T2I-CompBench++ scores in Table 1, and the operating points S and τ are explicitly chosen from those ablations ('set S at the saturation point', 'adopt τ=0.5'). Thus the benchmark superiority claim is not a prediction of a fixed method with pre-specified settings; it is the performance of settings that were selected by maximizing the same metric on the same prompts. This is a mild statistical-circularity burden on the 'significantly outperforms' claim, although the core mask-injection mechanism (Eqs. 3-4) is not definitionally tied to the evaluation metrics.

full rationale

No derivation-level circularity is present: Eq. (3) defines binary masks by normalizing and thresholding cross-attention maps, and Eq. (4) gates latent replacement with those masks; neither equation is defined in terms of BLIP-VQA, ImageReward, or the SCB CLIP scores, so the reported improvements are not true by construction. The centroid-alignment loss (Eq. 7) is a test-time objective on attention concentration, and the ablation in Table 3 and Figs. 10-13 empirically test its contribution rather than assuming it. The layout-sharing premise relies on external prior work cited as [5, 35, 41, 55, 54], not on a self-citation chain, and the paper honestly documents failure modes in Sec. 7.1 and Fig. A10 where masks or early layouts are inaccurate, which would be impossible if the outcome were definitionally forced. The only self-citation is [29] in the Introduction naming art-design applications; it is background and not load-bearing. The main circularity burden is evaluative rather than mathematical: S, τ, and the selected operating points are tuned on the same T2I-CompBench++ test portion whose scores are then presented as evidence of state-of-the-art performance, and the style benchmark is author-constructed. These practices weaken the external validity of the headline comparison but do not make the attribute-injection derivation circular.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The method rests on heuristic assumptions about attention maps rather than on proven mathematical statements. Its free parameters are empirical design choices, and some are tuned on the evaluation split. No new physical or conceptual entities are introduced.

free parameters (5)
  • tau_binarization_threshold = 0.5 (SDXL and Flux), 0.6 (SD3)
    Threshold in Eq. 3 for masking subject regions; selected from ablation on T2I-CompBench++ test set (Fig. 12).
  • lambda_entropy_weight = 1
    Weight for the entropy loss in L_total; chosen for balance, with no sensitivity analysis reported.
  • SA_propagation_steps_S = 0.5T for SDXL, 0.25T for SD3 and Flux
    Number of initial denoising steps with shared self-attention maps; selected from test-set ablations (Fig. 13, Table 5).
  • cross_attention_mask_layers = all blocks at 32x32 for SDXL; top-5 high-activation layers for DiT
    Layer selection for cross-attention mask extraction, described in Supplementary Sec. 7.7, affects mask reliability.
  • mask_extraction_blur_parameters = not specified
    Gaussian blur kernel size and standard deviation are referenced as mitigations in Fig. A10, but fixed values are not given, making exact reproduction difficult.
assumptions (5)
  • domain assumption Self-attention maps encode layout and can be shared across different sub-prompts to keep the same composition.
    Invoked in Sec. 4.2 for the shared self-attention map mechanism; taken from the editing literature [5,35,41,55] rather than derived.
  • domain assumption Thresholded cross-attention maps of subject tokens provide reliable subject masks.
    Used in Eq. 3 for Accumulative Latent Modification; the paper's Fig. A10 shows concrete failures of this assumption.
  • domain assumption Complex prompts can be decomposed into independent attribute branches without changing global semantics.
    The LLM-based prompt decomposition in Sec. 4.2 assumes attributes are separable; this may fail for interacting or relational attributes.
  • domain assumption Early denoising steps dominate layout, so sharing SA maps only for the first S steps is sufficient.
    Stated in Sec. 4.2; supported empirically by Fig. 13 but not proven.
  • domain assumption Centroid alignment makes cross-attention maps more focused and therefore masks more accurate.
    Sec. 4.3 introduces this loss; Fig. A10 shows it can be harmful for hollow subjects, so the assumption does not always hold.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Detail++: Training-Free Detail Enhancer for T2I Diffusion Models." pith.science (2026). https://pith.science/paper/4NO7LAZJ

@misc{pith2026250717853,
  author       = {Pith},
  title        = {Pith review of: Detail++: Training-Free Detail Enhancer for T2I Diffusion Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4NO7LAZJ}},
  note         = {Machine review of arXiv:2507.17853}
}
read the original abstract

Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, particularly those involving multiple subjects with distinct attributes. Inspired by the human drawing process, which first outlines the composition and then incrementally adds details, we propose Detail++, a training-free framework that introduces a novel Progressive Detail Injection (PDI) strategy to address this limitation. Specifically, we decompose a complex prompt into a sequence of simplified sub-prompts, guiding the generation process in stages. This staged generation leverages the inherent layout-controlling capacity of self-attention to first ensure global composition, followed by precise refinement. To achieve accurate binding between attributes and corresponding subjects, we exploit cross-attention mechanisms and further introduce a Centroid Alignment Loss at test time to reduce binding noise and enhance attribute consistency. Extensive experiments on T2I-CompBench and a newly constructed style composition benchmark demonstrate that Detail++ significantly outperforms existing methods, particularly in scenarios involving multiple objects and complex stylistic conditions.

Figures

Figures reproduced from arXiv: 2507.17853 by the authors.

Figure 1
Figure 1. A comparison between our method and current state-of-the-art generative models. The mainstream models often suffer from issues such as semantic overflow, complex attribute mismatching, and style blending. Even Flux, the leading generative model under the DiT framework, struggles to overcome these challenges. In contrast, our method, Detail++, based on SDXL, achieves highly accurate semantic binding in a training-fre… view at source ↗
Figure 2
Figure 2. The basic process of Detail++. As shown in the first column, generating complex prompts in a single branch often results in inaccurate or blended attribute assignments. For example, attributes such as “sunglasses” and “necklace” may be mistakenly applied to the wrong subject. Our method addresses this challenge through a progressive approach: we first ignore all complex modifiers to produce a rough generation base, … view at source ↗
Figure 3
Figure 3. Attention visualization. The visualization of the self￾attention map and cross cross-attention map of the prompt “man stand in front of car”. The self-attention map is visualized by displaying the top-6 components obtained after SVD [57]. The cross-attention map visualizations correspond to each token in the prompt. image feature space [35]. This attention map can be repre￾sented as: Mself = Softmax QimgK⊤ img √ dk … view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Overview of Detail++. 1) Firstly, all sub-prompts are passed to U-net for generation at the same time. In each denoising step, we first share the SA map of the first branch with all other branches to ensure a consistent layout base. Then, for each timestep generated la…
Figure 5
Figure 5. Figure 5: Efficient self-attention propagation for multi-branch [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparison of various methods on complex prompts with multiple attribute types (object, color, texture, style and [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Quantitative comparison of baselines under different numbers of attributes. Detail++ (Flux) and Detail++ (SD3) exhibit near [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Comparison of inference time and average VRAM usage across different methods. With the efficient self-attention propagation [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Qualitative ablation on the self-attention map propagation step [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Ablation study of different optimization terms in bi [PITH_FULL_IMAGE:figures/full_fig_p010_10.png]
Figure 11
Figure 11. Figure 11: Visual results of different optimization strategies for the [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: Ablation study of the mask binarization threshold [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: Ablation study of the self-attention propagation step [PITH_FULL_IMAGE:figures/full_fig_p012_13.png]
Figure 14
Figure 14. Figure 14: User study results evaluating different methods across [PITH_FULL_IMAGE:figures/full_fig_p013_14.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Model

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    STEDiff improves semantic alignment in text-to-image diffusion models via training-free embedding strengthening with the [EOT] token and a spatial semantic loss, showing gains on T2I-CompBench.

Reference graph

Works this paper leans on

83 extracted references · 51 canonical work pages · cited by 1 Pith paper

  1. [1]

    A-star: Test-time attention segregation and retention for text-to-image synthesis

    Aishwarya Agarwal, Srikrishna Karanam, KJ Joseph, Apoorv Saxena, Koustava Goswami, and Balaji Vasan Srini- vasan. A-star: Test-time attention segregation and retention for text-to-image synthesis. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 2283– 2293, 2023. 3

  2. [2]

    ediff-i: Text-to-image dif- fusion models with an ensemble of expert denoisers.arXiv preprint arXiv:2211.01324, 2022

    Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Ji- aming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, et al. ediff-i: Text-to-image dif- fusion models with an ensemble of expert denoisers.arXiv preprint arXiv:2211.01324, 2022. 3

  3. [3]

    In- structpix2pix: Learning to follow image editing instructions

    Tim Brooks, Aleksander Holynski, and Alexei A Efros. In- structpix2pix: Learning to follow image editing instructions. InProceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 18392–18402, 2023. 3

  4. [4]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz...

  5. [5]

    Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing

    Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xi- aohu Qie, and Yinqiang Zheng. Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing. InProceedings of the IEEE/CVF international con- ference on computer vision, pages 22560–22570, 2023. 3, 4

  6. [6]

    Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.ACM transactions on Graphics (TOG), 42(4):1–10, 2023

    Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.ACM transactions on Graphics (TOG), 42(4):1–10, 2023. 3

  7. [7]

    A cat is a cat (not a dog!): Unravel- ing information mix-ups in text-to-image encoders through causal analysis and embedding optimization.arXiv preprint arXiv:2410.00321, 2024

    Chieh-Yun Chen, Chiang Tseng, Li-Wu Tsao, and Hong- Han Shuai. A cat is a cat (not a dog!): Unravel- ing information mix-ups in text-to-image encoders through causal analysis and embedding optimization.arXiv preprint arXiv:2410.00321, 2024. 3

  8. [8]

    Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis.arXiv preprint arXiv:2310.00426, 2023

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis.arXiv preprint arXiv:2310.00426, 2023. 8, 1

Show all 83 references
  1. [9]

    Geodiffusion: Text- prompted geometric control for object detection data gen- eration.arXiv preprint arXiv:2306.04607, 2023

    Kai Chen, Enze Xie, Zhe Chen, Yibo Wang, Lanqing Hong, Zhenguo Li, and Dit-Yan Yeung. Geodiffusion: Text- prompted geometric control for object detection data gen- eration.arXiv preprint arXiv:2306.04607, 2023. 3

  2. [10]

    Visual pro- gramming for step-by-step text-to-image generation and evaluation.Advances in Neural Information Processing Sys- tems, 36:6048–6069, 2023

    Jaemin Cho, Abhay Zala, and Mohit Bansal. Visual pro- gramming for step-by-step text-to-image generation and evaluation.Advances in Neural Information Processing Sys- tems, 36:6048–6069, 2023. 3

  3. [11]

    What’s in a text-to-image prompt? the potential of stable diffusion in vi- sual arts education.Heliyon, 9(6), 2023

    Nassim Dehouche and Kullathida Dehouche. What’s in a text-to-image prompt? the potential of stable diffusion in vi- sual arts education.Heliyon, 9(6), 2023. 1

  4. [12]

    Diffusion models beat gans on image synthesis.Advances in neural informa- tion processing systems, 34:8780–8794, 2021

    Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis.Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 3

  5. [13]

    Scaling recti- fied flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. InForty-first international conference on machi...

  6. [14]

    Training-free structured diffusion guidance for compositional text-to-image synthesis.arXiv preprint arXiv:2212.05032, 2022

    Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang. Training-free structured diffusion guidance for compositional text-to-image synthesis.arXiv preprint arXiv:2212.05032, 2022. 3, 8, 11, 1

  7. [15]

    Ranni: Taming text-to-image diffu- sion for accurate instruction following

    Yutong Feng, Biao Gong, Di Chen, Yujun Shen, Yu Liu, and Jingren Zhou. Ranni: Taming text-to-image diffu- sion for accurate instruction following. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4744–4753, 2024. 2, 3, 8, 9, 1

  8. [16]

    Stylegan-nada: Clip- guided domain adaptation of image generators.ACM Trans- actions on Graphics (TOG), 41(4):1–13, 2022

    Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Stylegan-nada: Clip- guided domain adaptation of image generators.ACM Trans- actions on Graphics (TOG), 41(4):1–13, 2022. 3

  9. [17]

    Expressive text-to-image generation with rich text

    Songwei Ge, Taesung Park, Jun-Yan Zhu, and Jia-Bin Huang. Expressive text-to-image generation with rich text. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 7545–7556, 2023. 3

  10. [18]

    Prompt-to-prompt im- age editing with cross attention control.arXiv preprint arXiv:2208.01626, 2022

    Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt im- age editing with cross attention control.arXiv preprint arXiv:2208.01626, 2022. 3, 4 14

  11. [19]

    Style aligned image generation via shared atten- tion

    Amir Hertz, Andrey V oynov, Shlomi Fruchter, and Daniel Cohen-Or. Style aligned image generation via shared atten- tion. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 4775–4785,

  12. [20]

    spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing.(No Title), 2017

    Matthew Honnibal. spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing.(No Title), 2017. 4, 8

  13. [21]

    Token merging for training- free semantic binding in text-to-image synthesis.Advances in Neural Information Processing Systems, 37:137646– 137672, 2025

    Taihang Hu, Linxuan Li, Joost van de Weijer, Hongcheng Gao, Fahad Shahbaz Khan, Jian Yang, Ming-Ming Cheng, Kai Wang, and Yaxing Wang. Token merging for training- free semantic binding in text-to-image synthesis.Advances in Neural Information Processing Systems, 37:137646– 137...

  14. [22]

    Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135, 2024

    Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. Ella: Equip diffusion models with llm for enhanced semantic alignment.arXiv preprint arXiv:2403.05135, 2024. 2, 3, 8, 9, 1

  15. [23]

    T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.Advances in Neural Information Processing Systems, 36:78723–78747, 2023

    Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.Advances in Neural Information Processing Systems, 36:78723–78747, 2023. 2, 7, 8, 10, 11, 1

  16. [24]

    T2i-r1: Reinforcing image generation with col- laborative semantic-level and token-level cot.arXiv preprint arXiv:2505.00703, 2025

    Dongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong, Hao Li, Le Zhuo, Shilin Yan, Pheng-Ann Heng, and Hong- sheng Li. T2i-r1: Reinforcing image generation with col- laborative semantic-level and token-level cot.arXiv preprint arXiv:2505.00703, 2025. 3, 8, 9, 11, 13, 1

  17. [25]

    Comat: Aligning text-to-image diffusion model with image- to-text concept matching.Advances in Neural Information Processing Systems, 37:76177–76209, 2024

    Dongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang, Dazhong Shen, Zhuofan Zong, Yu Liu, and Hongsheng Li. Comat: Aligning text-to-image diffusion model with image- to-text concept matching.Advances in Neural Information Processing Systems, 37:76177–76209, 2024. 2, 3

  18. [26]

    Diffu- sionclip: Text-guided diffusion models for robust image ma- nipulation

    Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffu- sionclip: Text-guided diffusion models for robust image ma- nipulation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2426–2435,

  19. [27]

    Text embed- ding is not all you need: Attention control for text-to-image semantic alignment with text self-attention maps.arXiv preprint arXiv:2411.15236, 2024

    Jeeyung Kim, Erfan Esmaeili, and Qiang Qiu. Text embed- ding is not all you need: Attention control for text-to-image semantic alignment with text self-attention maps.arXiv preprint arXiv:2411.15236, 2024. 3

  20. [28]

    Flux.https://github.com/bla ck-forest-labs/flux, 2024

    Black Forest Labs. Flux.https://github.com/bla ck-forest-labs/flux, 2024. 1, 3, 6, 8, 9, 12

  21. [29]

    Stylestudio: Text-driven style transfer with selective control of style elements.arXiv preprint arXiv:2412.08503,

    Mingkun Lei, Xue Song, Beier Zhu, Hao Wang, and Chi Zhang. Stylestudio: Text-driven style transfer with selective control of style elements.arXiv preprint arXiv:2412.08503,

  22. [30]

    Manigan: Text-guided image manipulation

    Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip HS Torr. Manigan: Text-guided image manipulation. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7880–7889, 2020. 3

  23. [31]

    Unbounded: A generative infinite game of character life simulation.arXiv preprint arXiv:2410.18975, 2024

    Jialu Li, Yuanzhen Li, Neal Wadhwa, Yael Pritch, David E Jacobs, Michael Rubinstein, Mohit Bansal, and Nataniel Ruiz. Unbounded: A generative infinite game of character life simulation.arXiv preprint arXiv:2410.18975, 2024. 1

  24. [32]

    Gligen: Open-set grounded text-to-image generation

    Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jian- wei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. Gligen: Open-set grounded text-to-image generation. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22511–22521, 2023. 2

  25. [33]

    Divide & bind your attention for improved generative seman- tic nursing.arXiv preprint arXiv:2307.10864, 2023

    Yumeng Li, Margret Keuper, Dan Zhang, and Anna Khoreva. Divide & bind your attention for improved generative seman- tic nursing.arXiv preprint arXiv:2307.10864, 2023. 3

  26. [34]

    Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

    Long Lian, Boyi Li, Adam Yala, and Trevor Darrell. Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models. arXiv preprint arXiv:2305.13655, 2023. 2, 3

  27. [35]

    Towards understanding cross and self-attention in stable diffusion for text-guided image editing

    Bingyan Liu, Chengyu Wang, Tingfeng Cao, Kui Jia, and Jun Huang. Towards understanding cross and self-attention in stable diffusion for text-guided image editing. InProceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7817–7826, 2024. 3, 4, 5

  28. [36]

    Compositional visual generation with composable diffusion models

    Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. Compositional visual generation with composable diffusion models. InEuropean Conference on Computer Vision, pages 423–439. Springer, 2022. 8, 1

  29. [37]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision, pages 38–55. Springer...

  30. [38]

    Flow straight and fast: Learning to generate and transfer data with rectified flow

    Xingchao Liu, Chengyue Gong, et al. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Repre- sentations, 2023. 12

  31. [39]

    Re- thinking cross-modal interaction in multimodal diffusion transformers.arXiv preprint arXiv:2506.07986, 2025

    Zhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen, Wangmeng Zuo, Ziwei Liu, and Kwan-Yee K Wong. Re- thinking cross-modal interaction in multimodal diffusion transformers.arXiv preprint arXiv:2506.07986, 2025. 3, 8, 9, 11, 1

  32. [40]

    Sdedit: Guided image synthesis and editing with stochastic differential equa- tions.arXiv preprint arXiv:2108.01073, 2021

    Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jia- jun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Guided image synthesis and editing with stochastic differential equa- tions.arXiv preprint arXiv:2108.01073, 2021. 3

  33. [41]

    Dreammatcher: appearance matching self-attention for semantically-consistent text-to- image personalization

    Jisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin, Seungry- ong Kim, and Seunggyu Chang. Dreammatcher: appearance matching self-attention for semantically-consistent text-to- image personalization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,...

  34. [42]

    Text- adaptive generative adversarial networks: manipulating im- ages with natural language.Advances in neural information processing systems, 31, 2018

    Seonghyeon Nam, Yunji Kim, and Seon Joo Kim. Text- adaptive generative adversarial networks: manipulating im- ages with natural language.Advances in neural information processing systems, 31, 2018. 3

  35. [43]

    Energy-based cross attention for bayesian context update in text-to-image diffusion mod- els.Advances in Neural Information Processing Systems, 36:76382–76408, 2023

    Geon Yeong Park, Jeongsol Kim, Beomsu Kim, Sang Wan Lee, and Jong Chul Ye. Energy-based cross attention for bayesian context update in text-to-image diffusion mod- els.Advances in Neural Information Processing Systems, 36:76382–76408, 2023. 3

  36. [44]

    Styleclip: Text-driven manipulation of stylegan imagery

    Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. InProceedings of the IEEE/CVF inter- national conference on computer vision, pages 2085–2094,

  37. [45]

    Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis.arXiv preprint arXiv:2307.01952, 2023

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis.arXiv preprint arXiv:2307.01952, 2023. 1, 3, 8, 9, 11, 12

  38. [46]

    Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation

    Leigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie, and Tat- Seng Chua. Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation. InProceedings of the 31st ACM International Conference on Multimedia, pages 643– 654, 2023. 3

  39. [47]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. InInternational conference on machine learning, p...

  40. [48]

    Hierarchical text-conditional image gen- eration with clip latents.arXiv preprint arXiv:2204.06125, 1(2):3, 2022

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gen- eration with clip latents.arXiv preprint arXiv:2204.06125, 1(2):3, 2022. 3

  41. [49]

    Zero-shot text-to-image generation

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. InInternational confer- ence on machine learning, pages 8821–8831. Pmlr, 2021. 3

  42. [50]

    Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.Advances in Neural Infor- mation Processing Systems, 36:3536–3559, 2023

    Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Rav- fogel, Yoav Goldberg, and Gal Chechik. Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.Advances in Neural Infor- mation Processing Systems, 36:3536–3559, 2023. 3

  43. [51]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1, 3, 8

  44. [52]

    U- net: Convolutional networks for biomedical image segmen- tation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. InMedical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, par...

  45. [53]

    Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information ...

  46. [54]

    Exploring multimodal diffusion transform- ers for enhanced prompt-based image editing

    Joonghyuk Shin, Alchan Hwang, Yujin Kim, Daneul Kim, and Jaesik Park. Exploring multimodal diffusion transform- ers for enhanced prompt-based image editing. InProceed- ings of the IEEE/CVF International Conference on Com- puter Vision, 2025. 3, 13, 1

  47. [55]

    Plug-and-play diffusion features for text-driven image-to-image translation

    Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1921–1930, 2023. 3, 4

  48. [56]

    Attention is all you need.Advances in neural information processing systems, 30, 2017

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in neural information processing systems, 30, 2017. 3

  49. [57]

    Singular value decomposition and principal component anal- ysis

    Michael E Wall, Andreas Rechtsteiner, and Luis M Rocha. Singular value decomposition and principal component anal- ysis. InA practical approach to microarray data analysis, pages 91–109. Springer, 2003. 4

  50. [58]

    Enhancing mmdit-based text-to-image models for similar subject generation.arXiv preprint arXiv:2411.18301, 2024

    Tianyi Wei, Dongdong Chen, Yifan Zhou, and Xingang Pan. Enhancing mmdit-based text-to-image models for similar subject generation.arXiv preprint arXiv:2411.18301, 2024. 3

  51. [59]

    Qwen-image technical report.arXiv preprint arXiv:2508.02324, 2025

    Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng-ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al. Qwen-image technical report.arXiv preprint arXiv:2508.02324, 2025. 6

  52. [60]

    Janus: Decoupling visual encod- ing for unified multimodal understanding and generation

    Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, et al. Janus: Decoupling visual encod- ing for unified multimodal understanding and generation. In Proceedings of the Computer Vision and Pattern Recognitio...

  53. [61]

    Core: Context- regularized text embedding learning for text-to-image per- sonalization.arXiv preprint arXiv:2408.15914, 2024

    Feize Wu, Yun Pang, Junyi Zhang, Lianyu Pang, Jian Yin, Baoquan Zhao, Qing Li, and Xudong Mao. Core: Context- regularized text embedding learning for text-to-image per- sonalization.arXiv preprint arXiv:2408.15914, 2024. 3

  54. [62]

    Tedigan: Text-guided diverse face image generation and ma- nipulation

    Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu. Tedigan: Text-guided diverse face image generation and ma- nipulation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2256–2265,

  55. [63]

    Smartbrush: Text and shape guided object inpainting with diffusion model

    Shaoan Xie, Zhifei Zhang, Zhe Lin, Tobias Hinz, and Kun Zhang. Smartbrush: Text and shape guided object inpainting with diffusion model. InProceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 22428–22437, 2023. 3

  56. [64]

    Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023

    Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagere- ward: Learning and evaluating human preferences for text- to-image generation.Advances in Neural Information Pro- cessing Systems, 36:15903–15935, 2023. 7, 10

  57. [65]

    Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms

    Ling Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu, Ste- fano Ermon, and CUI Bin. Mastering text-to-image dif- fusion: Recaptioning, planning, and generating with multi- modal llms. InForty-first International Conference on Ma- chine Learning, 2024. 3, 8, 9, 1

  58. [66]

    Reco: Region-controlled text-to-image genera- tion

    Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, et al. Reco: Region-controlled text-to-image genera- tion. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 14246–14255,

  59. [67]

    R-bind: Unified enhance- ment of attribute and relation binding in text-to-image dif- fusion models

    Huixuan Zhang and Xiaojun Wan. R-bind: Unified enhance- ment of attribute and relation binding in text-to-image dif- fusion models. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025. 3, 8, 9, 11, 13, 1

  60. [68]

    Adding conditional control to text-to-image diffusion models

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In 16 Proceedings of the IEEE/CVF international conference on computer vision, pages 3836–3847, 2023. 5, 7

  61. [69]

    Controllable text-to-image generation with gpt- 4.arXiv preprint arXiv:2305.18583, 2023

    Tianjun Zhang, Yi Zhang, Vibhav Vineet, Neel Joshi, and Xin Wang. Controllable text-to-image generation with gpt- 4.arXiv preprint arXiv:2305.18583, 2023. 3

  62. [70]

    Realcompo: Balancing realism and compositionality improves text-to-image diffusion models.Advances in Neu- ral Information Processing Systems, 37:96963–96992, 2024

    Xinchen Zhang, Ling Yang, Yaqi Cai, Zhaochen Yu, Kai-Ni Wang, Ye Tian, Minkai Xu, Yong Tang, Yujiu Yang, Bin Cui, et al. Realcompo: Balancing realism and compositionality improves text-to-image diffusion models.Advances in Neu- ral Information Processing Systems, 37:96963–9699...

  63. [71]

    Itercomp: Iterative composition-aware feedback learning from model gallery for text-to-image generation

    Xinchen Zhang, Ling Yang, Guohao Li, YaQi Cai, Yong Tang, Yujiu Yang, Mengdi Wang, Bin CUI, et al. Itercomp: Iterative composition-aware feedback learning from model gallery for text-to-image generation. InThe Thirteenth In- ternational Conference on Learning Representations. ...

  64. [72]

    Enhancing semantic fidelity in text-to-image synthesis: Attention regulation in diffusion models

    Yang Zhang, Teoh Tze Tzun, Lim Wei Hern, and Kenji Kawaguchi. Enhancing semantic fidelity in text-to-image synthesis: Attention regulation in diffusion models. In European Conference on Computer Vision, pages 70–86. Springer, 2024. 3, 8, 9, 1

  65. [73]

    Object- conditioned energy-based attention map alignment in text-to- image diffusion models

    Yasi Zhang, Peiyu Yu, and Ying Nian Wu. Object- conditioned energy-based attention map alignment in text-to- image diffusion models. InEuropean Conference on Com- puter Vision, pages 55–71. Springer, 2024. 3

  66. [74]

    Loco: Locally constrained training-free layout-to-image synthesis

    Peiang Zhao, Han Li, Ruiyang Jin, and S Kevin Zhou. Loco: Locally constrained training-free layout-to-image synthesis. arXiv preprint arXiv:2311.12342, 2023. 2

  67. [75]

    Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

    Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li. Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 22490–22499, 2023. 2 17 Det...

  68. [76]

    a bench and a car

    Details and Analysis 7.1. Limitation and Discussion The proposed Progressive Detail Injection (PDI) framework addresses detail binding issues in text-to-image generation in a training-free manner, yet it has several limitations. First, the method relies heavily on the quality ...

  69. [77]

    7.3, and is also the formation we used in test

    Example of LLM Decomposing Prompts Here, we provide an example of prompt used in LLM de- composing, this is in Config.B formation we discussed in Sec. 7.3, and is also the formation we used in test. """**Detailed Instruction Prompt for Decomposing Image Descriptions**,→ You ar...

  70. [78]

    **Output Format Requirements:** - **First Line:** - Begin with`[original]`followed by a space and then the complete original prompt exactly as provided. ,→ ,→ - **Subsequent Lines:** - Each additional line must start with `[sub-index][subject]`where:,→ -`sub-index`is a sequent...

  71. [79]

    on one subject and a \blue tracksuit

    **Decomposition Rules:** - **Generic Version ([sub-0][None]):** - Create a version of the prompt that has all specific detailed attributes (e.g., color adjectives, style adjectives) removed. This produces a simplified, generic description of the scene. ,→ ,→ ,→ ,→ - **Attribut...

  72. [80]

    Only one attribute should be reintroduced per branch, while all other attribute details remain generic

    **General Guidelines:** - **Consistency:** - Ensure that the modified sub-prompts are logically consistent with the original description. Only one attribute should be reintroduced per branch, while all other attribute details remain generic. ,→ ,→ ,→ ,→ - **Precision:** - Foll...

  73. [81]

    variants

    **Example to Follow:** Given the original prompt: ``` a man wearing a red hat and blue tracksuit is standing in front of a green sports car,→ ``` The output should be: ``` {"variants": [ [original] a man wearing a red hat and blue tracksuit is standing in front of a green spor...

  74. [82]

    variants

    **Another Example to Follow:** Given the original prompt: ``` In a cyberpunk style city night, a VanGogh-style hound dog is standing in front of a lego-style sports car ,→ ,→ ``` The output should be: ``` {"variants": [ [original] In a cyberpunk style city night, a VanGogh-sty...

  75. [83]

    **Task Summary:** - Your task is to read the given original prompt and output a set of sub-prompts using the format above. ,→ ,→ - The first sub-prompt ([sub-0][None]) should be the fully generic version.,→ - Each subsequent sub-prompt should selectively reintroduce one detail...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.