Pith. sign in

REVIEW 3 major objections 5 minor 85 references

CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read CoMPaSS claims that spatial failures in text-to-image diffusion models are fixable by pairing curated spatial data with token-order reinjection, achieving up to +131% relative gains on GenEval Position across four open-weight models.

desk verdict A well-executed recipe for improving spatial compliance on COCO-style benchmarks, but the headline gains track the training distribution more than they prove a general spatial understanding. read the letter →

arxiv 2412.13195 v2 pith:2V7MGPF5 submitted 2024-12-17 cs.CV

classification cs.CV
keywords text-to-imagediffusionmodelsspatialrelationshipsdatacurationtextencoderlimitationstokenorderingcross-attentioninjectionCOCObenchmarkevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that text-to-image diffusion models fail at spatial relations for two fixable reasons: existing datasets contain ambiguous or contradictory spatial language, and standard text encoders (CLIP, T5) collapse prompts that differ only in token order into nearly identical embeddings. CoMPaSS attacks both causes with a data engine, SCOP, that mines only visually clear, size-balanced object pairs from COCO images, and a parameter-free module, TENOR, that re-injects positional ordering into every text-image attention layer. Applied to SD1.4, SD1.5, SD2.1, and FLUX.1, it reports state-of-the-art spatial accuracy, with relative gains of +98% on VISOR, +67% on T2I-CompBench Spatial, and +131% on GenEval Position, while also improving overall alignment and image fidelity. If the results hold, spatial control is a data-and-conditioning problem solvable by light fine-tuning of existing open-weight models rather than by architectural redesign.

What carries the argument

Two components carry the argument. SCOP (Spatial Constraints-Oriented Pairing) is a data engine that enumerates object pairs in an image and keeps only those passing five geometric constraints—visual significance, semantic distinction, spatial clarity, minimal overlap, and size balance—then decodes the surviving pairs into image crops paired with templated spatial captions. TENOR (Token ENcoding ORdering) is a parameter-free module that adds sinusoidal positional encodings to the key vectors in UNet cross-attention and to the text query/key vectors in MMDiT blocks, making token order visible at every attention step so structurally different prompts produce different conditioning signals.

What would settle it

Evaluate CoMPaSS against its base models on a held-out benchmark built from non-COCO object categories (for example, 'a wrench to the left of a screwdriver') or object-centric spatial language ('to the child's right hand side'); if accuracy returns to baseline levels, the gains are distribution matching and the claim of enhanced general spatial understanding fails.

Watch

Extended reading notes

Core claim

CoMPaSS establishes that injecting token-order information into the text-image attention of diffusion models, in combination with training on a small set of spatially unambiguous image-text pairs, makes both UNet-based and MMDiT-based text-to-image models substantially better at rendering left, right, above, and below relations. The paper's central claim is that the two interventions are complementary: SCOP supplies clean spatial supervision that was missing from web-scale training data, and TENOR provides the structural signal that lets the model tell 'A left of B' from 'B left of A', which standard encoders fail to preserve. On FLUX.1 the combination lifts VISOR from 37.96 to 75.17, T2I-CompBench Spatial from 0.18 to 0.30, and GenEval Position from 0.26 to 0.60, with no trainable parameters added at inference time.

Load-bearing premise

The reported gains generalize beyond the exact training distribution: SCOP pairs come only from COCO object categories with eight spatial tags, and the benchmarks test the same categories and the same binary left/right/above/below relations, so the improvements could reflect distribution matching rather than general spatial understanding.

Editorial extensions

If this is right

  • Any existing UNet- or MMDiT-based text-to-image model can be upgraded for spatial accuracy with a short fine-tuning phase that adds no parameters at inference and only about 3% latency.
  • A random 500-image subset of SCOP already lifts GenEval Position from 0.26 to 0.56 on FLUX.1, so the recipe is data-efficient enough for settings without access to web-scale datasets.
  • The model trained only on two-object pairs improves three-object spatial accuracy (e.g., FLUX.1 'any' accuracy from 30.12 to 52.44), indicating the token-order signal transfers beyond the training template.
  • The improvements are not confined to spatial metrics: overall GenEval, DPG-Bench, FID, and CMMD all improve, suggesting that cleaning spatial supervision also helps general prompt following.
  • The ablations assign distinct roles to the two components: SCOP alone raises spatial accuracy substantially, and TENOR adds generalization to unseen prompt structures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Every reported benchmark shares SCOP's own COCO vocabulary and binary relation set, so the true test of general spatial understanding would be an out-of-distribution probe with non-COCO objects or context-dependent spatial language; the paper does not provide one.
  • The same token-order blindness that scrambles left/right also plausibly degrades attribute binding and other order-sensitive compositions, so TENOR may transfer to color, size, and count tasks—an untested implication of the paper's analysis.
  • The authors' listed limitations (extreme size disparities, object-centric frames) suggest concrete next experiments: building SCOP-style pairs that include size-contrast or object-centric annotations should extend the method toward fuller Qualitative Spatial Relations coverage, and their Fig. 7 shows a preliminary positive result for size.
  • The 85.2% human-agreement check validates the SCOP captions, but the paper does not decompose how much of the benchmark gain comes from the crop-and-template decoding versus the geometric filtering itself.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper identifies two causes of poor spatial-relation generation in text-to-image (T2I) diffusion models: ambiguous spatial captions in existing image-text datasets and loss of token-ordering information in text encoders. It proposes CoMPaSS, composed of SCOP, a constraint-based data engine applied to the COCO training split that extracts object pairs satisfying visual significance, semantic distinction, spatial clarity, minimal overlap, and size balance, and TENOR, a parameter-free module that injects positional encodings into the text-image attention keys and queries. Experiments on SD1.4, SD1.5, SD2.1, and FLUX.1 report large gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%), together with improved overall scores on DPG-Bench and fidelity metrics, at low computational overhead and with promising data efficiency.

Significance. If the results hold, CoMPaSS is a practical and lightweight recipe for improving spatial compliance of open-weight T2I models: it adds no trainable parameters, requires only a brief fine-tuning phase, works across UNet and MMDiT architectures, and its data engine is simple and reproducible. The strengths of the paper are the systematic threshold-based curation pipeline, the sensible diagnostic in Table 1 showing that text encoders fail to rank logically equivalent spatial paraphrases as most similar, and the ablations in Tables 5 and 6 showing that both components contribute and that performance scales with training data. However, because SCOP and the headline benchmarks share the COCO category vocabulary and a small set of binary viewer-centric spatial relations, the evidence as presented is stronger for distribution-matched spatial compliance than for a general spatial-understanding capability. The broad claim of generalization needs an out-of-distribution evaluation and error bars before the conclusion is fully supported.

major comments (3)
  1. [Sec. 3.1, Sec. A.1, Tab. 2] The evidence for the paper's main claim that CoMPaSS enhances spatial understanding generally is currently confined to the same distribution used to build SCOP. SCOP curates pairs from the COCO training split and encodes only eight spatial tokens (<left>, <right>, <above>, <below>, and four diagonals), while the three headline benchmarks evaluate simple binary viewer-centric relations over essentially the same object vocabulary. The paper's own Sec. 5 lists context-dependent spatial language and object-centric frames as unsupported. Without an evaluation on categories and relation types outside this closed vocabulary, the large relative gains are equally consistent with distribution matching. Please add such an out-of-distribution test, or revise the conclusion to claim improved spatial compliance on this benchmark distribution rather than general spatial understanding.
  2. [Sec. 4.3, Tab. 5] The SCOP thresholds tau_v, tau_u, tau_o, and tau_s are selected by grid search on the same benchmarks that produce the headline SOTA numbers, and no repeated-seed or bootstrap intervals are reported for any accuracy result in Tabs. 2, 4, 5, 6, or A8-A11. This makes it impossible to quantify how much of the reported margin is selection bias and leaves the true improvement over baselines uncertain. Please provide confidence intervals and a validation split that is not used for threshold selection.
  3. [Tab. A11 vs. abstract/Sec. 4.2] The abstract and Sec. 4.2 state that gains are achieved 'without compromising general generation capabilities,' but the per-task breakdown in Tab. A11 shows several non-spatial tasks degrading: SD2.1+CoMPaSS drops GenEval Color from 0.85 to 0.71 and Count from 0.44 to 0.20, and SD1.5+CoMPaSS drops DPG-Bench Other from 67.81 to 60.80. Since overall scores can mask these trade-offs, the no-compromise claim should be made conditional on aggregate metrics or accompanied by a per-task analysis of which capabilities are preserved and which are not.
minor comments (5)
  1. [Sec. 3.1, Eqs. (1)-(5)] The thresholds are introduced as 'principled constraints' but are free parameters; please justify the chosen values or soften the terminology, and state how the resulting dataset size varies with each threshold.
  2. [Tab. 1] The proxy task tests nearest-neighbor ranking among four prompt variations; please clarify how ties are handled and report per-relation results, since 'above'/'below' may behave differently from 'left'/'right'.
  3. [Sec. 3.1] The human validation reports an 85.2% agreement rate but does not state the number of annotators, the number of items judged, or the exact instructions given; please add this information for reproducibility.
  4. [Sec. 5, Fig. 7] The size-disparity fine-tuning experiment is described only with one qualitative example; provide the training protocol and quantitative results, or label it explicitly as preliminary.
  5. [Appendix B] The latency overhead table reports mean +/- SD but not the number of measurement repetitions or the hardware conditions; please state the measurement protocol so the overhead numbers can be reproduced.

Circularity Check

1 steps flagged · score 4.0 of 10

No equation-level circularity and no load-bearing self-citation; the main residual circular element is that SCOP's thresholds are grid-searched on the GenEval Position benchmark that is then reported as a headline +131% 'prediction'.

  1. fitted input called prediction [Sec. 4.3 (Ablation Studies), Table 5; headline results in Sec. 1 and Table 2]
    "The SCOP data engine has four tunable hyperparameters. We empirically determine the optimal values to be {τv, τu, τo, τs} = {0.2, 2.0, 0.3, 0.5} via grid search. In Tab. 5, we report the model’s sensitivity to each of these four hyperparameters by evaluating performance on nearby values. While our chosen hyperparameters yield optimal results, the model’s performance remains high across a range of nearby values."

    Table 5's 'Ours' values are 0.54 for SD1.5 and 0.60 for FLUX.1, exactly the GenEval Position scores reported for SD1.5+CoMPaSS and FLUX.1+CoMPaSS in Table 2. Thus the SCOP threshold tuple was selected by maximizing the GenEval Position benchmark, and the same number is later presented as an independent +131% spatial-understanding prediction. The GenEval Position claim is therefore partly a fitted quantity rather than an out-of-sample result. This does not make the entire method circular: the VISOR and T2I-CompBench Spatial gains, and the TENOR ablation in Table 6, provide independent evidence. The issue is localized to one headline number and is a tuning-on-the-test-set loop, not an identity of equations.

full rationale

The paper's derivation chain is otherwise self-contained. SCOP is a data-curation engine that filters COCO pairs by explicit geometric constraints; TENOR adds absolute positional encodings to cross-attention key/query vectors, which is a parameter-free architectural intervention tested against the original models in Table 6. No claim is derived from a fitted parameter in the sense of a formula reducing to its inputs, and no load-bearing self-citation or imported uniqueness theorem appears. The main circularity concern is the SCOP hyperparameter grid search: the optimal thresholds in Table 5 coincide with the GenEval Position numbers later reported as SOTA, indicating selection on that benchmark. The reported VISOR (+98%) and T2I-CompBench Spatial (+67%) gains are not tied to that tuning loop, and the ablation shows large gains even for nearby threshold values, so the central claim retains substantial independent content. The broader worry that SCOP's COCO/8-token distribution matches the evaluation benchmarks is a generalization/overfitting concern rather than circularity and is therefore not scored as a circular step.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim rests on four hand-tuned thresholds, a bounding-box definition of spatial relations, a particular text-encoder diagnostic, and an unexamined assumption that evaluation benchmarks are out-of-distribution. No invented entities are introduced.

free parameters (4)
  • tau_v (visual significance threshold) = 0.2
    Chosen via grid search (Sec. 4.3); filters out object pairs whose combined area is too small for a meaningful spatial relation.
  • tau_u (spatial clarity threshold) = 2.0
    Chosen via grid search; upper-bounds centroid distance relative to object diagonal, keeping only spatially close pairs.
  • tau_o (minimal overlap threshold) = 0.3
    Chosen via grid search; caps overlap between objects to preserve visibility while allowing partial overlap.
  • tau_s (size balance threshold) = 0.5
    Chosen via grid search; requires comparable prominence of both objects so neither is too small to serve as a reference.
assumptions (4)
  • domain assumption Bounding-box geometry and category labels suffice to determine the spatial relation between two objects.
    SCOP assigns <left>, <above>, etc. purely from centroids and boxes (Sec. 3.1). This ignores viewer/object frame ambiguity and context-dependent spatial language, a limitation the authors acknowledge in Sec. 5.
  • domain assumption The proxy task in Table 1 (highest similarity to rephrased variation) measures how well a text encoder preserves spatial semantics.
    Used to conclude CLIP/T5 fail at spatial understanding and to motivate TENOR (Sec. 3.2). The conclusion depends on cosine similarity in pooled embedding space being a valid diagnostic.
  • domain assumption Benchmark prompts (VISOR, GenEval Position, T2I-CompBench Spatial) are unseen relative to SCOP training data and indicative of general spatial ability.
    The paper does not analyze train/eval category overlap; SCOP uses the COCO train split and these benchmarks use COCO categories, so the assumption is questionable (Secs. 3.1, 4.1).
  • domain assumption Fine-tuning on 28k spatial pairs does not degrade other generative abilities.
    The paper claims no compromise, but Tab. A11 shows drops on GenEval Count and Color for some models. The claim rests on aggregate overall scores rather than per-task guarantees.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models." pith.science (2026). https://pith.science/paper/2V7MGPF5

@misc{pith2026241213195,
  author       = {Pith},
  title        = {Pith review of: CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2V7MGPF5}},
  note         = {Machine review of arXiv:2412.13195}
}
read the original abstract

Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning spatial relationships in existing datasets, and 2) the inability of current text encoders to accurately interpret the spatial semantics of input descriptions. We propose CoMPaSS, a versatile framework that enhances spatial understanding in T2I models. It first addresses data ambiguity with the Spatial Constraints-Oriented Pairing (SCOP) data engine, which curates spatially-accurate training data via principled constraints. To leverage these priors, CoMPaSS also introduces the Token ENcoding ORdering (TENOR) module, which preserves crucial token ordering information lost by text encoders, thereby reinforcing the prompt's linguistic structure. Extensive experiments on four popular T2I models (UNet and MMDiT-based) show CoMPaSS sets a new state of the art on key spatial benchmarks, with substantial relative gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%). Code is available at https://github.com/blurgyy/CoMPaSS.

Figures

Figures reproduced from arXiv: 2412.13195 by the authors.

Figure 1
Figure 1. CoMPaSS enhances the spatial understanding of ex￾isting T2I diffusion models, enabling them to generate images that faithfully reflect spatial configurations specified in the text prompt. to correctly render spatial relationships described in text. For example, when given seemingly simple spatial config￾urations like “a motorcycle to the right of a bear”, or “a bird below a skateboard”, models that excel in realism … view at source ↗
Figure 2
Figure 2. Examples highlighting common ambiguities and er￾rors in spatial language annotations from COCO, LAION, and CC12M datasets. (a1, a2) inconsistent frame of reference; (b1) non-spatial usage of spatial terms; (b2, c1, c2) missing or incor￾rect reference objects. is meaningful: \frac { \text {Area}(\mathbf {B}_i \cup \mathbf {B}_j) }{ \text {Area}(\mathbf {I}) } > \tau _\text {v}. (1) A relatively small value of τv ensu… view at source ↗
Figure 3
Figure 3. Overview of the Spatial Constraints-Oriented Pairing (SCOP) data engine. SCOP first (1) reasons about all possible object pairs in an image, then (2) validates their spatial relationships using a set of carefully designed constraints. Finally, it (3) decodes the resulting unambiguous descriptors into image crops paired with accurate textual descriptions. sual separation: \frac { \text {Area}(\mathbf {B}_i \cap \math… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Overview of Token Encoding Ordering (TENOR). TENOR injects token ordering information into every text-image attention operation in both UNet- and MMDiT-based diffusion models. encodings to both the text query (Qtext) and key (Ktext) vec￾tors. Alternative approaches for…
Figure 5
Figure 5. Figure 5: Qualitative results of models enhanced with CoMPaSS. Our method improves the spatial understanding of both UNet-based models (SD1.4, SD1.5, SD2.1) and the MMDiT-based FLUX.1. More results are available in the supplementary material. 4.1. Evaluation Details Baselines. W…
Figure 7
Figure 7. Figure 7: Results for “a tiny phone above a large couch”. Left: FLUX.1. Middle: FLUX.1+CoMPaSS. Right: Fine-tuning (100 steps, batch size 4) on FLUX.1+CoMPaSS with only size-contrast data filtered from adapted SCOP. when the semantic representations of completely opposite text d…
Figure 6
Figure 6. Figure 6: Qualitative ablation study on the TENOR module. The TENOR module improves generalization to unseen prompts. (i) First row: original models. (ii) Second row: models trained only with the SCOP dataset. (iii) Third row: our full method. The left three columns show results…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 60 canonical work pages

  1. [1]

    Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models

    Eslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan, Li Erran Li, and Mohamed Elhoseiny. Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models. In IEEE/CVF International Confer- ence on Computer Vision, ICCV 2023, Paris, France, Octo- ber 1-6, 2023, pages 19984–19996. IEEE, 2023. 2

  2. [2]

    Improving Image Genera- tion with Better Captions

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, Wesam Manassra, Prafulla Dhariwal, Casey Chu, Yunxin Jiao, and Aditya Ramesh. Improving Image Genera- tion with Better Captions. 1

  3. [3]

    Training diffusion models with reinforce- ment learning

    Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine. Training diffusion models with reinforce- ment learning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. 2

  4. [4]

    FLUX.1-dev

    Black Forest Labs. FLUX.1-dev. https : / / huggingface . co / black - forest - labs / FLUX . 1-dev, 2024. Accessed: 2025-07-16. 1, 5

  5. [5]

    Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

    Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts. InIEEE Con- ference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , pages 3558–3568. Com- puter Vision Foundation / IEEE, 2021. 1, 2, 3, 4

  6. [6]

    Getting it right: Improving spatial consis- tency in text-to-image models

    Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, and Yezhou Yang. Getting it right: Improving spatial consis- tency in text-to-image models. In Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 20...

  7. [7]

    Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

    Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models. ACM Trans. Graph., 42(4):148:1–148:10, 2023. 2, 5

  8. [8]

    Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu

    Juan Chen, Anthony G. Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu. A survey of qual- itative spatial representations. pages 106–136, 2015. 8

Show all 85 references
  1. [9]

    Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu

    Juan Chen, Anthony G. Cohn, Dayou Liu, Sheng-Sheng Wang, Jihong Ouyang, and Qiangyuan Yu. A survey of qual- itative spatial representations. Knowl. Eng. Rev., 30(1):106– 136, 2015. 8

  2. [10]

    Pixart- Σ: Weak-to-strong training of dif- fusion transformer for 4k text-to-image generation

    Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- Σ: Weak-to-strong training of dif- fusion transformer for 4k text-to-image generation. In Com- puter Vision - ECCV 2024 - 18th European Conference...

  3. [11]

    Pixart- δ: Fast and controllable image generation with latent consistency mod- els

    Junsong Chen, Yue Wu, Simian Luo, Enze Xie, Sayak Paul, Ping Luo, Hang Zhao, and Zhenguo Li. Pixart- δ: Fast and controllable image generation with latent consistency mod- els. CoRR, abs/2401.05252, 2024

  4. [12]

    Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Zhongdao Wang, James T. Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart- alpha: Fast training of diffu- sion transformer for photorealistic text-to-image synthesis. In The Twelfth International Conference on Lear...

  5. [13]

    Training-free layout control with cross-attention guidance

    Minghao Chen, Iro Laina, and Andrea Vedaldi. Training-free layout control with cross-attention guidance. In IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2024, Waikoloa, HI, USA, January 3-8, 2024 , pages 5331–5341. IEEE, 2024. 2, 5, 6, 7

  6. [14]

    Reproducible scaling laws for contrastive language-image learning

    Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuh- mann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. In IEEE/CVF Conference on Computer Vision and Pattern Reco...

  7. [15]

    Visual pro- gramming for step-by-step text-to-image generation and evaluation

    Jaemin Cho, Abhay Zala, and Mohit Bansal. Visual pro- gramming for step-by-step text-to-image generation and evaluation. In Advances in Neural Information Process- ing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, U...

  8. [16]

    Coventry, Merc `e Prat-Sala, and Lynn Richards

    Kenny R. Coventry, Merc `e Prat-Sala, and Lynn Richards. The interplay between geometry and function in the compre- hension of over, under, above, and below.Journal of Memory and Language, 44(3):376–398, 2001. 8

  9. [17]

    Dall·e mini

    Boris Dayma, Suraj Patil, Pedro Cuenca, Khalid Saifullah, Tanishq Abraham, Ph ´uc L ˆe Khac, Luke Melas, and Rito- brata Ghosh. Dall·e mini. https://github.com/ borisdayma/dalle-mini, 2021. 8

  10. [18]

    Diffu- sion models beat gans on image synthesis

    Prafulla Dhariwal and Alexander Quinn Nichol. Diffu- sion models beat gans on image synthesis. In Advances in Neural Information Processing Systems 34: Annual Con- ference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual , pages 8780– 8...

  11. [19]

    Cogview2: Faster and better text-to-image generation via hierarchical transformers

    Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: Faster and better text-to-image generation via hierarchical transformers. In Advances in Neural Informa- tion Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New O...

  12. [20]

    Scaling rec- tified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rombach. Scaling rec- tified flow transformers for high-resolution image syn...

  13. [21]

    Re- inforcement learning for fine-tuning text-to-image diffusion models

    Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Moham- mad Ghavamzadeh, Kangwook Lee, and Kimin Lee. Re- inforcement learning for fine-tuning text-to-image diffusion models. In Advances in Neural Information Processing Sys- tems 36:...

  14. [22]

    Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang

    Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Ar- jun R. Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang. Training-free structured dif- fusion guidance for compositional text-to-image synthesis. In The Eleventh International Conference on Learn...

  15. [23]

    Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang

    Weixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani, Ar- jun R. Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang. Layoutgpt: Compositional visual plan- ning and generation with large language models. In Ad- vances in Neural Information Processing Systems 36: ...

  16. [24]

    Tc-bench: Benchmark- ing temporal compositionality in text-to-video and image-to- video generation

    Weixi Feng, Jiachen Li, Michael Saxon, Tsu-Jui Fu, Wenhu Chen, and William Yang Wang. Tc-bench: Benchmark- ing temporal compositionality in text-to-video and image-to- video generation. CoRR, abs/2406.08656, 2024. 2

  17. [25]

    Geneval: An object-focused framework for evaluating text- to-image alignment

    Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text- to-image alignment. In Advances in Neural Information Pro- cessing Systems 36: Annual Conference on Neural Informa- tion Processing Systems 2023, NeurIPS 2023, New ...

  18. [26]

    Benchmarking spatial relationships in text-to-image generation

    Tejas Gokhale, Hamid Palangi, Besmira Nushi, Vibhav Vi- neet, Eric Horvitz, Ece Kamar, Chitta Baral, and Yezhou Yang. Benchmarking spatial relationships in text-to-image generation. CoRR, abs/2212.10015, 2022. 2, 6, 7, 8

  19. [27]

    Diffusion-rpo: Aligning diffusion mod- els through relative preference optimization

    Yi Gu, Zhendong Wang, Yueqin Yin, Yujia Xie, and Mingyuan Zhou. Diffusion-rpo: Aligning diffusion mod- els through relative preference optimization. CoRR, abs/2406.06382, 2024. 2

  20. [28]

    Yagmur G ¨uc ¸l¨ut¨urk, Umut G ¨uc ¸l¨u, Rob van Lier, and Mar- cel A. J. van Gerven. Convolutional sketch inversion. In Computer Vision - ECCV 2016 Workshops - Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part I, pages 810–824. Springer, 2016. 2

  21. [29]

    Ganspace: Discovering interpretable GAN controls

    Erik H ¨ark¨onen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. Ganspace: Discovering interpretable GAN controls. In Advances in Neural Information Processing Sys- tems 33: Annual Conference on Neural Information Process- ing Systems 2020, NeurIPS 2020, December 6-12, 2...

  22. [30]

    Prompt-to-prompt image editing with cross-attention control

    Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross-attention control. InThe Eleventh Interna- tional Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, ...

  23. [31]

    Gans trained by a two time-scale update rule converge to a local nash equilib- rium

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Proc...

  24. [32]

    Denoising dif- fusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. In Advances in Neural Informa- tion Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, De- cember 6-12, 2020, virtual, 2020. 2

  25. [33]

    ELLA: equip diffusion models with LLM for en- hanced semantic alignment

    Xiwei Hu, Rui Wang, Yixiao Fang, Bin Fu, Pei Cheng, and Gang Yu. ELLA: equip diffusion models with LLM for en- hanced semantic alignment. CoRR, abs/2403.05135, 2024. 2, 5, 6, 7, 8

  26. [34]

    Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Os- tendorf, Ranjay Krishna, and Noah A. Smith. TIFA: accurate and interpretable text-to-image faithfulness evaluation with question answering. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, Fran...

  27. [35]

    T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

    Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xi- hui Liu. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation. In Ad- vances in Neural Information Processing Systems 36: An- nual Conference on Neural Information Processing Syste...

  28. [36]

    Re- thinking FID: towards a better evaluation metric for image generation

    Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar. Re- thinking FID: towards a better evaluation metric for image generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, ...

  29. [37]

    Comat: Aligning text-to-image diffusion model with image- to-text concept matching

    Dongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang, Dazhong Shen, Zhuofan Zong, Yu Liu, and Hongsheng Li. Comat: Aligning text-to-image diffusion model with image- to-text concept matching. InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural ...

  30. [38]

    Scalable ranked preference optimization for text-to-image generation

    Shyamgopal Karthik, Huseyin Coskun, Zeynep Akata, Sergey Tulyakov, Jian Ren, and Anil Kag. Scalable ranked preference optimization for text-to-image generation. CoRR, abs/2410.18013, 2024. 2

  31. [39]

    Evaluating and improving composi- tional text-to-visual generation

    Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li, Yixin Fei, Kewen Wu, Xide Xia, Pengchuan Zhang, Graham Neubig, and Deva Ramanan. Evaluating and improving composi- tional text-to-visual generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - W...

  32. [40]

    Photomaker: Customizing realistic human photos via stacked ID embedding

    Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming- Ming Cheng, and Ying Shan. Photomaker: Customizing realistic human photos via stacked ID embedding. CoRR, abs/2312.04461, 2023. 2

  33. [41]

    Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

    Long Lian, Boyi Li, Adam Yala, and Trevor Darrell. Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models. Trans. Mach. Learn. Res., 2024, 2024. 2, 6

  34. [42]

    Collins, Yiwen Luo, Yang Li, Kai J

    Youwei Liang, Junfeng He, Gang Li, Peizhao Li, Arseniy Klimovskiy, Nicholas Carolan, Jiao Sun, Jordi Pont-Tuset, Sarah Young, Feng Yang, Junjie Ke, Krishnamurthy Dj Dvijotham, Katherine M. Collins, Yiwen Luo, Yang Li, Kai J. Kohlhoff, Deepak Ramachandran, and Vidhya Naval- pak...

  35. [43]

    Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

    Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In Computer Vision - ECCV 2014 - 13th Eu- ropean Conference, Zurich, Switzerland, September 6-12, 2014, ...

  36. [44]

    Tenenbaum

    Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B. Tenenbaum. Compositional visual generation with composable diffusion models. In Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, Octo- ber 23-27, 2022, Proceedings, Part XVII , pages 423–439...

  37. [45]

    Repaint: Inpainting using denoising diffusion probabilistic models

    Andreas Lugmayr, Martin Danelljan, Andr´es Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022...

  38. [46]

    Pick-and-draw: Training-free semantic guidance for text-to-image person- alization

    Henglei Lv, Jiayu Xiao, and Liang Li. Pick-and-draw: Training-free semantic guidance for text-to-image person- alization. In Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Aus- tralia, 28 October 2024 - 1 November 2024 , pages 1053...

  39. [47]

    MidJourney

    Inc. MidJourney. Midjourney: Ai-powered image genera- tion. https://www.midjourney.com/ , 2023. Ac- cessed: 2024-04-27. 1

  40. [48]

    T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

    Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Thirty-Eighth AAAI Conference on Ar- tificial Intelligence, AAAI 2024, Thirty-Sixt...

  41. [49]

    GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models

    Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealis- tic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, I...

  42. [50]

    Drag your GAN: interactive point-based manipulation on the generative image manifold

    Xingang Pan, Ayush Tewari, Thomas Leimk”uhler, Lingjie Liu, Abhimitra Meka, and Christian Theobalt. Drag your GAN: interactive point-based manipulation on the generative image manifold. In ACM SIGGRAPH 2023 Conference Pro- ceedings, SIGGRAPH 2023, Los Angeles, CA, USA, August ...

  43. [51]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1- 6, 2023, pages 4172–4182. IEEE, 2023. 2, 5

  44. [52]

    Grounded text-to-image synthesis with attention refocusing

    Quynh Phung, Songwei Ge, and Jia-Bin Huang. Grounded text-to-image synthesis with attention refocusing. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 7932–7942. IEEE, 2024. 2, 6, 7

  45. [53]

    SDXL: improving latent diffusion models for high-resolution image synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M”uller, Joe Penna, and Robin Rombach. SDXL: improving latent diffusion models for high-resolution image synthesis. In The Twelfth Interna- tional Conference on Learning Representations, ICLR 2024,...

  46. [54]

    Barron, and Ben Milden- hall

    Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representa- tions, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenRe- view.net, 2023. 1

  47. [55]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of th...

  48. [56]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21: 140:1–140:67, 2020. 2, 5

  49. [57]

    Hierarchical text-conditional image gener- ation with CLIP latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gener- ation with CLIP latents. CoRR, abs/2204.06125, 2022. 1, 2, 8

  50. [58]

    High-resolution image syn- thesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 , pages 10674–...

  51. [59]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, CVPR 2023, Vancouver, BC, Ca...

  52. [60]

    Runway AI

    Inc. Runway AI. Runwayml: Creative ai tools for content creation. https://runwayml.com/, 2023. Accessed: 2024-04-27. 1

  53. [61]

    Dual caption preference optimization for diffusion models

    Amir Saeidi, Yiran Luo, Agneet Chatterjee, Shamanthak Hegde, Bimsara Pathiraja, Yezhou Yang, and Chitta Baral. Dual caption preference optimization for diffusion models. CoRR, abs/2502.06023, 2025. 2

  54. [62]

    Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Moham- mad Norouzi. Photorealistic text-to-image diffusion mod- els wit...

  55. [63]

    Scribbler: Controlling deep image synthesis with sketch and color

    Patsorn Sangkloy, Jingwan Lu, Chen Fang, Fisher Yu, and James Hays. Scribbler: Controlling deep image synthesis with sketch and color. In 2017 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2017, Hon- olulu, HI, USA, July 21-26, 2017 , pages 6836–6845. IEEE...

  56. [64]

    LAION- 400M: open dataset of clip-filtered 400 million image-text pairs

    Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION- 400M: open dataset of clip-filtered 400 million image-text pairs. CoRR, abs/2111.02114, 2021. 1, 2, 3, 4

  57. [65]

    LAION-5B: an open large-scale dataset for training next generation image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAI...

  58. [66]

    A picture is worth a thousand words: Principled recaptioning improves image generation

    Eyal Segalis, Dani Valevski, Danny Lumen, Yossi Matias, and Yaniv Leviathan. A picture is worth a thousand words: Principled recaptioning improves image generation. CoRR, abs/2310.16656, 2023. 2

  59. [67]

    Box it to bind it: Unified layout control and attribute binding in t2i diffusion models

    Ashkan Taghipour, Morteza Ghahremani, Mohammed Ben- namoun, Aref Miri Rekavandi, Hamid Laga, and Farid Bous- said. Box it to bind it: Unified layout control and attribute binding in t2i diffusion models. CoRR, abs/2402.17910,

  60. [68]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems 30, pages 5998–6008, 2017. 5

  61. [69]

    Diffusion model align- ment using direct preference optimization

    Bram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou, Aaron Lou, Senthil Purushwalkam, Stefano Ermon, Caiming Xiong, Shafiq Joty, and Nikhil Naik. Diffusion model align- ment using direct preference optimization. In IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  62. [70]

    Instantid: Zero-shot identity-preserving gener- ation in seconds

    Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, and An- thony Chen. Instantid: Zero-shot identity-preserving gener- ation in seconds. CoRR, abs/2401.07519, 2024. 2

  63. [71]

    Tokencompose: Text-to-image diffusion with token-level supervision

    Zirui Wang, Zhizhou Sha, Zheng Ding, Yilin Wang, and Zhuowen Tu. Tokencompose: Text-to-image diffusion with token-level supervision. In IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 8553–8564. IEEE, 2024. 2, 1

  64. [72]

    Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau

    Zijie J. Wang, Evan Montoya, David Munechika, Haoyang Yang, Benjamin Hoover, and Duen Horng Chau. Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models. In Proceedings of the 61st An- nual Meeting of the Association for Computational Linguis-...

  65. [73]

    Seesr: Towards semantics-aware real-world image super-resolution

    Rongyuan Wu, Tao Yang, Lingchen Sun, Zhengqiang Zhang, Shuai Li, and Lei Zhang. Seesr: Towards semantics-aware real-world image super-resolution. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024 , pages 25456–25467...

  66. [74]

    Paragraph-to-image gener- ation with information-enriched diffusion model

    Weijia Wu, Zhuang Li, Yefei He, Mike Zheng Shou, Chunhua Shen, Lele Cheng, Yan Li, Tingting Gao, Di Zhang, and Zhongyuan Wang. Paragraph-to-image gener- ation with information-enriched diffusion model. CoRR, abs/2311.14284, 2023. 2, 5

  67. [75]

    Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis

    Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. CoRR, abs/2306.09341, 2023. 2

  68. [76]

    Human preference score: Better aligning text-to- image models with human preference

    Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hong- sheng Li. Human preference score: Better aligning text-to- image models with human preference. In IEEE/CVF Inter- national Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023 , pages 2096–2105. IEEE, 2023. 2

  69. [77]

    Stylespace analysis: Disentangled controls for stylegan image genera- tion

    Zongze Wu, Dani Lischinski, and Eli Shechtman. Stylespace analysis: Disentangled controls for stylegan image genera- tion. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 , pages 12863–12872. Computer Vision Foundation / IEEE...

  70. [78]

    Freeman, Fr ´edo Durand, and Song Han

    Guangxuan Xiao, Tianwei Yin, William T. Freeman, Fr ´edo Durand, and Song Han. Fastcomposer: Tuning-free multi- subject image generation with localized attention. Int. J. Comput. Vis., 133(3):1175–1194, 2025. 2

  71. [79]

    R&b: Region and boundary aware zero-shot grounded text-to-image generation

    Jiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang, and Qing- ming Huang. R&b: Region and boundary aware zero-shot grounded text-to-image generation. In The Twelfth Interna- tional Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2...

  72. [80]

    Imagere- ward: Learning and evaluating human preferences for text- to-image generation

    Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagere- ward: Learning and evaluating human preferences for text- to-image generation. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Infor- m...

  73. [81]

    Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models

    Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models. CoRR, abs/2308.06721, 2023. 2

  74. [82]

    Scaling up to excellence: Practicing model scaling for photo- realistic image restoration in the wild

    Fanghua Yu, Jinjin Gu, Zheyuan Li, Jinfan Hu, Xiangtao Kong, Xintao Wang, Jingwen He, Yu Qiao, and Chao Dong. Scaling up to excellence: Practicing model scaling for photo- realistic image restoration in the wild. In IEEE/CVF Confer- ence on Computer Vision and Pattern Recognit...

  75. [83]

    Scaling autoregressive models for content-rich text-to-image generation

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin- fei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content...

  76. [84]

    Adding conditional control to text-to-image diffusion models

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 3813–

  77. [85]

    A horse to the left of a bottle

    Bo Zhao, Lili Meng, Weidong Yin, and Leonid Sigal. Image generation from layout. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019 , pages 8584–8593. Computer Vision Foundation / IEEE, 2019. 2 Appendix A. Additional...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.