Pith. sign in

REVIEW 3 minor 56 references

COLLAR refines object-level features in diffusion transformers through cascaded training-free steps to improve control and quality.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-28 17:49 UTC pith:3MJFCYE4

load-bearing objection COLLAR gives a training-free way to refine object features in diffusion transformers through cascaded FoV expansion with attention-based alignment and cyclic injection.

arxiv 2606.00954 v1 pith:3MJFCYE4 submitted 2026-05-31 cs.CV

COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation

classification cs.CV
keywords object-level controldiffusion transformerstraining-free generationconditional image synthesislatent refinementfield-of-view expansionsemantic alignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper proposes COLLAR as a training-free method that progressively refines object features in diffusion transformers by expanding the field of view. It uses two modules to close spatial-semantic gaps and feed context back into the global image without extra training. If correct, this would let users place and control specific objects more precisely while avoiding the artifacts common in current approaches that rely on depth or edge maps. Experiments on two COCO benchmarks show gains in alignment, quality, and spatial accuracy over prior methods.

Core claim

COLLAR achieves higher-fidelity object-level control by applying Cross-Scale Semantic Alignment to inject local features into extended-FoV branches and Cyclic Feature Injection to update the global backbone with frequency-adapted reciprocal feedback, all within a cascaded FoV-expansion process that keeps final image quality intact.

What carries the argument

Cascaded Object-Level Latent Refinement via Field-of-View expansion, which serves as the hub integrating the CSSA attention-based injection and CFI adaptive feedback modules.

Load-bearing premise

The two modules can fold object-level features into the global diffusion process through FoV expansion without creating new artifacts or lowering overall image quality.

What would settle it

Running the same COCO-MIG and COCO-POS evaluations and finding no consistent gains in semantic alignment, image quality, or spatial fidelity metrics, or finding visible new artifacts, would disprove the central claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Object control becomes possible at small localized scales without retraining the underlying diffusion transformer.
  • Structural priors such as depth or Canny maps can be supplemented rather than replaced to reduce visual artifacts.
  • The extended-FoV branch acts as an optimization hub that preserves global coherence while incorporating local detail.
  • Performance improves across semantic alignment, image quality, and spatial fidelity on the reported benchmarks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same FoV-expansion pattern might transfer to other conditional tasks like layout-guided or text-plus-region editing.
  • Removing the training-free constraint could allow further gains if the modules were made differentiable and fine-tuned end-to-end.
  • The frequency-based selection in the feedback loop may generalize to other modalities where local and global features must be balanced.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. The manuscript proposes COLLAR, a training-free framework for high-fidelity object-level conditional generation in Diffusion Transformers. It employs progressive Field-of-View (FoV) expansion with two core modules: Cross-Scale Semantic Alignment (CSSA) for injecting object-level features via attention to close spatial-semantic gaps, and Cyclic Feature Injection (CFI) for reciprocal background feedback using frequency-adaptive updates. The extended-FoV branch integrates these refinements into the global process. The central claim is consistent outperformance over state-of-the-art methods on the COCO-MIG and COCO-POS benchmarks in semantic alignment, image quality, and spatial fidelity.

Significance. If the benchmark results hold under the stated training-free protocol, the work provides a practical, modular refinement strategy that mitigates artifacts in localized object control without retraining. This is a meaningful contribution to controllable diffusion-based generation, particularly for applications requiring precise spatial fidelity on standard benchmarks. The explicit use of attention-based alignment and frequency-adaptive injection offers a clear, reproducible design that could be adopted or extended by others.

minor comments (3)
  1. [Abstract / §4] The abstract states quantitative outperformance but the provided text does not include the actual tables or figures reporting the metrics, error bars, or ablation studies on COCO-MIG and COCO-POS. Adding these (or confirming their presence in §4) would strengthen verifiability.
  2. [Method description of CFI] Notation for the frequency-based adaptive strategy in CFI is introduced without an explicit equation or pseudocode; a short algorithmic box or Eq. reference would clarify the selective update rule.
  3. [Discussion / Conclusion] The manuscript should include a brief limitations paragraph addressing potential failure cases when object regions are extremely small or when background complexity is high, to balance the positive benchmark claims.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive evaluation of our work on COLLAR and the recommendation for minor revision. The assessment accurately captures the core contributions of the training-free cascaded refinement approach using CSSA and CFI modules. As no specific major comments were provided in the report, we have no individual points requiring rebuttal or clarification at this stage.

Circularity Check

0 steps flagged

No significant circularity; method is a novel framework with independent experimental validation

full rationale

The paper introduces COLLAR as a new training-free framework with CSSA and CFI modules for object-level refinement in diffusion models. No equations, derivations, fitted parameters, or predictions appear in the abstract or task description. Claims of outperformance rest on external benchmarks (COCO-MIG, COCO-POS) rather than any reduction to self-defined inputs or self-citation chains. The approach is presented as an original proposal addressing stated limitations, with no load-bearing steps that equate outputs to inputs by construction.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

No free parameters, axioms, or invented entities are identifiable from the abstract; the work relies on standard diffusion transformer components and benchmark datasets without introducing new postulated entities or fitted constants.

pith-pipeline@v0.9.1-grok · 5756 in / 1191 out tokens · 20178 ms · 2026-06-28T17:49:34.272090+00:00 · methodology

0 comments
read the original abstract

Achieving high-fidelity object-level control in Diffusion Transformers remains a significant challenge despite the introduction of structural priors like depth and Canny maps. Current object-level conditional generation methods frequently suffer from visual artifacts and struggle to maintain precise control over objects within small localized regions. To address these limitations, we propose Cascaded Object-Level Latent Refinement (COLLAR), a training-free framework that progressively optimizes object-level features via the Field-of-View (FoV) expansion. First, we propose the Cross-Scale Semantic Alignment (CSSA) module to address spatial-semantic gaps by injecting object-level features into extended-FoV branches via attention mechanisms. To further optimize these features, the Cyclic Feature Injection (CFI) module introduces a reciprocal background feedback mechanism. It leverages a frequency-based adaptive strategy to selectively update the global backbone with context-aligned local information. Finally, the extended-FoV branch serves as a hub for feature optimization, ensuring that object-level features are integrated into the global generation process without compromising final image quality. Extensive experiments on the COCO-MIG and COCO-POS benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods across semantic alignment, image quality, and spatial fidelity.

Figures

Figures reproduced from arXiv: 2606.00954 by Chengyu Lin, Jia Wei, Teng Zhou, Xiaoyu Zhang, Xinlong Zhang, Yongchuan Tang.

Figure 1
Figure 1. Figure 1: Our proposed framework, COLLAR, is a plug-and-pla [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of COLLAR. We freeze the FLUX-Depth and us [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Qualitative comparison between our method and sta [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: (a) Users were asked to rate: text-image alignment, spatial fidelity, and image quality. (b) Ablation study of model designs. Our method can achieve spatial fidelity and image quality [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗
Figure 4
Figure 4. Figure 4: Visualization results of User Study and Ablation S [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Screenshot of the instruction page and evaluation [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Qualitative results demonstrating the generaliz [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

56 extracted references · 5 canonical work pages · 2 internal anchors

  1. [1]

    Y ogesh Balaji, Seungjun Nah, Xun Huang, Arash V ahdat, Ji aming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, and 1 others. 2022. ediff-i: Text-to-image diffusion models with an ensemble of expert d enoisers. arXiv preprint arXiv:2211.01324

  2. [2]

    BlackForest. 2024. Black forest labs; frontier ai lab

  3. [3]

    Nicolas Carion, Laura Gustafson, Y uan-Ting Hu, Shoubhi k Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan V asudev Alwala, Haitham Khe dr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Gr eer, Meng Wang, Peize Sun, Roman Rädle, and 19 others. 2025. Sam 3: Segment anythin g with concepts. Preprint, arXiv:2511.16719

  4. [4]

    Anthony Chen, Jianjin Xu, Wenzhao Zheng, Gaole Dai, Yida Wang, Renrui Zhang, Haofan Wang, and Shanghang Zhang. 2024. Training-free regional pr ompting for diffusion transform- ers. arXiv preprint arXiv:2411.02395

  5. [5]

    Hongyu Chen, Yiqi Gao, Min Zhou, Peng Wang, Xubin Li, Tiez heng Ge, and Bo Zheng. 2024. Enhancing prompt following with visual control through tra ining-free mask-guided diffusion. Preprint, arXiv:2404.14768

  6. [6]

    Junsong Chen, Jincheng Y u, Chongjian Ge, Lewei Y ao, Enze Xie, Y ue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. 2024. Pixar t-α : Fast training of diffu- sion transformer for photorealistic text-to-image synthe sis. In Proceedings of the International Conference on Learning Representations

  7. [7]

    Xi Chen, Lianghua Huang, Y u Liu, Y ujun Shen, Deli Zhao, and Hengshuang Zhao. 2024. Any- door: Zero-shot object-level image customization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 6593–6602

  8. [8]

    Yiyun Chen and Weikai Y ang. 2025. Refadgen: High-fidelit y advertising image generation. arXiv preprint arXiv:2508.11695

  9. [9]

    Zhennan Chen, Y ajie Li, Haofan Wang, Zhibo Chen, Zhengka i Jiang, Jun Li, Qian Wang, Jian Y ang, and Ying Tai. 2025. Ragd: Regional-aware diffusion model for text-to-image generation. In Proceedings of the IEEE/CVF International Conference on Co mputer Vision, pages 19331– 19341. 10

  10. [10]

    Bo Cheng, Y uhang Ma, Liebucha Wu, Shanyuan Liu, Ao Ma, Xi aoyu Wu, Dawei Leng, and Y uhui Yin. 2025. Hico: Hierarchical controllable diffusio n model for layout-to-image genera- tion. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY , USA. Curran Associates Inc

  11. [11]

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Y am Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, and 1 others. 2024. Scaling rectified flow transformers for high-resolution image synthesis. In Proceedings of the 41st International Conference on Machine Learning

  12. [12]

    Y uwei Guo, Ceyuan Y ang, Anyi Rao, Zhengyang Liang, Y aoh ui Wang, Y u Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2024. Animatediff: Animate your personalized text-to- image diffusion models without specific tuning. International Conference on Learning Repre- sentations

  13. [13]

    Runze He, Bo Cheng, Y uhang Ma, Qingxiang Jia, Shanyuan L iu, Ao Ma, Xiaoyu Wu, Liebucha Wu, Dawei Leng, and Y uhui Yin. 2025. Plangen: Towar ds unified layout plan- ning and image generation in auto-regressive vision langua ge models. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 18143–18154

  14. [14]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denois ing diffusion probabilistic models. Advances in neural information processing systems , 33:6840–6851

  15. [15]

    Edward J Hu, Y elong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Y uanzhi Li, Shean Wang, Liang Wang, Weizhu Chen, and 1 others. 2022. Lora: Low-rank adapta tion of large language models. Iclr, 1(2):3

  16. [16]

    Minghui Hu, Jianbin Zheng, Daqing Liu, Chuanxia Zheng, Chaoyue Wang, Dacheng Tao, and Tat-Jen Cham. 2023. Cocktail: Mixing multi-modality contr ols for text-conditional image gen- eration. In Proceedings of the 37th International Conference on NeuralInformation Processing Systems, NIPS ’23, Red Hook, NY , USA. Curran Associates Inc

  17. [17]

    Y uval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy

  18. [18]

    In Proceed- ings of the 37th International Conference on Neural Informa tion Processing Systems , NIPS ’23, Red Hook, NY , USA

    Pick-a-pic: an open dataset of user preferences for te xt-to-image generation. In Proceed- ings of the 37th International Conference on Neural Informa tion Processing Systems , NIPS ’23, Red Hook, NY , USA. Curran Associates Inc

  19. [19]

    Lee, Taehoon Y oon, and Minhyuk Sung

    Phillip Y . Lee, Taehoon Y oon, and Minhyuk Sung. 2024. Gr oundit: grounding diffusion trans- formers via noisy patch transplantation. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY , USA. Curran Associates Inc

  20. [20]

    Seung Hyun Lee, Jijun Jiang, Yiran Xu, Zhuofang Li, Junj ie Ke, Yinxiao Li, Junfeng He, Steven Hickson, Katie Datsenko, Sangpil Kim, Ming-Hsuan Y ang, Irfan Essa, and Feng Y ang

  21. [21]

    In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and P attern Recognition (CVPR) , pages 30010–30019

    Cropper: Vision-language model for image cropping th rough in-context learning. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and P attern Recognition (CVPR) , pages 30010–30019

  22. [22]

    Danfeng Li, Hui Zhang, Sheng Wang, Jiacheng Li, and Zuxu an Wu. 2025. Seg2any: Open-set segmentation-mask-to-image generation with precise shap e and semantic control

  23. [23]

    Ming Li, Taojiannan Y ang, Huafeng Kuang, Jie Wu, Zhaoning Wang, Xuefeng Xiao, and Chen Chen. 2024. Controlnet++: Improving conditional controls with efficient consistency feedback. In Proceedings of the European Conference on Computer Vision , pages 129–147. Springer

  24. [24]

    Y uheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianw ei Y ang, Jianfeng Gao, Chun- yuan Li, and Y ong Jae Lee. 2023. Gligen: Open-set grounded te xt-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22511–22521

  25. [25]

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hay s, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision , pages 740–755. Springer. 11

  26. [26]

    Cross-controlnet: Training- free fusion of multiple conditions for text-to-image generation

    Xiang Liu, Junjun Jiang, Wei Han, Kui Jiang, and Xianmin g Liu. Cross-controlnet: Training- free fusion of multiple conditions for text-to-image generation. In The F ourteenth International Conference on Learning Representations

  27. [27]

    Y uhang Ma, Xiaoshi Wu, Keqiang Sun, and Hongsheng Li. 20 25. Hpsv3: Towards wide- spectrum human preference score. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15086–15095

  28. [28]

    Chong Mou, Xintao Wang, Liangbin Xie, Y anze Wu, Jian Zhang, Zhongang Qi, and Ying Shan

  29. [29]

    In Proceedings of the AAAI conference on artificial intelligen ce, volume 38, pages 4296–4304

    T2i-adapter: Learning adapters to dig out more contro llable ability for text-to-image diffusion models. In Proceedings of the AAAI conference on artificial intelligen ce, volume 38, pages 4296–4304

  30. [30]

    William Peebles and Saining Xie. 2023. Scalable diffus ion models with transformers. In Proceedings of the IEEE/CVF international conference on co mputer vision, pages 4195–4205

  31. [31]

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blatt mann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024. Sdxl: Improving latent d iffusion models for high- resolution image synthesis. In Proceedings of the International Conference on Learning Re p- resentations, pages 1862–1874

  32. [32]

    Can Qin, Shu Zhang, Ning Y u, Yihao Feng, Xinyi Y ang, Ying bo Zhou, Huan Wang, Juan Car- los Niebles, Caiming Xiong, Silvio Savarese, Stefano Ermon , Y un Fu, and Ran Xu. 2023. Uni- control: A unified diffusion model for controllable visual g eneration in the wild. In Proceed- ings of the 37th International Conference on Neural Informa tion Processing Sys...

  33. [33]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Le e, Sharan Narang, Michael Matena, Y anqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limi ts of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. , 21(1)

  34. [34]

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray , Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image ge neration. In International confer- ence on machine learning , pages 8821–8831. Pmlr

  35. [35]

    Scott Reed, Zeynep Akata, Xinchen Y an, Lajanugen Loges waran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. In International conference on machine learning, pages 1060–1069. Pmlr

  36. [36]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Pat rick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 10684–10695

  37. [37]

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li , Jay Whang, Emily L Denton, Kam- yar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol A yan , Tim Salimans, and 1 others

  38. [38]

    Ad- vances in neural information processing systems , 35:36479–36494

    Photorealistic text-to-image diffusion models with deep language understanding. Ad- vances in neural information processing systems , 35:36479–36494

  39. [39]

    Takahiro Shirakawa and Seiichi Uchida. 2024. Noisecol lage: A layout-aware text-to-image diffusion model based on noise cropping and merging. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 8921–8930

  40. [40]

    Shuyuan Tu, Qi Dai, Zhi-Qi Cheng, Han Hu, Xintong Han, Zu xuan Wu, and Y u-Gang Jiang

  41. [41]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni tion, pages 7882–7891

    Motioneditor: Editing video motion via content-awar e diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni tion, pages 7882–7891

  42. [42]

    Xudong Wang, Trevor Darrell, Sai Saketh Rambhatla, Roh it Girdhar, and Ishan Misra

  43. [43]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni tion, pages 6232–6242

    Instancediffusion: Instance-level control for imag e generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni tion, pages 6232–6242

  44. [44]

    Qiang Xiang, Shuang Sun, Binglei Li, Dejia Song, Huaxia Li, Nemo Chen, Xu Tang, Y ao Hu, and Junping Zhang. 2025. Instanceassemble: Layout-awa re image generation via instance assembling attention. In Proceedings of the 39th Conference on Neural Information Processing Systems. 12

  45. [45]

    Jiayu Xiao, Henglei Lv, Liang Li, Shuhui Wang, and Qingm ing Huang. 2024. R&b: Region and boundary aware zero-shot grounded text-to- image generation. In The Twelfth International Conference on Learning Representat ions

  46. [46]

    Shitao Xiao, Y ueze Wang, Junjie Zhou, Huaying Y uan, Xingrun Xing, Ruiran Y an, Chaofan Li, Shuting Wang, Tiejun Huang, and Zheng Liu. 2025. Omnigen: Un ified image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13294–13304

  47. [47]

    Denis Zavadski, Johann-Friedrich Feiden, and Carsten Rother. 2024. Controlnet-xs: Rethinking the control of text-to-image di ffusion models as feedback-control systems. In Proceedings of the European Conference on Computer Vision , page 343–362, Berlin, Hei- delberg. Springer-V erlag

  48. [48]

    Hui Zhang, Dexiang Hong, Yitong Wang, Jie Shao, Xinglon g Wu, Zuxuan Wu, and Y u-Gang Jiang. 2025. Creatilayout: Siamese multimodal diffusion t ransformer for creative layout-to- image generation. In Proceedings of the IEEE/CVF International Conference on Co mputer Vision, pages 18487–18497

  49. [49]

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Addi ng conditional control to text- to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847

  50. [50]

    Xinlong Zhang, Zejian Li, Wei Li, Xiaoyu Zhang, Jia Wei, Chengyu Lin, and Y ongchuan Tang. 2025. Objctrl: Object-based control relaxation for c onditional text-to-image generation. In Proceedings of the 33rd ACM International Conference on Mul timedia, MM ’25, page 10064–10073, New Y ork, NY , USA. Association for Computing Machinery

  51. [51]

    Y uxuan Zhang, Yirui Y uan, Yiren Song, Haofan Wang, and J iaming Liu. 2025. Easycontrol: Adding efficient and flexible control for diffusion transfor mer. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 19513–19524

  52. [52]

    Shihao Zhao, Dongdong Chen, Y en-Chun Chen, Jianmin Bao , Shaozhe Hao, Lu Y uan, and Kwan-Y ee K. Wong. 2023. Uni-controlnet: All-in-one contro l to text-to-image diffusion mod- els. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23

  53. [53]

    Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li. 2023. Lay- outdiffusion: Controllable diffusion model for layout-to -image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco gnition, pages 22490–22499

  54. [54]

    Dewei Zhou, Mingwei Li, Zongxin Y ang, and Yi Y ang. 2025. Dreamrenderer: Taming multi- instance attribute control in large-scale text-to-image models. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 16712–16722

  55. [55]

    Dewei Zhou, Y ou Li, Fan Ma, Xiaoting Zhang, and Yi Y ang. 2 024. Migc: Multi-instance generation controller for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 6818–6828

  56. [56]

    Dewei Zhou, Ji Xie, Zongxin Y ang, and Yi Y ang. 2025. 3dis: Depth-driven decoupled instance synthesis for text-to-image generation. In Proceedings of the International Conference on Learning Representations. 13 Appendix A User Study Figure 4 (a) shows a user study with 20 participants with a bac kground in computer science, com- prising 12 Master’s stude...