Pith. sign in

REVIEW 45 references

How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.16679 v1 pith:NB7XOPQD submitted 2025-06-20 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords captionsmodeltext-to-imagetrainingdesignmodelsperformancesynthetic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inconsistency in web-scraped datasets, recent works shifted towards synthetic training captions. While this setup is generally believed to produce more capable models, current literature does not provide any insights into its design choices. This study closes this gap by systematically investigating how different synthetic captioning strategies impact the downstream performance of text-to-image models. Our experiments demonstrate that dense, high-quality captions enhance text alignment but may introduce trade-offs in output aesthetics and diversity. Conversely, captions of randomized lengths yield balanced improvements across aesthetics and alignment without compromising sample diversity. We also demonstrate that varying caption distributions introduce significant shifts in the output bias of a trained model. Our findings underscore the importance of caption design in achieving optimal model performance and provide practical insights for more effective training data strategies in text-to-image generation.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

45 extracted references · 23 canonical work pages

  1. [1]

    eDiff-I: Text-to-image diffusion models with an ensem- ble of expert denoisers.arXiv:2211.01324, 2022

    Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. eDiff-I: Text-to-image diffusion models with an ensem- ble of expert denoisers.arXiv:2211.01324, 2022. 4

  2. [2]

    Improving image generation with better captions, 2023

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, and Ope- nAI. Improving image generation with better captions, 2023. 1, 2

  3. [3]

    Easily Ac- cessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. Easily Ac- cessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale. InFAccT, 2023. 3

  4. [4]

    Se- mantics derived automatically from language corpora con- tain human-like biases.Science, 356(6334):183–186, 2017

    Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. Se- mantics derived automatically from language corpora con- tain human-like biases.Science, 356(6334):183–186, 2017. 3

  5. [5]

    How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.arXiv preprint arXiv:2404.16821, 2024

    Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.arXiv preprint arXiv:2404.16821, 2024. 5

  6. [6]

    Lmdeploy: A toolkit for com- pressing, deploying, and serving llm.https://github

    LMDeploy Contributors. Lmdeploy: A toolkit for com- pressing, deploying, and serving llm.https://github. com/InternLM/lmdeploy, 2023. 1

  7. [7]

    DeepFloyd-IF-I-XL-v1.0: DeepFloyd’s Image Generation Model, 2023

    DeepFloyd. DeepFloyd-IF-I-XL-v1.0: DeepFloyd’s Image Generation Model, 2023. Accessed: 2024-11-13. 1

  8. [8]

    Scaling rectified flow trans- formers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yan- nik Marek, and Robin Rombach. Scaling rectified flow trans- formers for high-resolution image synthesis. InProceedings of the 41st Inte...

Show all 45 references
  1. [9]

    Auditing and instructing text-to-image gener- ation models on fairness.AI and Ethics, 2024

    Felix Friedrich, Manuel Brack, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, Sasha Luccioni, and Kris- tian Kersting. Auditing and instructing text-to-image gener- ation models on fairness.AI and Ethics, 2024. 3, 5

  2. [10]

    Multilingual text-to-image generation magnifies gender stereotypes and prompt engineering may not help you, 2024

    Felix Friedrich, Katharina H ¨ammerl, Patrick Schramowski, Manuel Brack, Jindrich Libovicky, Kristian Kersting, and Alexander Fraser. Multilingual text-to-image generation magnifies gender stereotypes and prompt engineering may not help you, 2024. 5

  3. [11]

    Datacomp: In search of the next generation of multimodal datasets

    Samir Yitzhak Gadre*, Gabriel Ilharco*, Alex Fang*, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Muss...

  4. [12]

    Com- moncanvas: An open diffusion model trained with creative- commons images.arXiv preprint arXiv:2310.16825, 2023

    Aaron Gokaslan, A Feder Cooper, Jasmine Collins, Lan- dan Seguin, Austin Jacobson, Mihir Patel, Jonathan Fran- kle, Cory Stephenson, and V olodymyr Kuleshov. Com- moncanvas: An open diffusion model trained with creative- commons images.arXiv preprint arXiv:2310.16825, 2023. 1, 2

  5. [13]

    Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

    Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large visi...

  6. [14]

    Benchmarking of deep architectures for segmentation of medical images.Trans

    Daniel Gut, Zbislaw Tabor, Mateusz Szymkowski, Milosz Rozynek, Iwona Kucybala, and Wadim Wojciechowski. Benchmarking of deep architectures for segmentation of medical images.Trans. Med. Imaging, 2022. 4

  7. [15]

    CLIPScore: a reference-free evaluation met- ric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: a reference-free evaluation met- ric for image captioning. InEMNLP, 2021. 5

  8. [16]

    Gans trained by a two time-scale update rule converge to a local nash equilib- rium

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. InAdvances in Neural Information Processing Sys- tems, 2017. 5

  9. [17]

    Classifier-free diffusion guidance.arXiv:2207.12598, 2022

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance.arXiv:2207.12598, 2022. 4

  10. [18]

    Fairface: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation

    Kimmo Karkkainen and Jungseock Joo. Fairface: Face at- tribute dataset for balanced race, gender, and age for bias measurement and mitigation. InWACV, pages 1548–1558,

  11. [19]

    Transformers are minimax optimal nonparametric in-context learners

    Juno Kim, Tai Nakamaki, and Taiji Suzuki. Transformers are minimax optimal nonparametric in-context learners. In Proceedings of the Advances in Neural Information Process- ing Systems: Annual Conference on Neural Information Pro- cessing Systems (NeurIPS), 2024. 4

  12. [20]

    Pick-a-pic: An open dataset of user preferences for text-to-image generation

    Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Ma- tiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation. In arxiv, 2023. 5

  13. [21]

    Genai-bench: Evaluat- ing and improving compositional text-to-visual generation,

    Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li, Yixin Fei, Kewen Wu, Tiffany Ling, Xide Xia, Pengchuan Zhang, Gra- ham Neubig, and Deva Ramanan. Genai-bench: Evaluat- ing and improving compositional text-to-visual generation,

  14. [22]

    Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models.arXiv preprint arXiv:2407.07895, 2024

    Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models.arXiv preprint arXiv:2407.07895, 2024. 5

  15. [23]

    Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto

    Hao Li, Yang Zou, Ying Wang, Orchid Majumder, Yusheng Xie, R. Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, and Stefano Soatto. On the scalability of diffusion-based text-to-image generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  16. [24]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning (ICML), 2023. 1

  17. [25]

    MVPTR: multi- level semantic alignment for vision-language pre-training via multi-stage learning

    Zejun Li, Zhihao Fan, Huaixiao Tou, Jingjing Chen, Zhongyu Wei, and Xuanjing Huang. MVPTR: multi- level semantic alignment for vision-language pre-training via multi-stage learning. InMM ’22: The 30th ACM Interna- tional Conference on Multimedia, 2022. 4

  18. [26]

    Playground v3: Im- proving text-to-image alignment with deep-fusion large lan- guage models, 2024

    Bingchen Liu, Ehsan Akhgari, Alexander Visheratin, Aleks Kamko, Linmiao Xu, Shivam Shrirao, Chase Lambert, Joao Souza, Suhail Doshi, and Daiqing Li. Playground v3: Im- proving text-to-image alignment with deep-fusion large lan- guage models, 2024. 1, 2, 3

  19. [27]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InProceedings of the International Confer- ence on Learning Representations (ICLR), 2019. 4

  20. [28]

    Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu, Dirk Groeneveld, Oyvind Tafjord, Noah A

    Ian Magnusson, Nguyen Tai, Ben Bogin, David Heineman, Jena D. Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu, Dirk Groeneveld, Oyvind Tafjord, Noah A. Smith, Pang Wei Koh, and Jesse Dodge. Datadecide: How to predict best pretraining data with small experiments.arXiv prepri...

  21. [29]

    On aliased resizing and surprising subtleties in GAN evaluation

    Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On aliased resizing and surprising subtleties in GAN evaluation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 5

  22. [30]

    Hierarchical text-conditional image gener- ation with CLIP latents.arXiv preprint arXiv:2204.06125,

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gener- ation with CLIP latents.arXiv preprint arXiv:2204.06125,

  23. [31]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 3

  24. [32]

    Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion mo...

  25. [33]

    Fast high- resolution image synthesis with latent adversarial diffusion distillation.arXiv:2403.12015, 2024

    Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast high- resolution image synthesis with latent adversarial diffusion distillation.arXiv:2403.12015, 2024. 2

  26. [34]

    Safe latent diffusion: Mitigating inappro- priate degeneration in diffusion models

    Patrick Schramowski, Manuel Brack, Bj ¨orn Deiseroth, and Kristian Kersting. Safe latent diffusion: Mitigating inappro- priate degeneration in diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3

  27. [35]

    LAION- 400M: open dataset of clip-filtered 400 million image-text pairs

    Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION- 400M: open dataset of clip-filtered 400 million image-text pairs. Preprint athttps://arxiv.org/abs/2111. 02114, 2021. 1

  28. [36]

    Laion-5b: An open large-scale dataset for training next gen- eration image-text models

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. Laion-5b: An op...

  29. [37]

    The bias amplification paradox in text-to-image generation

    Preethi Seshadri, Sameer Singh, and Yanai Elazar. The bias amplification paradox in text-to-image generation. In NAACL, 2024. 3, 8

  30. [38]

    Scaling autoregressive models for content-rich text-to-image generation.Transactions on Machine Learn- ing Research, 2022

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gun- jan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yin- fei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content...

  31. [39]

    Sigmoid loss for language image pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2023. 2

  32. [40]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 5

  33. [41]

    The image

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Bar- rett, and Ying Sheng. Sglang: Efficient execution of structured language model programs.arXiv preprint. arXiv:2312.07104,...

  34. [42]

    Training on long, dense captions results in better prompt following of the downstream text-to-image model

  35. [43]

    Randomizing training caption length does not adversely affect text alignment while providing the benefits out- lined in Sec. 4.2

  36. [44]

    We observe no measurable benefit in additionally diver- sifying captions across epochs

  37. [45]

    psychologist

    Attempting to make caption diversity explicit through personas results in a performance drop D. Inter Epoch Diversity In addition to the experiment discussed in Sec. 4.2, we con- ducted a more rigorous evaluation of inter-epoch diversity for data-constrained scenarios. Experim...

Pith tools