REVIEW 4 major objections 6 minor 69 references
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Training a text-to-image model on only 20% of caption data, selected by a scene-graph detailness score, beats training on the full dataset.
desk verdict Practical caption-detailness metric for T2I data selection, but ICR's construct validity and missing statistical rigor keep the central claim from being fully convincing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the caption detailness score $\mathrm{CD}(x,c)=\mathrm{ICR}(x,c)\times \mathrm{AOD}(x,c)/\mathrm{Length}(c)$. ICR is computed by grounding every object mention with a segmentation model and taking the union mask area over image area; AOD counts scene-graph edges—attributes and relations—averaged over mentioned objects; and the length denominator penalizes verbose and non-visual wording. The scene graph, produced by an instructed large language model, turns raw text into countable object nodes and edges, making both terms computable without human annotation.
What would settle it
Compute, for a sample of images, both the proposed ICR and a recall score against all objects found by an open-vocabulary object detector; if high-ICR captions consistently miss objects the detector finds while low-ICR captions name more total objects, then the coverage ranking is not measuring completeness and the selection advantage should be re-examined.
Extended reading notes
Core claim
The paper's central discovery is that caption detailness for text-to-image training can be measured compositionally from a scene graph instead of being approximated by word count. A caption is parsed into objects, attributes, and relations; ICR (Eq. 2) is the union area of segmentation masks for mentioned objects divided by image area, AOD (Eq. 4) is the average number of attribute and relation edges per object, and the caption detailness score is $\mathrm{CD} = \mathrm{ICR} \times \mathrm{AOD} / \mathrm{Length}(c)$. The paper reports a positive correlation between each component and generation quality, and shows that sequential selection—keeping top-scoring captions by semantic correctness, then by CD—produces a 20,000-image subset that outperforms the full 113,162-image set and the length-based baselines on DPG, MSCOCO, and ImageInWords metrics.
Load-bearing premise
The whole ranking rests on the assumption that the area of the image covered by the objects a caption happens to name tells you how completely the caption covers the image's actual content, even though the metric never compares against the full set of objects in the image.
Editorial extensions
If this is right
- Selection by ICR and AOD with length normalization and a semantic-correctness gate is a usable data-curation pipeline for T2I training sets, not just an evaluation metric.
- Fine-tuning with about 20% of a detailed-caption dataset can match or exceed full-dataset training, lowering the data and compute needed for detailed-caption training.
- Caption annotation budgets should be spent on covering more image regions and adding per-object attributes and relations, and on correctness, rather than on making captions long.
- Length-based filtering heuristics for caption data are a worse proxy than scene-graph detailness when the goal is image-text alignment and image reconstruction.
- The positive ICR/AOD-performance trend is shown with controlled synthetic captions, so more coverage and more object detail per caption are part of the paper's claim, not only a selection heuristic.
Reading between the lines
- An implicit extension is that the same score could be applied during caption generation itself, as a reward or reranking signal for multimodal captioners, not just for post-hoc selection.
- Since ICR is defined against image area rather than against a complete inventory of image objects, a recall-style variant that divides by detected-object area would test whether area coverage tracks semantic coverage; that test is not in the paper.
- The 20% result was shown for one base model and one detailed-caption source; whether the advantage survives other backbones and other detailed caption datasets is an open corollary.
- The paper's inverted-U relationship between semantic correctness and detailness suggests an optimal detailness plateau, which could motivate adaptive detailness targets per dataset rather than a fixed more-is-better rule.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a new caption-detailness metric for text-to-image training data selection, combining an image coverage rate (ICR, Eq. 2), an average object detailness (AOD, Eq. 4), and a length-normalized caption detailness score (CD, Eq. 5). Using ShareGPT4V captions on MS COCO and retraining Lumina-Next-T2I, the authors report that (i) higher ICR/AOD synthetic caption subsets correlate with better DPG/FID/CLIP scores, and (ii) selecting only 20,000 high-CD captions outperforms full-dataset training and length-based filtering on DPG, MSCOCO, and IIW benchmarks. The core claim is that detailness, rather than length, drives data efficiency for text-to-image generation.
Significance. If the construct validity and experimental rigor concerns were addressed, the paper would make a useful contribution: it attacks a real problem (length is a poor proxy for visual detailness), and the proposed pipeline uses external frozen models (Llama3-70B for scene-graph parsing, LISA for segmentation, LLM2CLIP for image-text matching), so the metric is not fitted to the target benchmarks, which reduces circularity concerns. The correlation experiments in Section 4.3 are a sensible way to isolate the influence of ICR and AOD. However, the central construct-validity issue in Eq. (2), the matched-compute question, and the absence of statistical significance tests currently prevent the headline '20% data beats full data' claim from being fully supported.
major comments (4)
- [Section 3.3, Eq. (2)] The ICR as defined is not a coverage measure; it is an area-weighted mention score. Eq. (2) divides the union area of masks for caption-mentioned objects by the total image area, with no ground-truth object set or denominator. A caption naming only a large background region (e.g., sky, wall, road) can therefore receive a high ICR while omitting all small foreground objects. AOD in Eq. (4) has the same blind spot because it averages over mentioned objects only. Since CD in Eq. (5) is built on these two terms, the ranking used for data selection does not necessarily operationalize 'whether the caption covers all regions/objects in the image' as claimed in Section 3.3. Please validate ICR/AOD against human coverage judgments or an oracle object/region set (e.g., dense annotations from COCO) before drawing conclusions from the Section 4.4 filtering results.
- [Section 4.4, Tables 2 and 3] The central 'training on 20% of full data surpasses full-data training' claim is not supported with matched compute. Table 2 shows that random selection of 20,000 samples outperforms the full 113,162-sample run on DPG average (74.78 vs 73.18) and on several sub-scores, which strongly suggests the full-data run is undertrained (e.g., the same number of update steps with fewer samples per step would disadvantage the full-data model). Please report training iterations, epochs, batch sizes, and either match total compute across strategies or show the full-data result at a comparable compute budget. Without this, the headline result may reflect training schedule rather than caption quality.
- [Tables 2, 3, and 4] No error bars, multiple seeds, or significance tests are reported. The margins over the strongest baselines are small: DPG average 76.28 vs 75.62 for ITM-Len in Table 2, CLIP-IS 75.73 vs 75.53, and Table 4 CLIP-S 26.36 vs 26.22. Please report mean and standard deviation over at least three seeds and include a paired significance test (e.g., bootstrap or permutation test) for the key comparisons, because the practical value of a 1–2 point DPG improvement depends on its statistical reliability.
- [Section 4.3, Figure 5 and Table 1] The claim of a 'consistent positive correlation' is overstated. Table 1 shows non-monotonicities: ICR 0.8 yields Global=86.71 while ICR 1.0 yields 84.85, and AOD 0.8 yields Global=80.64 while AOD 0.6 yields 80.79. The text also admits that 'captions with a 60% ICR sometimes outperformed those with 80% ICR on certain metrics.' Please quantify the correlation with confidence intervals or a formal trend test, and temper the monotonicity claim accordingly.
minor comments (6)
- [Equation (5)] There is a typo: 'where is Length(·)is the word counting function' should read 'where Length(·) is the word counting function.'
- [Equation (2)] The notation 'Area(S|O(c)|j=1 {sj})' is ambiguous; please introduce a union operator, e.g., 'Area(∪_{j=1}^{|O(c)|} s_j)', and define what a mask s_j represents.
- [Section 4.1] The sentence 'We partition it into 5,000 image-text pairs for testing while allocating the remaining samples for ...' is incomplete; please specify the exact training split size and any validation data.
- [Figure 1] The caption's contrast between 'Caption A is more detailed than Caption B' and 'Length: Caption A is more detailed / Our metric: Caption A is less detailed' is confusing; clarify that this is the point being illustrated rather than a contradiction.
- [Table 3] The checkmark columns labeled 'ITM ICR AOD' are hard to read; please add an explicit legend or row labels indicating which metric is active in each row.
- [Throughout] Please use consistent dataset naming ('MS COCO 2017' instead of 'MSCOCO2017') and complete the reference [5], which currently lacks a full author list and title.
Circularity Check
The core data-selection claim is tested against external benchmarks and is not circular; only a minor, non-load-bearing same-author citation appears.
full rationale
The paper's central claim, that training on a 20% subset selected by its caption detailness metric beats both full-data training and length-based selection, is evaluated with an open-source base model (Lumina-Next-T2I) on external benchmarks (DPG, IIW, FID, CLIP-S, CLIP-IS). The ICR and AOD values are computed from external components (Llama3-70B scene-graph parsing and LISA segmentation) and are not fitted to the downstream benchmark outcomes, so the selection result is not statistically forced. Equation (2) does have a construct-validity gap: ICR measures the union area of masks for caption-mentioned objects divided by image area, so a caption naming only a large background region can receive a high score while omitting all small objects; this is a validity concern rather than a circularity concern, because the metric is not defined in terms of the target benchmark or the final result. Reference [61], which shares an author, appears only as one of several sources of inspiration for using scene graphs and is not load-bearing for any equation or conclusion. No step in the derivation reduces to its own inputs, and no fitted parameter is renamed as a prediction. The appropriate finding is therefore no significant circularity, with only a minor same-author citation that does not affect the independence of the experiments.
Assumptions & free parameters
free parameters (2)
- Selection sizes K and T =
K=30,000; T=20,000
- Metric composition weights =
ICR^1 x AOD^1 x Length^-1
assumptions (4)
- domain assumption Llama3-70B scene graph parser returns accurate object, relation, and attribute sets for captions
- domain assumption LISA segmentation masks accurately localize every object mentioned in a caption
- ad hoc to paper Union area of mentioned-object masks divided by image area is a valid proxy for semantic coverage of the image
- domain assumption All data selection strategies are compared under equal training budget and convergence
Cite this review
Pith. "Pith review of Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation." pith.science (2026). https://pith.science/paper/EXYWER24
@misc{pith2026250515172,
author = {Pith},
title = {Pith review of: Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/EXYWER24}},
note = {Machine review of arXiv:2505.15172}
}
read the original abstract
Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length to represent the detailness of the caption in the T2I training set. In this paper, we propose a new metric to estimate caption detailness based on two aspects: image coverage rate (ICR), which evaluates whether the caption covers all regions/objects in the image, and average object detailness (AOD), which quantifies the detailness of each object's description. Through experiments on the COCO dataset using ShareGPT4V captions, we demonstrate that T2I models trained on high-ICR and -AOD captions achieve superior performance on DPG and other benchmarks. Notably, our metric enables more effective data selection-training on only 20% of full data surpasses both full-dataset training and length-based selection method, improving alignment and reconstruction ability. These findings highlight the critical role of detail-aware metrics over length-based heuristics in caption selection for T2I tasks.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
SPICE: Semantic Propositional Image Cap- tion Evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. SPICE: Semantic Propositional Image Cap- tion Evaluation. InComputer Vision – ECCV 2016, pages 382–398. Springer International Publishing, Cham, 2016. 3
work page 2016
-
[3]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: a versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023. 2
work page 2023
-
[4]
METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments
Satanjeev Banerjee and Alon Lavie. METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments. InProceedings of the ACL Work- shop on Intrinsic and Extrinsic Evaluation Measures for Ma- chine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan, 2005. Association for Computational Lin- guistics. 2
work page 2005
-
[5]
Improving image generation with better captions.Computer Science
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, and others. Improving image generation with better captions.Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023. 2
work page 2023
-
[6]
PixLore: A Dataset-driven Approach to Rich Image Captioning, 2024
Diego Bonilla-Salvador, Marcelino Mart ´ınez-Sober, Joan Vila-Franc ´es, Antonio Jos ´e Serrano-L ´opez, Pablo Rodr´ıguez-Belenguer, and Fernando Mateo. PixLore: A Dataset-driven Approach to Rich Image Captioning, 2024. 2
work page 2024
-
[7]
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, S...
work page 1901
-
[8]
Junsong Chen, Jincheng YU, Chongjian GE, Lewei Yao, Enze Xie, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. PixArt- {\textbackslashalpha\: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthe- sis. InThe Twelfth International Conference on Learning Representations, 2024. 1
work page 2024
Show all 69 references
-
[9]
ShareGPT4V: Improving Large Multi-modal Models with Better Captions
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. ShareGPT4V: Improving Large Multi-modal Models with Better Captions. InComputer Vision – ECCV 2024, pages 370–387. Springer Nature Switzerland, Cham, 2025. 2, 4
2024
-
[10]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality, 2023. 2
2023
-
[11]
Davidsonian Scene Graph: Improving Relia- bility in Fine-grained Evaluation for Text-to-Image Genera- tion
Jaemin Cho, Yushi Hu, Jason Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang. Davidsonian Scene Graph: Improving Relia- bility in Fine-grained Evaluation for Text-to-Image Genera- tion. InICLR, 2024. 2, 3
2024
-
[12]
Instructblip: towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: towards general-purpose vision-language models with instruction tuning. InThirty-seventh Conference on Neural Information Processing Systems, 2023. 2
2023
-
[13]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shi- rong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei...
2025 arXiv
-
[14]
CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers. InAdvances in Neural Informa- tion Processing Systems, 2022. 2
2022
-
[15]
Benchmarking and Improv- ing Detail Image Caption, 2024
Hongyuan Dong, Jiawen Li, Bohong Wu, Jiacong Wang, Yuan Zhang, and Haoyuan Guo. Benchmarking and Improv- ing Detail Image Caption, 2024. arXiv:2405.19092. 2, 3
2024 arXiv
-
[16]
An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020
Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020. 2
2010 arXiv
-
[17]
Scaling rec- tified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, and others. Scaling rec- tified flow transformers for high-resolution image synthesis. InForty-first international conference on ...
-
[18]
Lumina-T2X: Scalable Flow- based Large Diffusion Transformer for Flexible Resolution Generation
Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, Tong He, Jingwen He, Junjun He, Yu Qiao, and Hongsheng Li. Lumina-T2X: Scalable Flow- base...
2025
-
[19]
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bun- ner, Ranjay Krishna, Jason Michael Baldridge, and Radu Soricut. ImageInWords: Unlocking Hyper-Detailed Image Descriptions. InProc. EMNLP 2024, pages 93–127, Miami, Flor...
2024
-
[20]
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools, 2024
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Jingyu Sun, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong...
2024 arXiv
-
[21]
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. InAdvances in Neural Information Processing Systems. Curran Associates, Inc., 2014. 2
2014
-
[22]
Generative adversarial networks.Commun
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks.Commun. ACM, 63(11):139–144, 2020. 2
2020
-
[23]
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. InProc. 2021 Conf. Empirical Methods NLP, pages 7514–7528, Online and Punta Cana, Dominican Republic, 2021. Association for Computation...
2021
-
[24]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017. 2
2017
-
[25]
Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans. Classifier-Free Diffusion Guidance. InNeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 2
2021
-
[26]
Denoising Dif- fusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Dif- fusion Probabilistic Models. InAdvances in Neural Infor- mation Processing Systems, pages 6840–6851. Curran Asso- ciates, Inc., 2020. 2
2020
-
[27]
Faith- Score: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models
Liqiang Jing, Ruosen Li, Yunmo Chen, and Xinya Du. Faith- Score: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 5042– 5063, Miami, Florida, USA, 2024. Association for Co...
2024
-
[28]
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
Kibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In, Jinyoung Moon, Donghyun Kim, and Chanyoung Park. LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 28306–28316, ...
2024
-
[29]
LISA: Reasoning Segmenta- tion via Large Language Model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: Reasoning Segmenta- tion via Large Language Model. In2024 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), pages 9579–9589, Seattle, W A, USA, 2024. IEEE. 3
2024
-
[30]
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
Zhengfeng Lai, Vasileios Saveris, Chen Chen, Hong-You Chen, Haotian Zhang, Bowen Zhang, Wenze Hu, Juan Lao Tebar, Zhe Gan, Peter Grasch, Meng Cao, and Yinfei Yang. Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models. InThe Thirteenth Interna- ti...
2025
-
[31]
From Pixels to Graphs: Open-V ocabulary Scene Graph Generation with Vision-Language Models
Rongjie Li, Songyang Zhang, Dahua Lin, Kai Chen, and Xuming He. From Pixels to Graphs: Open-V ocabulary Scene Graph Generation with Vision-Language Models. In 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 28076–28086, Seattle, W A, USA, 20...
2024
-
[32]
What If We Recaption Billions of Web Images with LLaMA-3?, 2024
Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, Yuyin Zhou, and Cihang Xie. What If We Recaption Billions of Web Images with LLaMA-3?, 2024. arXiv:2406.08478. 2
2024 arXiv
-
[33]
DenseFusion-1M: Merg- 10 ing Vision Experts for Comprehensive Multimodal Percep- tion
Xiaotong Li, Fan Zhang, Haiwen Diao, Yueze Wang, Xin- long Wang, and LINGYU DUAN. DenseFusion-1M: Merg- 10 ing Vision Experts for Comprehensive Multimodal Percep- tion. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. 2
2024
-
[34]
FACTUAL: A Benchmark for Faithful and Consistent Tex- tual Scene Graph Parsing
Zhuang Li, Yuyang Chai, Terry Yue Zhuo, Lizhen Qu, Gho- lamreza Haffari, Fei Li, Donghong Ji, and Quan Hung Tran. FACTUAL: A Benchmark for Faithful and Consistent Tex- tual Scene Graph Parsing. InFindings of the Association for Computational Linguistics: ACL 2023, pages 6377–6...
2023
-
[35]
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Under- standing, 2024
Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, Dayou Chen, Jiajun He, Jia- hao Li, Wenyue Li, Chen Zhang, Rongwei Quan, Jianxiang Lu, Jiabin Huang, Xiaoyan Yuan, Xiaoxiao Zheng, Yixuan Li, ...
2024 arXiv
-
[36]
ROUGE: A Package for Automatic Evalu- ation of Summaries
Chin-Yew Lin. ROUGE: A Package for Automatic Evalu- ation of Summaries. InText Summarization Branches Out, pages 74–81, Barcelona, Spain, 2004. Association for Com- putational Linguistics. 2
2004
-
[37]
Lawrence Zitnick
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft COCO: Common Objects in Context. In Computer Vision – ECCV 2014, pages 740–755. Springer In- ternational Publishing, Cham, 2014. 1, 2, 4
2014
-
[38]
Playground v3: Im- proving Text-to-Image Alignment with Deep-Fusion Large Language Models, 2024
Bingchen Liu, Ehsan Akhgari, Alexander Visheratin, Aleks Kamko, Linmiao Xu, Shivam Shrirao, Chase Lambert, Joao Souza, Suhail Doshi, and Daiqing Li. Playground v3: Im- proving Text-to-Image Alignment with Deep-Fusion Large Language Models, 2024. arXiv:2409.10695. 2
2024 arXiv
-
[39]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdvances in Neural Information Processing Systems, pages 34892–34916. Curran Associates, Inc., 2023. 2
2023
-
[40]
Improving Long-Text Alignment for Text-to-Image Diffusion Models
Luping Liu, Chao Du, Tianyu Pang, Zehan Wang, Chongx- uan Li, and Dong Xu. Improving Long-Text Alignment for Text-to-Image Diffusion Models. InThe Thirteenth Interna- tional Conference on Learning Representations, 2025. 2
2025
-
[41]
Tenenbaum, and Antonio Torralba
Nan Liu, Yilun Du, Shuang Li, Joshua B. Tenenbaum, and Antonio Torralba. Unsupervised Compositional Concepts Discovery with Text-to-Image Generative Models. In2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 2085–2095, Paris, France, 2023. IEEE. 2
2023
-
[42]
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning,
Fan Lu, Wei Wu, Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu, Wei Zhai, Yang Cao, Yujun Shen, and Zheng- Jun Zha. Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning,
-
[43]
DOCCI: Descriptions of Connected and Con- trasting Images
Yasumasa Onoe, Sunayana Rane, Zachary Berger, Yonatan Bitton, Jaemin Cho, Roopal Garg, Alexander Ku, Zarana Parekh, Jordi Pont-Tuset, Garrett Tanzer, Su Wang, and Ja- son Baldridge. DOCCI: Descriptions of Connected and Con- trasting Images. InComputer Vision – ECCV 2024, pages...
2024
-
[44]
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a Method for Automatic Evaluation of Machine Translation. InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311– 318, Philadelphia, Pennsylvania, USA, 2002. Associ...
2002
-
[45]
Scalable Diffusion Models with Transformers
William Peebles and Saining Xie. Scalable Diffusion Models with Transformers. In2023 IEEE/CVF International Confer- ence on Computer Vision (ICCV), pages 4172–4182, Paris, France, 2023. IEEE. 2
2023
-
[46]
Image Textualization: An Auto- matic Framework for Generating Rich and Detailed Image Descriptions
Renjie Pi, Jianshu Zhang, Jipeng Zhang, Rui Pan, Zhekai Chen, and Tong Zhang. Image Textualization: An Auto- matic Framework for Generating Rich and Detailed Image Descriptions. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Trac...
2024
-
[47]
SDXL: Improving Latent Diffusion Mod- els for High-Resolution Image Synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. SDXL: Improving Latent Diffusion Mod- els for High-Resolution Image Synthesis. InThe Twelfth In- ternational Conference on Learning Representations, 2024. 1
2024
-
[48]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and others. Learning transferable visual models from natural language supervision. InInternational conference on machine learn- i...
2021
-
[49]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.Journal of Machine Learning Research, 21(140):1–67, 2020. 2
2020
-
[50]
Zero-Shot Text-to-Image Generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-Shot Text-to-Image Generation. InProceedings of the 38th International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 1
2021
-
[51]
Generative Ad- versarial Text to Image Synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Lo- geswaran, Bernt Schiele, and Honglak Lee. Generative Ad- versarial Text to Image Synthesis. InProceedings of The 33rd International Conference on Machine Learning, pages 1060–1069, New York, New York, USA, 2016. PMLR
2016
-
[52]
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10674–10685, New Orleans, LA, USA, 2022. IEEE. 1, 2
2022
-
[53]
Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa R
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W. Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa R. Kundurthy, Kather- ine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and 11 Jenia ...
2022
-
[54]
Conceptual Captions: A Cleaned, Hypernymed, Im- age Alt-text Dataset For Automatic Image Captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual Captions: A Cleaned, Hypernymed, Im- age Alt-text Dataset For Automatic Image Captioning. In 56th Annual Meeting of the Association for Computational Linguistics, pages 2556–2565, Melbourne, Australia, 20...
2018
-
[55]
Denois- ing Diffusion Implicit Models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing Diffusion Implicit Models. InInternational Conference on Learning Representations, 2021. 2
2021
-
[56]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi`ere, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, L ´eonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, A...
2024 arXiv
-
[57]
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
Jack Urbanek, Florian Bordes, Pietro Astolfi, Mary Williamson, Vasu Sharma, and Adriana Romero-Soriano. A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pa...
2024
-
[58]
Lawrence Zitnick, and Devi Parikh
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. CIDEr: Consensus-based image description evalua- tion. In2015 IEEE Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 4566–4575, Boston, MA, USA, 2015. IEEE. 2
2015
-
[59]
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution, 2024
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Jun- yang Lin. Qwen2-VL: Enhancing Vision-Language Model’s ...
2024
-
[60]
Emu3: Next-Token Prediction is All You Need, 2024
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tie...
2024 arXiv
-
[61]
Detailed Object Description with Controllable Dimensions, 2025
Xinran Wang, Haiwen Zhang, Baoteng Li, Kongming Liang, Hao Sun, Zhongjiang He, Zhanyu Ma, and Jun Guo. Detailed Object Description with Controllable Dimensions, 2025. arXiv:2411.19106. 3
2025 arXiv
-
[62]
LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation
Aoqi Wu, weiquan Huang, Yifan Yang, Xufang Luo, Yuqing Yang, Chunyu Wang, Liang Hu, Xiyang Dai, Dongdong Chen, Chong Luo, and Lili Qiu. LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation. In NeurIPS 2024 Workshop: Self-Supervised Learning - The- ory and Prac...
2024
-
[63]
BoxDiff: Text-to-Image Synthesis with Training-Free Box- Constrained Diffusion
Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou. BoxDiff: Text-to-Image Synthesis with Training-Free Box- Constrained Diffusion. In2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7418–7427,
-
[64]
Choy, and Li Fei- Fei
Danfei Xu, Yuke Zhu, Christopher B. Choy, and Li Fei- Fei. Scene Graph Generation by Iterative Message Passing. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3097–3106, Honolulu, HI, 2017. IEEE. 3
2017
-
[65]
Peter Young, Alice Lai, Micah Hodosh, and Julia Hocken- maier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.Transactions of the Association for Computational Linguistics, 2:67–78, 2014. 1
2014
-
[66]
ITI-Gen: Inclusive Text-to-Image Generation
Cheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, and Fernando De La Torre. ITI-Gen: Inclusive Text-to-Image Generation. In2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3946–3957, 2023. 2
2023
-
[67]
OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding
Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, and Shuicheng Y AN. OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding. InThe Thirty-eighth Annual Conference on Neural Information Processing Sys- t...
2024
-
[68]
Lumina-Next : Mak- ing Lumina-T2X Stronger and Faster with Next-DiT
Le Zhuo, Ruoyi Du, Han Xiao, Yangguang Li, Dongyang Liu, Rongjie Huang, Wenze Liu, Xiangyang Zhu, Fu-Yun Wang, Zhanyu Ma, Xu Luo, Zehan Wang, Kaipeng Zhang, Lirui Zhao, Si Liu, Xiangyu Yue, Wanli Ouyang, Yu Qiao, Hongsheng Li, and Peng Gao. Lumina-Next : Mak- ing Lumina-T2X St...
2024
-
[2024]
arXiv:2412.08614. 2, 3
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.