Pith. sign in

REVIEW 4 major objections 6 minor 69 references

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Training a text-to-image model on only 20% of caption data, selected by a scene-graph detailness score, beats training on the full dataset.

desk verdict Practical caption-detailness metric for T2I data selection, but ICR's construct validity and missing statistical rigor keep the central claim from being fully convincing. read the letter →

arxiv 2505.15172 v1 pith:EXYWER24 submitted 2025-05-21 cs.CV

classification cs.CV
keywords captiondetailnesstext-to-imagegenerationdataselectionscenegraphimagecoveragerateaverageobjectdiffusiontransformerlongpromptfollowing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Text-to-image models improve when trained on richly detailed captions, but choosing which captions are genuinely detailed is usually done by caption length. This paper proposes a two-part detailness score: image coverage rate (ICR), the share of image area covered by segmented objects named in the caption, times average object detailness (AOD), the average number of attributes and relations per mentioned object, normalized by caption length. The claim is that this score identifies captions whose visual content is dense and concise, and that training a T2I model on just 20% of the dataset selected this way beats training on the full dataset and beats length-based selection on alignment and reconstruction benchmarks. If true, it makes detailed-caption training much cheaper and gives a principled alternative to length-based data curation.

What carries the argument

The load-bearing object is the caption detailness score $\mathrm{CD}(x,c)=\mathrm{ICR}(x,c)\times \mathrm{AOD}(x,c)/\mathrm{Length}(c)$. ICR is computed by grounding every object mention with a segmentation model and taking the union mask area over image area; AOD counts scene-graph edges—attributes and relations—averaged over mentioned objects; and the length denominator penalizes verbose and non-visual wording. The scene graph, produced by an instructed large language model, turns raw text into countable object nodes and edges, making both terms computable without human annotation.

What would settle it

Compute, for a sample of images, both the proposed ICR and a recall score against all objects found by an open-vocabulary object detector; if high-ICR captions consistently miss objects the detector finds while low-ICR captions name more total objects, then the coverage ranking is not measuring completeness and the selection advantage should be re-examined.

Watch

Extended reading notes

Core claim

The paper's central discovery is that caption detailness for text-to-image training can be measured compositionally from a scene graph instead of being approximated by word count. A caption is parsed into objects, attributes, and relations; ICR (Eq. 2) is the union area of segmentation masks for mentioned objects divided by image area, AOD (Eq. 4) is the average number of attribute and relation edges per object, and the caption detailness score is $\mathrm{CD} = \mathrm{ICR} \times \mathrm{AOD} / \mathrm{Length}(c)$. The paper reports a positive correlation between each component and generation quality, and shows that sequential selection—keeping top-scoring captions by semantic correctness, then by CD—produces a 20,000-image subset that outperforms the full 113,162-image set and the length-based baselines on DPG, MSCOCO, and ImageInWords metrics.

Load-bearing premise

The whole ranking rests on the assumption that the area of the image covered by the objects a caption happens to name tells you how completely the caption covers the image's actual content, even though the metric never compares against the full set of objects in the image.

Editorial extensions

If this is right

  • Selection by ICR and AOD with length normalization and a semantic-correctness gate is a usable data-curation pipeline for T2I training sets, not just an evaluation metric.
  • Fine-tuning with about 20% of a detailed-caption dataset can match or exceed full-dataset training, lowering the data and compute needed for detailed-caption training.
  • Caption annotation budgets should be spent on covering more image regions and adding per-object attributes and relations, and on correctness, rather than on making captions long.
  • Length-based filtering heuristics for caption data are a worse proxy than scene-graph detailness when the goal is image-text alignment and image reconstruction.
  • The positive ICR/AOD-performance trend is shown with controlled synthetic captions, so more coverage and more object detail per caption are part of the paper's claim, not only a selection heuristic.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implicit extension is that the same score could be applied during caption generation itself, as a reward or reranking signal for multimodal captioners, not just for post-hoc selection.
  • Since ICR is defined against image area rather than against a complete inventory of image objects, a recall-style variant that divides by detected-object area would test whether area coverage tracks semantic coverage; that test is not in the paper.
  • The 20% result was shown for one base model and one detailed-caption source; whether the advantage survives other backbones and other detailed caption datasets is an open corollary.
  • The paper's inverted-U relationship between semantic correctness and detailness suggests an optimal detailness plateau, which could motivate adaptive detailness targets per dataset rather than a fixed more-is-better rule.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a new caption-detailness metric for text-to-image training data selection, combining an image coverage rate (ICR, Eq. 2), an average object detailness (AOD, Eq. 4), and a length-normalized caption detailness score (CD, Eq. 5). Using ShareGPT4V captions on MS COCO and retraining Lumina-Next-T2I, the authors report that (i) higher ICR/AOD synthetic caption subsets correlate with better DPG/FID/CLIP scores, and (ii) selecting only 20,000 high-CD captions outperforms full-dataset training and length-based filtering on DPG, MSCOCO, and IIW benchmarks. The core claim is that detailness, rather than length, drives data efficiency for text-to-image generation.

Significance. If the construct validity and experimental rigor concerns were addressed, the paper would make a useful contribution: it attacks a real problem (length is a poor proxy for visual detailness), and the proposed pipeline uses external frozen models (Llama3-70B for scene-graph parsing, LISA for segmentation, LLM2CLIP for image-text matching), so the metric is not fitted to the target benchmarks, which reduces circularity concerns. The correlation experiments in Section 4.3 are a sensible way to isolate the influence of ICR and AOD. However, the central construct-validity issue in Eq. (2), the matched-compute question, and the absence of statistical significance tests currently prevent the headline '20% data beats full data' claim from being fully supported.

major comments (4)
  1. [Section 3.3, Eq. (2)] The ICR as defined is not a coverage measure; it is an area-weighted mention score. Eq. (2) divides the union area of masks for caption-mentioned objects by the total image area, with no ground-truth object set or denominator. A caption naming only a large background region (e.g., sky, wall, road) can therefore receive a high ICR while omitting all small foreground objects. AOD in Eq. (4) has the same blind spot because it averages over mentioned objects only. Since CD in Eq. (5) is built on these two terms, the ranking used for data selection does not necessarily operationalize 'whether the caption covers all regions/objects in the image' as claimed in Section 3.3. Please validate ICR/AOD against human coverage judgments or an oracle object/region set (e.g., dense annotations from COCO) before drawing conclusions from the Section 4.4 filtering results.
  2. [Section 4.4, Tables 2 and 3] The central 'training on 20% of full data surpasses full-data training' claim is not supported with matched compute. Table 2 shows that random selection of 20,000 samples outperforms the full 113,162-sample run on DPG average (74.78 vs 73.18) and on several sub-scores, which strongly suggests the full-data run is undertrained (e.g., the same number of update steps with fewer samples per step would disadvantage the full-data model). Please report training iterations, epochs, batch sizes, and either match total compute across strategies or show the full-data result at a comparable compute budget. Without this, the headline result may reflect training schedule rather than caption quality.
  3. [Tables 2, 3, and 4] No error bars, multiple seeds, or significance tests are reported. The margins over the strongest baselines are small: DPG average 76.28 vs 75.62 for ITM-Len in Table 2, CLIP-IS 75.73 vs 75.53, and Table 4 CLIP-S 26.36 vs 26.22. Please report mean and standard deviation over at least three seeds and include a paired significance test (e.g., bootstrap or permutation test) for the key comparisons, because the practical value of a 1–2 point DPG improvement depends on its statistical reliability.
  4. [Section 4.3, Figure 5 and Table 1] The claim of a 'consistent positive correlation' is overstated. Table 1 shows non-monotonicities: ICR 0.8 yields Global=86.71 while ICR 1.0 yields 84.85, and AOD 0.8 yields Global=80.64 while AOD 0.6 yields 80.79. The text also admits that 'captions with a 60% ICR sometimes outperformed those with 80% ICR on certain metrics.' Please quantify the correlation with confidence intervals or a formal trend test, and temper the monotonicity claim accordingly.
minor comments (6)
  1. [Equation (5)] There is a typo: 'where is Length(·)is the word counting function' should read 'where Length(·) is the word counting function.'
  2. [Equation (2)] The notation 'Area(S|O(c)|j=1 {sj})' is ambiguous; please introduce a union operator, e.g., 'Area(∪_{j=1}^{|O(c)|} s_j)', and define what a mask s_j represents.
  3. [Section 4.1] The sentence 'We partition it into 5,000 image-text pairs for testing while allocating the remaining samples for ...' is incomplete; please specify the exact training split size and any validation data.
  4. [Figure 1] The caption's contrast between 'Caption A is more detailed than Caption B' and 'Length: Caption A is more detailed / Our metric: Caption A is less detailed' is confusing; clarify that this is the point being illustrated rather than a contradiction.
  5. [Table 3] The checkmark columns labeled 'ITM ICR AOD' are hard to read; please add an explicit legend or row labels indicating which metric is active in each row.
  6. [Throughout] Please use consistent dataset naming ('MS COCO 2017' instead of 'MSCOCO2017') and complete the reference [5], which currently lacks a full author list and title.

Circularity Check

0 steps flagged · score 1.0 of 10

The core data-selection claim is tested against external benchmarks and is not circular; only a minor, non-load-bearing same-author citation appears.

full rationale

The paper's central claim, that training on a 20% subset selected by its caption detailness metric beats both full-data training and length-based selection, is evaluated with an open-source base model (Lumina-Next-T2I) on external benchmarks (DPG, IIW, FID, CLIP-S, CLIP-IS). The ICR and AOD values are computed from external components (Llama3-70B scene-graph parsing and LISA segmentation) and are not fitted to the downstream benchmark outcomes, so the selection result is not statistically forced. Equation (2) does have a construct-validity gap: ICR measures the union area of masks for caption-mentioned objects divided by image area, so a caption naming only a large background region can receive a high score while omitting all small objects; this is a validity concern rather than a circularity concern, because the metric is not defined in terms of the target benchmark or the final result. Reference [61], which shares an author, appears only as one of several sources of inspiration for using scene graphs and is not load-bearing for any equation or conclusion. No step in the derivation reduces to its own inputs, and no fitted parameter is renamed as a prediction. The appropriate finding is therefore no significant circularity, with only a minor same-author citation that does not affect the independence of the experiments.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the validity of ICR and AOD as detailness measures, on the parser and segmenter accuracy, and on a fair training comparison. The metric composition and selection thresholds are hand-chosen free parameters. No new physical or architectural entities are introduced.

free parameters (2)
  • Selection sizes K and T = K=30,000; T=20,000
    Section 4.4 chooses these thresholds by hand. T=20k is 17.7% of 113,162, though the paper calls it 20%. Results could shift with different K/T choices.
  • Metric composition weights = ICR^1 x AOD^1 x Length^-1
    Eq. 5 multiplies ICR and AOD and divides by word count with no sensitivity analysis. The unweighted product and inverse-length penalty are ad hoc choices that are not derived or validated.
assumptions (4)
  • domain assumption Llama3-70B scene graph parser returns accurate object, relation, and attribute sets for captions
    Section 3.2 uses Llama3-70B to extract O(c), R(c), A(c), which feed both AOD and ICR. Parser errors directly change the metric scores.
  • domain assumption LISA segmentation masks accurately localize every object mentioned in a caption
    Section 3.3, Eq. 2 computes ICR from the union area of these masks. Missed, merged, or imprecise masks change ICR and therefore the data ranking.
  • ad hoc to paper Union area of mentioned-object masks divided by image area is a valid proxy for semantic coverage of the image
    No ground-truth object inventory is used. A caption naming only a large background object scores high ICR despite omitting many objects, so the metric may not measure what its name claims.
  • domain assumption All data selection strategies are compared under equal training budget and convergence
    Section 4.4 compares full-data training against 20% subsets without stating step or epoch counts. Random selection beating full-data training suggests this assumption may be violated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation." pith.science (2026). https://pith.science/paper/EXYWER24

@misc{pith2026250515172,
  author       = {Pith},
  title        = {Pith review of: Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EXYWER24}},
  note         = {Machine review of arXiv:2505.15172}
}
read the original abstract

Training text-to-image (T2I) models with detailed captions can significantly improve their generation quality. Existing methods often rely on simplistic metrics like caption length to represent the detailness of the caption in the T2I training set. In this paper, we propose a new metric to estimate caption detailness based on two aspects: image coverage rate (ICR), which evaluates whether the caption covers all regions/objects in the image, and average object detailness (AOD), which quantifies the detailness of each object's description. Through experiments on the COCO dataset using ShareGPT4V captions, we demonstrate that T2I models trained on high-ICR and -AOD captions achieve superior performance on DPG and other benchmarks. Notably, our metric enables more effective data selection-training on only 20% of full data surpasses both full-dataset training and length-based selection method, improving alignment and reconstruction ability. These findings highlight the critical role of detail-aware metrics over length-based heuristics in caption selection for T2I tasks.

Figures

Figures reproduced from arXiv: 2505.15172 by the authors.

Figure 1
Figure 1. Caption length is not a perfect indicator of caption detailness. In this paper, we propose a new metric to more effectively estimate [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Illustration of calculating average object detailness (AOD) and image coverage rate (ICR). The image caption is first parsed into [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. An illustration of how scene graphs are sampled to pro [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Generation examples of models trained by captions of different ICR ratios and AOD ratios using DPG prompts. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Text-to-image performance of Lumina-Next-T2I fine [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 7
Figure 7. Figure 7: Impact of AOD and ICR on ITM Score: Captions with [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 6
Figure 6. Figure 6: Comparison of Pearson correlation coefficients between [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 8
Figure 8. Figure 8: Generated image examples from DPG benchmark. The blue text highlights matching details generated by our method. [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

69 extracted references · 59 canonical work pages

  1. [1]

    Llama 3 Model Card

    AI@Meta. Llama 3 Model Card. 2024. 2, 3

  2. [2]

    SPICE: Semantic Propositional Image Cap- tion Evaluation

    Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. SPICE: Semantic Propositional Image Cap- tion Evaluation. InComputer Vision – ECCV 2016, pages 382–398. Springer International Publishing, Cham, 2016. 3

  3. [3]

    Qwen-vl: a versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: a versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023. 2

  4. [4]

    METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments

    Satanjeev Banerjee and Alon Lavie. METEOR: An Auto- matic Metric for MT Evaluation with Improved Correlation with Human Judgments. InProceedings of the ACL Work- shop on Intrinsic and Extrinsic Evaluation Measures for Ma- chine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan, 2005. Association for Computational Lin- guistics. 2

  5. [5]

    Improving image generation with better captions.Computer Science

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, and others. Improving image generation with better captions.Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023. 2

  6. [6]

    PixLore: A Dataset-driven Approach to Rich Image Captioning, 2024

    Diego Bonilla-Salvador, Marcelino Mart ´ınez-Sober, Joan Vila-Franc ´es, Antonio Jos ´e Serrano-L ´opez, Pablo Rodr´ıguez-Belenguer, and Fernando Mateo. PixLore: A Dataset-driven Approach to Rich Image Captioning, 2024. 2

  7. [7]

    Language Models are Few-Shot Learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, S...

  8. [8]

    PixArt- {\textbackslashalpha\: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthe- sis

    Junsong Chen, Jincheng YU, Chongjian GE, Lewei Yao, Enze Xie, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. PixArt- {\textbackslashalpha\: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthe- sis. InThe Twelfth International Conference on Learning Representations, 2024. 1

Show all 69 references
  1. [9]

    ShareGPT4V: Improving Large Multi-modal Models with Better Captions

    Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. ShareGPT4V: Improving Large Multi-modal Models with Better Captions. InComputer Vision – ECCV 2024, pages 370–387. Springer Nature Switzerland, Cham, 2025. 2, 4

  2. [10]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality, 2023. 2

  3. [11]

    Davidsonian Scene Graph: Improving Relia- bility in Fine-grained Evaluation for Text-to-Image Genera- tion

    Jaemin Cho, Yushi Hu, Jason Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang. Davidsonian Scene Graph: Improving Relia- bility in Fine-grained Evaluation for Text-to-Image Genera- tion. InICLR, 2024. 2, 3

  4. [12]

    Instructblip: towards general-purpose vision-language models with instruction tuning

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: towards general-purpose vision-language models with instruction tuning. InThirty-seventh Conference on Neural Information Processing Systems, 2023. 2

  5. [13]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shi- rong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei...

  6. [14]

    CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

    Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers. InAdvances in Neural Informa- tion Processing Systems, 2022. 2

  7. [15]

    Benchmarking and Improv- ing Detail Image Caption, 2024

    Hongyuan Dong, Jiawen Li, Bohong Wu, Jiacong Wang, Yuan Zhang, and Haoyuan Guo. Benchmarking and Improv- ing Detail Image Caption, 2024. arXiv:2405.19092. 2, 3

  8. [16]

    An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020

    Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020. 2

  9. [17]

    Scaling rec- tified flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, and others. Scaling rec- tified flow transformers for high-resolution image synthesis. InForty-first international conference on ...

  10. [18]

    Lumina-T2X: Scalable Flow- based Large Diffusion Transformer for Flexible Resolution Generation

    Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xie, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, Tong He, Jingwen He, Junjun He, Yu Qiao, and Hongsheng Li. Lumina-T2X: Scalable Flow- base...

  11. [19]

    ImageInWords: Unlocking Hyper-Detailed Image Descriptions

    Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bun- ner, Ranjay Krishna, Jason Michael Baldridge, and Radu Soricut. ImageInWords: Unlocking Hyper-Detailed Image Descriptions. InProc. EMNLP 2024, pages 93–127, Miami, Flor...

  12. [20]

    ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools, 2024

    Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Jingyu Sun, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong...

  13. [21]

    Generative Adversarial Nets

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. InAdvances in Neural Information Processing Systems. Curran Associates, Inc., 2014. 2

  14. [22]

    Generative adversarial networks.Commun

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks.Commun. ACM, 63(11):139–144, 2020. 2

  15. [23]

    CLIPScore: A Reference-free Evaluation Metric for Image Captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. InProc. 2021 Conf. Empirical Methods NLP, pages 7514–7528, Online and Punta Cana, Dominican Republic, 2021. Association for Computation...

  16. [24]

    Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems, 30, 2017. 2

  17. [25]

    Classifier-Free Diffusion Guidance

    Jonathan Ho and Tim Salimans. Classifier-Free Diffusion Guidance. InNeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 2

  18. [26]

    Denoising Dif- fusion Probabilistic Models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Dif- fusion Probabilistic Models. InAdvances in Neural Infor- mation Processing Systems, pages 6840–6851. Curran Asso- ciates, Inc., 2020. 2

  19. [27]

    Faith- Score: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

    Liqiang Jing, Ruosen Li, Yunmo Chen, and Xinya Du. Faith- Score: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 5042– 5063, Miami, Florida, USA, 2024. Association for Co...

  20. [28]

    LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation

    Kibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In, Jinyoung Moon, Donghyun Kim, and Chanyoung Park. LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 28306–28316, ...

  21. [29]

    LISA: Reasoning Segmenta- tion via Large Language Model

    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: Reasoning Segmenta- tion via Large Language Model. In2024 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR), pages 9579–9589, Seattle, W A, USA, 2024. IEEE. 3

  22. [30]

    Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models

    Zhengfeng Lai, Vasileios Saveris, Chen Chen, Hong-You Chen, Haotian Zhang, Bowen Zhang, Wenze Hu, Juan Lao Tebar, Zhe Gan, Peter Grasch, Meng Cao, and Yinfei Yang. Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models. InThe Thirteenth Interna- ti...

  23. [31]

    From Pixels to Graphs: Open-V ocabulary Scene Graph Generation with Vision-Language Models

    Rongjie Li, Songyang Zhang, Dahua Lin, Kai Chen, and Xuming He. From Pixels to Graphs: Open-V ocabulary Scene Graph Generation with Vision-Language Models. In 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 28076–28086, Seattle, W A, USA, 20...

  24. [32]

    What If We Recaption Billions of Web Images with LLaMA-3?, 2024

    Xianhang Li, Haoqin Tu, Mude Hui, Zeyu Wang, Bingchen Zhao, Junfei Xiao, Sucheng Ren, Jieru Mei, Qing Liu, Huangjie Zheng, Yuyin Zhou, and Cihang Xie. What If We Recaption Billions of Web Images with LLaMA-3?, 2024. arXiv:2406.08478. 2

  25. [33]

    DenseFusion-1M: Merg- 10 ing Vision Experts for Comprehensive Multimodal Percep- tion

    Xiaotong Li, Fan Zhang, Haiwen Diao, Yueze Wang, Xin- long Wang, and LINGYU DUAN. DenseFusion-1M: Merg- 10 ing Vision Experts for Comprehensive Multimodal Percep- tion. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. 2

  26. [34]

    FACTUAL: A Benchmark for Faithful and Consistent Tex- tual Scene Graph Parsing

    Zhuang Li, Yuyang Chai, Terry Yue Zhuo, Lizhen Qu, Gho- lamreza Haffari, Fei Li, Donghong Ji, and Quan Hung Tran. FACTUAL: A Benchmark for Faithful and Consistent Tex- tual Scene Graph Parsing. InFindings of the Association for Computational Linguistics: ACL 2023, pages 6377–6...

  27. [35]

    Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Under- standing, 2024

    Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, Dayou Chen, Jiajun He, Jia- hao Li, Wenyue Li, Chen Zhang, Rongwei Quan, Jianxiang Lu, Jiabin Huang, Xiaoyan Yuan, Xiaoxiao Zheng, Yixuan Li, ...

  28. [36]

    ROUGE: A Package for Automatic Evalu- ation of Summaries

    Chin-Yew Lin. ROUGE: A Package for Automatic Evalu- ation of Summaries. InText Summarization Branches Out, pages 74–81, Barcelona, Spain, 2004. Association for Com- putational Linguistics. 2

  29. [37]

    Lawrence Zitnick

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C. Lawrence Zitnick. Microsoft COCO: Common Objects in Context. In Computer Vision – ECCV 2014, pages 740–755. Springer In- ternational Publishing, Cham, 2014. 1, 2, 4

  30. [38]

    Playground v3: Im- proving Text-to-Image Alignment with Deep-Fusion Large Language Models, 2024

    Bingchen Liu, Ehsan Akhgari, Alexander Visheratin, Aleks Kamko, Linmiao Xu, Shivam Shrirao, Chase Lambert, Joao Souza, Suhail Doshi, and Daiqing Li. Playground v3: Im- proving Text-to-Image Alignment with Deep-Fusion Large Language Models, 2024. arXiv:2409.10695. 2

  31. [39]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdvances in Neural Information Processing Systems, pages 34892–34916. Curran Associates, Inc., 2023. 2

  32. [40]

    Improving Long-Text Alignment for Text-to-Image Diffusion Models

    Luping Liu, Chao Du, Tianyu Pang, Zehan Wang, Chongx- uan Li, and Dong Xu. Improving Long-Text Alignment for Text-to-Image Diffusion Models. InThe Thirteenth Interna- tional Conference on Learning Representations, 2025. 2

  33. [41]

    Tenenbaum, and Antonio Torralba

    Nan Liu, Yilun Du, Shuang Li, Joshua B. Tenenbaum, and Antonio Torralba. Unsupervised Compositional Concepts Discovery with Text-to-Image Generative Models. In2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 2085–2095, Paris, France, 2023. IEEE. 2

  34. [42]

    Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning,

    Fan Lu, Wei Wu, Kecheng Zheng, Shuailei Ma, Biao Gong, Jiawei Liu, Wei Zhai, Yang Cao, Yujun Shen, and Zheng- Jun Zha. Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning,

  35. [43]

    DOCCI: Descriptions of Connected and Con- trasting Images

    Yasumasa Onoe, Sunayana Rane, Zachary Berger, Yonatan Bitton, Jaemin Cho, Roopal Garg, Alexander Ku, Zarana Parekh, Jordi Pont-Tuset, Garrett Tanzer, Su Wang, and Ja- son Baldridge. DOCCI: Descriptions of Connected and Con- trasting Images. InComputer Vision – ECCV 2024, pages...

  36. [44]

    Bleu: a Method for Automatic Evaluation of Machine Translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a Method for Automatic Evaluation of Machine Translation. InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311– 318, Philadelphia, Pennsylvania, USA, 2002. Associ...

  37. [45]

    Scalable Diffusion Models with Transformers

    William Peebles and Saining Xie. Scalable Diffusion Models with Transformers. In2023 IEEE/CVF International Confer- ence on Computer Vision (ICCV), pages 4172–4182, Paris, France, 2023. IEEE. 2

  38. [46]

    Image Textualization: An Auto- matic Framework for Generating Rich and Detailed Image Descriptions

    Renjie Pi, Jianshu Zhang, Jipeng Zhang, Rui Pan, Zhekai Chen, and Tong Zhang. Image Textualization: An Auto- matic Framework for Generating Rich and Detailed Image Descriptions. InThe Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Trac...

  39. [47]

    SDXL: Improving Latent Diffusion Mod- els for High-Resolution Image Synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. SDXL: Improving Latent Diffusion Mod- els for High-Resolution Image Synthesis. InThe Twelfth In- ternational Conference on Learning Representations, 2024. 1

  40. [48]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and others. Learning transferable visual models from natural language supervision. InInternational conference on machine learn- i...

  41. [49]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.Journal of Machine Learning Research, 21(140):1–67, 2020. 2

  42. [50]

    Zero-Shot Text-to-Image Generation

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-Shot Text-to-Image Generation. InProceedings of the 38th International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 1

  43. [51]

    Generative Ad- versarial Text to Image Synthesis

    Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Lo- geswaran, Bernt Schiele, and Honglak Lee. Generative Ad- versarial Text to Image Synthesis. InProceedings of The 33rd International Conference on Machine Learning, pages 1060–1069, New York, New York, USA, 2016. PMLR

  44. [52]

    High-Resolution Image Synthesis with Latent Diffusion Models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10674–10685, New Orleans, LA, USA, 2022. IEEE. 1, 2

  45. [53]

    Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa R

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W. Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa R. Kundurthy, Kather- ine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and 11 Jenia ...

  46. [54]

    Conceptual Captions: A Cleaned, Hypernymed, Im- age Alt-text Dataset For Automatic Image Captioning

    Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual Captions: A Cleaned, Hypernymed, Im- age Alt-text Dataset For Automatic Image Captioning. In 56th Annual Meeting of the Association for Computational Linguistics, pages 2556–2565, Melbourne, Australia, 20...

  47. [55]

    Denois- ing Diffusion Implicit Models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing Diffusion Implicit Models. InInternational Conference on Learning Representations, 2021. 2

  48. [56]

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi`ere, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, L ´eonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, A...

  49. [57]

    A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions

    Jack Urbanek, Florian Bordes, Pietro Astolfi, Mary Williamson, Vasu Sharma, and Adriana Romero-Soriano. A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pa...

  50. [58]

    Lawrence Zitnick, and Devi Parikh

    Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. CIDEr: Consensus-based image description evalua- tion. In2015 IEEE Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 4566–4575, Boston, MA, USA, 2015. IEEE. 2

  51. [59]

    Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution, 2024

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Jun- yang Lin. Qwen2-VL: Enhancing Vision-Language Model’s ...

  52. [60]

    Emu3: Next-Token Prediction is All You Need, 2024

    Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, Bo Zhao, Bowen Zhang, Liangdong Wang, Guang Liu, Zheqi He, Xi Yang, Jingjing Liu, Yonghua Lin, Tie...

  53. [61]

    Detailed Object Description with Controllable Dimensions, 2025

    Xinran Wang, Haiwen Zhang, Baoteng Li, Kongming Liang, Hao Sun, Zhongjiang He, Zhanyu Ma, and Jun Guo. Detailed Object Description with Controllable Dimensions, 2025. arXiv:2411.19106. 3

  54. [62]

    LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation

    Aoqi Wu, weiquan Huang, Yifan Yang, Xufang Luo, Yuqing Yang, Chunyu Wang, Liang Hu, Xiyang Dai, Dongdong Chen, Chong Luo, and Lili Qiu. LLM2CLIP: Powerful Language Model Unlock Richer Visual Representation. In NeurIPS 2024 Workshop: Self-Supervised Learning - The- ory and Prac...

  55. [63]

    BoxDiff: Text-to-Image Synthesis with Training-Free Box- Constrained Diffusion

    Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou. BoxDiff: Text-to-Image Synthesis with Training-Free Box- Constrained Diffusion. In2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7418–7427,

  56. [64]

    Choy, and Li Fei- Fei

    Danfei Xu, Yuke Zhu, Christopher B. Choy, and Li Fei- Fei. Scene Graph Generation by Iterative Message Passing. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3097–3106, Honolulu, HI, 2017. IEEE. 3

  57. [65]

    Peter Young, Alice Lai, Micah Hodosh, and Julia Hocken- maier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.Transactions of the Association for Computational Linguistics, 2:67–78, 2014. 1

  58. [66]

    ITI-Gen: Inclusive Text-to-Image Generation

    Cheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, and Fernando De La Torre. ITI-Gen: Inclusive Text-to-Image Generation. In2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 3946–3957, 2023. 2

  59. [67]

    OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding

    Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, and Shuicheng Y AN. OMG-LLaV A: Bridging Image-level, Object-level, Pixel- level Reasoning and Understanding. InThe Thirty-eighth Annual Conference on Neural Information Processing Sys- t...

  60. [68]

    Lumina-Next : Mak- ing Lumina-T2X Stronger and Faster with Next-DiT

    Le Zhuo, Ruoyi Du, Han Xiao, Yangguang Li, Dongyang Liu, Rongjie Huang, Wenze Liu, Xiangyang Zhu, Fu-Yun Wang, Zhanyu Ma, Xu Luo, Zehan Wang, Kaipeng Zhang, Lirui Zhao, Si Liu, Xiangyu Yue, Wanli Ouyang, Yu Qiao, Hongsheng Li, and Peng Gao. Lumina-Next : Mak- ing Lumina-T2X St...

  61. [2024]

    arXiv:2412.08614. 2, 3

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.