Pith. sign in

REVIEW 3 major objections 8 minor 62 references

AI-generated Image Quality Assessment in Visual Communication

T0 review · 3 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper introduces AIGI-VC, a 2,500-image benchmark arguing that current image-quality models cannot judge whether AI-generated ads communicate clearly and emotionally.

desk verdict A genuinely useful new benchmark for AI-generated image communicability, but the ground-truth preference labels are validated at a different sampling density than the one used, and the interpretation benchmark has an avoidable circularity. read the letter →

arxiv 2412.15677 v1 pith:PW5OG7BZ submitted 2024-12-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords AI-generatedimagequalityassessmentvisualcommunicationadvertisinghumanpreferencedatasetinformationclarityemotionalinteractionlargemultimodalmodelsbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that the quality of an AI-generated advertisement cannot be judged by how realistic or aesthetically pleasing it is: what matters is whether it clearly delivers the intended message and triggers the intended emotion. To make this measurable, the paper introduces AIGI-VC, a 2,500-image database of AI-generated ads spanning 14 topics and 8 emotions, with human pairwise preferences on those two dimensions and written explanations of the preferences. On this database, the paper reports that existing image-quality metrics and open-source multimodal models rank far below human judgment, while the proprietary multimodal model GPT-4o comes closest across preference prediction, interpretation, and reasoning. If the benchmark is reliable, it gives automated image evaluation a concrete new target: communicative effectiveness rather than low-level fidelity.

What carries the argument

The annotation protocol carries the argument. Raters compare pairs of images generated from the same prompt and choose which one conveys the text more clearly or evokes the intended emotion more strongly; a Thurstone Case V model with maximum-a-posteriori estimation turns the sparse pairwise votes into a global preference score, allowing 2,000 labeled pairs to stand in for the full set of comparisons. For fine-grained labels, the lower-ranked image is explained against a pseudo-reference: a large multimodal model drafts reasons using prescribed visual cues, and human experts verify and supplement those drafts into golden descriptions. The evaluation layer then measures whether automated models can predict, explain, and reason about these human preferences using accuracy, correlation, order-consistency, and GPT-assisted completeness, preciseness, and relevance scores.

What would settle it

Re-run a random sample of the 500 prompts with exhaustive pairwise comparisons among the five generated images using new raters, then compare the full ranking to the paper's MAP-estimated preferences; if accuracy drops well below the level reported in the pilot, the ground-truth labels are not stable enough to support the benchmark conclusions.

Watch

Extended reading notes

Core claim

In the paper's own telling, the core discovery is that communicability is a measurable quality axis that current IQA tools miss. AIGI-VC contains 2,500 images generated from 500 advertisement prompts by five text-to-image models, covering 14 ad topics and 8 emotion categories, and it annotates each image pair for information clarity and emotional interaction separately. Benchmarking 14 metrics and 7 LMMs, the paper finds that CLIP-based preference metrics reach only about 0.75 accuracy on clarity and 0.69 on emotion, that open-source LMMs are close to chance or inconsistent when the order of the two images is flipped, and that GPT-4o reaches roughly 0.79 to 0.88 accuracy with much higher consistency. The paper also reports a strong correlation between the two dimensions (SRCC 0.9371, PLCC 0.9360), which it reads as evidence that clarity and emotional impact are closely linked. The conclusion is that automated evaluators are not yet effective for the communicability of AIGIs in visual communication.

Load-bearing premise

The benchmark's ground-truth labels are only as trustworthy as the assumption that preference scores estimated from a sampled subset of pairwise comparisons match the preferences people would give in exhaustive comparison, and that assumption was tested on only 250 images.

Editorial extensions

If this is right

  • Ad teams can use AIGI-VC to filter generated images for clarity and emotional impact before publication, since the benchmark provides human preference ground truth for both axes.
  • Future quality metrics for AI-generated content should be tested on preference prediction and on interpretation and reasoning, because the paper shows these abilities are not the same.
  • The high correlation between clarity and emotion preferences implies that a single ranking score captures much of what humans care about in ads, although the two dimensions remain separately labeled.
  • Open-source multimodal models need order-consistency improvements before they can serve as dependable automatic judges of visual communication.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A practical extension the authors do not draw: merging the two axes into one composite communicability score would likely preserve most ranking information given the 0.94 correlation, and could halve future annotation costs.
  • The same pairwise-protocol design could be reused for news illustrations, educational graphics, or social-media creatives, where clarity and emotion also determine whether an image works.
  • If a future open-source model matches GPT-4o's accuracy and consistency on AIGI-VC, that would show the current gap is a modeling problem rather than a missing benchmark.
  • The order-consistency failure of open-source LMMs suggests that any deployable ad-quality judge should be required to pass a counterbalanced presentation test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The paper introduces AIGI-VC, a dataset of 2,500 AI-generated advertisement images produced from 500 prompts by five text-to-image models, spanning 14 ad topics and 8 emotion categories. Human annotators provided pairwise preferences along two dimensions, information clarity and emotional interaction; the paper then applies a Thurstone-based MAP model (Eq. 1) to estimate preference probabilities from a sampled subset of the exhaustive pairs. The dataset also includes fine-grained natural-language descriptions generated by GPT-4o and verified by human experts, intended to explain the reasons behind human preferences. The paper benchmarks 14 baselines (IQA metrics and large multimodal models, including GPT-4o) on three tasks: preference prediction, interpretation of human choices, and reasoning about image pairs. The main empirical claims are that existing IQA methods and open-source LMMs perform poorly on this task, while GPT-4o significantly outperforms all competitors, and that the dataset is the first to study communicability of AIGIs in visual communication.

Significance. If the ground-truth preferences are reliable, AIGI-VC fills a genuine gap: existing AIGI datasets focus on general quality or aesthetics, whereas this one targets information clarity and emotional interaction in an advertising context. The dataset construction covers diverse topics and emotions, and the authors provide code, three challenge subsets (human-object interaction, fantastical ads, positive/negative emotions), and a broad set of baseline models. These are strengths. However, the central empirical claims depend on two load-bearing assumptions: that MAP estimates from a 40%-coverage pair sample reproduce exhaustive human preferences, and that the GPT-assisted interpretation/reasoning evaluation is not biased toward GPT-4o. Both need stronger support before the dataset and the model rankings can be fully trusted.

major comments (3)
  1. [Dataset Construction / Human Preference Annotation (Fig. 3)] The validation of MAP-estimated preferences is performed at M=4 rounds, which corresponds to roughly 80-100% of the exhaustive within-prompt pairs for the pilot (50 prompts with 5 images each imply 10 pairs per prompt, i.e., about 400-500 comparisons at M=4), yet the actual annotation protocol samples only 2,000 of the 5,000 possible pairs, i.e., 40% coverage, which is closer to M=2 density. The claim that labeling 2,000 pairs reduces exhaustive comparisons by 60% 'while producing the same preferences' is therefore an extrapolation from a denser regime. Please report the MAP accuracy at the actual sampling density (e.g., the M=2 point in Fig. 3, or a direct simulation on the pilot data that samples exactly 4 pairs per prompt and compares against exhaustive preferences). Without this, the preference probabilities used in all benchmark tables are not demonstrated to be reliable.
  2. [Experimental Settings / Tables 2-5] Model comparisons are reported only as point estimates with no confidence intervals or significance tests. For example, the claim that GPT-4o 'significantly outperforms all other competing models' rests on accuracy differences such as 0.7928 vs. 0.7518 (Table 3, IC Dall), with no measure of variability. Because the ground-truth labels are themselves MAP estimates from a sparse sample, label uncertainty propagates into the model rankings. Please provide bootstrap confidence intervals or paired significance tests (e.g., McNemar's test for accuracy) for α, ρ, and κ, so that 'significantly outperforms' is statistically supported.
  3. [Performance on Interpretation and Reasoning / Fine-grained Descriptions (Tables 6-7)] The interpretation and reasoning benchmarks score LMM outputs against golden descriptions that were initially generated by GPT-4o and then verified by human experts, and the evaluation itself is described as 'GPT-assisted' following Q-Bench, without identifying the judge model. This design risks favoring GPT-4o in two ways: (i) the golden descriptions may retain GPT-4o's style and priorities even after human verification, and (ii) if the judge is also GPT-4o, the scoring is not independent. Please identify the judge model and either use a different model as the judge or provide a human-evaluated subset to demonstrate that the rankings in Tables 6-7 are not an artifact of self-scoring.
minor comments (8)
  1. [Data Collection] The PSA topics are listed as 'environment protection, animal rights, social welfare, safety, healthcare, and self-esteem,' which is six topics, while the dataset is described as having 14 topics and Figure 1 includes 'Bullying & violence' as the seventh PSA topic; please correct the enumeration.
  2. [Table 1] The AIGI-VC row contains the typo 'Commnunication' for 'Communication'; the table header has the same typo.
  3. [Experimental Settings / Eq. (2)] Equation (2) defines ρ as PLCC between ground-truth and predicted preference probabilities, but the conditioning notation P(X,Y)|Z is not defined clearly; please clarify that it denotes the preference probability distribution over pairs given the reference text or emotion.
  4. [Table 4 discussion] The text states that 'HPSv2 achieves higher λ and ρ values,' but no λ is defined in the paper; presumably α is meant.
  5. [Experimental Settings] The paper says 'we employ 14 objective metrics' but then lists LMMs (e.g., GPT-4o) which are not typically called objective metrics; consider renaming to '14 baseline models.'
  6. [Figure 2 and model references] Figure 2 does not label the columns with the generative model names; adding column labels would improve readability. Also, Stable Diffusion 2.0 and Dreamlike Photoreal 2.0 are both cited to Rombach et al. 2022, which is the original latent diffusion paper; more specific references would be appropriate.
  7. [Tables 2-3] The abbreviations 'Dall' and 'Dsub' are used without definition in the table captions; please define them when first used in the main text.
  8. [Conclusion] The conclusion does not mention any limitations of the dataset or the evaluation; a short limitations paragraph would help readers interpret the claims.

Circularity Check

1 steps flagged · score 5.0 of 10

The interpretation/reasoning benchmark is partially circular: golden descriptions are initially authored by GPT-4o, and GPT-4o is then scored against those same descriptions.

  1. self definitional [Dataset Construction, Fine-grained Descriptions; Evaluation on AIGI-VC, Performance on Interpretation and Reasoning]
    "we treat the top-ranked image in each evaluation dimension for each prompt as a pseudo-reference and incorporate GPT-4o ... to identify why the other image under the same prompt is worse than the pseudo-reference. ... We collect responses from GPT-4o as the initial descriptions and recruit human experts to verify and supplement each GPT-generated description, creating a golden standard description. ... We evaluate the interpretation and reasoning abilities of the LMMs using golden descriptions."

    The 'golden standard description' used as ground truth for the interpretation and reasoning benchmarks is initially produced by GPT-4o, then human-verified and supplemented. The same model, GPT-4o, is later evaluated by how closely its responses align with these golden descriptions (Tables 6 and 7). This gives GPT-4o a built-in advantage: the benchmark's reference text inherits GPT-4o's phrasing, emphasis, and selected visual cues, so its top scores partly measure self-consistency with its own generated content rather than independent communicative quality. Human verification mitigates but does not remove this overlap, because verification starts from GPT-4o's wording and framing.

full rationale

The central dataset contribution is grounded in human pairwise preference judgments, and the preference-prediction experiments (Tables 2-5) compare models against those human choices, so that part is not circular. The MAP-estimation validation concern raised by the reader is a data-reliability issue, not a circularity issue. The one genuine circularity burden is the fine-grained description benchmark: the golden descriptions used to score interpretation and reasoning are initially generated by GPT-4o, then human-verified, and GPT-4o is subsequently evaluated against those same descriptions. This makes GPT-4o's reported superiority in interpretation and reasoning partially self-referential. The paper's other citations, including self-citations to prior work on pairwise prompting and Thurstone scaling, are not load-bearing because they are supported by external references (Prashnani et al., Q-Bench) and by the paper's own human annotation protocol. Overall, the core preference data and the majority of benchmark comparisons are independent; the circularity is localized to the interpretation/reasoning evaluation, so a moderate score is appropriate.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central dataset claim rests on standard statistical modeling and several domain assumptions about advertising communication. The most fragile choices are the Thurstone Case V prior, the equation of pairwise preference with communicability, the 512x512 resizing, and the use of GPT-4o-generated, human-verified gold descriptions. No new physical or conceptual entities are introduced.

free parameters (2)
  • Preference probability calibration mapping per model = not specified
    The paper states that 'all predicted scores by the model are fitted before computing the preference probabilities', but the fitting procedure and its parameters are not described. This mapping affects all reported correlation values for every model.
  • Strong preference threshold = [0.3, 0.7]
    The authors restrict the reliability validation and some analyses to pairs whose estimated preference probability falls outside [0.3, 0.7]. This hand-chosen threshold defines the D_sub subset and excludes ambiguous pairs.
assumptions (4)
  • standard math Preference data follow the Thurstone Case V model with a zero-mean unit-variance Gaussian prior over quality scores.
    The MAP estimation in Eq. (1) assumes the probability of preferring image i over j is Phi(q_i - q_j) with a unit-variance Gaussian prior. This is a standard but unverified statistical assumption for the collected human judgments.
  • domain assumption Pairwise preference under a fixed prompt is a valid measure of advertising communicability.
    The authors equate human preference between two images sharing the same text and intended emotion with the effectiveness of visual communication, without validating against actual advertising outcomes or expert advertisers.
  • domain assumption Resizing all images to 512x512 preserves information clarity and emotional impact of advertisements.
    The dataset construction resizes every image to 512x512 to standardize resolution. For ads that contain text, downsampling can degrade legibility and alter communicative content, yet the paper does not test for this effect.
  • domain assumption GPT-4o-generated descriptions, after human expert verification, are a valid golden standard for preference interpretation and reasoning.
    The fine-grained descriptions are initially produced by GPT-4o and then verified and supplemented by human experts. This makes the reference standard partially authored by the model that is later scored against it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-generated Image Quality Assessment in Visual Communication." pith.science (2026). https://pith.science/paper/PW5OG7BZ

@misc{pith2026241215677,
  author       = {Pith},
  title        = {Pith review of: AI-generated Image Quality Assessment in Visual Communication},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PW5OG7BZ}},
  note         = {Machine review of arXiv:2412.15677}
}
read the original abstract

Assessing the quality of artificial intelligence-generated images (AIGIs) plays a crucial role in their application in real-world scenarios. However, traditional image quality assessment (IQA) algorithms primarily focus on low-level visual perception, while existing IQA works on AIGIs overemphasize the generated content itself, neglecting its effectiveness in real-world applications. To bridge this gap, we propose AIGI-VC, a quality assessment database for AI-Generated Images in Visual Communication, which studies the communicability of AIGIs in the advertising field from the perspectives of information clarity and emotional interaction. The dataset consists of 2,500 images spanning 14 advertisement topics and 8 emotion types. It provides coarse-grained human preference annotations and fine-grained preference descriptions, benchmarking the abilities of IQA methods in preference prediction, interpretation, and reasoning. We conduct an empirical study of existing representative IQA methods and large multi-modal models on the AIGI-VC dataset, uncovering their strengths and weaknesses.

Figures

Figures reproduced from arXiv: 2412.15677 by the authors.

Figure 1
Figure 1. Outline of the AIGI-VC dataset. reflect (Holbrook and O’Shaughnessy 1984; Hussain et al. 2017; Yang et al. 2023). However, due to hardware limita￾tions and technical proficiency, the quality of AI-generated images (AIGIs) varies widely, necessitating refinement and filtering before distributing them to practical applications. There have been substantial efforts in establishing benchmarks to facilitate research on AI… view at source ↗
Figure 2
Figure 2. Sample images from the AIGI-VC database, where the first to fifth columns show images generated by Dall [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Accuracy of preference choices via MAP estima [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: The process of description generation. Given two images with preference choices collected from human users, GPT [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 2
Figure 2. Figure 2: Human Preference Annotation Coarse-grained Preference Choices We collect human opinions via the pairwise image comparison method, di￾rectly asking participants to choose their preferred image from a pair. Generally speaking, global ranking results of N test stimuli are…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 49 canonical work pages

  1. [1]

    L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al

    Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774

  2. [2]

    H.; and Ramkumar, J

    Akhtar, M. H.; and Ramkumar, J. 2023. AI in Visual Communication: AI Taking Down Graphic Designers. Scary? In AI for Designers, 85--103. Springer

  3. [3]

    Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023. Qwen-VL : A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv preprint arXiv:2308.12966

  4. [4]

    Bao, Q.; Hui, Z.; Zhu, R.; Ren, P.; Xie, X.; and Yang, W. 2024. Improving Diffusion-Based Image Restoration with Error Contraction and Error Correction. In AAAI, volume 38, 756--764

  5. [5]

    Campbell, C.; Plangger, K.; Sands, S.; and Kietzmann, J. 2022. Preparing for An Era of Deepfakes and AI -generated Ads: A Framework for Understanding Responses to Manipulated Advertising. Journal of Advertising, 51(1): 22--38

  6. [6]

    Chen, B.; Zhu, H.; Zhu, L.; Wang, S.; and Kwong, S. 2024 a . Deep Feature Statistics Mapping for Generalized Screen Content Image Quality Assessment. IEEE Trans. Image Process., 33: 3227--3241

  7. [7]

    Chen, C.; Zhou, S.; Liao, L.; Wu, H.; Sun, W.; Yan, Q.; and Lin, W. 2024 b . Iterative Token Evaluation and Refinement for Real-world Super-Resolution. In AAAI, volume 38, 1010--1018

  8. [8]

    Chen, J.; An, J.; Lyu, H.; Kanan, C.; and Luo, J. 2024 c . Learning to Evaluate The Artness of AI-generated Images. IEEE Trans. Multimedia, 1--10

Show all 62 references
  1. [9]

    Chen, Z.; Sun, W.; Wu, H.; Zhang, Z.; Jia, J.; Min, X.; Zhai, G.; and Zhang, W. 2023. Exploring the Naturalness of AI -generated Images. arXiv preprint arXiv:2312.05476

  2. [10]

    Chen, Z.; Zhou, Q.; Shen, Y.; Hong, Y.; Sun, Z.; Gutfreund, D.; and Gan, C. 2024 d . Visual Chain-of-thought Prompting for Knowledge-Based Visual Reasoning. In AAAI, volume 38, 1254--1262

  3. [11]

    Cho, J.; Hu, Y.; Baldridge, J.; Garg, R.; Anderson, P.; Krishna, R.; Bansal, M.; Pont-Tuset, J.; and Wang, S. 2024. Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-image Generation. In ICLR

  4. [12]

    Cho, J.; Zala, A.; and Bansal, M. 2023. Dall-eval: Probing The Reasoning Skills and Social Biases of Text-to-image Generation Models. In ICCV, 3043--3054

  5. [13]

    Dong, X.; Zhang, P.; Zang, Y.; Cao, Y.; Wang, B.; Ouyang, L.; Wei, X.; Zhang, S.; Duan, H.; Cao, M.; Zhang, W.; Li, Y.; Yan, H.; Gao, Y.; Zhang, X.; Li, W.; Li, J.; Chen, K.; He, C.; Zhang, X.; Qiao, Y.; Lin, D.; and Wang, J. 2024. InternLM-XComposer2 : Mastering Free-form Tex...

  6. [14]

    Duan, X.; Ma, S.; Liu, H.; and Jia, C. 2024. PKU-AIGI-500K : A Neural Compression Benchmark and Model for AI-Generated Images. IEEE J. Emerg. Sel. Topics Circuits Syst

  7. [15]

    Ford, J.; Jain, V.; Wadhwani, K.; and Gupta, D. G. 2023. AI Advertising: An Overview and Guidelines. Journal of Business Research, 166: 114124

  8. [16]

    Gao, Y.; Min, X.; Zhu, Y.; Zhang, X.-P.; and Zhai, G. 2024. Blind Image Quality Assessment: A Fuzzy Neural Network for Opinion Score Distribution Prediction. IEEE Trans. Circuits Syst. Video Technol., 34(3): 1641--1655

  9. [17]

    Gu, S.; Bao, J.; Chen, D.; and Wen, F. 2020. GIQA : Generated Image Quality Assessment. In ECCV, 369--385

  10. [18]

    L.; and Choi, Y

    Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021. CLIPScore : A Reference-free Evaluation Metric for Image Captioning. arXiv preprint arXiv:2104.08718

  11. [19]

    Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising Diffusion Probabilistic Models. NeurIPS, 33: 6840--6851

  12. [20]

    B.; and O'Shaughnessy, J

    Holbrook, M. B.; and O'Shaughnessy, J. 1984. The Role of Emotion in Advertising. Psychology & Marketing, 1(2): 45--64

  13. [21]

    Huang, K.; Sun, K.; Xie, E.; Li, Z.; and Liu, X. 2024 a . T2I -compbench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation. NeurIPS, 36

  14. [22]

    Huang, Y.; Sheng, X.; Yang, Z.; Yuan, Q.; Duan, Z.; Chen, P.; Li, L.; Lin, W.; and Shi, G. 2024 b . AesExpert : Towards Multi-modality Foundation Model for Image Aesthetics Perception. arXiv preprint arXiv:2404.09624

  15. [23]

    Hussain, Z.; Zhang, M.; Zhang, X.; Ye, K.; Thomas, C.; Agha, Z.; Ong, N.; and Kovashka, A. 2017. Automatic Understanding of Image and Video Advertisements. In CVPR, 1705--1715

  16. [24]

    Jiang-Lin, J.-Y.; Huang, K.-Y.; Lo, L.; Huang, Y.-N.; Lin, T.; Wu, J.-C.; Shuai, H.-H.; and Cheng, W.-H. 2024. ReCorD : Reasoning and Correcting Diffusion for HOI Generation. In ACM MM

  17. [25]

    Kirstain, Y.; Polyak, A.; Singer, U.; Matiana, S.; Penna, J.; and Levy, O. 2024. Pick-a-pic: An Open Dataset of User Preferences for Text-to-image Generation. NeurIPS, 36

  18. [26]

    Lauren c on, H.; Tronchon, L.; Cord, M.; and Sanh, V. 2024. What Matters when Building Vision-language models? arXiv preprint arXiv:2405.02246

  19. [27]

    S.; Gupta, A.; Zhang, Y.; Narayanan, D.; Teufel, H.; Bellagente, M.; et al

    Lee, T.; Yasunaga, M.; Meng, C.; Mai, Y.; Park, J. S.; Gupta, A.; Zhang, Y.; Narayanan, D.; Teufel, H.; Bellagente, M.; et al. 2024. Holistic Evaluation of Text-to-image Models. NeurIPS, 36

  20. [28]

    Li, C.; Zhang, Z.; Wu, H.; Sun, W.; Min, X.; Liu, X.; Zhai, G.; and Lin, W. 2023. AGIQA-3K : An Open Database for AI -Generated Image Quality Assessment. IEEE Trans. Circuits Syst. Video Technol

  21. [29]

    Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024. Visual Instruction Tuning. NeurIPS, 36

  22. [30]

    E.; and Wang, W

    Lu, Y.; Yang, X.; Li, X.; Wang, X. E.; and Wang, W. Y. 2024. LLMscore : Unveiling the Power of Large Language Models in Text-to-image Synthesis Evaluation. NeurIPS, 36

  23. [31]

    Ma, K.; Liu, W.; Liu, T.; Wang, Z.; and Tao, D. 2017. dipIQ : Blind Image Quality Assessment by Learning-to-rank Discriminable Image Pairs. IEEE Trans. Image Process., 26(8): 3951--3964

  24. [32]

    A.; Fredrickson, B

    Mikels, J. A.; Fredrickson, B. L.; Larkin, G. R.; Lindberg, C. M.; Maglio, S. J.; and Reuter-Lorenz, P. A. 2005. Emotional Category Data on Images from The International Affective Picture System. Behavior research methods, 37: 626--630

  25. [33]

    Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2022. GLIDE : Towards Photorealistic Image Generation and Editing with Text-guided Diffusion Models. In ICML, 16784--16804

  26. [34]

    OpenAI. 2023. GPT-4V(ision) System Card. https://cdn.openai.com/papers/GPTV_System_Card.pdf

  27. [35]

    Prashnani, E.; Cai, H.; Mostofi, Y.; and Sen, P. 2018. PieAPP : Perceptual Image-error Assessment through Pairwise Preference. In CVPR, 1808--1817

  28. [36]

    D.; Crowson, K.; and Contributors, S

    Pressman, J. D.; Crowson, K.; and Contributors, S. C. 2022. Simulacra Aesthetic Captions. Technical Report Version 1.0, Stability AI. \ url https://github.com/JD-P/simulacra-aesthetic-captions

  29. [37]

    Quan, H.; Li, S.; Zeng, C.; Wei, H.; and Hu, J. 2023. Big Data and AI -driven Product Design: A Survey. Applied Sciences, 13(16): 9433

  30. [38]

    Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022. Hierarchical Text-conditional Image Generation with CLIP Latents. arXiv preprint arXiv:2204.06125

  31. [39]

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022. High-resolution Image Synthesis with Latent Diffusion Models. In CVPR, 10684--10695

  32. [40]

    Sagar, A.; Srivastava, R.; Rakshitha; Kesav, V.; and Kiran, R. 2024. MAdVerse : A Hierarchical Dataset of Multi-Lingual Ads from Diverse Sources and Categories. In IEEE Winter Conf. App. Comput. Vis

  33. [41]

    Schramowski, P.; Brack, M.; Deiseroth, B.; and Kersting, K. 2023. Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. In CVPR, 22522--22531

  34. [42]

    Series, B. T. 2012. Methodology for The Subjective Assessment of The Quality of Television Pictures. Recommendation ITU-R BT, 500(13)

  35. [43]

    L.; and Wang, L

    She, D.; Yang, J.; Cheng, M.-M.; Lai, Y.-K.; Rosin, P. L.; and Wang, L. 2020. WSCNet: Weakly Supervised Coupled Networks for Visual Sentiment Classification and Detection. IEEE Trans. Multimedia, 22(5): 1358--1371

  36. [44]

    SkunkworksAI. 2024. BakLLaVA. https://github.com/SkunkworksAI/BakLLaVA

  37. [45]

    Su, S.; Yan, Q.; Zhu, Y.; Zhang, C.; Ge, X.; Sun, J.; and Zhang, Y. 2020. Blindly Assess Image Quality in The Wild Guided by A Self-adaptive Hyper Network. In CVPR, 3667--3676

  38. [46]

    Teo, C.; Abdollahzadeh, M.; and Cheung, N.-M. M. 2024. On Measuring Fairness in Generative Models. NeurIPS, 36

  39. [47]

    Tian, Y.; Wang, S.; Chen, B.; and Kwong, S. 2024. Causal Representation Learning for GAN-Generated Face Image Quality Assessment. IEEE Trans. Circuits Syst. Video Technol., 1--1

  40. [48]

    R.; et al

    Tsukida, K.; Gupta, M. R.; et al. 2011. How to Analyze Paired Comparison Data. Department of Electrical Engineering University of Washington, Tech. Rep. UWEETR-2011-0004, 1

  41. [49]

    Wu, H.; Zhang, Z.; Zhang, E.; Chen, C.; Liao, L.; Wang, A.; Li, C.; Sun, W.; Yan, Q.; Zhai, G.; and Lin, W. 2024 a . Q-Bench : A Benchmark for General-Purpose Foundation Models on Low-level Vision. In ICLR

  42. [50]

    Wu, H.; Zhang, Z.; Zhang, E.; Chen, C.; Liao, L.; Wang, A.; Xu, K.; Li, C.; Hou, J.; Zhai, G.; et al. 2024 b . Q-instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models. In CVPR, 25490--25500

  43. [51]

    Wu, X.; Hao, Y.; Sun, K.; Chen, Y.; Zhu, F.; Zhao, R.; and Li, H. 2023. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis. arXiv preprint arXiv:2306.09341

  44. [52]

    Wu, X.; Xiao, L.; Sun, Y.; Zhang, J.; Ma, T.; and He, L. 2022. A Survey of Human-in-the-loop for Machine Learning. Future Gener. Comp. Sy., 135: 364--381

  45. [53]

    Xu, J.; Liu, X.; Wu, Y.; Tong, Y.; Li, Q.; Ding, M.; Tang, J.; and Dong, Y. 2024. ImageReward : Learning and Evaluating Human Preferences for Text-to-image Generation. NeurIPS, 36

  46. [54]

    Yang, J.; Huang, Q.; Ding, T.; Lischinski, D.; Cohen-Or, D.; and Huang, H. 2023. EmoSet : A Large-scale Visual Emotion Dataset with Rich Attributes. In ICCV, 20383--20394

  47. [55]

    Yarom, M.; Bitton, Y.; Changpinyo, S.; Aharoni, R.; Herzig, J.; Lang, O.; Ofek, E.; and Szpektor, I. 2024. What You See is What You Read? Improving Text-image Alignment Evaluation. NeurIPS, 36

  48. [56]

    Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023. mPLUG-Owl : Modularization Empowers Large Language Models with Multimodality. arXiv preprint arXiv:2304.14178

  49. [57]

    Zhang, W.; Ma, K.; Zhai, G.; and Yang, X. 2021. Uncertainty-aware Blind Image Quality Assessment in The Laboratory and Wild. IEEE Trans. Image Process., 30: 3474--3486

  50. [58]

    Zhao, S.; Gao, Y.; Jiang, X.; Yao, H.; Chua, T.-S.; and Sun, X. 2014. Exploring Principles-of-Art Features For Image Emotion Recognition. In ACM MM, 47–56

  51. [59]

    Zhu, H.; Sui, X.; Chen, B.; Liu, X.; Chen, P.; Fang, Y.; and Wang, S. 2024 a . 2AFC Prompting of Large Multimodal Models for Image Quality Assessment. arXiv preprint arXiv:2402.01162

  52. [60]

    Zhu, H.; Wu, H.; Li, Y.; Zhang, Z.; Chen, B.; Zhu, L.; Fang, Y.; Zhai, G.; Lin, W.; and Wang, S. 2024 b . Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare. NeurIPS

  53. [61]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  54. [62]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.