REVIEW 3 major objections 8 minor 62 references
AI-generated Image Quality Assessment in Visual Communication
T0 review · 3 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper introduces AIGI-VC, a 2,500-image benchmark arguing that current image-quality models cannot judge whether AI-generated ads communicate clearly and emotionally.
desk verdict A genuinely useful new benchmark for AI-generated image communicability, but the ground-truth preference labels are validated at a different sampling density than the one used, and the interpretation benchmark has an avoidable circularity. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The annotation protocol carries the argument. Raters compare pairs of images generated from the same prompt and choose which one conveys the text more clearly or evokes the intended emotion more strongly; a Thurstone Case V model with maximum-a-posteriori estimation turns the sparse pairwise votes into a global preference score, allowing 2,000 labeled pairs to stand in for the full set of comparisons. For fine-grained labels, the lower-ranked image is explained against a pseudo-reference: a large multimodal model drafts reasons using prescribed visual cues, and human experts verify and supplement those drafts into golden descriptions. The evaluation layer then measures whether automated models can predict, explain, and reason about these human preferences using accuracy, correlation, order-consistency, and GPT-assisted completeness, preciseness, and relevance scores.
What would settle it
Re-run a random sample of the 500 prompts with exhaustive pairwise comparisons among the five generated images using new raters, then compare the full ranking to the paper's MAP-estimated preferences; if accuracy drops well below the level reported in the pilot, the ground-truth labels are not stable enough to support the benchmark conclusions.
Extended reading notes
Core claim
In the paper's own telling, the core discovery is that communicability is a measurable quality axis that current IQA tools miss. AIGI-VC contains 2,500 images generated from 500 advertisement prompts by five text-to-image models, covering 14 ad topics and 8 emotion categories, and it annotates each image pair for information clarity and emotional interaction separately. Benchmarking 14 metrics and 7 LMMs, the paper finds that CLIP-based preference metrics reach only about 0.75 accuracy on clarity and 0.69 on emotion, that open-source LMMs are close to chance or inconsistent when the order of the two images is flipped, and that GPT-4o reaches roughly 0.79 to 0.88 accuracy with much higher consistency. The paper also reports a strong correlation between the two dimensions (SRCC 0.9371, PLCC 0.9360), which it reads as evidence that clarity and emotional impact are closely linked. The conclusion is that automated evaluators are not yet effective for the communicability of AIGIs in visual communication.
Load-bearing premise
The benchmark's ground-truth labels are only as trustworthy as the assumption that preference scores estimated from a sampled subset of pairwise comparisons match the preferences people would give in exhaustive comparison, and that assumption was tested on only 250 images.
Editorial extensions
If this is right
- Ad teams can use AIGI-VC to filter generated images for clarity and emotional impact before publication, since the benchmark provides human preference ground truth for both axes.
- Future quality metrics for AI-generated content should be tested on preference prediction and on interpretation and reasoning, because the paper shows these abilities are not the same.
- The high correlation between clarity and emotion preferences implies that a single ranking score captures much of what humans care about in ads, although the two dimensions remain separately labeled.
- Open-source multimodal models need order-consistency improvements before they can serve as dependable automatic judges of visual communication.
Reading between the lines
- A practical extension the authors do not draw: merging the two axes into one composite communicability score would likely preserve most ranking information given the 0.94 correlation, and could halve future annotation costs.
- The same pairwise-protocol design could be reused for news illustrations, educational graphics, or social-media creatives, where clarity and emotion also determine whether an image works.
- If a future open-source model matches GPT-4o's accuracy and consistency on AIGI-VC, that would show the current gap is a modeling problem rather than a missing benchmark.
- The order-consistency failure of open-source LMMs suggests that any deployable ad-quality judge should be required to pass a counterbalanced presentation test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AIGI-VC, a dataset of 2,500 AI-generated advertisement images produced from 500 prompts by five text-to-image models, spanning 14 ad topics and 8 emotion categories. Human annotators provided pairwise preferences along two dimensions, information clarity and emotional interaction; the paper then applies a Thurstone-based MAP model (Eq. 1) to estimate preference probabilities from a sampled subset of the exhaustive pairs. The dataset also includes fine-grained natural-language descriptions generated by GPT-4o and verified by human experts, intended to explain the reasons behind human preferences. The paper benchmarks 14 baselines (IQA metrics and large multimodal models, including GPT-4o) on three tasks: preference prediction, interpretation of human choices, and reasoning about image pairs. The main empirical claims are that existing IQA methods and open-source LMMs perform poorly on this task, while GPT-4o significantly outperforms all competitors, and that the dataset is the first to study communicability of AIGIs in visual communication.
Significance. If the ground-truth preferences are reliable, AIGI-VC fills a genuine gap: existing AIGI datasets focus on general quality or aesthetics, whereas this one targets information clarity and emotional interaction in an advertising context. The dataset construction covers diverse topics and emotions, and the authors provide code, three challenge subsets (human-object interaction, fantastical ads, positive/negative emotions), and a broad set of baseline models. These are strengths. However, the central empirical claims depend on two load-bearing assumptions: that MAP estimates from a 40%-coverage pair sample reproduce exhaustive human preferences, and that the GPT-assisted interpretation/reasoning evaluation is not biased toward GPT-4o. Both need stronger support before the dataset and the model rankings can be fully trusted.
major comments (3)
- [Dataset Construction / Human Preference Annotation (Fig. 3)] The validation of MAP-estimated preferences is performed at M=4 rounds, which corresponds to roughly 80-100% of the exhaustive within-prompt pairs for the pilot (50 prompts with 5 images each imply 10 pairs per prompt, i.e., about 400-500 comparisons at M=4), yet the actual annotation protocol samples only 2,000 of the 5,000 possible pairs, i.e., 40% coverage, which is closer to M=2 density. The claim that labeling 2,000 pairs reduces exhaustive comparisons by 60% 'while producing the same preferences' is therefore an extrapolation from a denser regime. Please report the MAP accuracy at the actual sampling density (e.g., the M=2 point in Fig. 3, or a direct simulation on the pilot data that samples exactly 4 pairs per prompt and compares against exhaustive preferences). Without this, the preference probabilities used in all benchmark tables are not demonstrated to be reliable.
- [Experimental Settings / Tables 2-5] Model comparisons are reported only as point estimates with no confidence intervals or significance tests. For example, the claim that GPT-4o 'significantly outperforms all other competing models' rests on accuracy differences such as 0.7928 vs. 0.7518 (Table 3, IC Dall), with no measure of variability. Because the ground-truth labels are themselves MAP estimates from a sparse sample, label uncertainty propagates into the model rankings. Please provide bootstrap confidence intervals or paired significance tests (e.g., McNemar's test for accuracy) for α, ρ, and κ, so that 'significantly outperforms' is statistically supported.
- [Performance on Interpretation and Reasoning / Fine-grained Descriptions (Tables 6-7)] The interpretation and reasoning benchmarks score LMM outputs against golden descriptions that were initially generated by GPT-4o and then verified by human experts, and the evaluation itself is described as 'GPT-assisted' following Q-Bench, without identifying the judge model. This design risks favoring GPT-4o in two ways: (i) the golden descriptions may retain GPT-4o's style and priorities even after human verification, and (ii) if the judge is also GPT-4o, the scoring is not independent. Please identify the judge model and either use a different model as the judge or provide a human-evaluated subset to demonstrate that the rankings in Tables 6-7 are not an artifact of self-scoring.
minor comments (8)
- [Data Collection] The PSA topics are listed as 'environment protection, animal rights, social welfare, safety, healthcare, and self-esteem,' which is six topics, while the dataset is described as having 14 topics and Figure 1 includes 'Bullying & violence' as the seventh PSA topic; please correct the enumeration.
- [Table 1] The AIGI-VC row contains the typo 'Commnunication' for 'Communication'; the table header has the same typo.
- [Experimental Settings / Eq. (2)] Equation (2) defines ρ as PLCC between ground-truth and predicted preference probabilities, but the conditioning notation P(X,Y)|Z is not defined clearly; please clarify that it denotes the preference probability distribution over pairs given the reference text or emotion.
- [Table 4 discussion] The text states that 'HPSv2 achieves higher λ and ρ values,' but no λ is defined in the paper; presumably α is meant.
- [Experimental Settings] The paper says 'we employ 14 objective metrics' but then lists LMMs (e.g., GPT-4o) which are not typically called objective metrics; consider renaming to '14 baseline models.'
- [Figure 2 and model references] Figure 2 does not label the columns with the generative model names; adding column labels would improve readability. Also, Stable Diffusion 2.0 and Dreamlike Photoreal 2.0 are both cited to Rombach et al. 2022, which is the original latent diffusion paper; more specific references would be appropriate.
- [Tables 2-3] The abbreviations 'Dall' and 'Dsub' are used without definition in the table captions; please define them when first used in the main text.
- [Conclusion] The conclusion does not mention any limitations of the dataset or the evaluation; a short limitations paragraph would help readers interpret the claims.
Circularity Check
The interpretation/reasoning benchmark is partially circular: golden descriptions are initially authored by GPT-4o, and GPT-4o is then scored against those same descriptions.
-
self definitional
[Dataset Construction, Fine-grained Descriptions; Evaluation on AIGI-VC, Performance on Interpretation and Reasoning]
"we treat the top-ranked image in each evaluation dimension for each prompt as a pseudo-reference and incorporate GPT-4o ... to identify why the other image under the same prompt is worse than the pseudo-reference. ... We collect responses from GPT-4o as the initial descriptions and recruit human experts to verify and supplement each GPT-generated description, creating a golden standard description. ... We evaluate the interpretation and reasoning abilities of the LMMs using golden descriptions."
The 'golden standard description' used as ground truth for the interpretation and reasoning benchmarks is initially produced by GPT-4o, then human-verified and supplemented. The same model, GPT-4o, is later evaluated by how closely its responses align with these golden descriptions (Tables 6 and 7). This gives GPT-4o a built-in advantage: the benchmark's reference text inherits GPT-4o's phrasing, emphasis, and selected visual cues, so its top scores partly measure self-consistency with its own generated content rather than independent communicative quality. Human verification mitigates but does not remove this overlap, because verification starts from GPT-4o's wording and framing.
full rationale
The central dataset contribution is grounded in human pairwise preference judgments, and the preference-prediction experiments (Tables 2-5) compare models against those human choices, so that part is not circular. The MAP-estimation validation concern raised by the reader is a data-reliability issue, not a circularity issue. The one genuine circularity burden is the fine-grained description benchmark: the golden descriptions used to score interpretation and reasoning are initially generated by GPT-4o, then human-verified, and GPT-4o is subsequently evaluated against those same descriptions. This makes GPT-4o's reported superiority in interpretation and reasoning partially self-referential. The paper's other citations, including self-citations to prior work on pairwise prompting and Thurstone scaling, are not load-bearing because they are supported by external references (Prashnani et al., Q-Bench) and by the paper's own human annotation protocol. Overall, the core preference data and the majority of benchmark comparisons are independent; the circularity is localized to the interpretation/reasoning evaluation, so a moderate score is appropriate.
Assumptions & free parameters
free parameters (2)
- Preference probability calibration mapping per model =
not specified
- Strong preference threshold =
[0.3, 0.7]
assumptions (4)
- standard math Preference data follow the Thurstone Case V model with a zero-mean unit-variance Gaussian prior over quality scores.
- domain assumption Pairwise preference under a fixed prompt is a valid measure of advertising communicability.
- domain assumption Resizing all images to 512x512 preserves information clarity and emotional impact of advertisements.
- domain assumption GPT-4o-generated descriptions, after human expert verification, are a valid golden standard for preference interpretation and reasoning.
Cite this review
Pith. "Pith review of AI-generated Image Quality Assessment in Visual Communication." pith.science (2026). https://pith.science/paper/PW5OG7BZ
@misc{pith2026241215677,
author = {Pith},
title = {Pith review of: AI-generated Image Quality Assessment in Visual Communication},
year = {2026},
howpublished = {\url{https://pith.science/paper/PW5OG7BZ}},
note = {Machine review of arXiv:2412.15677}
}
read the original abstract
Assessing the quality of artificial intelligence-generated images (AIGIs) plays a crucial role in their application in real-world scenarios. However, traditional image quality assessment (IQA) algorithms primarily focus on low-level visual perception, while existing IQA works on AIGIs overemphasize the generated content itself, neglecting its effectiveness in real-world applications. To bridge this gap, we propose AIGI-VC, a quality assessment database for AI-Generated Images in Visual Communication, which studies the communicability of AIGIs in the advertising field from the perspectives of information clarity and emotional interaction. The dataset consists of 2,500 images spanning 14 advertisement topics and 8 emotion types. It provides coarse-grained human preference annotations and fine-grained preference descriptions, benchmarking the abilities of IQA methods in preference prediction, interpretation, and reasoning. We conduct an empirical study of existing representative IQA methods and large multi-modal models on the AIGI-VC dataset, uncovering their strengths and weaknesses.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[2]
Akhtar, M. H.; and Ramkumar, J. 2023. AI in Visual Communication: AI Taking Down Graphic Designers. Scary? In AI for Designers, 85--103. Springer
work page 2023
-
[3]
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023. Qwen-VL : A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv preprint arXiv:2308.12966
arXiv 2023
-
[4]
Bao, Q.; Hui, Z.; Zhu, R.; Ren, P.; Xie, X.; and Yang, W. 2024. Improving Diffusion-Based Image Restoration with Error Contraction and Error Correction. In AAAI, volume 38, 756--764
work page 2024
-
[5]
Campbell, C.; Plangger, K.; Sands, S.; and Kietzmann, J. 2022. Preparing for An Era of Deepfakes and AI -generated Ads: A Framework for Understanding Responses to Manipulated Advertising. Journal of Advertising, 51(1): 22--38
work page 2022
-
[6]
Chen, B.; Zhu, H.; Zhu, L.; Wang, S.; and Kwong, S. 2024 a . Deep Feature Statistics Mapping for Generalized Screen Content Image Quality Assessment. IEEE Trans. Image Process., 33: 3227--3241
work page 2024
-
[7]
Chen, C.; Zhou, S.; Liao, L.; Wu, H.; Sun, W.; Yan, Q.; and Lin, W. 2024 b . Iterative Token Evaluation and Refinement for Real-world Super-Resolution. In AAAI, volume 38, 1010--1018
work page 2024
-
[8]
Chen, J.; An, J.; Lyu, H.; Kanan, C.; and Luo, J. 2024 c . Learning to Evaluate The Artness of AI-generated Images. IEEE Trans. Multimedia, 1--10
work page 2024
Show all 62 references
-
[9]
Chen, Z.; Sun, W.; Wu, H.; Zhang, Z.; Jia, J.; Min, X.; Zhai, G.; and Zhang, W. 2023. Exploring the Naturalness of AI -generated Images. arXiv preprint arXiv:2312.05476
2023 arXiv
-
[10]
Chen, Z.; Zhou, Q.; Shen, Y.; Hong, Y.; Sun, Z.; Gutfreund, D.; and Gan, C. 2024 d . Visual Chain-of-thought Prompting for Knowledge-Based Visual Reasoning. In AAAI, volume 38, 1254--1262
2024
-
[11]
Cho, J.; Hu, Y.; Baldridge, J.; Garg, R.; Anderson, P.; Krishna, R.; Bansal, M.; Pont-Tuset, J.; and Wang, S. 2024. Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-image Generation. In ICLR
2024
-
[12]
Cho, J.; Zala, A.; and Bansal, M. 2023. Dall-eval: Probing The Reasoning Skills and Social Biases of Text-to-image Generation Models. In ICCV, 3043--3054
2023
-
[13]
Dong, X.; Zhang, P.; Zang, Y.; Cao, Y.; Wang, B.; Ouyang, L.; Wei, X.; Zhang, S.; Duan, H.; Cao, M.; Zhang, W.; Li, Y.; Yan, H.; Gao, Y.; Zhang, X.; Li, W.; Li, J.; Chen, K.; He, C.; Zhang, X.; Qiao, Y.; Lin, D.; and Wang, J. 2024. InternLM-XComposer2 : Mastering Free-form Tex...
2024 arXiv
-
[14]
Duan, X.; Ma, S.; Liu, H.; and Jia, C. 2024. PKU-AIGI-500K : A Neural Compression Benchmark and Model for AI-Generated Images. IEEE J. Emerg. Sel. Topics Circuits Syst
2024
-
[15]
Ford, J.; Jain, V.; Wadhwani, K.; and Gupta, D. G. 2023. AI Advertising: An Overview and Guidelines. Journal of Business Research, 166: 114124
2023
-
[16]
Gao, Y.; Min, X.; Zhu, Y.; Zhang, X.-P.; and Zhai, G. 2024. Blind Image Quality Assessment: A Fuzzy Neural Network for Opinion Score Distribution Prediction. IEEE Trans. Circuits Syst. Video Technol., 34(3): 1641--1655
2024
-
[17]
Gu, S.; Bao, J.; Chen, D.; and Wen, F. 2020. GIQA : Generated Image Quality Assessment. In ECCV, 369--385
2020
-
[18]
L.; and Choi, Y
Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021. CLIPScore : A Reference-free Evaluation Metric for Image Captioning. arXiv preprint arXiv:2104.08718
2021 arXiv
-
[19]
Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising Diffusion Probabilistic Models. NeurIPS, 33: 6840--6851
2020
-
[20]
B.; and O'Shaughnessy, J
Holbrook, M. B.; and O'Shaughnessy, J. 1984. The Role of Emotion in Advertising. Psychology & Marketing, 1(2): 45--64
1984
-
[21]
Huang, K.; Sun, K.; Xie, E.; Li, Z.; and Liu, X. 2024 a . T2I -compbench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation. NeurIPS, 36
2024
-
[22]
Huang, Y.; Sheng, X.; Yang, Z.; Yuan, Q.; Duan, Z.; Chen, P.; Li, L.; Lin, W.; and Shi, G. 2024 b . AesExpert : Towards Multi-modality Foundation Model for Image Aesthetics Perception. arXiv preprint arXiv:2404.09624
2024 arXiv
-
[23]
Hussain, Z.; Zhang, M.; Zhang, X.; Ye, K.; Thomas, C.; Agha, Z.; Ong, N.; and Kovashka, A. 2017. Automatic Understanding of Image and Video Advertisements. In CVPR, 1705--1715
2017
-
[24]
Jiang-Lin, J.-Y.; Huang, K.-Y.; Lo, L.; Huang, Y.-N.; Lin, T.; Wu, J.-C.; Shuai, H.-H.; and Cheng, W.-H. 2024. ReCorD : Reasoning and Correcting Diffusion for HOI Generation. In ACM MM
2024
-
[25]
Kirstain, Y.; Polyak, A.; Singer, U.; Matiana, S.; Penna, J.; and Levy, O. 2024. Pick-a-pic: An Open Dataset of User Preferences for Text-to-image Generation. NeurIPS, 36
2024
-
[26]
Lauren c on, H.; Tronchon, L.; Cord, M.; and Sanh, V. 2024. What Matters when Building Vision-language models? arXiv preprint arXiv:2405.02246
2024 arXiv
-
[27]
S.; Gupta, A.; Zhang, Y.; Narayanan, D.; Teufel, H.; Bellagente, M.; et al
Lee, T.; Yasunaga, M.; Meng, C.; Mai, Y.; Park, J. S.; Gupta, A.; Zhang, Y.; Narayanan, D.; Teufel, H.; Bellagente, M.; et al. 2024. Holistic Evaluation of Text-to-image Models. NeurIPS, 36
2024
-
[28]
Li, C.; Zhang, Z.; Wu, H.; Sun, W.; Min, X.; Liu, X.; Zhai, G.; and Lin, W. 2023. AGIQA-3K : An Open Database for AI -Generated Image Quality Assessment. IEEE Trans. Circuits Syst. Video Technol
2023
-
[29]
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024. Visual Instruction Tuning. NeurIPS, 36
2024
-
[30]
E.; and Wang, W
Lu, Y.; Yang, X.; Li, X.; Wang, X. E.; and Wang, W. Y. 2024. LLMscore : Unveiling the Power of Large Language Models in Text-to-image Synthesis Evaluation. NeurIPS, 36
2024
-
[31]
Ma, K.; Liu, W.; Liu, T.; Wang, Z.; and Tao, D. 2017. dipIQ : Blind Image Quality Assessment by Learning-to-rank Discriminable Image Pairs. IEEE Trans. Image Process., 26(8): 3951--3964
2017
-
[32]
A.; Fredrickson, B
Mikels, J. A.; Fredrickson, B. L.; Larkin, G. R.; Lindberg, C. M.; Maglio, S. J.; and Reuter-Lorenz, P. A. 2005. Emotional Category Data on Images from The International Affective Picture System. Behavior research methods, 37: 626--630
2005
-
[33]
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2022. GLIDE : Towards Photorealistic Image Generation and Editing with Text-guided Diffusion Models. In ICML, 16784--16804
2022
-
[34]
OpenAI. 2023. GPT-4V(ision) System Card. https://cdn.openai.com/papers/GPTV_System_Card.pdf
2023
-
[35]
Prashnani, E.; Cai, H.; Mostofi, Y.; and Sen, P. 2018. PieAPP : Perceptual Image-error Assessment through Pairwise Preference. In CVPR, 1808--1817
2018
-
[36]
D.; Crowson, K.; and Contributors, S
Pressman, J. D.; Crowson, K.; and Contributors, S. C. 2022. Simulacra Aesthetic Captions. Technical Report Version 1.0, Stability AI. \ url https://github.com/JD-P/simulacra-aesthetic-captions
2022
-
[37]
Quan, H.; Li, S.; Zeng, C.; Wei, H.; and Hu, J. 2023. Big Data and AI -driven Product Design: A Survey. Applied Sciences, 13(16): 9433
2023
-
[38]
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022. Hierarchical Text-conditional Image Generation with CLIP Latents. arXiv preprint arXiv:2204.06125
2022 arXiv
-
[39]
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022. High-resolution Image Synthesis with Latent Diffusion Models. In CVPR, 10684--10695
2022
-
[40]
Sagar, A.; Srivastava, R.; Rakshitha; Kesav, V.; and Kiran, R. 2024. MAdVerse : A Hierarchical Dataset of Multi-Lingual Ads from Diverse Sources and Categories. In IEEE Winter Conf. App. Comput. Vis
2024
-
[41]
Schramowski, P.; Brack, M.; Deiseroth, B.; and Kersting, K. 2023. Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. In CVPR, 22522--22531
2023
-
[42]
Series, B. T. 2012. Methodology for The Subjective Assessment of The Quality of Television Pictures. Recommendation ITU-R BT, 500(13)
2012
-
[43]
L.; and Wang, L
She, D.; Yang, J.; Cheng, M.-M.; Lai, Y.-K.; Rosin, P. L.; and Wang, L. 2020. WSCNet: Weakly Supervised Coupled Networks for Visual Sentiment Classification and Detection. IEEE Trans. Multimedia, 22(5): 1358--1371
2020
-
[44]
SkunkworksAI. 2024. BakLLaVA. https://github.com/SkunkworksAI/BakLLaVA
2024
-
[45]
Su, S.; Yan, Q.; Zhu, Y.; Zhang, C.; Ge, X.; Sun, J.; and Zhang, Y. 2020. Blindly Assess Image Quality in The Wild Guided by A Self-adaptive Hyper Network. In CVPR, 3667--3676
2020
-
[46]
Teo, C.; Abdollahzadeh, M.; and Cheung, N.-M. M. 2024. On Measuring Fairness in Generative Models. NeurIPS, 36
2024
-
[47]
Tian, Y.; Wang, S.; Chen, B.; and Kwong, S. 2024. Causal Representation Learning for GAN-Generated Face Image Quality Assessment. IEEE Trans. Circuits Syst. Video Technol., 1--1
2024
-
[48]
R.; et al
Tsukida, K.; Gupta, M. R.; et al. 2011. How to Analyze Paired Comparison Data. Department of Electrical Engineering University of Washington, Tech. Rep. UWEETR-2011-0004, 1
2011
-
[49]
Wu, H.; Zhang, Z.; Zhang, E.; Chen, C.; Liao, L.; Wang, A.; Li, C.; Sun, W.; Yan, Q.; Zhai, G.; and Lin, W. 2024 a . Q-Bench : A Benchmark for General-Purpose Foundation Models on Low-level Vision. In ICLR
2024
-
[50]
Wu, H.; Zhang, Z.; Zhang, E.; Chen, C.; Liao, L.; Wang, A.; Xu, K.; Li, C.; Hou, J.; Zhai, G.; et al. 2024 b . Q-instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models. In CVPR, 25490--25500
2024
-
[51]
Wu, X.; Hao, Y.; Sun, K.; Chen, Y.; Zhu, F.; Zhao, R.; and Li, H. 2023. Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis. arXiv preprint arXiv:2306.09341
2023 arXiv
-
[52]
Wu, X.; Xiao, L.; Sun, Y.; Zhang, J.; Ma, T.; and He, L. 2022. A Survey of Human-in-the-loop for Machine Learning. Future Gener. Comp. Sy., 135: 364--381
2022
-
[53]
Xu, J.; Liu, X.; Wu, Y.; Tong, Y.; Li, Q.; Ding, M.; Tang, J.; and Dong, Y. 2024. ImageReward : Learning and Evaluating Human Preferences for Text-to-image Generation. NeurIPS, 36
2024
-
[54]
Yang, J.; Huang, Q.; Ding, T.; Lischinski, D.; Cohen-Or, D.; and Huang, H. 2023. EmoSet : A Large-scale Visual Emotion Dataset with Rich Attributes. In ICCV, 20383--20394
2023
-
[55]
Yarom, M.; Bitton, Y.; Changpinyo, S.; Aharoni, R.; Herzig, J.; Lang, O.; Ofek, E.; and Szpektor, I. 2024. What You See is What You Read? Improving Text-image Alignment Evaluation. NeurIPS, 36
2024
-
[56]
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023. mPLUG-Owl : Modularization Empowers Large Language Models with Multimodality. arXiv preprint arXiv:2304.14178
2023 arXiv
-
[57]
Zhang, W.; Ma, K.; Zhai, G.; and Yang, X. 2021. Uncertainty-aware Blind Image Quality Assessment in The Laboratory and Wild. IEEE Trans. Image Process., 30: 3474--3486
2021
-
[58]
Zhao, S.; Gao, Y.; Jiang, X.; Yao, H.; Chua, T.-S.; and Sun, X. 2014. Exploring Principles-of-Art Features For Image Emotion Recognition. In ACM MM, 47–56
2014
-
[59]
Zhu, H.; Sui, X.; Chen, B.; Liu, X.; Chen, P.; Fang, Y.; and Wang, S. 2024 a . 2AFC Prompting of Large Multimodal Models for Image Quality Assessment. arXiv preprint arXiv:2402.01162
2024 arXiv
-
[60]
Zhu, H.; Wu, H.; Li, Y.; Zhang, Z.; Chen, B.; Zhu, L.; Fang, Y.; Zhai, G.; Lin, W.; and Wang, S. 2024 b . Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare. NeurIPS
2024
-
[61]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[62]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.