REVIEW 4 major objections 6 minor 56 references
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper reports that GPT-4o and Claude 3.5 Sonnet agree with crowd-rated ground truth in about 70 percent of pairwise chart comparisons, while their direct 7-point experiential scores are compressed and only weakly-to-moderately…
desk verdict A genuinely new benchmark dataset, but the headline claim about reliable pairwise comparisons needs confidence intervals before I'd trust the 0.70/0.69 numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing artifact is Chart-to-Experience itself: 36 charts selected across three topics (COVID-19, house prices, global warming), with 12 charts per topic spanning common chart types, journalistic visuals, and detailed infographics. Each chart-factor combination carries approximately 36 crowd-sourced 7-point Likert ratings (216 workers per topic in a rotating design, with an 'insufficient information' option), and the seven factors are memorability, interest, trustworthiness, empathy, aesthetic pleasure, intuitiveness, and comfort. The evaluation machinery is the two task types built on those ratings: direct score prediction, where each model rates charts under 216 personas so its variance and rank correlation can be compared with human variance, and pairwise comparison, where human mean scores define the correct answer for each chart pair and accuracy is analyzed within bins of human-score difference. The explanations the models gave for their pairwise choices serve as the material for post hoc failure analysis.
What would settle it
Resample the existing crowd ratings per chart-factor by bootstrapping the 36 ratings and recompute the pairwise labels; if a substantial fraction of pairs change their winner, or if a held-out set of charts from new topics drops GPT-4o and Claude 3.5 accuracy near chance, the claim that these models are accurate and reliable pairwise judges fails.
Extended reading notes
Core claim
On Chart-to-Experience, the central empirical claim is two-sided. First, state-of-the-art MLLMs are poor absolute judges: across the seven factors, their score distributions are compressed (standard deviations 0.45–1.22 versus human 1.63–2.20) and their rank correlation with human mean scores is weak to moderate, with the clearest alignment on aesthetic pleasure and intuitiveness. Second, MLLMs are better relative judges: GPT-4o and Claude 3.5 Sonnet select the chart with the higher human mean in about 70% of pairwise comparisons (0.70 and 0.69 overall), while Llama-3.2-11B-Vision-Instruct hovers near chance at 0.52. Accuracy rises with the gap between human means, so the models perform best on pairs humans also find easy and fall toward chance on near-ties, with notable dips for visually busy charts that models overvalue as interesting and beautiful.
Load-bearing premise
The load-bearing premise is that the average of about 36 crowd ratings per chart and factor is a stable enough measure of a chart's true experiential impact; if sampling noise, topic bias, or the authors' chart selection can flip which chart wins a pairwise comparison, the reported accuracy rates would not generalize.
Editorial extensions
If this is right
- Pairwise chart comparison with GPT-4o or Claude 3.5 Sonnet can serve as a usable proxy for human preference in visualization evaluation, provided the charts being compared are not near-ties in human ratings.
- Direct 7-point experiential scores from MLLMs should not be used as evaluation metrics, since their compressed range would systematically understate differences between chart designs.
- Evaluations of chart experience should report factor-specific results, because alignment with human ratings varies from moderate (aesthetic pleasure, intuitiveness) to weak or absent for other factors.
- The accuracy-versus-rating-gap trend means an MLLM judge is most trustworthy for choosing among clearly distinct design candidates and least trustworthy for fine-grained design iteration.
- Failure analysis of pairwise choices can pinpoint specific chart styles that MLLMs systematically misjudge, such as complex, color-vivid infographics that models find appealing but humans find chaotic.
Reading between the lines
- A practical acceptance test for any chart-judging MLLM would be its accuracy in the 1.4–1.6 score-gap bin; models that cannot beat chance there should not be used to rank visual design candidates.
- Because the benchmark uses only 36 charts and three topics, the reported 70% accuracies are likely optimistic for deployment, and a broader chart distribution could be expected to pull them toward chance.
- The compressed variance of direct scores suggests that a simple recalibration, scaling model scores to match human variance, might recover some absolute-score utility, a test the paper does not run.
- The pairwise paradigm transfers naturally to new experiential factors such as joy, surprise, or persuasiveness, but each new factor would need its own human norms rather than assuming the same accuracy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Chart-to-Experience, a benchmark of 36 charts across three topics (COVID-19, House Prices, Global Warming) with crowdsourced 7-point Likert ratings on seven experiential factors (memorability, interest, trustworthiness, empathy, aesthetic pleasure, intuitiveness, comfort) collected from 216 Prolific workers, yielding roughly 36 ratings per chart-factor along with textual explanations. It then evaluates three multimodal LLMs (GPT-4o, Claude 3.5 Sonnet, Llama-3.2-11B-Vision-Instruct) on two tasks: direct Likert score prediction using generated personas (Section 4.1) and pairwise comparison of charts (Section 4.2). The main reported findings are that MLLM score distributions are compressed relative to humans (SD 0.45-1.22 vs 1.63-2.20) with weak-to-moderate Kendall tau correlations, while GPT-4o and Claude achieve overall pairwise accuracies of 0.70 and 0.69 that rise with human-rating gaps. The authors conclude that MLLMs are not sensitive enough for absolute scoring but are accurate and reliable in pairwise comparisons.
Significance. If the pairwise result is robust, the paper offers a useful benchmark and a clear practical message: MLLM judges can be used for relative ranking of charts' experiential impact but not for absolute score prediction. The public release of chart images, human ratings, explanations, and prompts is a concrete strength, as is the inclusion of two commercial and one open-weight MLLM and the factor-level breakdown. The dataset is small (36 charts, three topics), the seven factors are reasonable but not exhaustive, and the sampling is not described as stratified beyond topic, so the benchmark is best viewed as an initial probe rather than a definitive evaluation. The core contribution—the dataset plus the observed score-vs-comparison dissociation—is interesting, but the pairwise claim needs substantially stronger statistical support before it can be taken as established.
major comments (4)
- [Section 4.2, Table 3] The central pairwise-accuracy claim rests on labels derived from means of about 36 ratings per chart-factor (Section 3.4). With human SDs of 1.63-2.20, the standard error of each mean is roughly 0.27-0.37 and the standard error of a difference between two chart means is roughly 0.38-0.52, so many pairwise labels, especially in the small-difference regime, are likely to be incorrect relative to the true population ordering. This measurement noise attenuates observed accuracy toward 0.5. The manuscript reports no confidence intervals, cluster bootstrap, intraclass correlation, split-half reliability, chance baseline, or repeated MLLM sampling, so the point estimates 0.70 and 0.69 do not by themselves establish that GPT-4o and Claude are accurate and reliable at pairwise comparison. Please add uncertainty quantification and a comparison against a 0.5 chance baseline, or a noise-corrected estimate of true agreement.
- [Section 4.2, Figure 3] The upward accuracy trend in Figure 3 is partly an artifact of binning comparisons by the observed human-rating difference. Because the observed difference contains measurement error, larger observed differences are more likely to reflect genuinely large differences (and thus easier comparisons), while smaller observed differences are more likely to have their sign flipped by noise. Without correcting for this regression-to-the-mean effect or reporting per-bin intervals, the figure does not cleanly establish that MLLM accuracy tracks true task difficulty. Please provide a formal trend test, error bars, and/or an analysis that models measurement error in the binning variable.
- [Section 4.1, Table 2] The conclusion that MLLMs are not as sensitive as human evaluators is based on comparing the SD of MLLM Likert outputs to the SD of human ratings, but the MLLM outputs were generated under an unvalidated persona protocol: 216 author-written persona profiles are assumed to capture human response variability, with no evidence that MLLMs actually vary across personas, no no-persona baseline, and no comparison to the actual demographic or response distributions of the crowd workers. The compressed SD range (0.45-1.22) could be an artifact of prompt design or model behavior rather than a stable property of MLLMs. Please validate the persona manipulation (e.g., inter-persona variance, distributional similarity to human responses, or a condition without personas) before drawing the sensitivity conclusion.
- [Section 4.2, pairwise task protocol] The paper's related-work section notes that LLM judges exhibit an ordering bias in pairwise comparisons [54], but the manuscript does not state whether chart-pair order was randomized or whether accuracy differs by position, nor does it report temperature or seed settings. Since the pairwise claim is the paper's main positive result, an order-bias check (e.g., first-position vs. second-position accuracy) is needed. Additionally, the reported one-way ANOVA (p<0.001) supports that not all models have equal accuracy, but it does not establish that each of GPT-4o and Claude is significantly above chance or above Llama; please add post-hoc comparisons with multiple-comparison correction and account for the non-independence of the multiple pairs evaluated per factor and per chart.
minor comments (6)
- [Abstract, Section 6] The wording 'accurate and reliable in pairwise comparisons' is stronger than the factor-level results in Table 3 support, where several factor accuracies for GPT-4o and Claude fall between 0.62 and 0.66; please qualify the claim or report confidence intervals that justify it.
- [Table 2] Kendall's tau values are reported without significance tests or confidence intervals; with only 36 charts, many weak correlations (e.g., tau below 0.2) may not be reliably different from zero.
- [Section 3.4] The paper reports that 4.1% of responses were ignored due to 'Insufficient information to answer,' but it does not state whether the analysis was repeated with different missing-data treatments or whether the missingness pattern could bias the factor means; please clarify.
- [Figure 3] Figure 3 has no error bars and appears to use variable bin widths; the binning rule is not described, and the number of pairs per bin is not reported, which makes the apparent trends and the 1.4-1.6 dips hard to evaluate.
- [Section 4.2] The manuscript does not state whether ties in human mean ratings were possible and how they were handled in computing pairwise accuracy; if ties were assigned to one label, this should be reported.
- [General] There are minor typographical issues: 'we did not explore' in the Conclusion should begin with a capital letter, 'One-way ANOV A' has a stray space, and the author list in the header contains spacing artifacts that should be cleaned.
Circularity Check
No significant circularity: the benchmark ground truth and model outputs are independent, and no fitted parameter or self-citation is load-bearing.
full rationale
Chart-to-Experience constructs its ground truth from 216 crowdsourced workers' 7-point Likert ratings (Section 3.3-3.4), independent of the MLLM outputs. The two evaluation tasks (Section 4.1 and 4.2) compare model outputs to these external human ratings; no model parameter is fitted to the ground truth, and no target quantity is defined in terms of model output. The pairwise comparisons in Section 4.2 use the mean human rating as the label, but the model's choices are made without access to those means, so the accuracy measure is an external agreement statistic, not a self-fulfilling construction. The personas used in Task 1 are author-generated prompt conditions, not calibrated to the human data, so the finding that MLLM score distributions are narrower than human distributions is an empirical observation rather than a circular derivation. The limitations acknowledged in Section 6 (no Chain-of-Thought/few-shot prompting, limited demographic modeling, and potential additional factors) concern scope and generalizability, not circularity. No self-citation is used as load-bearing evidence: references to prior benchmark work are contextual. The skeptical concern about small per-chart sample sizes (about 36 ratings per chart-factor) and absent confidence intervals is a statistical robustness issue, not a circularity issue, and therefore does not raise the circularity score.
Assumptions & free parameters
assumptions (4)
- domain assumption Self-reported Likert ratings from crowd workers accurately measure the experiential impact of charts.
- domain assumption The mean of roughly 36 ratings per chart-factor is a stable estimate of true experiential impact.
- ad hoc to paper Generated persona profiles capture enough human response variability to allow a comparison of MLLM and human sensitivity.
- domain assumption The 36 charts selected by the authors are representative of charts relevant to experiential impact.
invented entities (1)
-
Generated persona profiles for MLLM evaluators
Cite this review
Pith. "Pith review of Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts." pith.science (2026). https://pith.science/paper/O66F2HYN
@misc{pith2026250517374,
author = {Pith},
title = {Pith review of: Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts},
year = {2026},
howpublished = {\url{https://pith.science/paper/O66F2HYN}},
note = {Machine review of arXiv:2505.17374}
}
read the original abstract
The field of Multimodal Large Language Models (MLLMs) has made remarkable progress in visual understanding tasks, presenting a vast opportunity to predict the perceptual and emotional impact of charts. However, it also raises concerns, as many applications of LLMs are based on overgeneralized assumptions from a few examples, lacking sufficient validation of their performance and effectiveness. We introduce Chart-to-Experience, a benchmark dataset comprising 36 charts, evaluated by crowdsourced workers for their impact on seven experiential factors. Using the dataset as ground truth, we evaluated capabilities of state-of-the-art MLLMs on two tasks: direct prediction and pairwise comparison of charts. Our findings imply that MLLMs are not as sensitive as human evaluators when assessing individual charts, but are accurate and reliable in pairwise comparisons.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[54]
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. Judging llm-as-a-judge with mt- bench and chatbot arena. Advances in Neural Information Processing Systems, 36:46595–46623, 2023
work page 2023
-
[1]
Y . Abe, T. Daikoku, and Y . Kuniyoshi. Assessing the aesthetic evalu- ation capabilities of gpt-4 with vision: Insights from group and indi- vidual assessments. pp. 2Q1IS301–2Q1IS301, 2024
work page 2024
- [2]
-
[3]
L. Bartram, A. Patra, and M. Stone. Affective color in visualization. In Proceedings of the 2017 CHI conference on human factors in com- puting systems, pp. 1364–1374, 2017
work page 2017
-
[4]
M. Behrisch, M. Blumenschein, N. W. Kim, L. Shao, M. El-Assady, J. Fuchs, D. Seebacher, A. Diehl, U. Brandes, H. Pfister, et al. Quality metrics for information visualization. In Computer Graphics F orum, vol. 37, pp. 625–662. Wiley Online Library, 2018
work page 2018
- [5]
-
[6]
M. A. Borkin, A. A. V o, Z. Bylinskii, P. Isola, S. Sunkavalli, A. Oliva, and H. Pfister. What makes a visualization memorable? IEEE trans- actions on visualization and computer graphics , 19(12):2306–2315, 2013
work page 2013
-
[7]
J. Boy, A. V . Pandey, J. Emerson, M. Satterthwaite, O. Nov, and E. Bertini. Showing people behind data: Does anthropomorphizing visualizations elicit more empathy for human rights data? In Pro- ceedings of the 2017 CHI conference on human factors in computing systems, pp. 5462–5474, 2017
work page 2017
Show all 56 references
-
[8]
D. Chen, R. Chen, S. Zhang, Y . Liu, Y . Wang, H. Zhou, Q. Zhang, Y . Wan, P. Zhou, and L. Sun. Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark. arXiv preprint arXiv:2402.04788, 2024
2024 arXiv
-
[9]
Datta, D
R. Datta, D. Joshi, J. Li, and J. Z. Wang. Studying aesthetics in photographic images using a computational approach. In Computer Vision–ECCV 2006: 9th European Conference on Computer Vision, Graz, Austria, May 7-13, 2006, Proceedings, Part III 9 , pp. 288–301. Springer, 2006
2006
-
[10]
A. Dave, A. Saxena, and A. Jha. Understanding user comfort and expectations in ai-based systems. 2023
2023
-
[11]
K. Deng, A. Ray, R. Tan, S. Gabriel, B. A. Plummer, and K. Saenko. Socratis: Are large multimodal models emotionally aware? arXiv preprint arXiv:2308.16741, 2023
2023 arXiv
-
[12]
P. Duan, J. Warner, Y . Li, and B. Hartmann. Generating automatic feedback on ui mockups with large language models. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pp. 1–20, 2024
2024
-
[13]
Errey, J
N. Errey, J. Liang, T. W. Leong, and D. Zowghi. Evaluating narrative visualization: a survey of practitioners. International Journal of Data Science and Analytics, 18(1):19–34, 2024. 4available at http://chart2experience.github.io
2024
-
[14]
Few and P
S. Few and P. Edge. Data visualization effectiveness profile. Percep- tual Edge, 10:12, 2017
2017
-
[15]
Y . Guo, F. Siddiqui, Y . Zhao, R. Chellappa, and S.-Y . Lo. Stimu- var: Spatiotemporal stimuli-aware video affective reasoning with mul- timodal large language models. arXiv preprint arXiv:2409.00304 , 2024
2024 arXiv
-
[16]
Y . Han, C. Zhang, X. Chen, X. Yang, Z. Wang, G. Yu, B. Fu, and H. Zhang. Chartllama: A multimodal llm for chart understanding and generation. arXiv preprint arXiv:2311.16483, 2023
2023 arXiv
-
[17]
Huang, H
K.-H. Huang, H. P. Chan, Y . R. Fung, H. Qiu, M. Zhou, S. Joty, S.-F. Chang, and H. Ji. From pixels to insights: A survey on automatic chart understanding in the era of large foundation models. arXiv preprint arXiv:2403.12027, 2024
2024 arXiv
-
[18]
Huang, P
W. Huang, P. Eades, and S.-H. Hong. Measuring effectiveness of graph visualizations: A cognitive load perspective. Information Vi- sualization, 8(3):139–152, 2009
2009
-
[19]
Huddy and A
L. Huddy and A. H. Gunnthorsdottir. The persuasive effects of emo- tive visual imagery: Superficial manipulation or the product of pas- sionate reason? Political Psychology, 21(4):745–778, 2000
2000
-
[20]
Isola, J
P. Isola, J. Xiao, D. Parikh, A. Torralba, and A. Oliva. What makes a photograph memorable? IEEE transactions on pattern analysis and machine intelligence, 36(7):1469–1482, 2013
2013
-
[21]
Kafle, B
K. Kafle, B. Price, S. Cohen, and C. Kanan. Dvqa: Understanding data visualizations via question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 5648– 5656, 2018
2018
-
[22]
Kantharaj, R
S. Kantharaj, R. T. K. Leong, X. Lin, A. Masry, M. Thakkar, E. Hoque, and S. Joty. Chart-to-text: A large-scale benchmark for chart summa- rization. arXiv preprint arXiv:2203.06486, 2022
2022 arXiv
-
[23]
X. Lan, Y . Shi, Y . Wu, X. Jiao, and N. Cao. Kineticharts: Augment- ing affective expressiveness of charts in data stories with animation design. IEEE Transactions on Visualization and Computer Graphics , 28(1):933–943, 2021
2021
-
[24]
X. Lan, Y . Shi, Y . Zhang, and N. Cao. Smile or scowl? looking at infographic design through the affective lens. IEEE Transactions on Visualization and Computer Graphics, 27(6):2796–2807, 2021
2021
-
[25]
X. Lan, Y . Wu, and N. Cao. Affective visualization design: Leveraging the emotional impact of data. IEEE Transactions on Visualization and Computer Graphics, 2023
2023
-
[26]
S. Lee, S. Kim, S. H. Park, G. Kim, and M. Seo. Prometheusvision: Vision-language model as a judge for fine-grained evaluation. arXiv preprint arXiv:2401.06591, 2024
2024 arXiv
-
[27]
Lee-Robbins and E
E. Lee-Robbins and E. Adar. Affective learning objectives for com- municative visualizations. IEEE Transactions on Visualization and Computer Graphics, 29(1):1–11, 2022
2022
-
[28]
D. Li, B. Jiang, L. Huang, A. Beigi, C. Zhao, Z. Tan, A. Bhat- tacharjee, Y . Jiang, C. Chen, T. Wu, et al. From generation to judg- ment: Opportunities and challenges of llm-as-a-judge. arXiv preprint arXiv:2411.16594, 2024
2024
-
[29]
Z. Lian, L. Sun, H. Sun, K. Chen, Z. Wen, H. Gu, B. Liu, and J. Tao. Gpt-4v with emotion: A zero-shot benchmark for generalized emotion recognition. Information Fusion, 108:102367, 2024
2024
-
[30]
J. Liem, C. Perin, and J. Wood. Structure and empathy in visual data storytelling: Evaluating their influence on attitude. In Computer Graphics F orum, vol. 39, pp. 277–289. Wiley Online Library, 2020
2020
-
[31]
F. Liu, J. M. Eisenschlos, F. Piccinno, S. Krichene, C. Pang, K. Lee, M. Joshi, W. Chen, N. Collier, and Y . Altun. Deplot: One-shot vi- sual language reasoning by plot-to-table translation. arXiv preprint arXiv:2212.10505, 2022
2022 arXiv
-
[32]
Machajdik and A
J. Machajdik and A. Hanbury. Affective image classification using features inspired by psychology and art theory. In Proceedings of the 18th ACM international conference on Multimedia , pp. 83–92, 2010
2010
-
[33]
Mehrabian
A. Mehrabian. An approach to environmental psychology. Mas- sachusetts Institute of Technology, 1974
1974
-
[34]
Micallef, G
L. Micallef, G. Palmas, A. Oulasvirta, and T. Weinkauf. Towards per- ceptual optimization of the visual design of scatterplots. IEEE trans- actions on visualization and computer graphics , 23(6):1588–1599, 2017
2017
-
[35]
Obeid and E
J. Obeid and E. Hoque. Chart-to-text: Generating natural language de- scriptions for charts by adapting the transformer model.arXiv preprint arXiv:2010.09142, 2020
2010 arXiv
-
[36]
E. M. Peck, S. E. Ayuso, and O. El-Etr. Data is personal: Attitudes and perceptions of data visualization in rural pennsylvania. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pp. 1–12, 2019
2019
-
[37]
Sabini and M
J. Sabini and M. Silver. Ekman’s basic emotions: Why not love and jealousy? Cognition & Emotion, 19(5):693–712, 2005
2005
-
[38]
Saket, A
B. Saket, A. Endert, and J. Stasko. Beyond usability and performance: A review of user experience-focused evaluations in visualization. In Proceedings of the Sixth Workshop on Beyond Time and Errors on Novel Evaluation Methods for Visualization, pp. 133–142, 2016
2016
-
[39]
J. Stasko. Value-driven evaluation of visualizations. In Proceedings of the Fifth Workshop on Beyond Time and Errors: Novel Evaluation Methods for Visualization, pp. 46–53, 2014
2014
-
[40]
T. Sun, Y . Shao, H. Qian, X. Huang, and X. Qiu. Black-box tuning for language-model-as-a-service. In International Conference on Ma- chine Learning, pp. 20841–20855. PMLR, 2022
2022
-
[41]
B. J. Tang, A. Boggust, and A. Satyanarayan. Vistext: A benchmark for semantically rich chart captioning. arXiv preprint arXiv:2307.05356, 2023
2023 arXiv
-
[42]
J. Tang, Q. Liu, Y . Ye, J. Lu, S. Wei, C. Lin, W. Li, M. F. F. B. Mahmood, H. Feng, Z. Zhao, et al. Mtvqa: Benchmarking multilingual text-centric visual question answering. arXiv preprint arXiv:2405.11985, 2024
2024 arXiv
-
[43]
A. Tatu, P. Bak, E. Bertini, D. Keim, and J. Schneidewind. Visual qual- ity metrics and human perception: an initial study on 2d projections of large multidimensional data. In Proceedings of the international conference on advanced visual interfaces , pp. 49–56, 2010
2010
-
[44]
Valdez and A
P. Valdez and A. Mehrabian. Effects of color on emotions. Journal of experimental psychology: General, 123(4):394, 1994
1994
-
[45]
Y . Wang, A. Segal, R. Klatzky, D. F. Keefe, P. Isenberg, J. Hurtienne, E. Hornecker, T. Dwyer, and S. Barrass. An emotional response to the value of visualization. IEEE computer graphics and applications , 39(5):8–17, 2019
2019
-
[46]
Willigen
T. Willigen. Measuring the user experience of data visualization. Mas- ter’s thesis, University of Twente, 2019
2019
-
[47]
Y . Wu, C. Bauckhage, and C. Thurau. The good, the bad, and the ugly: Predicting aesthetic image labels. In 2010 20th International Conference on Pattern Recognition, pp. 1586–1589. IEEE, 2010
2010
-
[48]
J. Ye, Y . Wang, Y . Huang, D. Chen, Q. Zhang, N. Moniz, T. Gao, W. Geyer, C. Huang, P.-Y . Chen, et al. Justice or prejudice? quantify- ing biases in llm-as-a-judge. arXiv preprint arXiv:2410.02736, 2024
2024 arXiv
-
[49]
Zamfirescu-Pereira, R
J. Zamfirescu-Pereira, R. Y . Wong, B. Hartmann, and Q. Yang. Why johnny can’t prompt: how non-ai experts try (and fail) to design llm prompts. In Proceedings of the 2023 CHI Conference on Human Fac- tors in Computing Systems , pp. 1–21, 2023
2023
-
[50]
Zhang, E
H. Zhang, E. Augilius, T. Honkela, J. Laaksonen, H. Gamper, and H. Alene. Analyzing emotional semantics of abstract art using low- level image features. In Advances in Intelligent Data Analysis X: 10th International Symposium, IDA 2011, Porto, Portugal, October 29-31,
2011
-
[51]
Zhang, M
Y . Zhang, M. Wang, P. Tiwari, Q. Li, B. Wang, and J. Qin. Dialoguellm: Context and emotion knowledge-tuned llama mod- els for emotion recognition in conversations. arXiv preprint arXiv:2310.11374, 2023
2023 arXiv
-
[52]
S. Zhao, G. Ding, Q. Huang, T.-S. Chua, B. Schuller, and K. Keutzer. Affective image content analysis: A comprehensive survey. 2018
2018
-
[53]
S. Zhao, Y . Gao, X. Jiang, H. Yao, T.-S. Chua, and X. Sun. Exploring principles-of-art features for image emotion recognition. In Proceed- ings of the 22nd ACM international conference on Multimedia , pp. 47–56, 2014
2014
-
[55]
M. Zhou, Y . R. Fung, L. Chen, C. Thomas, H. Ji, and S.-F. Chang. Enhanced chart understanding in vision and language task via cross-modal pre-training on plot table pairs. arXiv preprint arXiv:2305.18641, 2023
2023 arXiv
-
[2011]
Proceedings 10, pp. 413–423. Springer, 2011
2011
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.