REVIEW 3 major objections 6 minor 55 references
Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Large language model agents can rate visualizations like human users, but this paper shows their agreement with human conclusions is largely limited to experiments where expert hypotheses are confident.
desk verdict A useful empirical study of LLM agents as rating proxies, but the central expert-confidence correlation is a marginal post-hoc pattern—closer to a hypothesis than a demonstration. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a comparison table in which every conclusion from the six replications carries three labels: whether the original study's hypothesis (H) matched its conclusion (C), whether the agent feedback (A) matched the hypothesis, and whether it matched the conclusion. Onto these rows the paper grafts a confidence score per conclusion, obtained by asking five external visualization experts to form their own predictions for each experiment and rate their certainty on a three-point scale, with opposed expert hypotheses downgraded to low; the five scores are then averaged and rounded. The agent runs themselves use a fixed replication protocol—the original instructions and stimuli fed to a multimodal language model, with between-subject designs converted to within-subject batches because the model cannot hold a stable rating standard across separate sessions.
What would settle it
Take the six replicated experiments and have two independent coders re-label, from the raw transcripts, whether agent feedback matches each hypothesis and each conclusion, then recompute the match-rate difference between high- and low-confidence rows; if inter-coder agreement is low or the gap disappears, the paper's central correlation does not hold. A complementary test would pre-register expert hypotheses and confidence before a new set of visualization user studies and check whether agent-human agreement rises monotonically with pre-registered confidence.
Extended reading notes
Core claim
The paper claims that alignment between language-model agents and human raters in visualization experiments is predictable from expert confidence. In six replications, the match rate between agent and human conclusions was 3 out of 5 for high-confidence expert hypotheses, but 1 out of 12 elsewhere; the one high-confidence case that failed was a magnitude-judgement task where the agent consistently inverted the human interpretation by reading axis labels as text. The authors interpret this pattern as evidence that the agent behaves like an aggregator of historically trained knowledge: it can reproduce basic, well-established perceptual findings, but it does not possess the visual and cognitive machinery to track the subtle or contested effects that human experiments are usually run to discover.
Load-bearing premise
The correlation between expert confidence and agent alignment is only as strong as the hand-assigned labels in the comparison table and the experts' self-reported confidence; if those labels or confidence ratings are unreliable, the correlation may be an artifact.
Editorial extensions
If this is right
- When experts have low confidence in a hypothesis, agent ratings are unlikely to reproduce the human conclusion, so low-confidence conditions still require human participants.
- For basic, high-confidence findings—noise changes perceived fit, horizontal timelines are most readable—an agent can serve as a quick pre-check before a small pilot study.
- Prompt and input choices are consequential: explicitly comparing variables or removing images pushes the agent toward textual stereotypes, so agent-based evaluations must report and validate their exact prompt configuration.
- Web-retrieved knowledge injection can repair a specific reasoning error, such as aggregating icicle children, but the retrieved knowledge is uncertain in relevance and cannot yet be treated as a general fix.
- If an agent already matches a validated user study, changing only experimental parameters, such as a more extreme decentering condition, can provide a preliminary preview of the new outcome and speed up iterative design.
Reading between the lines
- Because the replication protocol converted between-subject designs into within-subject batches, the reported alignment may be an upper bound for how an agent would do when each condition is judged alone, as in most field use.
- The high-confidence alignments are consistent with the agent reproducing patterns it has seen during training; a decisive test would run the same protocol on experiments published after the model's knowledge cutoff.
- The five-expert confidence coding suggests a practical gate: pre-register expert confidence for a new visualization question and only trust agent simulation where confidence is high.
- The failure to model demographic profiles and aesthetic/texture judgments implies agent ratings describe an average reader, not a population; studies in which individual differences drive ratings are the least suitable for simulation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether GPT-4V-based agents can simulate human ratings in visualization experiments across three studies. Study I replicates a published CHI time-series user study and finds that agent ratings can mimic human-like reasoning but fail to capture user diversity. Study II replicates six human-subject experiments from five OSF-sourced papers, introduces external expert confidence coding, and reports that agent-human alignment positively correlates with experts' pre-experiment confidence. Study III tests input preprocessing and knowledge injection, finding that these techniques can improve or bias agent ratings. The paper concludes that agents can potentially simulate human ratings when guided by high-confidence expert hypotheses, while explicitly cautioning that agents cannot replace user studies.
Significance. If the central claim holds, the paper makes a useful contribution to visualization evaluation methodology by identifying when LLM agents can serve as low-cost proxies in formative rating studies. The work uses external published human studies as benchmarks, avoiding circular validation, and the authors provide open code, detailed replication protocols, and candid discussions of limitations such as data pollution and modified experimental procedures. The proposed fast-prototyping scenario in Sec. 6 is a concrete, falsifiable use case. However, the significance is bounded by the small dataset (five papers, six experiments) and the fragility of the headline confidence-alignment correlation, which is supported by a small non-independent sample and coding procedures that need better verification.
major comments (3)
- [Sec. 4.6, Table 1] The central claim that agent-human alignment positively correlates with expert confidence is supported by a single 2x2 split: 3/5 match in high-confidence rows versus 1/12 elsewhere. This split is not statistically significant (one-sided Fisher exact test yields p about 0.053, and a two-sided test is larger), and the rows are not independent because C1.1, C3.1, and C4.2 are duplicated for different rating types, while all rows come from only five papers. The abstract and conclusion state that the study 'demonstrates' this correlation; given the evidence, this is an overstatement. Please report an analysis that accounts for clustering by paper, or explicitly reframe the finding as exploratory and hypothesis-generating.
- [Sec. 4.6, Confidence Coding] The expert confidence judgments were elicited after the agent experiments, and the paper does not describe any blinding of experts to the published outcomes, so the label 'pre-experiment' is asserted rather than verified. Furthermore, the protocol of lowering confidence to 'low' for experts who proposed opposing hypotheses is a post hoc recoding that can inflate the apparent agreement between high confidence and agent-human alignment. In addition, the Y/N/P alignment labels in Table 1 come from a single coding pass without inter-rater reliability. Please provide the full elicitation protocol, blinding details, inter-rater reliability statistics, or at minimum a sensitivity analysis that treats the recoding alternative.
- [Sec. 7, Data pollution; Sec. 4.6] The data-pollution discussion in Sec. 7 acknowledges that GPT-4V's training data may inflate alignment for broad, well-known findings, but the same confound is not applied to the expert-confidence explanation. The three high-confidence matches (C1.1, C2.1, C4.1) are broad basic findings ('noise affects fit', 'horizontal timelines are preferred') that are likely overrepresented both in LLM training data and in expert prior knowledge, while the low-confidence rows tend to be detailed pattern claims. The paper does not distinguish the expert-confidence account from a training-data-familiarity account; this needs to be addressed in the revised discussion and ideally in the analysis.
minor comments (6)
- [Appendix 1.1, Prompt 3] The prompt text contains a typo: 'answe' should be 'answer'.
- [Appendix 1.2] The code block for the Imputation for Uncertainty execution is duplicated verbatim; one copy should be removed.
- [Appendix 2.5 and Sec. 5.2] The model name is written inconsistently as 'GPT-4V' and 'GPT-4v'; please standardize to 'GPT-4V' throughout.
- [Table 1] Several rows (C2.4, C4.3, C5.1, C5.2, C6.2) have blank entries under H-C, H-A, or C-A; please clarify whether these denote 'no hypothesis' or 'not applicable' in the table caption.
- [Abstract and Sec. 4.1] The abstract says the second study 'repeated six human-subject studies' while Sec. 4.1 states that 'we identified five suitable papers'; please clarify that six experiments were drawn from five papers.
- [Sec. 4.6, Implications] The sentence 'Our response to RQ2 suggests...' should read 'Our answer to RQ2 suggests...' for consistency with the RQ phrasing used elsewhere.
Circularity Check
No significant circularity; the comparison against published human studies is self-contained.
full rationale
The paper's core comparison is anchored to external, published human-subject studies (Adnan et al. 2016; Reimann et al. 2021; Sarma et al. 2023; Di Bartolomeo et al. 2020; He et al. 2024; Bradley et al. 2025) with open-source stimuli and human ratings. Agent ratings are generated from those fixed stimuli and prompts, and alignment is scored against the externally fixed human conclusions. No equation in the paper fits a parameter to the target outcome; the 3/5 versus 1/12 summary in Sec. 4.6 is a post-hoc descriptive comparison rather than a fitted prediction. The related-work self-citations (e.g., refs. 12, 41, 46, 56) are background context and are not load-bearing for the central claim. The weakest point is the confidence-coding procedure in Sec. 4.6: the authors consulted experts after the fact, recoded opposing expert hypotheses to low confidence, and the Table 1 Y/N/P alignment labels come from a single coding pass. These are legitimate threats to the independence and validity of the claimed correlation, and the paper's own Sec. 7 data-pollution discussion reinforces that the correlation may be confounded with task triviality or training-data overlap. However, confounding and weak statistical support are not constructional circularity: the predictor (expert confidence) and the outcome (agent-human conclusion alignment) are not the same variable by definition, and the claimed correlation is an empirical, testable observation rather than a result forced by the paper's definitions. There is no demonstrated step in which a prediction is equivalent to its input by construction, and no load-bearing self-citation chain. Therefore the appropriate circularity finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Conclusion-level comparison is a valid proxy for alignment; absolute rating baselines can be ignored if conclusions match.
- domain assumption The five selected OSF papers adequately represent visualization evaluation studies with subjective ratings.
- domain assumption Five external experts' post hoc confidence ratings are independent and valid measures of pre-experiment confidence.
- domain assumption GPT-4V is treated as a representative agent for LLM-based rating behavior.
Cite this review
Pith. "Pith review of Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study." pith.science (2026). https://pith.science/paper/KOSSBM6P
@misc{pith2026250506702,
author = {Pith},
title = {Pith review of: Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/KOSSBM6P}},
note = {Machine review of arXiv:2505.06702}
}
read the original abstract
Large language models encode knowledge in various domains and demonstrate the ability to understand visualizations. They may also capture visualization design knowledge and potentially help reduce the cost of formative studies. However, it remains a question whether large language models are capable of predicting human feedback on visualizations. To investigate this question, we conducted three studies to examine whether large model-based agents can simulate human ratings in visualization tasks. The first study, replicating a published study involving human subjects, shows agents are promising in conducting human-like reasoning and rating, and its result guides the subsequent experimental design. The second study repeated six human-subject studies reported in literature on subjective ratings, but replacing human participants with agents. Consulting with five human experts, this study demonstrates that the alignment of agent ratings with human ratings positively correlates with the confidence levels of the experts before the experiments. The third study tests commonly used techniques for enhancing agents, including preprocessing visual and textual inputs, and knowledge injection. The results reveal the issues of these techniques in robustness and potential induction of biases. The three studies indicate that language model-based agents can potentially simulate human ratings in visualization experiments, provided that they are guided by high-confidence hypotheses from expert evaluators. Additionally, we demonstrate the usage scenario of swiftly evaluating prototypes with agents. We discuss insights and future directions for evaluating and improving the alignment of agent ratings with human ratings. We note that simulation may only serve as complements and cannot replace user studies.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
M. Adnan, M. Just, and L. Baillie. Investigating Time Series Visual- isations to Improve the User Experience. InProceedings of the CHI Conference on Human Factors in Computing Systems, pp. 5444–5455. Association for Computing Machinery, New York, NY , USA, 2016. doi: 10.1145/2858036.2858300 1, 2, 6
arXiv 2016
-
[2]
K. Andrews. Evaluation Comes in Many Guises. InAVI Workshop on BEyond time and errors (BELIV) Position Paper, pp. 7–8, 2008. 2
work page 2008
-
[3]
L. Barkhuus and J. A. Rode. From Mice to Men - 24 Years of Eval- uation in CHI. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems. Association for Computing Machin- ery, New York, NY , USA, 2007. doi: 10.1145/1240624.2180963 2
-
[4]
D. Bradley, G. Strain, C. Jay, and A. J. Stewart. Magnitude Judgements Are Influenced by the Relative Positions of Data Points Within Axis Limits.IEEE Transactions on Visualization and Com- puter Graphics, 31(2):1414–1421, 2025. doi: 10.1109/TVCG.2024. 3364069 3, 6, 13
-
[5]
M. Chen and H. Jäenicke. An Information-theoretic Framework for Visualization.IEEE Transactions on Visualization and Computer Graphics, 16(6):1206–1215, 2010. doi: 10.1109/TVCG.2010.132 2
-
[7]
J. W. Creswell and J. D. Creswell.Research design: Qualitative, quantitative, and mixed methods approaches. Sage publications, 2017. 2
work page 2017
-
[9]
Z. Cui, L. Chen, Y . Wang, D. Haehn, Y . Wang, and H. Pfister. Gener- alization of CNNs on Relational Reasoning with Bar Charts.IEEE Transactions on Visualization and Computer Graphics, pp. 1–15,
-
[10]
S. Di Bartolomeo, A. Pandey, A. Leventidis, D. Saffo, U. H. Syeda, E. Carstensdottir, M. Seif El-Nasr, M. A. Borkin, and C. Dunne. Eval- uating the Effect of Timeline Shape on Visualization Task Perfor- mance. InProceedings of the CHI Conference on Human Factors in Computing Systems, pp. 1–12. Association for Computing Machin- ery, New York, NY , USA, 202...
arXiv 2020
Show all 55 references
- [11]
-
[12]
L. Gao, J. Lu, Z. Shao, Z. Lin, S. Yue, C. Ieong, Y . Sun, R. J. Za- uner, Z. Wei, and S. Chen. Fine-Tuned Large Language Model for Visualization System: A Study on Self-Regulated Learning in Edu- cation.IEEE Transactions on Visualization and Computer Graphics, 31(1):514–524, ...
2025
-
[13]
Y . Gao, Y . Xiong, X. Gao, K. Jia, J. Pan, Y . Bi, Y . Dai, J. Sun, Q. Guo, M. Wang, and H. Wang. Retrieval-Augmented Generation for Large Language Models: A Survey, 2024. 2, 7
2024
-
[14]
L. W. Ge, Y . Cui, and M. Kay. CALVI: Critical Thinking Assess- ment for Literacy in Visualizations. InProceedings of the CHI Con- ference on Human Factors in Computing Systems, pp. 1–18. Associa- tion for Computing Machinery, New York, NY , USA, 2023. doi: 10. 1145/3544548.3581406 2
2023
-
[15]
Grunde-McLaughlin, M
M. Grunde-McLaughlin, M. S. Lam, R. Krishna, D. S. Weld, and J. Heer. Designing LLM Chains by Adapting Techniques from Crowd- sourcing Workflows, 2023. 2, 9
2023
-
[16]
H. Guo, S. R. Gomez, C. Ziemkiewicz, and D. H. Laidlaw. A Case Study Using Visualization Interaction Logs and Insight Metrics to Un- derstand How Analysts Arrive at Insights.IEEE Transactions on Visu- alization and Computer Graphics, 22(1):51–60, 2016. doi: 10.1109/ TVCG.2015....
2016
-
[17]
Haehn, J
D. Haehn, J. Tompkin, and H. Pfister. Evaluating ‘graphical percep- tion’ with cnns.IEEE Transactions on Visualization and Computer Graphics, 25(1):641–650, 2019. doi: 10.1109/TVCG.2018.2865138 2
2019
-
[19]
T. He, Y . Zhong, P. Isenberg, and T. Isenberg. Design Characterization for Black-and-White Textures in Visualization.IEEE Transactions on Visualization and Computer Graphics, 30(1):1019–1029, 2024. doi: 10.1109/TVCG.2023.3326941 2, 3, 6, 13
2024
-
[20]
Hullman, X
J. Hullman, X. Qiao, M. Correll, A. Kale, and M. Kay. In Pursuit of Er- ror: A Survey of Uncertainty Visualization Evaluation.IEEE Transac- tions on Visualization and Computer Graphics, 25(1):903–913, 2019. doi: 10.1109/TVCG.2018.2864889 2
2019
-
[21]
Hämäläinen, M
P. Hämäläinen, M. Tavast, and A. Kunnari. Evaluating Large Lan- guage Models in Generating Synthetic HCI Research Data: a Case Study. InProceedings of the CHI Conference on Human Factors in Computing Systems, pp. 1–19. Association for Computing Machinery, New York, NY , USA, 2...
2023
-
[22]
Isenberg, F
P. Isenberg, F. Heimerl, S. Koch, T. Isenberg, P. Xu, C. D. Stolper, M. Sedlmair, J. Chen, T. Möller, and J. Stasko. Vispubdata.org: A Metadata Collection About IEEE Visualization (VIS) Publica- tions.IEEE Transactions on Visualization and Computer Graphics, 23(9):2199–2206, 2...
2017
-
[23]
Isenberg, P
T. Isenberg, P. Isenberg, J. Chen, M. Sedlmair, and T. Möller. A Sys- tematic Review on the Practice of Evaluating Visualization.IEEE Transactions on Visualization and Computer Graphics, 19(12):2818– 2827, 2013. doi: 10.1109/TVCG.2013.126 2
2013 doi
-
[24]
Johansson and C
J. Johansson and C. Forsell. Evaluation of Parallel Coordinates: Overview, Categorization and Guidelines for Future Research.IEEE Transactions on Visualization and Computer Graphics, 22(1):579– 588, 2016. doi: 10.1109/TVCG.2015.2466992 2
2016
-
[25]
Kahng and D
M. Kahng and D. H. P. Chau. How Does Visualization Help Peo- ple Learn Deep Learning? Evaluating GAN Lab with Observational Study and Log Analysis. InIEEE Visualization Conference (VIS), pp. 266–270. IEEE, 2020. doi: 10.1109/VIS47514.2020.00060 2
2020
-
[26]
Y .-a. Kang, C. Gorg, and J. Stasko. Evaluating visual analytics sys- tems for investigative analysis: Deriving design principles from a case study. InIEEE Symposium on Visual Analytics Science and Technol- ogy, pp. 139–146. IEEE, 2009. doi: 10.1109/V AST.2009.5333878 2
2009
-
[27]
Y .-a. Kang, C. Görg, and J. Stasko. How Can Visual Analytics As- sist Investigative Analysis? Design Implications from an Evalua- tion.IEEE Transactions on Visualization and Computer Graphics, 17(5):570–583, 2011. doi: 10.1109/TVCG.2010.84 2
2011 doi
-
[28]
H.-K. Ko, H. Jeon, G. Park, D. H. Kim, N. W. Kim, J. Kim, and J. Seo. Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models, 2024. 2
2024
-
[29]
H. Lam, E. Bertini, P. Isenberg, C. Plaisant, and S. Carpendale. Em- pirical Studies in Information Visualization: Seven Scenarios.IEEE Transactions on Visualization and Computer Graphics, 18(9):1520– 1536, 2012. doi: 10.1109/TVCG.2011.279 2
2012 doi
-
[30]
G. Li, X. Wang, G. Aodeng, S. Zheng, Y . Zhang, C. Ou, S. Wang, and C. H. Liu. Visualization Generation with Large Language Models: An Evaluation. doi: 10.48550/arXiv.2401.11255 2
-
[31]
Z. Li, H. Miao, V . Pascucci, and S. Liu. Visualization Literacy of Multimodal Large Language Models: A Comparative Study.arXiv preprint arXiv:2407.10996, 2024. 1, 2, 8
2024 arXiv
-
[32]
H. Liu, C. Li, Y . Li, and Y . J. Lee. Improved Baselines with Visual Instruction Tuning, 2023. 5
2023
-
[33]
S. Liu, H. Miao, Z. Li, M. Olson, V . Pascucci, and P.-T. Bremer.AVA: Towards Autonomous Visualization Agents through Visual Perception- Driven Decision-Making. 2023. 1, 2
2023
-
[35]
J. Lála, O. O’Donoghue, A. Shtedritski, S. Cox, S. G. Rodriques, and A. D. White. PaperQA: Retrieval-Augmented Generative Agent for Scientific Research, 2023. 2
2023
-
[36]
L. E. Matzen, M. J. Haass, K. M. Divis, Z. Wang, and A. T. Wil- son. Data Visualization Saliency Model: A Tool for Evaluating Ab- stract Data Visualizations.IEEE Transactions on Visualization and Computer Graphics, 24(1):563–573, 2018. doi: 10.1109/TVCG.2017 .2743939 2
2018 doi
-
[37]
GPT-4 Technical Report, 2023
OpenAI. GPT-4 Technical Report, 2023. 1, 2
2023
-
[38]
Open Science Framework - OSF, 2024
OSF. Open Science Framework - OSF, 2024. 1
2024
-
[39]
Reimann, C
D. Reimann, C. Blech, N. Ram, and R. Gaschler. Visual Model Fit Estimation in Scatterplots: Influence of Amount and Decentering of Noise.IEEE Transactions on Visualization and Computer Graphics, 27(9):3834–3838, 2021. doi: 10.1109/TVCG.2021.3051853 3, 4, 6, 7
2021
-
[40]
Sarma, S
A. Sarma, S. Guo, J. Hoffswell, R. Rossi, F. Du, E. Koh, and M. Kay. Evaluating the Use of Uncertainty Visualisations for Imputations of Data Missing At Random in Scatterplots.IEEE Transactions on Vi- sualization and Computer Graphics, 29(1):602–612, 2023. doi: 10. 1109/TVCG.2...
2023
-
[41]
Z. Shao, L. Shen, H. Li, Y . Shan, H. Qu, Y . Wang, and S. Chen. Nar- rative Player: Reviving Data Narratives with Visuals.IEEE Transac- tions on Visualization and Computer Graphics, pp. 1–15, 2025. doi: 10.1109/TVCG.2025.3530512 2
2025
-
[42]
L. Shen, Y . Zhang, H. Zhang, and Y . Wang. Data Player: Automatic Generation of Data Videos with Narration-Animation Interplay.IEEE Transactions on Visualization and Computer Graphics, 30(1):109– 119, 2024. doi: 10.1109/TVCG.2023.3327197 2
2024
-
[43]
D. A. Szafir, R. Borgo, M. Chen, D. J. Edwards, B. Fisher, and L. Padilla.Visualization Psychology. Springer Nature, 2023. 8
2023
-
[44]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. LLaMA: Open and Efficient Foundation Language Models, 2023. 2
2023
-
[45]
Verma, K
A. Verma, K. Mukherjee, C. Potts, E. Kreiss, and J. E. Fan. Evaluating human and machine understanding of data visualizations. InProceed- ings of the Annual Meeting of the Cognitive Science Society, vol. 46,
-
[46]
F. Wang, B. Wang, X. Shu, Z. Liu, Z. Shao, C. Liu, and S. Chen. ChartInsighter: An Approach for Mitigating Hallucination in Time- series Chart Summary Generation with A Benchmark Dataset.IEEE Transactions on Visualization and Computer Graphics, pp. 1–11,
-
[48]
H. W. Wang, J. Hoffswell, S. M. T. Thane, V . S. Bursztyn, and C. X. Bearfield. How Aligned are Human Chart Takeaways and LLM Pre- dictions? A Case Study on Bar Charts with Varying Layouts.IEEE Transactions on Visualization and Computer Graphics, 31(1):536– 546, 2025. doi: 10....
2025
-
[49]
L. Wang, S. Zhang, Y . Wang, E.-P. Lim, and Y . Wang. LLM4Vis: Explainable Visualization Recommendation using ChatGPT, 2023. 1, 2
2023
-
[50]
J. Xia, Y . Zhang, J. Song, Y . Chen, Y . Wang, and S. Liu. Revisiting Di- mensionality Reduction Techniques for Visual Cluster Analysis: An Empirical Study.IEEE Transactions on Visualization and Computer Graphics, 28(1):529–539, 2021. doi: 10.1109/TVCG.2021.3114694 2
2021
-
[51]
Xu and E
Z. Xu and E. Wall. Exploring the Capability of LLMs in Performing Low-Level Visual Analytic Tasks on SVG Data Visualizations.arXiv preprint arXiv:2404.19097, 2024. 2
2024 arXiv
-
[52]
F. Yang, Y . Ma, L. Harrison, J. Tompkin, and D. H. Laidlaw. How Can Deep Neural Networks Aid Visualization Perception Research? Three Studies on Correlation Judgments in Scatterplots. InProceedings of the CHI Conference on Human Factors in Computing Systems, pp. 1–17. Associa...
-
[53]
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao. Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V,
-
[54]
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision),
-
[55]
A. Zeng, M. Liu, R. Lu, B. Wang, X. Liu, Y . Dong, and J. Tang. AgentTuning: Enabling Generalized Agent Abilities for LLMs, 2023. 2
2023
-
[56]
Y . Zhao, Y . Zhang, Y . Zhang, X. Zhao, J. Wang, Z. Shao, C. Turkay, and S. Chen. LEV A: Using Large Language Models to Enhance Visual Analytics.IEEE Transactions on Visualization and Computer Graph- ics, 31(3):1830–1847, 2025. doi: 10.1109/TVCG.2024.3368060 2
2025
-
[57]
ROLE" refers to the desired background role for the large model,
B. Zheng, B. Gou, J. Kil, H. Sun, and Y . Su. GPT-4V(ision) is a Generalist Web Agent, if Grounded, 2024. 2 APPENDIX In the first section, we present the prompt and codes for two ex- periments (Fit Estimation and Imputation for Uncertainty) as refer- ences. In the second secti...
2024
-
[2023]
doi: 10.1145/3544548.3581111 1, 2
-
[2024]
doi: 10.1109/TVCG.2024.3463800 2
2024
-
[2025]
doi: 10.1109/TVCG.2025.3567122 2
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.