REVIEW 4 major objections 4 minor 70 references
AnnoBench: A Benchmark for Visualization Annotation Generation
T0 review · 4 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read AnnoBench makes chart annotation quality measurable and shows current AI models fail most on raster inputs.
desk verdict AnnoBench is a real, usable benchmark for chart annotation, and the paper is mostly honest about where its own VLM-judge evidence is weak; the abstract overstates that part, but the artifact and experiments deserve a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the five-dimensional evaluation rubric—precise, relevant, visible, preservant, and consistent—instantiated in AnnoBench. The rubric turns qualitative annotation design principles into verifiable per-instance questions, and the benchmark pairs each chart with multiple representations (raster, SVG vector, Vega grammar, D3 code), five description levels (from no description to domain-specific context), and two task prompt levels (intent and execution). A VLM-as-a-judge pipeline scores outputs automatically, and the paper uses one-factor-at-a-time experiments to attribute failures to specific dimensions.
What would settle it
Score a new held-out sample of raster annotation outputs with at least two human raters and all three VLM judges used in the paper. If the best VLM judge's Spearman correlation with human scores fails to exceed 0.4 on the precise and preservant dimensions—as the paper's Table 2 already suggests for raster—then the claim that VLM-as-a-judge is aligned with manual human assessment does not hold for one of the four core representations.
Extended reading notes
Core claim
The paper's central claim is that the three necessary conditions for a correct annotation—target correctness, task accuracy, and spatial feasibility—plus two holistic qualities, chart preservation and stylistic coherence, can be defined as per-instance criteria and used to evaluate any annotation system. On that basis AnnoBench provides a dataset and an end-to-end pipeline that renders model outputs, scores them automatically with vision-language models as judges, and supports human evaluation. Using it, the paper reports systematic failure modes in current models: annotations that point at the wrong mark while using plausible text, charts whose data marks shift to make room for labels, and
Load-bearing premise
The benchmark's automated execution assumes that vision-language-model judges can reliably stand in for human judgment of annotation quality; the paper's own data show this assumption is weakest for raster outputs, where human–VLM agreement is weak (ρ≈0.3).
Editorial extensions
If this is right
- Annotation automation should prioritize structured, editable chart representations like code over raster images, which currently suffer from chart-preservation failures.
- Providing explicit execution-level instructions materially improves annotation targeting, placement, and chart preservation compared with intent-level prompts.
- VLM-as-a-judge is a viable scaling tool only for structured outputs; raster evaluation still needs human review or substantially better judges.
- Model rankings appear stable across representations, prompt types, and description levels, suggesting annotation ability is a general model competence rather than a configuration-specific skill.
- The benchmark provides a reusable instrument for comparing future annotation tools, grammars, and generation pipelines under consistent evaluation conditions.
Reading between the lines
- If the benchmark's dimensions are sound, the same rubric could be extended to chart editing and chart explanation tasks, which share the need to verify that an output preserves the original chart's data encoding.
- The VLM judge's tendency to over-credit visually plausible raster outputs suggests a testable fix: inject deliberate data corruption into otherwise correct annotations and measure whether judges downgrade them; current evidence says they would not.
- The finding that models avoid enclosure annotations under intent prompts implies that progress in annotation automation depends less on stronger chart reading and more on strategic design inference—an area that could be probed with targeted intent-level tasks across diverse annotation types.
- Because model outputs are deliberately excluded from the dataset, AnnoBench remains useful as models improve; one could re-run the same configurations yearly to track whether raster annotation weaknesses close over time.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AnnoBench, a benchmark for evaluating the generation of chart annotations by LLMs/VLMs. The benchmark spans four chart representations (raster, vector, grammar, code), five semantic description levels (plus a no-description baseline), and two task-prompt levels (intent and execution). A five-dimensional rubric (precise, visible, relevant, preservant, consistent) is proposed, and a browser-based pipeline supports human evaluation and VLM-as-a-judge scoring. The authors report four one-factor-at-a-time experiments varying representation, prompt specificity, chart description, and model selection, and claim the benchmark reveals systematic failure modes. The central claim is that VLM-as-a-judge is 'aligned with manual human assessment,' enabling scalable automated execution of the benchmark.
Significance. If the headline claims held, AnnoBench would be a valuable first instrument for a task that has lacked a testable benchmark: chart annotation. The five-dimensional rubric, the multi-representation dataset (342 charts, 650+ tasks), and the open browser/pipeline are concrete, reusable contributions. The paper also reports human inter-rater reliability and human–VLM agreement, which is more transparent than many benchmark papers. However, the paper's own results weaken the central evaluation claim, and the automated scoring mechanism is load-bearing for the benchmark's usefulness at scale. The core idea is defensible, but the presentation and evidence need substantial revision before the claims can be accepted.
major comments (4)
- [Abstract, §6.1, Table 2, §7] The abstract states the benchmark is 'executed via VLM-as-a-judge, using models aligned with manual human assessment,' but Table 2 shows weak human–VLM agreement on raster outputs (ρ = 0.32 precise, 0.22 visible, 0.48 relevant, 0.39 preservant, 0.47 consistent for the best-performing VLM judge). Section 6.1 itself reports 'much weaker agreement' for raster, and §7 says VLM judging 'should not be treated as a substitute for human assessment' and 'tends to overestimate quality on raster outputs.' Since raster is one of the four core representations and chart preservation is one of the claimed failure modes, a judge with ρ≈0.39 on preservant cannot reliably detect the data-distortion failures the benchmark claims to reveal. The abstract and Figure 1 should be revised to state the conditional reliability of the automated judge, and Table 2 should report all three judges plus confidence inter
- [§6] The three VLM judges (GPT-5.2, Claude Sonnet 4, Gemini 2.5 Flash) are from the same model families as the generation models under test. This creates an in-family evaluation loop: the judges may systematically favor outputs from models of their own family, which would bias Experiment 4's model ranking and, more generally, any fully automated use of AnnoBench. The paper should report agreement stratified by judge-family versus generation-family, or at least show that the model rankings do not change when judges are restricted to a family different from the generation model. Without such analysis, the automated execution claim is incomplete.
- [§6.2–§6.4] Human–VLM agreement is only quantified in Experiment 1. Experiments 2–4 also use VLM-as-a-judge to draw conclusions (e.g., 'VLM-Judge tends to score slightly lower on tasks with execution-level prompts'), but no correlation or agreement statistics are reported for those experiments. Given that Experiment 2 already shows a qualitative divergence between human and VLM evaluators, the reader cannot assess how much of the later results reflect the judge rather than the annotation quality. Report per-experiment, per-rubric human–VLM agreement, or explicitly limit VLM-based conclusions to representations where agreement is established.
- [§6, §6.1] The sample-size reporting is internally inconsistent. Section 6 states 'n=15 tasks per experiment condition' and then 'we use AnnoBench’s VLM-as-a-Judge feature to score an additional sample of equal size (n=30).' If the additional sample is of equal size, it should be n=15, making the total n=30; the parenthetical 'n=30' is not consistent with the preceding sentence. This ambiguity matters because the confidence intervals and effect sizes in Sections 6.1–6.4 depend on the actual sample size. The authors should correct the description and, where possible, report the per-condition cell sizes explicitly.
minor comments (4)
- [Table 1] The chart-type counts do not sum to the stated dataset sizes. For the Vega Gallery, listed counts sum to 101 while # Charts is 92; for Vega-Lite, listed counts sum to 191 while # Charts is 192. Please clarify whether chart types are overlapping categories or correct the numbers.
- [Fig. 4] The figure contains repeated 'Vue' labels in the experiment charts. This appears to be a typo (likely for 'Vega' or a generic axis label) and should be corrected for clarity.
- [Fig. 2, §6.3] There are typos: 'erronous' in the level-3 description example in Fig. 2, and 'visualizaiton' in §6.3. Minor language edits throughout would improve readability.
- [References [45] and [46]] References [45] and [46] appear to be the same paper (same title, venue, and authors). If they are the same work, the duplicate reference number should be removed and the in-text citations should be updated.
Circularity Check
No significant circularity: AnnoBench's results are empirical measurements, not derivations from their inputs.
full rationale
AnnoBench is an empirical benchmark/measurement paper rather than a derivation from first principles. The central outputs—representation-condition score distributions, prompt-level differences, caption-level effects, and model rankings—are measured on generated outputs and scored by human raters (with high IRR, kappa > 0.93) plus VLM judges; they are not algebraically or statistically forced by the benchmark definitions. The self-citations to Rahman et al. [44-46] supply the annotation taxonomy and the prior 'open problem' framing, but the benchmark's load-bearing measurements do not reduce to those citations: the task distribution is drawn from the taxonomy, and the rubric is operationalized from prior annotation systems, yet the experimental results are independent observations reported in raw form (e.g., Table 2, Fig. 4). The in-family use of GPT/Claude/Gemini as both generators and judges is an evaluation-validity concern, not a circular derivation, because VLM scores are explicitly compared against human scores and the paper reports weak raster agreement (rho about 0.3) and states in Limitations that VLM methods 'should not be treated as a substitute for human assessment.' The abstract's 'aligned with manual human assessment' claim is internally inconsistent with Table 2 for raster outputs, but that is a correctness/consistency weakness, not a construction-level circularity. No equation or fitted parameter is renamed as a prediction; no uniqueness theorem is imported from the authors' prior work; no prediction reduces by definition to its inputs. Therefore no significant circularity.
Assumptions & free parameters
free parameters (4)
- retry_limit =
3
- token_threshold =
16384
- human_sample_size =
15
- vlm_sample_size =
30
assumptions (5)
- domain assumption VLM-as-a-judge output scores are a reliable proxy for human annotation-quality judgments.
- domain assumption The five-dimension rubric (precise, relevant, visible, preservant, consistent) fully operationalizes annotation correctness.
- domain assumption Lundgard et al.'s four-level semantic hierarchy, extended with level 0, is appropriate for chart descriptions used in annotation.
- domain assumption Reconstructed/synthesized data for professional charts preserves the perceptual and structural properties required for annotation tasks.
- domain assumption The annotation taxonomy of Rahman et al. [45] is a valid and complete basis for task generation.
Cite this review
Pith. "Pith review of AnnoBench: A Benchmark for Visualization Annotation Generation." pith.science (2026). https://pith.science/paper/5EWPSXPT
@misc{pith2026260725911,
author = {Pith},
title = {Pith review of: AnnoBench: A Benchmark for Visualization Annotation Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/5EWPSXPT}},
note = {Machine review of arXiv:2607.25911}
}
read the original abstract
Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints. Failure to meet any of these conditions severely undermines the utility of an annotation, rendering it challenging to read, inaccurate, or visually discordant. Despite a growing body of annotation tools and automations, no existing benchmark or evaluation framework tests whether these conditions are met because of their scope and annotation not being the focus of their studies. We introduce AnnoBench, a benchmark for visualization annotation that materializes the inherent challenges of this domain in a structured and testable manner. AnnoBench pairs visualizations from professional data journalism and visualization galleries with annotation tasks, spanning four representation formats, five chart description conditions, and two prompt specification levels. The benchmark is executed via VLM-as-a-judge, using models aligned with manual human assessment. We evaluate the benchmark via four one-factor-at-a-time experiments, exploring the effects of input representation, semantic context, and prompt specificity, and model selection on annotation quality. This work provides a foundation for advancing annotation automation, tooling, and visualization-generation pipelines.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Claude api pricing: Model costs, prompt caching, and batch processing
Anthropic. Claude api pricing: Model costs, prompt caching, and batch processing. https://platform.claude.com, 2026. Accessed: 2026- 03-30. 6
2026
-
[2]
M. Bancilhon, A. Wright, S. Ha, R. J. Crouser, and A. Ottley. Why combining text and visualization could improve bayesian reasoning: A cognitive load perspective. InACM CHI, pp. 1–15, 2023. doi: 10.1145/ 3544548.3581218 2
arXiv 2023
-
[3]
S. Bateman, R. L. Mandryk, C. Gutwin, A. Genest, D. McDine, and C. Brooks. Useful junk? the effects of visual embellishment on compre- hension and memorability of charts. InACM CHI, pp. 2573–2582, 2010. doi: 10.1145/1753326.1753716 2
arXiv 2010
-
[5]
M. Borkin, A. A. V o, Z. Bylinskii, P. Isola, S. Sunkavalli, A. Oliva, and H. Pfister. What makes a visualization memorable?IEEE Transactions on Visualization and Computer Graphics, 19(12):2306–2315, 2013. doi: 10.1109/TVCG.2013.234 2
-
[6]
M. A. Borkin, Z. Bylinskii, N. W. Kim, C. M. Bainbridge, C. S. Yeh, D. Borkin, H. Pfister, and A. Oliva. Beyond memorability: Visualization recognition and recall.IEEE Transactions on Visualization and Computer Graphics, 22(1):519–528, 2015. doi: 10.1109/TVCG.2015.2467732 2
arXiv 2015
-
[7]
C. Bryan, A. Mishra, H. Shidara, and K.-L. Ma. Analyzing gaze behavior for text-embellished narrative visualizations under different task scenarios. Visual Informatics, 4(3):41–50, 2020. doi: 10.1016/j.visinf.2020.08.001 2
-
[8]
N. Chen, Y . Zhang, J. Xu, K. Ren, and Y . Yang. Viseval: A benchmark for data visualization in the era of large language models.IEEE Transactions on Visualization and Computer Graphics, 2024. doi: 10.1109/TVCG.2024 .3456320 2
-
[9]
Y . Chen, S. Barlowe, and J. Yang. Click2annotate: Automated insight externalization with rich semantics. InIEEE VAST, pp. 155–162, 2010. doi: 10.1109/V AST.2010.5652885 1
arXiv 2010
Show all 70 references
-
[10]
Y . Chen, Y . Wu, S. Shen, Y . Xie, L. Shen, H. Xiong, and Y . Luo. Chartmark: A structured grammar for chart annotation.arXiv preprint arXiv:2507.21810, 2025. doi: 10.1109/VIS60296.2025.00068 1, 2, 3
2025 arXiv
-
[11]
R. Chun. Giving guidance to graphs: Evaluating annotations of data visualizations for the news.Visual Communication Quarterly, 27(2):84– 97, 2020. doi: 10.1080/15551393.2020.1749842 2
2020
-
[12]
Cutler, J
Z. Cutler, J. Wilburn, H. Shrestha, Y . Ding, B. Bollen, K. A. Nadib, T. He, A. McNutt, L. Harrison, and A. Lex. Revisit 2: A full experiment life cycle user study framework.IEEE Transactions on Visualization and Computer Graphics, 32(1):13–23, 2026. doi: 10.1109/TVCG.2025.3633896 6
2026
-
[13]
A. Fan, F. Lei, M. Mancenido, A. M. Maceachren, and R. Maciejewski. Understanding reader takeaways in thematic maps under varying text, detail, and spatial autocorrelation. InACM CHI, pp. 1–17, 2024. doi: 10. 1145/3613904.3642132 1, 2
2024
-
[14]
A. Fan, Y . Ma, M. Mancenido, and R. Maciejewski. Annotating line charts for addressing deception. InACM CHI, 12 pages, 2022. doi: 10. 1145/3491102.3502138 1, 2
2022
-
[15]
T. Gao, J. R. Hullman, E. Adar, B. Hecht, and N. Diakopoulos. Newsviews: an automated pipeline for creating custom geovisualizations for news. In ACM CHI, pp. 3005–3014, 2014. doi: 10.1145/2556288.2557228 1
2014
-
[16]
L. W. Ge, Y . Cui, and M. Kay. Calvi: Critical thinking assessment for literacy in visualizations. InACM CHI, pp. 1–18, 2023. doi: 10.1145/ 3544548.3581406 2
2023
-
[17]
G. Guo, J. Stasko, and A. Endert. What we augment when we augment visualizations: A design elicitation study of how we visually express data relationships. InACM Advanced Visual Interfaces, pp. 1–5, 2024. doi: 10. 1145/3656650.3656666 2
2024
-
[18]
J. Hong, C. Seto, A. Fan, and R. Maciejewski. Do llms have visualization literacy? an evaluation on modified visualizations to test generalization in data interpretation.IEEE Transactions on Visualization and Computer Graphics, 31(10):7004–7018, 2025. doi: 10.1109/TVCG.2025.35...
2025
-
[19]
Huang, W
L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Information Systems, 43(2):1–55, 2025. doi: 10.1145/3703155 4
2025 doi
-
[20]
Hullman, N
J. Hullman, N. Diakopoulos, and E. Adar. Contextifier: automatic genera- tion of annotated stock visualizations. InACM CHI, pp. 2707–2716, 2013. doi: 10.1145/2470654.2481374 1, 2, 3
2013
-
[21]
Kafle, B
K. Kafle, B. Price, S. Cohen, and C. Kanan. Dvqa: Understanding data visualizations via question answering. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5648–5656, 2018. doi: 10. 1109/CVPR.2018.00592 2
2018
-
[22]
Kantharaj, R
S. Kantharaj, R. Ramesh, A. Shetty, M. M. Chandrasekaran, and A. Modi. Chart-to-text: A large-scale benchmark for chart summarization. InFind- ings of the Association for Computational Linguistics (ACL), pp. 2694–
-
[23]
D. H. Kim, E. Hoque, and M. Agrawala. Answering questions about charts and generating visual explanations. InACM CHI, pp. 1–13, 2020. doi: 10. 1145/3313831.3376467 5
2020
-
[24]
Kittivorawong, D
C. Kittivorawong, D. Moritz, and J. Heer. Fastlabels: Fast and flexible overlap detection for chart labeling with occupancy bitmaps.IEEE Visu- alization Conference (VIS), 2020. doi: 10.1109/VIS47514.2020.00027 1
2020
-
[25]
H.-K. Kong, Z. Liu, and K. Karahalios. Internal and external visual cue preferences for visualizations in presentations.Computer Graphics Forum, 36(3):515–525, 2017. doi: 10.1111/cgf.13207 2
2017 doi
-
[26]
H.-K. Kong, W. Zhu, Z. Liu, and K. Karahalios. Understanding visual cues in visualizations accompanied by audio narrations. InACM CHI, pp. 1–13, 2019. doi: 10.1145/3290605.3300280 2
2019
-
[27]
Kong and M
N. Kong and M. Agrawala. Graphical overlays: Using layered elements to aid chart reading.IEEE Transactions on Visualization and Computer Graphics, 18(12):2631–2638, 2012. doi: 10.1109/TVCG.2012.229 1, 2, 3, 4
2012 doi
-
[28]
Lee, S.-H
S. Lee, S.-H. Kim, and B. C. Kwon. Vlat: Development of a visualization literacy assessment test.IEEE Transactions on Visualization and Computer Graphics, 23(1):551–560, 2016. doi: 10.1109/TVCG.2016.2598920 1, 2
2016
-
[29]
H. Li, Y . Wang, and H. Qu. Where are we so far? understanding data storytelling tools from the perspective of human-ai collaboration. InACM CHI, pp. 1–19, 2024. doi: 10.1145/3613904.3642726 2
2024
-
[30]
Lisnic, C
M. Lisnic, C. Polychronis, A. Lex, and M. Kogan. Misleading beyond visual tricks: How people actually lie with charts. InACM CHI, pp. 1–21,
-
[31]
Lundgard and A
A. Lundgard and A. Satyanarayan. Accessible visualization via natural language descriptions: A four-level model of semantic content.IEEE Transactions on Visualization and Computer Graphics, 28(1):1073–1083,
- [32]
-
[33]
Masry, X
A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguistics (ACL), pp. 2263–2279, 2022. doi: 10.18653/v1/2022.findings-acl.177 1, 2
2022 doi
-
[34]
A. M. McNutt and R. Chugh. Integrated visualization editing via param- eterized declarative templates. InACM CHI, pp. 1–14, 2021. doi: 10. 1145/3411764.3445356 5
2021
-
[35]
Methani, P
N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar. Plotqa: Reasoning over scientific plots. InIEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1527–1536, 2020. doi: 10.1109/W ACV45572.2020. 9093523 1, 2
2020
-
[36]
Mukherjee, D
K. Mukherjee, D. Ren, D. Moritz, and Y . Assogba. Encqa: Benchmarking vision-language models on visual encodings for charts.IEEE Transactions on Visualization and Computer Graphics, 2025. Early Access; also on arXiv:2508.04650. doi: 10.1109/TVCG.2025.3634249 2
2025
-
[37]
J. Neyman. On the two different aspects of the representative method: the method of stratified sampling and the method of purposive selection. In Breakthroughs in statistics: Methodology and distribution, pp. 123–150. Springer, 1992. doi: 10.1007/978-1-4612-4380-9_12 5, 6
1992 doi
-
[38]
D3 gallery, 2024
Observable. D3 gallery, 2024. 1
2024
-
[39]
Ottley, A
A. Ottley, A. Kaszowska, R. J. Crouser, and E. M. Peck. The curious case of combining text and visualization. InEuroVis (Short Papers). Eurographics, 2019. doi: 10.2312/evs.20191181 1, 2
2019 doi
-
[40]
Pandey and A
S. Pandey and A. Ottley. Mini-vlat: A short and effective measure of visualization literacy.Computer Graphics Forum, 42(3):1–11, 2023. doi: 10.1111/cgf.14809 2
2023 doi
-
[41]
Pandey and A
S. Pandey and A. Ottley. Benchmarking visual language models on stan- dardized visualization literacy tests. InComputer Graphics Forum, p. e70137, 2025. doi: 10.1111/cgf.70137 2
2025 doi
-
[42]
P. Parsons. Understanding data visualization design practice.IEEE Trans- actions on Visualization and Computer Graphics, 28(1):665–675, 2021. doi: 10.1109/TVCG.2021.3114959 5
2021
-
[43]
Rahman, M
M. Rahman, M. T. R. Laskar, S. Joty, and E. Hoque. Text2vis: A challeng- ing and diverse benchmark for generating multimodal visualizations from text. InEmpirical Methods in Natural Language Processing (EMNLP), pp. 31837–31862, 2025. doi: 10.18653/v1/2025.emnlp-main.1622 2
2025 doi
-
[44]
M. D. Rahman, B. Doppalapudi, G. J. Quadri, and P. Rosen. A survey on annotations in information visualization: Empirical studies, applica- tions and challenges.IEEE Transactions on Visualization and Computer Graphics, 2025. doi: 10.1109/TVCG.2025.3600957 1, 2, 3
2025
-
[46]
M. D. Rahman, G. J. Quadri, B. Doppalapudi, D. A. Szafir, and P. Rosen. A qualitative analysis of common practices in annotations: A taxonomy and design space.IEEE Transactions on Visualization and Computer Graphics, 2024. doi: 10.1109/TVCG.2024.3456359 3, 6
2024
-
[47]
M. D. Rahman, G. J. Quadri, and P. Rosen. Exploring annotation strategies in professional visualizations: Insights from prominent us news portals. VisComm Workshop at IEEE VIS, 2023. doi: 10.31219/osf.io/fd8zj 2, 3, 5
2023 doi
-
[48]
M. D. Rahman, G. J. Quadri, D. A. Szafir, and P. Rosen. Exploring annotation taxonomy in grouped bar charts: A qualitative classroom study.Inf. Visualization, p. 14738716241270247, 2024. doi: 10.1177/ 14738716241270247 1
2024
-
[49]
M. D. Rahman, M. R.-u. Zaman, A. McNutt, and P. Rosen. Annogram: An annotative grammar of graphics extension.IEEE Visualization Conference (VIS), 2025. doi: 10.1109/VIS60296.2025.00053 1, 2, 3
2025
-
[50]
D. Ren, M. Brehmer, B. Lee, T. Höllerer, and E. K. Choe. Chartaccent: Annotation for data-driven storytelling. InIEEE PacificVis, pp. 230–239,
-
[51]
Satyanarayan, D
A. Satyanarayan, D. Moritz, K. Wongsuphasawat, and J. Heer. Vega-lite: A grammar of interactive graphics.IEEE Transactions on Visualization and Computer Graphics, 23(1):341–350, 2016. doi: 10.1109/TVCG.2016. 2599030 5
2016 doi
-
[52]
Satyanarayan, R
A. Satyanarayan, R. Russell, J. Hoffswell, and J. Heer. Reactive vega: A streaming dataflow architecture for declarative interactive visualization. IEEE Transactions on Visualization and Computer Graphics, 2016. doi: 10.1109/TVCG.2015.2467091 5
2016
-
[53]
S. C. Spivak and M. Tory. Fiction vs friction: Challenges in evaluating llms on data visualization tasks. InHuman-centered Evaluation and Auditing of Language Models (HEAL) Workshop at CHI 2025, 2025. Workshop paper. 1, 2
2025
-
[54]
Stokes, C
C. Stokes, C. X. Bearfield, and M. A. Hearst. The role of text in visualiza- tions: How annotations shape perceptions of bias and influence predictions. IEEE Transactions on Visualization and Computer Graphics, 2023. doi: 10.1109/TVCG.2023.3338451 1, 2
2023
-
[55]
Stokes, V
C. Stokes, V . Setlur, B. Cogley, A. Satyanarayan, and M. A. Hearst. Strik- ing a balance: Reader takeaways and preferences when integrating text and charts.IEEE Transactions on Visualization and Computer Graphics, 29(1):1233–1243, 2022. doi: 10.1109/TVCG.2022.3209383 1, 2
2022
-
[56]
B. Tang, A. Boggust, and A. Satyanarayan. Vistext: A benchmark for semantically rich chart captioning. InFindings of the Association for Computational Linguistics (ACL), pp. 7268–7298, 2023. doi: 10.18653/ v1/2023.acl-long.401 1, 2, 3
2023
-
[57]
Valmeekam, M
K. Valmeekam, M. Marquez, S. Sreedharan, and S. Kambhampati. On the planning abilities of large language models: A critical investigation. Advances in Neural Information Processing Systems, 36:75993–76005,
-
[58]
C. Wang, B. Lee, S. M. Drucker, D. Marshall, and J. Gao. Data formulator 2: Iterative creation of data visualizations, with ai transforming data along the way. InACM CHI, pp. 1–17. Association for Computing Machinery,
-
[59]
C. Yang, Y . Shi, Q. Ma, M. X. Liu, C. Kästner, and T. Wu. What prompts don’t say: Understanding and managing underspecification in llm prompts,
-
[60]
D. Yang, L. Zhang, Z. Yue, L. Chen, Y . Xu, W. Wang, and Q. Jin. Chartm3: Benchmarking chart editing with multimodal instructions. InACM In- ternational Conference on Multimedia, pp. 5001–5009, 2025. doi: 10. 1145/3746027.3755714 2
2025
-
[61]
J. Yang, A. M. McNutt, and L. Battle. Considering visualization example galleries. InIEEE Visual Languages and Human-Centric Computing (VL/HCC), pp. 329–343, 2024. doi: 10.1109/VL/HCC60511.2024.00043 5
2024
- [62]
-
[63]
Q. Zhi, A. Ottley, and R. Metoyer. Linking and layout: Exploring the integration of text and visualization in storytelling.Computer Graphics Forum, 38(3):675–685, 2019. doi: 10.1111/cgf.13719 2
2019 doi
-
[64]
Z. Zhu, M. Jia, Z. Zhang, L. Li, and M. Jiang. Multichartqa: Bench- marking vision-language models on multi-chart problems. InFindings of the Association for Computational Linguistics (ACL), pp. 11341–11359. Association for Computational Linguistics, 2025. doi: 10.18653/v1/202...
2025 doi
-
[65]
B. Zou, M. Cai, J. Zhang, and Y . J. Lee. Vgbench: Evaluating large language models on vector graphics understanding and generation. In Empirical Methods in Natural Language Processing (EMNLP), pp. 3647– 3659, 2024. doi: 10.18653/v1/2024.emnlp-main.213 2 Appendix S1 ANNOBENCHB...
2024 doi
- [66]
-
[69]
X. Zhao, X. Liu, Y . Haoyue, X. Luo, F. Zeng, J. Li, Q. Shi, and C. Chen. Chartedit: How far are mllms from automating chart analysis? evaluating mllms’ capability via chart editing. InFindings of the Association for Computational Linguistics (ACL), pp. 3616–3630, 2025. doi: 1...
2025
-
[2017]
doi: 10.1109/PACIFICVIS.2017.8031599 1, 2, 3, 4
2017
-
[2022]
doi: 10.1109/TVCG.2021.3114770 2, 3
2021
-
[2023]
doi: 10.1145/3544548.3580910 1, 2
-
[2025]
doi: 10.1145/3706598.3713296 2
-
[2706]
doi: 10.18653/v1/ 2022.acl-long.277 1
Association for Computational Linguistics, 2022. doi: 10.18653/v1/ 2022.acl-long.277 1
2022 doi
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.