{"id":"9573de49-1b61-418c-b68e-541fc9a08bcc","arxiv_id":"2501.16277","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"On a contamination-controlled version of the VLAT, GPT-4 and Gemini show lower visualization literacy than reported human norms and rely more on prior knowledge than on the chart itself.","lead":"This study tested GPT-4 and Gemini on a modified 53-question visualization literacy test and found that both models score below the general public and often answer from memory. The result matters for researchers who hoped to use LLMs as cheap stand-ins for human evaluators of charts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline claim that both LLMs 'lack visualization literacy' compared to humans rests on comparing modified-test LLM scores to original VLAT human scores without establishing item difficulty equivalence; removing labels and the Omit option likely disadvantages LLMs.","rationale":"The reader's weakest_assumption correctly identifies the unvalidated comparability between the modified LLM test and the original human VLAT baseline. This is the single most load-bearing concern because the paper's strongest claim is explicitly a comparison to humans, and Table 2 is the primary evidence for it. The concern is concrete: removing value labels, randomizing data, and removing the Omit option are procedural changes that can systematically alter item difficulty for any test-taker, human or machine. Without a human baseline on the modified instrument, the observed LLM accuracy cannot be attributed to a lack of visualization literacy relative to humans. I do not see an internal inconsistency in the LLM-vs-LLM or visualization-present-vs-absent analyses; those comparisons use the same modified items and are informative about relative performance and knowledge reliance. The paper also provides reproducible scripts and a large trial count, which are strengths. The multiple-comparison issue noted by the reader is real but secondary, since the headline claim does not depend on the 49 interaction tests. Because the CONDITIONAL verdict already requires a same-test human baseline or a moderated claim, my read does not change the verdict; the requested condition should be implemented before the 'compared to humans' wording is accepted as established.","tokens_in":37520,"tokens_out":3160,"duration_ms":33361,"concrete_test":"Administer the modified 53-item test from Section 4.1.1 (modified charts, same questions, no value labels, no Omit option) to a sample of human participants on Prolific, e.g., N=100, screened comparably to the VLAT norming sample, and score with the same exact-answer rubric. If human accuracy on the modified test is significantly lower than the original VLAT human accuracy in Table 2, then the LLM-vs-human comparison cannot support the claim that LLMs underperform humans. A sharper version would counterbalance original and modified VLAT items within subjects and estimate an item-difficulty offset; if such an offset exists, the headline comparison should be recomputed after controlling for it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract and Section 4.3.1 ('compared to humans, both GPT-4 and Gemini lack visualization literacy') is operationalized in Table 2 as a direct comparison between LLM accuracy on modified charts and human accuracy on the original VLAT charts from Lee et al. (2017). But Section 4.1.1 changes the test in at least three ways that can shift difficulty independently of literacy: (1) all data values are randomized, so semantic priors become misleading rather than helpful; (2) data value labels are removed, so Retrieve Value tasks require interpolating from axes instead of reading labels; (3) the Omit option is removed and the LLMs are forced to guess. The original human VLAT included value labels and an Omit option. No human data were collected on the modified instrument, so the size and direction of the difficulty shift are unknown. If the modified items are also harder for humans, the low LLM accuracy in Table 2 reflects test difficulty rather than a deficiency relative to humans. This assumption enters in Section 4.1.1 and is load-bearing for the strongest claim; the within-LLM comparisons for H1-H4 do not depend on it, but the 'compared to humans' claim does. The paper's own internal evidence, such as decontextualization improving GPT-4's accuracy, is interesting but does not substitute for a same-instrument human baseline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a template for assessing the visualization literacy of GPT-4 and Gemini by administering a modified 53-item VLAT in which charts are redrawn with randomized data, value labels are removed, the Omit option is suppressed, and answer choices are counterbalanced over 120 trials per question. The authors compare LLM accuracy to the original VLAT human accuracy (Table 2), analyze performance by visualization/task/model via bootstrap logistic regression (H1-H4), run follow-up experiments without answer choices and with decontextualized charts, and estimate cost/time differences between LLMs and human participants. The paper reports that the LLMs lack visualization literacy relative to the general public and that they rely on pre-existing knowledge rather than on the information in the visualizations.","tokens_in":37910,"tokens_out":7056,"duration_ms":62123,"significance":"The paper's methodological core is strong: 25,440 trials with counterbalanced option order, separate sessions per question, bootstrap logistic regression with hyperparameter tuning, and a decontextualization experiment are careful and reproducible, with code and data publicly released. The template of re-rendering VLAT charts with randomized data to avoid memorization is a useful contribution to LLM evaluation in visualization. The within-LLM comparisons (H1-H4) and the decontextualization results are interesting and informative. However, the headline human-comparison claim is not supported by the evidence as presented, because the modified test was never administered to human participants. If a same-instrument human baseline is added or the claims are reframed, the internal findings would still be valuable to the visualization community.","major_comments":[{"comment":"The headline claim that both GPT-4 and Gemini 'lack visualization literacy compared to humans' (abstract and Section 4.3.1) is operationalized in Table 2 by comparing LLM accuracy on the modified VLAT to the original VLAT human accuracies from Lee et al. (2017). Section 4.1.1 changes the instrument in at least three consequential ways: all data values are randomized, value labels are removed, and the Omit option is replaced by a forced guess. These changes plausibly shift difficulty independently of visualization literacy: randomized values make semantic priors misleading, removing labels turns Retrieve Value tasks into axis interpolation, and removing Omit eliminates a safe response. The original human data were collected on the original instrument, with labels and Omit. No human participants took the modified instrument, so the size and direction of the difficulty shift are unknown. The paper assumes, but does not validate, item difficulty equivalence between the two instruments. This assumption is load-bearing for the strongest claim in the abstract and Section 4.3.1; the within-LLM comparisons for H1-H4 do not depend on it. The authors should either collect human data on the modified test (e.g., via crowdsourcing) or explicitly reframe the claim as 'LLMs perform poorly on a modified VLAT' rather than 'compared to humans, LLMs lack visualization literacy.'","section":"Section 4.1.1, Table 2"},{"comment":"The claim that LLMs 'heavily relied on their pre-existing knowledge' is not established by the H4 analysis as written. H4 tests whether accuracy is higher with a visualization than without; Section 4.3.4 reports that the majority of visualization/task interactions were not significant. The text then concludes 'suggesting that LLMs mostly rely on their knowledge base.' This is a non-sequitur: failing to find a benefit of the visualization is consistent with several alternatives, including that the models attempt to use the visual information but fail to extract it, or that the visual information actively misleads them. The comparison is also confounded by model version: Experiment 2 uses gpt-4-turbo-preview and gemini-pro, while Experiment 1 uses gpt-4-vision-preview and gemini-pro-vision (Section 4.1.3). The more direct evidence for prior-knowledge reliance comes from Experiment 5, where decontextualization improved GPT-4's accuracy (Section 5.3.2), but those results are reported as descriptive comparisons without statistical tests, and the paper itself notes that the number of examples per task is too small for statistical conclusions. The abstract's causal claim should either be supported by a statistical test of the decontextualization effect or softened to say that LLMs appear to rely on context when available and perform poorly whether or not the visualization is present.","section":"Section 4.3.4 and Section 5.1.3"}],"minor_comments":[{"comment":"The statement that 'GPT-4 performed better than random in 25 of the 53 questions' should be supported by a per-question binomial test or bootstrap confidence intervals, since 120 trials are available and the random baseline varies by the number of answer options.","section":"Section 4.3.1"},{"comment":"The entry 'Make Comperison' is a typo for 'Make Comparison.'","section":"Table 5"},{"comment":"The time comparison against the 25-second per-question limit in VLAT is not apples-to-apples because human participants answered all 53 questions in a single session, whereas LLMs were tested per question in separate sessions; the paper acknowledges this, but the cost table should include session overhead or explicitly state the limitation.","section":"Section 6"},{"comment":"The subfigure labels skip (f), (g), and (k) because three charts were excluded; renumbering or adding placeholders would avoid confusion.","section":"Figure 8"},{"comment":"The statement that providing answer choices 'guides' LLMs is a causal claim; the design compares different prompt conditions, so consider softening to 'is associated with more answerable responses.'","section":"Section 5.3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for TVCG and the proposed template is useful. The main obstacle is the unvalidated human baseline; the authors should be encouraged either to collect a small human sample on the modified test or to rewrite the abstract and Section 4 so that the headline claim does not rely on the original VLAT human accuracies. The internal contrasts and the decontextualization experiment are solid and can be preserved. The paper is repairable without discarding the core work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, the genuinely useful contribution is the modified VLAT template and the with/without visualization contrast: 25,000+ trials, randomized data, labels removed, Omit removed, a decontextualized follow-up, and code and data released. That internal contrast shows both GPT-4 and Gemini answer roughly as well without the chart as with it, which is strong evidence they lean on priors. This is new, and it is a reusable benchmark design. Second, the headline claim that the models 'lack visualization literacy compared to humans' does not hold as stated, because humans never took the modified test. The comparison uses original VLAT human accuracies from Lee et al. against LLMs on charts with randomized values, no data labels, and forced guessing. Those changes plausibly make the items harder for humans too, and the paper offers no same-test human baseline or item-difficulty calibration. The deficiency is real relative to random chance, but not demonstrably human-relative. The authors call it a 'qualitative comparison' in Section 4.2, yet the abstract and Section 4.3.1 go further than that evidence supports.\n\nThe internal hypotheses (H1–H4) are much sounder. The logistic regression with bootstrapped coefficients is appropriate, and the decontextualization experiment is a clever way to probe knowledge reliance. Two smaller concerns: the 629-variable model with 623 significant coefficients and no multiple-comparison correction will produce false positives, so the per-visualization/task conclusions should be read as exploratory; and the 'real truth' discussion in Section 7 is interesting but speculative, though it fits the data.\n\nThis paper is for anyone considering LLMs as cheap evaluators in visualization and anyone building contamination-resistant benchmarks. The template is the lasting artifact. I would send it to review with a request for either a human baseline on the modified items or a moderated claim; the data and code are already released, so a same-test baseline is feasible. The authors are honest about limitations, and the work is a real step forward.","headline":"A solid contamination-controlled LLM evaluation whose real finding is that the models ignore the charts; the human-literacy comparison is the weak link and needs a same-test baseline.","tokens_in":38283,"tokens_out":2596,"would_cite":true,"duration_ms":25872,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Both GPT-4 and Gemini fall short of the general public's visualization literacy, and they answer from prior knowledge rather than reading the charts.","keywords":["visualization literacy","large language models","GPT-4","Gemini","VLAT","chart question answering","knowledge reliance","evaluation benchmark"],"falsifier":"Give the same human participants both the original VLAT and this paper's modified version and show that the modified version is substantially harder for people; that would break the comparability assumption on which the 'below the general public' claim rests. Alternatively, run a newer multimodal model on the modified test and observe accuracy at or above the human baseline that drops sharply when the charts are removed.","tokens_in":37301,"feed_emoji":"📊","tokens_out":5549,"duration_ms":50358,"temperature":0.7,"pith_summary":"The paper asks whether today's multimodal large language models can read data visualizations well enough to serve as stand-in evaluators for human users. To avoid the problem that models may have memorized the standard Visualization Literacy Assessment Test, the authors rebuilt all 53 VLAT charts with randomized data, removed value labels and the 'omit' option, and re-keyed the answers. They then tested GPT-4 and Gemini across 25,440 trials, with and without the charts present. Their central claim is that both models fall short of the general-public baseline reported for VLAT, and that the models are often answering from pre-existing world knowledge rather than from the visual information. The paper also proposes this modified-test template as a reusable benchmark for future LLM literacy evaluations.","feed_headline":"GPT-4 and Gemini flunk a chart-literacy test","feed_subtitle":"Both models answer from memory instead of reading the charts, a 53-question benchmark shows.","key_machinery":"The central instrument is a modified version of the 53-item VLAT in which every chart is regenerated with randomized data values, value labels are stripped from marks, and all correct answers are re-derived so they differ from the original test's answers. Around this, the paper builds a paired protocol: the same questions are asked with the chart present and with the chart absent, across 120 counterbalanced trials per item and model, so that any accuracy that persists without the chart can be attributed to the model's prior knowledge rather than to visual reading. A logistic-regression model with bootstrapped coefficients over the factors visualization type, task type, model, and visualization presence, plus a decontextualized follow-up experiment that anonymizes proper nouns, supplies the quantitative evidence that the charts themselves contribute little to the answers.","core_discovery":"On the paper's own terms, the discovery is that visualization literacy, measured as the ability to read, understand, and interpret information from a chart, is not yet possessed by GPT-4 or Gemini at the level of the general public. Across 53 question stems and 12 chart types, the models beat the human VLAT accuracy on only 14 and 15 questions respectively, and they exceeded random guessing on only about half the items. The decisive evidence for knowledge substitution comes from the paired no-visualization condition: for the majority of chart/task combinations, showing the chart did not significantly change the models' accuracy, even though the charts contained randomized data whose correct answers could not be known in advance. Decontextualizing the charts—replacing real city, country, company, and party names with placeholders—raised GPT-4's accuracy from about 31 percent to 42 percent, which the paper reads as further confirmation that contextual labels trigger memorized answers instead of chart reading.","pith_inferences":["Because the charts used randomized data, the near-zero accuracy of both models on several 'find extremum' and 'retrieve value' items suggests those tasks are currently under-served by multimodal LLMs; one could test whether chain-of-thought prompting or higher-resolution inputs close that gap.","The knowledge-substitution failure likely extends to real-world chart QA benchmarks, so high scores on those benchmarks may overstate true visual reasoning; a direct comparison using the same protocol on such benchmarks would reveal how much is memorization.","The paper's human baseline comes from the original VLAT, but no human data was collected on the modified test; recruiting a human sample on the modified charts would let the authors convert their qualitative comparison into a proper head-to-head literacy test.","The decontextualization effect suggests a testable design rule: strip real-world labels before feeding charts to LLMs when the goal is to assess the chart itself—and the same trick might improve LLM performance in downstream applications like automated chart captioning or data extraction."],"forward_implications":["Current GPT-4 and Gemini cannot yet replace human raters for visualization readability assessment; human validation remains necessary for evaluation studies.","The paired with/without-visualization protocol offers a reusable, contamination-resistant template for benchmarking future LLM literacy.","Anonymizing or decontextualizing chart labels can push LLMs toward genuine chart reading, at least for GPT-4, suggesting a practical prompt or interface design for chart question answering.","LLM evaluation remains far cheaper and faster than crowdworker evaluation — roughly $0.53 versus $1.80 per 53-item run for GPT-4, and near zero for Gemini — so even imperfect models may be economical screening tools for flagging problematic visualizations.","Model choice can be tuned per task: GPT-4 did comparatively well at trend-finding and hierarchical structure, while Gemini won several comparison and retrieval tasks, so practitioners can mix models rather than relying on one."],"supporting_citations":[{"why":"Supplies the original 53-item VLAT, its questions and charts, its general-public accuracy baseline, and the definition of visualization literacy used throughout the paper.","marker":"[24]"},{"why":"Prior empirical evaluation that used VLAT directly on GPT-4, which the paper argues may be contaminated by training data and therefore motivates the modified, re-keyed test.","marker":"[4]"},{"why":"Provides the Mini-VLAT abridged literacy assessment, referenced as part of the measurement landscape that the modified test extends.","marker":"[31]"},{"why":"Shows that LLMs are sensitive to the order of options in multiple-choice questions, justifying the study's counterbalanced option-order design.","marker":"[32]"},{"why":"Supplies the beta-difference distribution used to test whether probability differences between models and between visualization-present and visualization-absent conditions are statistically significant.","marker":"[33]"},{"why":"Provides the bootstrap resampling procedure used to estimate coefficient and probability distributions for the logistic-regression hypothesis tests.","marker":"[40]"}],"fun_headline_variants":["LLMs fail chart literacy: they guess from memory","GPT-4 and Gemini can't read charts, study finds","Chart literacy test: LLMs score below general public","LLMs rely on prior knowledge, not charts, in VLAT","Visualization literacy gap: LLMs vs. humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The human accuracy scores from the original 2017 VLAT paper are treated as a valid baseline, even though the LLMs faced modified charts with randomized data, no value labels, and no 'omit' option; if those changes made the LLM version harder than the human version, the conclusion that LLMs lack literacy would be overstated.","fun_headline_variants_meta":{"raw":{"variants":["LLMs fail chart literacy: they guess from memory","GPT-4 and Gemini can't read charts, study finds","Chart literacy test: LLMs score below general public","LLMs rely on prior knowledge, not charts, in VLAT","Visualization literacy gap: LLMs vs. humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1407,"prompt_tokens":968,"completion_tokens":439,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":372}},"tokens_in":584,"tokens_out":439,"duration_ms":4431,"temperature":1.0,"reasoning_tokens":372,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:33:53.605919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same human participants both the original VLAT and this paper's modified version and show that the modified version is substantially harder for people; that would break the comparability assumption on which the 'below the general public' claim rests. Alternatively, run a newer multimodal model on the modified test and observe accuracy at or above the human baseline that drops sharply when the charts are removed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the beta-difference distribution used to test whether probability differences between models and between visualization-present and visualization-absent conditions are statistically significant."}],"review_version":1}