Pith. sign in

REVIEW 4 major objections 4 minor 55 references

Open-ended sketching and critique assessments separate visualization skill levels that multiple-choice tests cannot, especially among experienced viewers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 00:42 UTC pith:BZBNGAJF

load-bearing objection Worth engaging: the first real shot at open-ended sketching/critique assessments for visualization literacy, but the headline claim leans on single-rater, non-blind grading. the 4 major comments →

arxiv 2608.00330 v1 pith:BZBNGAJF submitted 2026-07-31 cs.HC

Read, Critique, or Sketch? Investigating Alternative Visualization Literacy Assessment Modalities

classification cs.HC
keywords visualization literacyqualitative assessmentsketchingcritiquethink-alouditem response theoryceiling effectshigher-order skills
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to show that visualization literacy is not fully captured by multiple-choice comprehension tests. It introduces two web-based qualitative assessments—thinking aloud while critiquing flawed charts, and sketching a visualization for a given dataset—and compares them with two established tests across crowdworkers, students, and visualization researchers. The results suggest both qualitative modalities differentiate between experience levels better than the established tests, with Mini-VLAT showing a ceiling even among crowdworkers. Because the new assessments correlate only moderately with existing ones, they appear to measure distinct, higher-order skills such as contextual judgment and design reasoning. A sympathetic reader would conclude that qualitative assessments are a promising complement for measuring high-end visualization literacy.

Core claim

The paper claims that both sketching and critique succeed at differentiating between groups of varying visualization skill and exhibit improved sensitivity over both Mini-VLAT and CALVI. Under the authors' rubrics—four categories for sketching and three for critique—sketching separates all three pairwise comparisons (experts vs students, students vs crowdworkers, experts vs crowdworkers), while critique separates experts from each other group but not students from crowdworkers. Mini-VLAT separates no pair and CALVI separates only experts from crowdworkers. The qualitative assessments correlate moderately with each other and with CALVI, and weakly with Mini-VLAT, which the authors read as evi

What carries the argument

The carrying mechanism is the pair of five-point holistic rubrics applied to think-aloud transcripts and final sketches. Unlike early analytic checklists, the rubrics credit the reasoning behind a judgment—a participant can defend a 3D treemap well and score well—rather than ticking off expected features. For sketching, rubric categories cover visual hierarchy, base encoding, semantic validity, and text/guides; for critique, reading, contextualization, and judgment. A Bayesian ordinal item-response model then converts the rubric scores into latent ability estimates used for group comparisons and correlation analysis.

Load-bearing premise

The load-bearing premise is that the single grader's rubric scores on the main dataset were not biased toward the expert group; inter-rater reliability was checked only on a small pilot, and the grader could recognize many experts' voices.

What would settle it

Rescore all 80 participants' sketches and critiques with two independent raters who are blind to group and hear only anonymized transcripts; if the expert–student and student–crowdworker gaps no longer reach significance, the paper's central sensitivity claim would not hold.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Multiple-choice-only assessment of visualization literacy should be reconsidered when the goal is to distinguish experienced individuals; Mini-VLAT's ceiling makes it unsuitable for that purpose.
  • The critique rubric appears transferable: similar discrimination across three very different charts suggests the method, not just the specific stimuli, carries the signal.
  • Sketching and critique can be used as a complement after a short screening test, an adaptive design the authors explicitly propose.
  • Researchers can reduce grading burden by selecting a subset of sketching tasks, such as keeping the tree task for expert populations, without much loss of discrimination.
  • Qualitative assessments open the door to studying how people construct transformations and design choices, not just whether they can read a chart.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: If baseline visualization literacy has risen since VLAT's introduction, as the paper tentatively suggests, norms for older multiple-choice instruments may need revalidation in current populations before being used as study filters.
  • Editorial inference: The authors' acknowledged grader-bias risk implies a concrete test: if expert-student differences persist when grading is blind and voice-anonymized, the construct-validity case would be substantially stronger.
  • Editorial inference: Because correlation between reading- and writing-style assessments around 0.5 mirrors text literacy research, one might predict that combining sketching and critique with multiple-choice scores gives a more complete latent visualization-literacy factor than any single modality.
  • Editorial inference: The rubric's treatment of ambiguous sketches, where audio is used to infer implied labels, suggests an automated pipeline could grade transcripts plus final canvas state; a pilot study comparing automated grading to human rubric scores would test feasibility.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper introduces two qualitative visualization literacy assessments—freeform sketching and think-aloud critique—and compares them against CALVI and Mini-VLAT in a within-subjects online study with 80 participants from three expertise groups (crowdworkers, students, and researchers). Using Bayesian IRT, the authors estimate item discrimination and latent abilities and report pairwise group contrasts. They find that both new assessments differentiate experts from non-experts and that they capture skills weakly to moderately correlated with existing tests. The paper argues that qualitative assessments complement multiple-choice measures, especially for distinguishing highly skilled individuals.

Significance. If the findings hold, the paper provides a useful step toward measuring higher-order visualization literacy, with practical guidance for instrument selection. Strengths include a transparent study design, Bayesian IRT with reported credible intervals, public study materials and code, and a sensitivity analysis for prior treemap exposure. The candid acknowledgment of grading bias in Section 7 is commendable. However, the central claim of improved sensitivity is weakened by single-rater, non-blind grading of the main study and the absence of a formal cross-assessment statistical comparison. The paper is likely to stimulate further research on qualitative literacy assessments, but the current evidence does not fully support the strong conclusions in the abstract.

major comments (4)
  1. [§4 Grading; §7 Limitations] The primary group-level results (Figs. 5, 8) are based on scores assigned by the first author alone on the main experiment. Pilot inter-rater reliability (κ=0.81 critique, 0.76 sketching) was computed on a small pilot subset (69 critique tasks, 15 sketching tasks) during rubric alignment, not on the main data. Section 7 explicitly acknowledges that 'many participants’ voices were recognizable by the first author during grading,' particularly in the expert condition. Because the rubric is holistic and the rater knew the hypotheses and group membership, the expert-vs-student and expert-vs-crowd gaps could be inflated. This is load-bearing for the paper’s central claim. Please blind the grading (e.g., use transcripts with voice masking), add a second independent rater on a sample of main data, or provide a sensitivity analysis bounding the possible bias.
  2. [§5 Results; Fig. 8; Abstract] The abstract and Section 1 claim that both sketching and critique 'succeed at differentiating between groups of varying skill level.' However, the critique assessment does not significantly differentiate students from crowdworkers (Fig. 8: 0.39 [−0.05, 0.82]). Only expert vs. student and expert vs. crowdworker contrasts are significant. The claim should be qualified to avoid overstating the evidence. Similarly, the phrase 'improved sensitivity over both Mini-VLAT and CALVI' is only supported by a qualitative comparison of which CIs exclude zero; no direct statistical comparison of sensitivity across assessments is reported.
  3. [§5 Will Sketching...; Fig. 8] The conclusion that the new assessments exhibit improved sensitivity is inferred from separate Welch t-tests per assessment. No model or test directly compares the magnitude of group differences across assessments (e.g., an assessment-by-group interaction or a posterior distribution of the contrast difference). A formal comparison would strengthen the claim and provide a quantitative basis for the title and abstract.
  4. [§6 Participant Group Differences; §7 Limitations; Appendix A.1] The sensitivity analysis in Appendix A.1 only removes four students with prior exposure to the treemap. However, the critique stimuli include other widely publicized charts (the line chart, the map), and Section 7 notes that experts 'likely have seen some of the examples before.' Prior familiarity with specific charts could differentially boost expert scores. Please measure or ask about prior exposure for all stimuli, or at least broaden the sensitivity analysis, so that the assessments measure generalizable literacy rather than recognition of famous charts.
minor comments (4)
  1. [Abstract] The abstract contains 'Y et' with a stray space; should be 'Yet.'
  2. [Fig. 8 caption] The caption reads 'If a test confidence interval does not overlap with the associated mean, then the test is significant.' This is inaccurate; significance is determined by whether the CI excludes zero, not the mean. Please correct.
  3. [§4 Study Methods] The test name is spelled inconsistently: 'Calvi/Mini-VLAT' and 'CALVI' appear interchangeably. Use 'CALVI' consistently.
  4. [§5 Results] The reference to 'Fig. 6' in the text occurs before the reader has context for the figure's layout; consider adding a pointer to the rubric table (Table 1) when discussing item discrimination.

Circularity Check

0 steps flagged

No significant circularity: the central comparison is empirical and benchmarked against external instruments; identified limitations are validity threats, not definitional circularity.

full rationale

The paper's central claim is that two newly developed qualitative assessments (sketching and critique) differentiate between expertise groups and capture skills distinct from Mini-VLAT and CALVI. This is an empirical claim supported by a user study, Bayesian IRT models, and correlation analyses. The rubrics were developed from prior literature, teaching experience, and pilot iteration (Section 3.1), not from the outcome variable being predicted. The comparison instruments (CALVI, Mini-VLAT) are external and not fit to the new assessments. Self-citations (e.g., Trrack [14], reVISit [16], crowdsourced think-aloud methods [15]) concern study infrastructure and methodology, not the central construct or its confirmation. Citation [22] (Ge et al.) is by overlapping authors and motivates the qualitative-modality direction, but it is background motivation, not the load-bearing evidence for the empirical results. The acknowledged non-blind, single-rater grading of the main dataset (Section 7) is a genuine internal-validity limitation and could affect the magnitude of group differences, but it is not circularity in the derivation: the assessment scores are not defined in terms of the claimed comparisons, nor are the conclusions forced by a fitted parameter renamed as a prediction. The IRT discriminations and pairwise differences are computed from the rubric scores as collected, and the paper transparently reports the extent of rater agreement on pilot data. Thus no circular step can be exhibited with quoted text; the appropriate finding is low score reflecting one minor self-citation that is not load-bearing.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No numeric free parameters are fitted to make the central claim; the hand-built rubrics and IRT modeling choices are the main analytical assumptions. There are no invented theoretical entities—the new instruments are measurement procedures, not postulates.

axioms (4)
  • domain assumption The rubric categories validly operationalize higher-order visualization literacy (sketching: Visual Design/Hierarchy, Base Encoding, Semantic Validity, Guides/Text/Annotations; critique: Reading, Contextualization, Judgment).
    Section 3.1: rubrics were developed from co-author teaching experience and arts-education rubrics; the paper explicitly forgoes formal content-validity indexing (Section 6), so construct validity is assumed rather than externally verified.
  • domain assumption A 2PL Bayesian ordinal IRT model with local independence and unidimensionality fits the four short assessments and produces comparable latent ability estimates.
    Section 5: no dimensionality checks or model-fit comparisons are reported; item counts are small (9 critique rubric items, 20 sketch rubric items across 5 tasks), making IRT parameter estimates relatively fragile.
  • domain assumption The recruited groups (Prolific crowdworkers, students who took a visualization course, visualization researchers) correspond to distinct expertise levels and are otherwise comparable in effort and motivation.
    Section 4.2: groups were recruited via different channels, compensation differed ($20 for crowdworkers/researchers vs. $50 for students), and group labels were not verified; crowdworkers who reported relevant courses or PhDs were retained in the crowd group.
  • domain assumption Inter-rater reliability measured on pilot data (critique: 69 tasks, sketching: 15 tasks) transfers to the main dataset, which was coded by a single author.
    Section 4.2 Grading and Section 7: the first author coded all main data and could recognize expert voices; this is acknowledged as a source of potential grader bias.

pith-pipeline@v1.3.0-alltime-deepseek · 18690 in / 9912 out tokens · 94760 ms · 2026-08-04T00:42:21.194028+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Read, Critique, or Sketch? Investigating Alternative Visualization Literacy Assessment Modalities." pith.science (2026). https://pith.science/paper/BZBNGAJF

@misc{pith2026260800330,
  author       = {Pith},
  title        = {Pith review of: Read, Critique, or Sketch? Investigating Alternative Visualization Literacy Assessment Modalities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BZBNGAJF}},
  note         = {Machine review of arXiv:2608.00330}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Visualization literacy is a multifaceted construct encompassing skills and competencies, such as decoding data, constructing charts, and identifying design flaws. Yet, assessments of these competencies has been primarily constrained to multiple choice assessments that target lower-order skills, such as chart comprehension. As a result, they often exhibit ceiling effects (i.e., even modestly skilled individuals commonly score near the top of the scale), and do not provide enough information about an individual's higher-order skills (e.g., applying external knowledge, formulating critiques, and designing visualizations). To close these gaps, we develop and investigate two web-based qualitative assessments for testing the critique and design aspects of visualization literacy through online think-aloud critique and sketching of visualization designs based on data and a prompt. We compare performance on our assessments to two established visualization literacy assessments, CALVI and Mini-VLAT, by administering them to three groups that represent three experience levels: crowdworkers, students who have taken a relevant course, and researchers. We find that our critique and sketching assessments capture skills distinct from existing measures and that they differentiate between experienced individuals better than multiple choice-based alternatives. Although administering and grading qualitative assessments can be challenging, our findings suggest qualitative, multimodal assessments are a promising complement to existing visualization literacy assessments, in particular when high visualization skills need to be distinguished.

Figures

Figures reproduced from arXiv: 2608.00330 by Alexander Lex, Andrew McNutt, Lane Harrison, Lily W. Ge, Matthew Kay, Zach Cutler.

Figure 1
Figure 1. Figure 1: Examples of the drawings that participants generated in different sketching tasks, based on the data snippets shown on the [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The visualizations used in the critique-based assessment of visualization literacy. (A) A pair of maps comparing windy days in the summer to [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: shows the self-assessed visualization skills in interpreting graphs, designing graphs, and interpreting numbers between groups. This study was marked exempt by the University of Utah IRB (IRB#00193994). We conducted a number of pilots, primarily to evalu￾ate and iterate on rubrics, as described in Sec. 3.1. Initially, we intended to run our study on tablets to improve the sketching experience. How￾ever, ea… view at source ↗
Figure 5
Figure 5. Figure 5: Sketching and critique scores strongly differentiate between experts, students and crowdworkers. Raw scores of participants for the four different assessments used in our study, with group means for each assessment marked. was κ = 0.81 for critique and κ = 0.76 for sketching, indicating sub￾stantial agreement in both cases [31]. The first author coded the data drawn from the main experiment. Answers that w… view at source ↗
Figure 6
Figure 6. Figure 6: Sketching tasks show similar discriminability as Critique, with a wider range. Tree, 100M, and all Critique tasks show strong discriminability, while Ranking and Rainfall tasks show slightly lower discriminability. Item discriminability calculated from a 2PL IRT model for every question and rubric item in Sketching and Critique assessments. Each rubric item displays the mean and 66% / 95% credible interval… view at source ↗
Figure 7
Figure 7. Figure 7: Sketching, critique, and CALVI show moderate correlations with each other, while Mini-VLAT correlates weakly. differentiating between participant groups. Given that sketching and critique are commonly learned through formal visualization training, observing that our assessments differentiate between experts and non￾experts better than past assessments is a promising sign of the tests’ content validity (how… view at source ↗
Figure 8
Figure 8. Figure 8: Latent ability scores derived from Bayesian IRT confirm sketching and critique differentiate between visualization expertise levels better than CALVI and Mini-VLAT. Distribution of participants’ latent abilities by test and group (with corresponding means marked). Pairwise comparisons are made between each group using Welch’s two-sample t-test and displayed below the Students and Experts groups. If a test … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

55 extracted references · 10 canonical work pages · 2 internal anchors

  1. [1]

    Adelberger, O

    P. Adelberger, O. Lesota, K. Eckelt, M. Schedl, and M. Streit. Iguan- odon: A Code-Breaking Game for Improving Visualization Construction Literacy.IEEE Transactions on Visualization and Computer Graphics, 31(9):5713–5725, Sept. 2025. doi: 10.1109/TVCG.2024.3468948 2

  2. [2]

    Beasley, A

    Z. Beasley, A. Friedman, L. Pieg, and P. Rosen. Leveraging Peer Feed- back to Improve Visualization Education. InIEEE Pacific Visualization Symposium, pp. 146–155. IEEE, Tianjin, China, June 2020. doi: 10.1109/ PacificVis48177.2020.1261 4

  3. [3]

    A. M. Bell and T. Gift. Fraud in Online Surveys: Evidence from a Nonprobability, Subpopulation Sample.Journal of Experimental Political Science, pp. 148–153, 2023. doi: 10.1017/XPS.2022.8 5

  4. [4]

    M. L. Blommel and M. A. Abate. A Rubric to Assess Critical Litera- ture Evaluation Skills.American Journal of Pharmaceutical Education, 71(4):63, Aug. 2007. doi: 10.5688/aj710463 3

  5. [5]

    J. Boy, R. A. Rensink, E. Bertini, and J.-D. Fekete. A Principled Way of Assessing Visualization Literacy.IEEE Transactions on Visualization and Computer Graphics, 20(12):1963–1972, 2014. doi: 10.1109/TVCG.2014. 2346984 2, 4

  6. [6]

    Brath and E

    R. Brath and E. Banissi. Evaluation of Visualization by Critiques. In Proceedings of the Sixth Workshop on Beyond Time and Errors on Novel Evaluation Methods for Visualization, BELIV ’16, pp. 19–26. Association for Computing Machinery, New York, NY , USA, Oct. 2016. doi: 10. 1145/2993901.2993904 3

  7. [7]

    H. M. Breland, B. Bridgeman, and M. E. Fowles. Writing Assessment in Admission to Higher Education: Review and Framework.ETS Re- search Report Series, 1999(1):i–37, 1999. doi: 10.1002/j.2333-8504.1999 .tb01801.x 2, 9

  8. [8]

    Bressa, J

    N. Bressa, J. Louis, W. Willett, and S. Huron. Input Visualization: Collect- ing and Modifying Data with Visual Representations. InProceedings of the CHI Conference on Human Factors in Computing Systems, pp. 1–18. ACM, Honolulu HI USA, May 2024. doi: 10.1145/3613904.3642808 9

  9. [9]

    P.-C. Bürkner. Bayesian Item Response Modeling in R with brms and Stan. Journal of Statistical Software, 100:1–54, Nov. 2021. doi: 10.18637/jss. v100.i05 6

  10. [11]

    P. L. Cooper. The Assessment of Writing Ability: A Review of Research. ETS Research Report Series, 1984(1):i–46, 1984. doi: 10.1002/j.2330 -8516.1984.tb00052.x 2

  11. [12]

    Cope and M

    B. Cope and M. Kalantzis. The Things You Do to Know: An Introduction to the Pedagogy of Multiliteracies. InA Pedagogy of Multiliteracies. Palgrave Macmillan, 2015. doi: 10.1057/9781137539724_1 2

  12. [13]

    Y . Cui, L. W. Ge, Y . Ding, F. Yang, L. Harrison, and M. Kay. Adaptive Assessment of Visualization Literacy.IEEE Transactions on Visualization and Computer Graphics, 30(1):628–637, Jan. 2024. doi: 10.1109/TVCG. 2023.3327165 2, 5, 8

  13. [14]

    Cutler, K

    Z. Cutler, K. Gadhave, and A. Lex. Trrack: A Library for Provenance Tracking in Web-Based Visualizations. InProc. VIS, pp. 116–120. IEEE,

  14. [15]

    Cutler, L

    Z. Cutler, L. Harrison, C. Nobre, and A. Lex. Crowdsourced Think-Aloud Studies. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–23. ACM, Yokohama Japan, Apr. 2025. doi: 10.1145/3706598.3714305 5

  15. [16]

    Cutler, J

    Z. Cutler, J. Wilburn, H. Shrestha, Y . Ding, B. Bollen, K. A. Nadib et al. ReVISit 2: A Full Experiment Life Cycle User Study Framework.IEEE Transactions on Visualization and Computer Graphics, 32(1):13–23, Jan

  16. [17]

    DeMars.Item Response Theory

    C. DeMars.Item Response Theory. Series in Understanding Statistics. Oxford University Press, Oxford New York, N.Y , 2010. doi: 10.1093/ acprof:oso/9780195377033.001.0001 2, 6

  17. [18]

    S. Few. One of Bill Gates’ favorite graphs redesigned, 2014. 3

  18. [19]

    Fitzgerald and T

    J. Fitzgerald and T. Shanahan. Reading and Writing Relations and Their Development.Educational Psychologist, 35(1):39–50, Jan. 2000. doi: 10. 1207/S15326985EP3501_5 8

  19. [20]

    A. R. Fox, M. Morgenstern, G. M. Jones, and A. Satyanarayan. Quanti- fying Visualization Vibes: Measuring Socio-Indexicality at Scale.IEEE Transactions on Visualization and Computer Graphics (VIS), 32:1273– 1283, 2026. doi: 10.1109/TVCG.2025.3634819 9

  20. [21]

    Leveraging Peer Review in Visualization Education: A Proposal for a New Model

    A. Friedman and P. Rosen. Leveraging Peer Review in Visualization Education: A Proposal for a New Model, Jan. 2021. doi: 10.48550/arXiv. 2101.07708 4

  21. [23]

    L. W. Ge, Y . Cui, and M. Kay. CALVI: Critical Thinking Assessment for Literacy in Visualizations. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, pp. 1–18. Association for Computing Machinery, New York, NY , USA, Apr. 2023. doi: 10. 1145/3544548.3581406 1, 2, 3, 4, 5, 6, 8

  22. [24]

    L. W. Ge, Y . Cui, and M. Kay. A VEC: An Assessment of Visual Encoding Ability in Visualization Construction. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, pp. 1–16. Association for Computing Machinery, New York, NY , USA, Apr. 2025. doi: 10.1145/3706598.3713364 1, 2, 4, 6, 8

  23. [25]

    A. Gelman. Bill Gates’s favorite graph of the year, 2013. 3

  24. [26]

    Gratzl, A

    S. Gratzl, A. Lex, N. Gehlenborg, H. Pfister, and M. Streit. LineUp: Visual Analysis of Multi-Attribute Rankings.IEEE Transactions on Visualization and Computer Graphics, 19(12):2277–2286, Dec. 2013. doi: 10.1109/ TVCG.2013.173 6

  25. [27]

    Hedayati, A

    M. Hedayati, A. Hunt, and M. Kay. From pixels to practices: Reconceptu- alizing visualization literacy, May 2024. doi: 10.31219/osf.io/6mq42 1, 2, 8

  26. [28]

    Hedayati and M

    M. Hedayati and M. Kay. What University Students Learn In Visualization Classes.IEEE Transactions on Visualization and Computer Graphics, 31(1):1072–1082, Jan. 2025. doi: 10.1109/TVCG.2024.3456291 2, 3, 6

  27. [29]

    Hofmann, P

    V . Hofmann, P. R. Kalluri, D. Jurafsky, and S. King. AI generates covertly racist decisions about people based on their dialect.Nature, 633(8028):147–154, 2024. doi: 10.1038/s41586-024-07856-5 9

  28. [30]

    R. Kosara. Visualization Criticism - The Missing Link Between Infor- mation Visualization and Art. InInformation Visualization, 2007. IV ’07. 11th International Conference, pp. 631–636, 2007. doi: 10.1109/IV.2007. 130 3

  29. [31]

    J. R. Landis and G. G. Koch. The Measurement of Observer Agreement for Categorical Data.Biometrics, 33(1):159–174, 1977. doi: 10.2307/ 2529310 6

  30. [32]

    Lee, S.-H

    S. Lee, S.-H. Kim, and B. C. Kwon. VLAT: Development of a Visual- ization Literacy Assessment Test.IEEE Transactions on Visualization and Computer Graphics, 23(1):551–560, 2017. doi: 10.1109/TVCG.2016. 2598920 1, 2, 8

  31. [33]

    S. A. Livingston. Constructed-Response Test Questions: Why We Use Them; How We Score Them.Educational Testing Service, 2009. 3

  32. [34]

    H. E. Merzdorf, D. Jaison, M. B. Weaver, J. Linsey, T. Hammond, and K. A. Douglas. Sketching assessment in engineering education: A systematic literature review.Journal of Engineering Education, 113(4):872–893,

  33. [35]

    Molina León, B

    G. Molina León, B. Bach, M. Valentim, and N. Elmqvist. A Multiliteracy Model for Interactive Visualization Literacy: Definitions, Literacies, and Steps for Future Research. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, CHI ’26, pp. 1–16. Association for Computing Machinery, New York, NY , USA, Apr. 2026. doi: 10. 1145/377...

  34. [36]

    Munzner.Visualization Analysis and Design

    T. Munzner.Visualization Analysis and Design. CRC Press, Taylor & Francis Group, 2014. doi: 10.1201/b17511 3

  35. [37]

    Nobre, K

    C. Nobre, K. Zhu, E. Mörth, H. Pfister, and J. Beyer. Reading Between the Pixels: Investigating the Barriers to Visualization Literacy. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24, pp. 1–17. Association for Computing Machinery, New York, NY , USA, May 2024. doi: 10.1145/3613904.3642760 9

  36. [38]

    Y . Okan, E. Janssen, M. Galesic, and E. A. Waters. Using the Short Graph Literacy Scale to Predict Precursors of Health Behavior Change. Medical Decision Making, 39(3):183–195, Apr. 2019. doi: 10.1177/ 0272989X19829728 2 10 © 2026 IEEE. This is the author’s version of the article that has been published in IEEE Transactions on Visualization and Computer ...

  37. [39]

    Pandey and A

    S. Pandey and A. Ottley. Mini-VLAT: A Short and Effective Measure of Visualization Literacy.Computer Graphics Forum, 42(3):1–11, 2023. doi: 10.1111/cgf.14809 1, 2, 4, 5

  38. [40]

    E. M. Peck, S. E. Ayuso, and O. El-Etr. Data is Personal: Attitudes and Perceptions of Data Visualization in Rural Pennsylvania. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI ’19, pp. 1–12. Association for Computing Machinery, New York, NY , USA, 2019. doi: 10.1145/3290605.3300474 7

  39. [41]

    Quispel and A

    A. Quispel and A. Maes. Would you prefer pie or cupcakes? Preferences for data visualization designs of professionals and laypeople in graphic design.Journal of Visual Languages & Computing, 25(2):107–116, Apr

  40. [42]

    R. A. K. Richardson, S. S. Hong, J. A. Byrne, T. Stoeger, and L. A. N. Amaral. The entities enabling scientific fraud at scale are large, resilient, and growing rapidly.Proceedings of the National Academy of Sciences, 122(32):e2420092122, Aug. 2025. doi: 10.1073/pnas.2420092122 3

  41. [43]

    J. C. Roberts, H. Alnjar, A. E. Owen, and P. D. Ritsos. Critical Design Strategy: A Method for Heuristically Evaluating Visualisation Designs. IEEE Transactions on Visualization and Computer Graphics, 32(1):1383– 1393, Jan. 2026. doi: 10.1109/TVCG.2025.3634783 3

  42. [44]

    J. C. Roberts, C. Headleand, and P. D. Ritsos. Sketching Designs Using the Five Design-Sheet Methodology.IEEE Transactions on Visualization and Computer Graphics, 22(1):419–428, Jan. 2016. doi: 10.1109/TVCG. 2015.2467271 2

  43. [45]

    Saske, L

    A. Saske, L. Koesten, T. Möller, J. Staudner, and S. Kritzinger. A Multidi- mensional Assessment Method for Visualization Understanding (MdamV). IEEE Transactions on Visualization and Computer Graphics, 32(3):2695– 2708, Mar. 2026. doi: 10.1109/TVCG.2026.3653265 2

  44. [46]

    J. Su, Y . Yan, F. Fu, Z. Han, J. Ye, X. Liu et al. EssayJudge: A Multi- Granular Benchmark for Assessing Automated Essay Scoring Capabil- ities of Multimodal Large Language Models. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, eds.,Findings of the Association for Computational Linguistics: ACL 2025, pp. 6363–6389. Association for Computational L...

  45. [47]

    C. O. Tam. Evaluating Students’ Performance in Responding to Art: The Development and Validation of an Art Criticism Assessment Rubric. International Journal of Art & Design Education, 37(3):519–529, 2018. doi: 10.1111/jade.12154 3, 4

  46. [48]

    Tldraw. Tldraw. https://github.com/tldraw/tldraw, 2026. 5

  47. [49]

    The State of the Art in Visualization Literacy

    M. Varona, K. Bonilla, M. Hedayati, A. Joshi, L. Harrison, M. Kay et al. The State of the Art in Visualization Literacy, Aug. 2025. doi: 10. 48550/arXiv.2509.01018 1, 2, 3, 4, 7

  48. [50]

    Walny, S

    J. Walny, S. Huron, and S. Carpendale. An Exploratory Study of Data Sketching for Visual Representation.Computer Graphics Forum, 34(3):231–240, 2015. doi: 10.1111/cgf.12635 2, 6

  49. [51]

    Y . Wang, J. Huang, L. Du, Y . Guo, Y . Liu, and R. Wang. Evaluating large language models as raters in large-scale writing assessments: A psycho- metric framework for reliability and validity.Computers and Education: Artificial Intelligence, 9:100481, Dec. 2025. doi: 10.1016/j.caeai.2025. 100481 9

  50. [52]

    Bill Gates’s graph of the year.The Washington Post, Dec

    Wonkborg. Bill Gates’s graph of the year.The Washington Post, Dec

  51. [2013]

    As with the original analysis, we ran an ordinal regression model run with 4 chains of 20,000 iterations

    3 11 A APPENDIX A.1 Critique re-analysis In order to ensure the results in the critique assessment were not influenced by the prior-exposure of participants to thetreemapplot, we re-performed the analysis without the four students who indicated prior exposure, including a new Bayesian IRT model. As with the original analysis, we ran an ordinal regression ...

  52. [2014]

    doi: 10.1016/j.jvlc.2013.11.007 9

  53. [2020]

    doi: 10.1109/VIS47514.2020.00030 5

  54. [2024]

    doi: 10.1002/jee.20560 2

  55. [2026]

    doi: 10.1109/TVCG.2025.3633896 5