Pith. sign in

REVIEW 3 major objections 7 minor 26 references

Five Levels of Reasoning Predict Where Science Figures Lose Readers

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-07 22:42 UTC pith:MI5IDVK5

load-bearing objection The R1–R5 typology for chart-image reasoning gaps is a genuine conceptual contribution, but its central empirical claim rests on a non-independent evaluation that the paper does not acknowledge. the 3 major comments →

arxiv 2607.05222 v1 pith:MI5IDVK5 submitted 2026-07-06 cs.CV

A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication

classification cs.CV
keywords multimodal reasoningchart-image coherencegrounding theoryscience communicationvision-language modelsscientific figuresinterpretive efforttypology
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that a chart, its accompanying image, and the caption linking them form a single multimodal unit, and that the inferential work required to read this unit varies systematically. To capture this variation, the authors develop a five-level typology of reasoning gaps (R1 through R5) that characterizes how chart, image, and text jointly convey a scientific claim. At R1 (Translation), the chart and image are alternative representations of the same data, requiring only recognition. At R2 (Quantification), the chart measures structures the image localizes. At R3 (Projection), the chart relates the image's content to an external variable not present in the image. At R4 (Evaluation), the chart audits the validity of the image's finding. At R5 (Framing), the panels share no data and only the reader's contextual frame links them. The typology was derived bottom-up with a neuroscience expert from a corpus of 79 traumatic brain injury papers. The authors test a falsifiable prediction: as the reasoning gap widens from R1 to R5, the figure does less of the linking work and the reader must supply more contextual knowledge. In a grounding study where a domain expert and three non-experts judged vision-language model descriptions of 25 chart-image pairs, the R-level predicted where their judgments aligned and where they diverged. Agreement decreased from R1 to R4, isolating the points where contextual knowledge rather than the figure carries coherence. At R5, the pattern reversed: the expert accepted thinner descriptions because her knowledge silently filled the gaps, while non-experts could not reconstruct the claim at all.

Core claim

The central object is the R1 to R5 typology itself, which operationalizes interpretive effort as a measurable gradient from figure-carried to reader-carried coherence. The key empirical finding is that this gradient predicts the specific points at which expert and non-expert readers diverge when interpreting chart-image pairs: agreement drops from R1 through R4, then reverses at R5. This reversal is the most diagnostic result, showing that at the highest reasoning gap, the expert's domain knowledge substitutes for the figure's missing link while the non-expert is stranded, unable to see what is absent from the description.

What carries the argument

The R1 to R5 typology, grounded in Clark's theory of communication where understanding requires shared common ground. The levels are: R1 Translation (same data, different encoding), R2 Quantification (chart measures what image localizes), R3 Projection (chart adds an external variable), R4 Evaluation (chart audits the image's finding), and R5 Framing (no shared data; reader supplies the linking frame). The typology is paired with Lundgard and Satyanarayan's four-level semantic model of chart descriptions as a complementary instrument: L1 to L4 measure description depth, while R1 to R5 measure the reach of the inference linking chart and image.

Load-bearing premise

The typology was derived from a single domain (traumatic brain injury neuroscience) with one expert, and its predictive claim rests on 25 chart-image pairs judged by three non-experts with no statistical significance testing reported. The authors state the five levels are domain-general but note that transferring the protocol to another field requires re-specifying the decision points, so the generality claim is not yet empirically established.

What would settle it

If expert and non-expert agreement patterns across chart-image pairs showed no systematic relationship to the R-level assignments, or if the R5 reversal did not replicate with additional experts and non-experts from other domains, the typology's predictive value would be undermined.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Figure designers can use the R-levels to diagnose where their chart-image pair demands contextual knowledge the reader may not have, and adjust captions or accompanying text to bridge the gap.
  • Benchmarks for vision-language models on scientific figures should account for reasoning gap levels rather than treating all chart-image pairs as equally recoverable from panels alone.
  • The R1 to R5 typology crossed with the L1 to L4 semantic depth model creates a grid that shows not only how deep a description goes but whether that depth meets the relational claim the pairing makes.
  • The R5 reversal finding suggests that expert review of figure descriptions may systematically overestimate accessibility, since experts unconsciously fill gaps that non-experts cannot cross.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. This paper introduces a typology of five reasoning-gap levels (R1–R5) characterizing the interpretive effort required to integrate chart-image pairs in scientific publications. The typology is grounded in Clark's grounding theory of communication and was derived bottom-up through a six-month collaboration with a neuroscience expert on a corpus of 79 TBI papers yielding 32 explicit chart-image pairs. A preliminary evaluation tests whether the R-levels predict where expert and non-expert readers align or diverge when judging VLM-generated descriptions of 25 implicit pairs. The central finding is that expert–non-expert agreement generally decreases from R1 to R4 and reverses at R5, where the expert's domain knowledge fills gaps that non-experts cannot cross. The work bridges chart-understanding research (Lundgard & Satyanarayan's L1–L4 model) with multimodal scientific figure analysis, proposing a crossed R×L grid as a design tool.

Significance. The paper addresses a genuine gap in computational chart understanding: most prior work treats single charts or assumes panel relationships are given, whereas this work models the inferential work a reader performs across paired panels. The grounding-theory anchor is well-motivated and distinguishes the contribution from purely taxonomic work. The falsifiable prediction—that VLM performance should decrease with R-level—is a strength, as is the learnability cross-check showing that zero-shot VLM agreement is low (34%) and plateaus at 65.6% even with caption augmentation. The R×L grid proposal (Section 5) is a concrete, actionable contribution for figure designers. However, the empirical base is narrow: the typology is derived from a single domain with one expert, and the predictive claim rests on 25 pairs with no statistical testing. The non-independence between the typology's co-deriver and its sole expert evaluator is a structural concern that the manuscript does not address.

major comments (3)
  1. Section 3.1–3.2 and Section 4: The neuroscience expert who co-derived the R1–R5 typology over six months is the same person who served as sole expert evaluator in the grounding study. Her judgments about whether VLM descriptions adequately ground each pair's relational claim are therefore not independent of the framework she helped construct. This concern is most acute for the R5 reversal—the paper's most distinctive diagnostic claim—because the typology itself predicts that at R5 the expert should rely on her own knowledge rather than the description. The manuscript should either (a) acknowledge this non-independence explicitly as a limitation and temper the R5 reversal claim accordingly, or (b) provide evidence that the expert's R-level assignments for the 25 implicit pairs were fixed prior to the evaluation (i.e., she did not both assign levels and evaluate descriptions for the same 5
  2. Figure 3b and Section 4: The agreement percentages (R1=50%, R2=57%, R3=40%, R4=29%, R5=25%) are described as showing that disagreement 'increases as the level rises,' but R2 (57%) exceeds R1 (50%), which is not monotonically decreasing. With approximately 5 pairs per level and no statistical significance test reported, it is unclear whether the observed trend differs from chance. A simple test (e.g., Cochran-Armitage trend test or a permutation test) would clarify whether the R-level genuinely predicts agreement patterns. Without it, the central predictive claim is unsupported by standard statistical evidence. The R5 reversal, resting on roughly 5 pairs, is especially fragile.
  3. Section 3.2: The claim that the five levels are 'domain-general' while the decision-point protocol is 'instantiated for TBI figure conventions' is internally consistent but empirically untested. The manuscript acknowledges that transferring the protocol requires re-specifying decision points for another field, yet the abstract and Section 5 frame the typology as a general contribution. The domain-generality claim should be either supported by at least a pilot application in a second domain or explicitly scoped as a hypothesis pending cross-domain validation.
minor comments (7)
  1. Section 3.2: The LLM-augmented annotation of 47 additional pairs is described as demonstrating scalability, but only 30% (14 of 47) were verified. The manuscript should report the agreement rate on this audited subset to support the scalability claim.
  2. Section 3.3: The VLM prompt iterations are described across four rounds, but the specific prompts used at each round are not provided in the main text or clearly referenced in the supplemental materials. Including or clearly linking to the prompt sequence would improve reproducibility.
  3. Figure 2: The triangle icons are described as using solid vs. dashed lines to indicate 'degree of interaction,' but the visual distinction is subtle and the semantics are not defined precisely. A brief legend explanation would help.
  4. Section 4: The three non-expert annotators are described as doctoral students with CS/DS backgrounds who took a data visualization course. Their specific levels of neuroscience exposure (even if 'limited') should be reported, as any prior exposure to TBI or neuroimaging could confound the expert–non-expert comparison.
  5. Section 4: The expert's evaluation took approximately 2 hours for 25 pairs, while non-experts took 60–70 minutes each. The difference in time-on-task could itself influence judgments and should be noted as a potential confound.
  6. The supplemental materials link (https://tinyurl.com/yzyp35rc) should be verified to ensure it resolves to the intended content for a published version.
  7. Section 7: The generative AI disclosure mentions Claude (Version 4.8, Anthropic). The version number may be a placeholder and should be verified.

Simulated Author's Rebuttal

3 responses · 1 unresolved

We thank the referee for a careful and constructive reading of our manuscript. The report identifies three substantive concerns: (1) non-independence between the typology's co-deriver and its sole expert evaluator, (2) the absence of statistical significance testing for the agreement trend, and (3) the unsupported domain-generality claim. We agree that all three points identify genuine weaknesses that the revision must address. We will (a) add an explicit limitation acknowledging the non-independence and temper the R5 reversal claim, (b) conduct and report a permutation test on the agreement data, and (c) reframe the domain-generality claim as a hypothesis pending cross-domain validation. One point—the R5 reversal resting on approximately five pairs—cannot be fully resolved without collecting additional data, and we acknowledge this as a standing limitation.

read point-by-point responses
  1. Referee: Section 3.1–3.2 and Section 4: The neuroscience expert who co-derived the R1–R5 typology over six months is the same person who served as sole expert evaluator in the grounding study. Her judgments about whether VLM descriptions adequately ground each pair's relational claim are therefore not independent of the framework she helped construct. This concern is most acute for the R5 reversal—the paper's most distinctive diagnostic claim—because the typology itself predicts that at R5 the expert should rely on her own knowledge rather than the description. The manuscript should either (a) acknowledge this non-independence explicitly as a limitation and temper the R5 reversal claim accordingly, or (b) provide evidence that the expert's R-level assignments for the 25 implicit pairs were fixed prior to the evaluation.

    Authors: The referee is correct that the non-independence between the typology's co-deriver and its sole expert evaluator is a structural concern that the manuscript does not currently address. We will take option (a): we will add an explicit limitation subsection acknowledging this non-independence and temper the R5 reversal claim accordingly. To be transparent about the timeline: the expert's R-level assignments for the 25 implicit pairs were indeed fixed prior to the evaluation phase—she assigned levels during the typology derivation (Section 3.2) and did not reassign or adjust them during the grounding study. We will state this explicitly in the revision. However, we agree that even with this temporal separation, the deeper concern remains: the expert constructed the framework that generates the predictions she then evaluates, and at R5 specifically, the typology predicts that her domain knowledge (not the description) carries coherence, which is precisely what she reported observing. This is a real circularity risk. In the revision, we will (1) add a dedicated limitation paragraph in Section 4, (2) reframe the R5 reversal as a pattern consistent with the typology's prediction rather than as independent confirmation of it, and (3) note that a fully independent test requires expert evaluators who did not participate in typology construction, which we commit to in future work. revision: yes

  2. Referee: Figure 3b and Section 4: The agreement percentages (R1=50%, R2=57%, R3=40%, R4=29%, R5=25%) are described as showing that disagreement 'increases as the level rises,' but R2 (57%) exceeds R1 (50%), which is not monotonically decreasing. With approximately 5 pairs per level and no statistical significance test reported, it is unclear whether the observed trend differs from chance. A simple test (e.g., Cochran-Armitage trend test or a permutation test) would clarify whether the R-level genuinely predicts agreement patterns. Without it, the central predictive claim is unsupported by standard statistical evidence. The R5 reversal, resting on roughly 5 pairs, is especially fragile.

    Authors: The referee is correct on both points. First, the trend is not monotonically decreasing—R2 (57%) exceeds R1 (50%)—and our language characterizing it as 'increasing disagreement' is imprecise. We will revise the description to acknowledge the non-monotonicity at R1–R2 and frame the pattern as a general downward trend with a local reversal at R2, rather than a strict monotonic decrease. Second, we will conduct a permutation test: treating the 25 pairs as fixed with their observed agreement outcomes, we will randomly permute R-level labels across pairs 10,000 times and compute the correlation between R-level and agreement rate for each permutation, building a null distribution against which we test the observed correlation. We will report the p-value and effect size. We expect this test to clarify whether the observed trend is distinguishable from chance given the small sample. We acknowledge upfront that with approximately 5 pairs per level, the test will have limited power, and we will state this explicitly. If the result is not significant at conventional thresholds, we will reframe the claim as an observed pattern that is consistent with the typology's prediction but not yet statistically confirmed, and we will note that a larger-scale study is needed to draw firm conclusions. The R5 reversal claim will be further tempered in light of both the small sample and the non-independence concern raised in the first comment. revision: yes

  3. Referee: Section 3.2: The claim that the five levels are 'domain-general' while the decision-point protocol is 'instantiated for TBI figure conventions' is internally consistent but empirically untested. The manuscript acknowledges that transferring the protocol requires re-specifying decision points for another field, yet the abstract and Section 5 frame the typology as a general contribution. The domain-generality claim should be either supported by at least a pilot application in a second domain or explicitly scoped as a hypothesis pending cross-domain validation.

    Authors: The referee is correct that the domain-generality claim is empirically untested and that the abstract and Section 5 overstate the current evidence. We do not have a pilot application in a second domain to offer at this time. We will therefore scope the claim explicitly as a hypothesis pending cross-domain validation. Specifically, we will (1) revise the abstract to replace language implying established generality with language framing the five levels as proposed to be domain-general, (2) add a sentence in Section 3.2 clarifying that the levels are theoretically motivated as domain-general (because they track the inferential structure of chart-image pairs, which is not field-specific) but that this has not been empirically tested beyond TBI, and (3) add cross-domain validation as an explicit next step in Section 5. We believe the theoretical argument for domain-generality is reasonable—the progression from shared data through external projection to reader-supplied framing is not specific to neuroscience—but we agree that this is a hypothesis, not a demonstrated result, and the revision will reflect that distinction. revision: yes

standing simulated objections not resolved
  • The R5 reversal rests on approximately 5 pairs and cannot be statistically strengthened without collecting additional data. We will temper the claim and report the permutation test, but we cannot make the evidence stronger than it is.

Circularity Check

2 steps flagged

The typology's central predictive claim is partially circular: the expert who co-derived the R1–R5 levels from explicit pairs is the sole expert evaluator in the grounding study, and the 'falsifiable prediction' that VLM performance decreases with R-level was confirmed by iteratively refining prompts against the same expert's labels.

specific steps
  1. fitted input called prediction [Section 3.3 (learnability cross-check) and Section 4 (grounding study)]
    "Treating the expert's R-level assignments as a reference, we prompted a frontier vision-language model to label 32 explicit pairs across four iterative rounds, refining the prompt after each round from the disagreement pattern with the expert. Zero-shot agreement was low (11/32, 34%), confirming that the levels cannot be recovered from the panels alone. Successive refinements raised agreement but plateaued; the largest gain came from adding the source paper's captions to the prompt (21/32, 65.6%). This is not validation of the typology, which was derived through the process above; it is a下游信号…"

    The paper claims the 'falsifiable prediction' that VLM performance decreases with R-level is confirmed by the learnability cross-check. But the cross-check fits the VLM prompt iteratively against the expert's R-level assignments (the same assignments used to derive the typology). The 'prediction' that performance decreases with R-level is thus confirmed by construction: the prompt was refined until agreement improved, and the improvement is attributed to adding captions — exactly the contextual material the typology says is needed. The paper acknowledges this is 'not validation,' but then presents it as support for the prediction in the grounding study.

  2. self definitional [Section 3.1–3.2 (typology derivation) and Section 4 (expert assessment)]
    "The typology emerged from sustained collaboration through a six-month, multi-step methodology (Figure 1) with a neuroscience expert… We derived the typology through a structured cognitive walkthrough… Working with the expert, three visualization-researcher co-authors read each chart-image pair and, for each pair, recorded the reasoning connecting the two panels… As an evaluation, we then ran a grounding study on 25 implicit pairs… with the expert and three non-experts."

    The expert who co-derived the typology (assigning R-levels, defining decision points, fitting level boundaries to her own reasoning) is the same person who serves as sole expert evaluator in the grounding study. Her judgments of whether VLM descriptions 'ground' each pair's relational claim are not independent of the typology she constructed. At R5 specifically, the typology predicts the expert will fill gaps with her own knowledge — and as the person who defined R5 as 'context-dependent,' her evaluation behavior is shaped by the framework itself. The R5 reversal (expert more lenient than non-experts) rests on ~5 pairs judged by the same person who defined what R5 means.

full rationale

The paper has genuine independent content: the 25 implicit pairs used in the grounding study are distinct from the 32 explicit pairs used in derivation, and the non-expert judgments are independent of the typology. However, the expert's dual role as co-deriver and sole evaluator creates non-independence for the most distinctive claim (the R5 reversal), and the learnability cross-check is a fitted-input-called-prediction pattern where prompts were iteratively refined against the expert's labels. The paper is transparent about some of these limitations ('This is not validation of the typology'), which mitigates the circularity somewhat. Score 4 reflects partial circularity: the central claim has independent empirical content from non-experts, but the expert-side evaluation and the VLM prediction are partially constructed from the typology's own inputs.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 1 invented entities

The typology has 5 levels and associated decision points derived empirically from one domain. The main axioms are the applicability of Clark's theory to multimodal figures and the domain-generality of the levels. No new physical entities or forces are introduced. The R-levels have a falsifiable handle (the grounding study prediction).

free parameters (3)
  • Number of typology levels (5) = 5
    The number of levels was determined bottom-up from the TBI corpus with the expert; it is not derived from theory.
  • Decision-point protocol thresholds = not specified
    The five decision points used to route pairs were 'instantiated for TBI figure conventions' and are domain-specific.
  • VLM prompt iterations (4 rounds) = 4
    The learnability cross-check used four iterative prompt refinement rounds; the stopping point and refinement criteria are not formalized.
axioms (4)
  • domain assumption Clark's grounding theory of communication applies to chart-image pairs in scientific figures.
    Section 1: The paper extends Clark's face-to-face dialogue account to multimodal scientific figures without independent validation that the mapping holds.
  • domain assumption The chart-image pair (with caption) is the correct unit of analysis for multimodal scientific communication.
    Section 1: This framing choice structures the entire typology but is asserted, not derived.
  • ad hoc to paper The R-levels are domain-general despite being derived from TBI only.
    Section 3.2: 'The five levels are domain-general, but transferring the protocol would require re-specifying the decision points for another field.' This claim is not tested.
  • domain assumption Three CS/DS doctoral students are sufficient non-expert controls for visualization literacy.
    Section 4: The paper assumes visualization literacy is ruled out as a confound because the non-experts took a data visualization course.
invented entities (1)
  • R1–R5 reasoning gap levels independent evidence
    purpose: Classify the inferential work required to connect chart and image panels in a scientific figure.
    The levels make a falsifiable prediction (expert–non-expert agreement decreases with R-level), which was tested in the grounding study with directional support.

pith-pipeline@v1.1.0-glm · 13393 in / 2629 out tokens · 159640 ms · 2026-07-07T22:42:57.943440+00:00 · methodology

0 comments
read the original abstract

Charts and images appear together throughout scientific publications, yet most computational work does not characterize their coherence. We argue that a chart, its accompanying image, and the caption that links them form a multimodal unit, and that the inferential work required to read it varies systematically. To capture this variation, we develop a typology of reasoning gaps, R1 through R5, that characterizes how chart, image, and text jointly convey a scientific claim, and the interpretive work this demands of the reader. Some pairs restate the same data, while in other pairs, charts are used to quantify a structure the image localizes, project image content onto an external variable, audit an image-based claim, or jointly construct a frame that neither panel can establish alone. The typology is anchored in the grounding theory of communication and was derived bottom-up, with a neuroscience expert, from a corpus of 79 traumatic brain injury papers and 32 chart-image pairs. Crucially, the levels provide a systematic mechanism for identifying where grounding succeeds or breaks down, rather than leaving it to subjective inference. We show this in a study in which a domain expert and three non-experts judge vision-language model (VLM) descriptions of 25 pairs: the level predicts where their judgments align and where they diverge, isolating the points at which contextual knowledge, not the figure, carries coherence. This typology thus offers figure designers a systematic way to balance text against chart-image pairs, bridging the expert-to-non-expert divide in reading a scientific takeaway.

Figures

Figures reproduced from arXiv: 2607.05222 by Aritra Dasgupta, Avina Nakarmi, Sohom Sen, Sreyashi Samaddar, Xun Song.

Figure 1
Figure 1. Figure 1: Methodology pipeline. The typology was derived bottom￾up, not imposed. A neuroscience expert framed the corpus; three visualization-researcher co-authors iterated with her over explicit pairs to surface the R1–R5 levels; three separate annotators then tested them on implicit pairs. Successive rounds of prompt refine￾ment ensured expert–VLM agreement is 65.6%. malized or operationalized systematically [6]. … view at source ↗
Figure 2
Figure 2. Figure 2: The R1–R5 typology illustrated with chart-image pairs from the TBI corpus. Each panel shows a level’s triangle icon (solid vs dashed lines indicate the degree of interaction) and one corpus example. Interpretive effort grows from (a) to (e) as the link between chart and image moves outside the figure, and so does the contextual knowledge a reader must bring to ground it. At (a) R1 Translation [19] and (b) … view at source ↗
Figure 3
Figure 3. Figure 3: Disagreement in grounding of VLM-generated descrip￾tion. Panel (a) shows that more non-experts considered the VLM capable of correctly describing the relational claim between chart￾image pairs than the domain expert did, highlighting the impact of prior knowledge on filling reasoning gaps. Panel (b) shows agree￾ment between experts and non-experts tends to decrease with the progression in reasoning gap lev… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 26 canonical work pages · 3 internal anchors

  1. [1]

    Bonnelle, R

    V . Bonnelle, R. Leech, K. M. Kinnunen, T. E. Ham, C. F. Beckmann, X. De Boissezon, R. J. Greenwood, and D. J. Sharp. Default mode net- work connectivity predicts sustained attention deficits after traumatic brain injury.Journal of Neuroscience, 31(38):13442–13451, 2011. 3

  2. [2]

    Bosak, P

    N. Bosak, P. Branco, P. Kuperman, C. Buxbaum, R. M. Cohen, S. Fadel, R. Zubeidat, R. Hadad, A. Lawen, N. Saadon-Grosman, et al. Brain connectivity predicts chronic pain in acute mild traumatic brain injury.Annals of Neurology, 92(5):819–833, 2022. 3

  3. [3]

    H. H. Clark.Using Language. Cambridge University Press, 1996. 1, 2

  4. [4]

    H. H. Clark and S. E. Brennan. Grounding in communication. In L. B. Resnick, J. M. Levine, and S. D. Teasley, eds.,Perspectives on Socially Shared Cognition, pp. 127–149. American Psychological As- sociation, 1991. 1, 2

  5. [5]

    H. H. Clark and T. B. Carlson. Hearers and speech acts.Language, 58(2):332–373, 1982. doi: 10.2307/414102 2

  6. [6]

    Dasgupta

    A. Dasgupta. Analytical reasoning and visualization: Estranged bed- fellows? reclaiming a missed opportunity in the human-ai era. pp. 1–7, 2026. 2

  7. [7]

    S. L. Franconeri, L. M. Padilla, P. Shah, J. M. Zacks, and J. Hull- man. The science of visual data communication: What works.Psy- chological Science in the Public Interest, 22(3):110–161, 2021. doi: 10.1177/15291006211051956 2

  8. [8]

    H. P. Grice. Logic and conversation. In P. Cole and J. L. Morgan, eds.,Syntax and Semantics, Vol. 3: Speech Acts, pp. 41–58. Academic Press, 1975. 2

  9. [9]

    Huang, H

    K.-H. Huang, H. P. Chan, M. Fung, H. Qiu, M. Zhou, S. Joty, S.-F. Chang, and H. Ji. From pixels to insights: A survey on automatic chart understanding in the era of large foundation models.IEEE Transac- tions on Knowledge and Data Engineering, 37(5):2550–2568, 2024. 2

  10. [10]

    Hullman and N

    J. Hullman and N. Diakopoulos. Visualization rhetoric: Framing effects in narrative visualization.IEEE Transactions on Visualiza- tion and Computer Graphics, 17(12):2231–2240, 2011. doi: 10.1109/ TVCG.2011.255 2

  11. [11]

    M. S. Islam, R. Rahman, A. Masry, M. T. R. Laskar, M. T. Nayeem, and E. Hoque. Are large vision language models up to the chal- lenge of chart comprehension and reasoning? an extensive investi- gation into the capabilities and limitations of LVLMs.arXiv preprint arXiv:2406.00257, 2024. doi: 10.48550/arXiv.2406.00257 1

  12. [12]

    S. E. Kahou, V . Michalski, A. Atkinson, ´A. K´ad´ar, A. Trischler, and Y . Bengio. Figureqa: An annotated figure dataset for visual reasoning. arXiv preprint arXiv:1710.07300, 2017. 2

  13. [13]

    Kembhavi, M

    A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi. A diagram is worth a dozen images. InEuropean Con- ference on Computer Vision (ECCV), pp. 235–251. Springer, 2016. 1, 2

  14. [14]

    D. H. Kim, E. Hoque, J. Kim, and M. Agrawala. Facilitating document reading by linking text and tables. InProceedings of the 31st Annual ACM Symposium on User Interface Software and Technology (UIST), pp. 423–434, 2018. doi: 10.1145/3242587.3242617 2

  15. [15]

    S. M. Kosslyn. Understanding charts and graphs.Applied Cognitive Psychology, 3(3):185–225, 1989. doi: 10.1002/acp.2350030302 2

  16. [16]

    Z. Li, X. Yang, K. Choi, W. Zhu, R. Hsieh, H. Kim, J. H. Lim, S. Ji, B. Lee, X. Yan, L. R. Petzold, S. D. Wilson, W. Lim, and W. Y . Wang. MMSci: A dataset for graduate-level multi-discipline multimodal sci- entific understanding.arXiv preprint arXiv:2407.04903, 2024. doi: 10 .48550/arXiv.2407.04903 2

  17. [17]

    Lundgard and A

    A. Lundgard and A. Satyanarayan. Accessible visualization via nat- ural language descriptions: A four-level model of semantic con- tent.IEEE Transactions on Visualization and Computer Graphics, 28(1):1073–1083, 2021. doi: 10.1109/TVCG.2021.3114770 1, 4

  18. [18]

    Masry, J

    A. Masry, J. Q. Tan, S. Joty, E. Hoque, et al. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the association for computational linguistics: ACL 2022, pp. 2263–2279, 2022. 1, 2

  19. [19]

    Raizman, I

    R. Raizman, I. Tavor, A. Biegon, S. Harnof, C. Hoffmann, G. Tsar- faty, E. Fruchter, L. Tatsa-Laur, M. Weiser, and A. Livny. Traumatic brain injury severity in a network perspective: A diffusion mri based connectome study.Scientific reports, 10(1):9121, 2020. 3

  20. [20]

    S. Sen, A. Nakarmi, X. Song, and A. Dasgupta. Pointers at uzh shared task 2026: Reasoning probes for argumentation mining in un resolu- tions. In13th Workshop on Argument Mining and Reasoning, p. 145,

  21. [21]

    Siegel, Z

    N. Siegel, Z. Horvitz, R. Levin, S. Divvala, and A. Farhadi. Figure- seer: Parsing result-figures in research papers. InEuropean Confer- ence on Computer Vision, pp. 664–680. Springer, 2016. 1, 2

  22. [22]

    M. S. F. Sorg, L. Delano-Wood, M. N. Luc, D. M. Schiehser, K. L. Hanson, D. A. Nation, M. E. Lanni, A. J. Jak, K. Lu, M. Meloy, et al. White matter integrity in veterans with mild traumatic brain injury: associations with executive function and loss of consciousness.The Journal of head trauma rehabilitation, 29(1):21, 2014. 3

  23. [23]

    Vaidya and A

    S. Vaidya and A. Dasgupta. Knowing what to look for: A fact-evidence reasoning framework for decoding communicative visu- alization. In2020 IEEE Visualization Conference (VIS), pp. 231–235. IEEE, 2020. 2

  24. [24]

    Z. Wang, M. Xia, L. He, H. Chen, Y . Liu, R. Zhu, K. Liang, X. Wu, H. Liu, S. Malladi, et al. Charxiv: Charting gaps in realistic chart understanding in multimodal llms. vol. 37, pp. 113569–113697, 2024. 2

  25. [25]

    Q. Zhi, A. Ottley, and R. Metoyer. Linking and layout: Exploring the integration of text and visualization in storytelling. 38(3):675–685,

  26. [26]

    J. Zong, C. Lee, A. Lundgard, J. Jang, D. Hajas, and A. Satya- narayan. Rich screen reader experiences for accessible data visual- ization. 41(3):15–27, 2022. 1 5