REVIEW 3 major objections 7 minor 26 references
Five Levels of Reasoning Predict Where Science Figures Lose Readers
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-07 22:42 UTC pith:MI5IDVK5
load-bearing objection The R1–R5 typology for chart-image reasoning gaps is a genuine conceptual contribution, but its central empirical claim rests on a non-independent evaluation that the paper does not acknowledge. the 3 major comments →
A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central object is the R1 to R5 typology itself, which operationalizes interpretive effort as a measurable gradient from figure-carried to reader-carried coherence. The key empirical finding is that this gradient predicts the specific points at which expert and non-expert readers diverge when interpreting chart-image pairs: agreement drops from R1 through R4, then reverses at R5. This reversal is the most diagnostic result, showing that at the highest reasoning gap, the expert's domain knowledge substitutes for the figure's missing link while the non-expert is stranded, unable to see what is absent from the description.
What carries the argument
The R1 to R5 typology, grounded in Clark's theory of communication where understanding requires shared common ground. The levels are: R1 Translation (same data, different encoding), R2 Quantification (chart measures what image localizes), R3 Projection (chart adds an external variable), R4 Evaluation (chart audits the image's finding), and R5 Framing (no shared data; reader supplies the linking frame). The typology is paired with Lundgard and Satyanarayan's four-level semantic model of chart descriptions as a complementary instrument: L1 to L4 measure description depth, while R1 to R5 measure the reach of the inference linking chart and image.
Load-bearing premise
The typology was derived from a single domain (traumatic brain injury neuroscience) with one expert, and its predictive claim rests on 25 chart-image pairs judged by three non-experts with no statistical significance testing reported. The authors state the five levels are domain-general but note that transferring the protocol to another field requires re-specifying the decision points, so the generality claim is not yet empirically established.
What would settle it
If expert and non-expert agreement patterns across chart-image pairs showed no systematic relationship to the R-level assignments, or if the R5 reversal did not replicate with additional experts and non-experts from other domains, the typology's predictive value would be undermined.
If this is right
- Figure designers can use the R-levels to diagnose where their chart-image pair demands contextual knowledge the reader may not have, and adjust captions or accompanying text to bridge the gap.
- Benchmarks for vision-language models on scientific figures should account for reasoning gap levels rather than treating all chart-image pairs as equally recoverable from panels alone.
- The R1 to R5 typology crossed with the L1 to L4 semantic depth model creates a grid that shows not only how deep a description goes but whether that depth meets the relational claim the pairing makes.
- The R5 reversal finding suggests that expert review of figure descriptions may systematically overestimate accessibility, since experts unconsciously fill gaps that non-experts cannot cross.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces a typology of five reasoning-gap levels (R1–R5) characterizing the interpretive effort required to integrate chart-image pairs in scientific publications. The typology is grounded in Clark's grounding theory of communication and was derived bottom-up through a six-month collaboration with a neuroscience expert on a corpus of 79 TBI papers yielding 32 explicit chart-image pairs. A preliminary evaluation tests whether the R-levels predict where expert and non-expert readers align or diverge when judging VLM-generated descriptions of 25 implicit pairs. The central finding is that expert–non-expert agreement generally decreases from R1 to R4 and reverses at R5, where the expert's domain knowledge fills gaps that non-experts cannot cross. The work bridges chart-understanding research (Lundgard & Satyanarayan's L1–L4 model) with multimodal scientific figure analysis, proposing a crossed R×L grid as a design tool.
Significance. The paper addresses a genuine gap in computational chart understanding: most prior work treats single charts or assumes panel relationships are given, whereas this work models the inferential work a reader performs across paired panels. The grounding-theory anchor is well-motivated and distinguishes the contribution from purely taxonomic work. The falsifiable prediction—that VLM performance should decrease with R-level—is a strength, as is the learnability cross-check showing that zero-shot VLM agreement is low (34%) and plateaus at 65.6% even with caption augmentation. The R×L grid proposal (Section 5) is a concrete, actionable contribution for figure designers. However, the empirical base is narrow: the typology is derived from a single domain with one expert, and the predictive claim rests on 25 pairs with no statistical testing. The non-independence between the typology's co-deriver and its sole expert evaluator is a structural concern that the manuscript does not address.
major comments (3)
- Section 3.1–3.2 and Section 4: The neuroscience expert who co-derived the R1–R5 typology over six months is the same person who served as sole expert evaluator in the grounding study. Her judgments about whether VLM descriptions adequately ground each pair's relational claim are therefore not independent of the framework she helped construct. This concern is most acute for the R5 reversal—the paper's most distinctive diagnostic claim—because the typology itself predicts that at R5 the expert should rely on her own knowledge rather than the description. The manuscript should either (a) acknowledge this non-independence explicitly as a limitation and temper the R5 reversal claim accordingly, or (b) provide evidence that the expert's R-level assignments for the 25 implicit pairs were fixed prior to the evaluation (i.e., she did not both assign levels and evaluate descriptions for the same 5
- Figure 3b and Section 4: The agreement percentages (R1=50%, R2=57%, R3=40%, R4=29%, R5=25%) are described as showing that disagreement 'increases as the level rises,' but R2 (57%) exceeds R1 (50%), which is not monotonically decreasing. With approximately 5 pairs per level and no statistical significance test reported, it is unclear whether the observed trend differs from chance. A simple test (e.g., Cochran-Armitage trend test or a permutation test) would clarify whether the R-level genuinely predicts agreement patterns. Without it, the central predictive claim is unsupported by standard statistical evidence. The R5 reversal, resting on roughly 5 pairs, is especially fragile.
- Section 3.2: The claim that the five levels are 'domain-general' while the decision-point protocol is 'instantiated for TBI figure conventions' is internally consistent but empirically untested. The manuscript acknowledges that transferring the protocol requires re-specifying decision points for another field, yet the abstract and Section 5 frame the typology as a general contribution. The domain-generality claim should be either supported by at least a pilot application in a second domain or explicitly scoped as a hypothesis pending cross-domain validation.
minor comments (7)
- Section 3.2: The LLM-augmented annotation of 47 additional pairs is described as demonstrating scalability, but only 30% (14 of 47) were verified. The manuscript should report the agreement rate on this audited subset to support the scalability claim.
- Section 3.3: The VLM prompt iterations are described across four rounds, but the specific prompts used at each round are not provided in the main text or clearly referenced in the supplemental materials. Including or clearly linking to the prompt sequence would improve reproducibility.
- Figure 2: The triangle icons are described as using solid vs. dashed lines to indicate 'degree of interaction,' but the visual distinction is subtle and the semantics are not defined precisely. A brief legend explanation would help.
- Section 4: The three non-expert annotators are described as doctoral students with CS/DS backgrounds who took a data visualization course. Their specific levels of neuroscience exposure (even if 'limited') should be reported, as any prior exposure to TBI or neuroimaging could confound the expert–non-expert comparison.
- Section 4: The expert's evaluation took approximately 2 hours for 25 pairs, while non-experts took 60–70 minutes each. The difference in time-on-task could itself influence judgments and should be noted as a potential confound.
- The supplemental materials link (https://tinyurl.com/yzyp35rc) should be verified to ensure it resolves to the intended content for a published version.
- Section 7: The generative AI disclosure mentions Claude (Version 4.8, Anthropic). The version number may be a placeholder and should be verified.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive reading of our manuscript. The report identifies three substantive concerns: (1) non-independence between the typology's co-deriver and its sole expert evaluator, (2) the absence of statistical significance testing for the agreement trend, and (3) the unsupported domain-generality claim. We agree that all three points identify genuine weaknesses that the revision must address. We will (a) add an explicit limitation acknowledging the non-independence and temper the R5 reversal claim, (b) conduct and report a permutation test on the agreement data, and (c) reframe the domain-generality claim as a hypothesis pending cross-domain validation. One point—the R5 reversal resting on approximately five pairs—cannot be fully resolved without collecting additional data, and we acknowledge this as a standing limitation.
read point-by-point responses
-
Referee: Section 3.1–3.2 and Section 4: The neuroscience expert who co-derived the R1–R5 typology over six months is the same person who served as sole expert evaluator in the grounding study. Her judgments about whether VLM descriptions adequately ground each pair's relational claim are therefore not independent of the framework she helped construct. This concern is most acute for the R5 reversal—the paper's most distinctive diagnostic claim—because the typology itself predicts that at R5 the expert should rely on her own knowledge rather than the description. The manuscript should either (a) acknowledge this non-independence explicitly as a limitation and temper the R5 reversal claim accordingly, or (b) provide evidence that the expert's R-level assignments for the 25 implicit pairs were fixed prior to the evaluation.
Authors: The referee is correct that the non-independence between the typology's co-deriver and its sole expert evaluator is a structural concern that the manuscript does not currently address. We will take option (a): we will add an explicit limitation subsection acknowledging this non-independence and temper the R5 reversal claim accordingly. To be transparent about the timeline: the expert's R-level assignments for the 25 implicit pairs were indeed fixed prior to the evaluation phase—she assigned levels during the typology derivation (Section 3.2) and did not reassign or adjust them during the grounding study. We will state this explicitly in the revision. However, we agree that even with this temporal separation, the deeper concern remains: the expert constructed the framework that generates the predictions she then evaluates, and at R5 specifically, the typology predicts that her domain knowledge (not the description) carries coherence, which is precisely what she reported observing. This is a real circularity risk. In the revision, we will (1) add a dedicated limitation paragraph in Section 4, (2) reframe the R5 reversal as a pattern consistent with the typology's prediction rather than as independent confirmation of it, and (3) note that a fully independent test requires expert evaluators who did not participate in typology construction, which we commit to in future work. revision: yes
-
Referee: Figure 3b and Section 4: The agreement percentages (R1=50%, R2=57%, R3=40%, R4=29%, R5=25%) are described as showing that disagreement 'increases as the level rises,' but R2 (57%) exceeds R1 (50%), which is not monotonically decreasing. With approximately 5 pairs per level and no statistical significance test reported, it is unclear whether the observed trend differs from chance. A simple test (e.g., Cochran-Armitage trend test or a permutation test) would clarify whether the R-level genuinely predicts agreement patterns. Without it, the central predictive claim is unsupported by standard statistical evidence. The R5 reversal, resting on roughly 5 pairs, is especially fragile.
Authors: The referee is correct on both points. First, the trend is not monotonically decreasing—R2 (57%) exceeds R1 (50%)—and our language characterizing it as 'increasing disagreement' is imprecise. We will revise the description to acknowledge the non-monotonicity at R1–R2 and frame the pattern as a general downward trend with a local reversal at R2, rather than a strict monotonic decrease. Second, we will conduct a permutation test: treating the 25 pairs as fixed with their observed agreement outcomes, we will randomly permute R-level labels across pairs 10,000 times and compute the correlation between R-level and agreement rate for each permutation, building a null distribution against which we test the observed correlation. We will report the p-value and effect size. We expect this test to clarify whether the observed trend is distinguishable from chance given the small sample. We acknowledge upfront that with approximately 5 pairs per level, the test will have limited power, and we will state this explicitly. If the result is not significant at conventional thresholds, we will reframe the claim as an observed pattern that is consistent with the typology's prediction but not yet statistically confirmed, and we will note that a larger-scale study is needed to draw firm conclusions. The R5 reversal claim will be further tempered in light of both the small sample and the non-independence concern raised in the first comment. revision: yes
-
Referee: Section 3.2: The claim that the five levels are 'domain-general' while the decision-point protocol is 'instantiated for TBI figure conventions' is internally consistent but empirically untested. The manuscript acknowledges that transferring the protocol requires re-specifying decision points for another field, yet the abstract and Section 5 frame the typology as a general contribution. The domain-generality claim should be either supported by at least a pilot application in a second domain or explicitly scoped as a hypothesis pending cross-domain validation.
Authors: The referee is correct that the domain-generality claim is empirically untested and that the abstract and Section 5 overstate the current evidence. We do not have a pilot application in a second domain to offer at this time. We will therefore scope the claim explicitly as a hypothesis pending cross-domain validation. Specifically, we will (1) revise the abstract to replace language implying established generality with language framing the five levels as proposed to be domain-general, (2) add a sentence in Section 3.2 clarifying that the levels are theoretically motivated as domain-general (because they track the inferential structure of chart-image pairs, which is not field-specific) but that this has not been empirically tested beyond TBI, and (3) add cross-domain validation as an explicit next step in Section 5. We believe the theoretical argument for domain-generality is reasonable—the progression from shared data through external projection to reader-supplied framing is not specific to neuroscience—but we agree that this is a hypothesis, not a demonstrated result, and the revision will reflect that distinction. revision: yes
- The R5 reversal rests on approximately 5 pairs and cannot be statistically strengthened without collecting additional data. We will temper the claim and report the permutation test, but we cannot make the evidence stronger than it is.
Circularity Check
The typology's central predictive claim is partially circular: the expert who co-derived the R1–R5 levels from explicit pairs is the sole expert evaluator in the grounding study, and the 'falsifiable prediction' that VLM performance decreases with R-level was confirmed by iteratively refining prompts against the same expert's labels.
specific steps
-
fitted input called prediction
[Section 3.3 (learnability cross-check) and Section 4 (grounding study)]
"Treating the expert's R-level assignments as a reference, we prompted a frontier vision-language model to label 32 explicit pairs across four iterative rounds, refining the prompt after each round from the disagreement pattern with the expert. Zero-shot agreement was low (11/32, 34%), confirming that the levels cannot be recovered from the panels alone. Successive refinements raised agreement but plateaued; the largest gain came from adding the source paper's captions to the prompt (21/32, 65.6%). This is not validation of the typology, which was derived through the process above; it is a下游信号…"
The paper claims the 'falsifiable prediction' that VLM performance decreases with R-level is confirmed by the learnability cross-check. But the cross-check fits the VLM prompt iteratively against the expert's R-level assignments (the same assignments used to derive the typology). The 'prediction' that performance decreases with R-level is thus confirmed by construction: the prompt was refined until agreement improved, and the improvement is attributed to adding captions — exactly the contextual material the typology says is needed. The paper acknowledges this is 'not validation,' but then presents it as support for the prediction in the grounding study.
-
self definitional
[Section 3.1–3.2 (typology derivation) and Section 4 (expert assessment)]
"The typology emerged from sustained collaboration through a six-month, multi-step methodology (Figure 1) with a neuroscience expert… We derived the typology through a structured cognitive walkthrough… Working with the expert, three visualization-researcher co-authors read each chart-image pair and, for each pair, recorded the reasoning connecting the two panels… As an evaluation, we then ran a grounding study on 25 implicit pairs… with the expert and three non-experts."
The expert who co-derived the typology (assigning R-levels, defining decision points, fitting level boundaries to her own reasoning) is the same person who serves as sole expert evaluator in the grounding study. Her judgments of whether VLM descriptions 'ground' each pair's relational claim are not independent of the typology she constructed. At R5 specifically, the typology predicts the expert will fill gaps with her own knowledge — and as the person who defined R5 as 'context-dependent,' her evaluation behavior is shaped by the framework itself. The R5 reversal (expert more lenient than non-experts) rests on ~5 pairs judged by the same person who defined what R5 means.
full rationale
The paper has genuine independent content: the 25 implicit pairs used in the grounding study are distinct from the 32 explicit pairs used in derivation, and the non-expert judgments are independent of the typology. However, the expert's dual role as co-deriver and sole evaluator creates non-independence for the most distinctive claim (the R5 reversal), and the learnability cross-check is a fitted-input-called-prediction pattern where prompts were iteratively refined against the expert's labels. The paper is transparent about some of these limitations ('This is not validation of the typology'), which mitigates the circularity somewhat. Score 4 reflects partial circularity: the central claim has independent empirical content from non-experts, but the expert-side evaluation and the VLM prediction are partially constructed from the typology's own inputs.
Axiom & Free-Parameter Ledger
free parameters (3)
- Number of typology levels (5) =
5
- Decision-point protocol thresholds =
not specified
- VLM prompt iterations (4 rounds) =
4
axioms (4)
- domain assumption Clark's grounding theory of communication applies to chart-image pairs in scientific figures.
- domain assumption The chart-image pair (with caption) is the correct unit of analysis for multimodal scientific communication.
- ad hoc to paper The R-levels are domain-general despite being derived from TBI only.
- domain assumption Three CS/DS doctoral students are sufficient non-expert controls for visualization literacy.
invented entities (1)
-
R1–R5 reasoning gap levels
independent evidence
read the original abstract
Charts and images appear together throughout scientific publications, yet most computational work does not characterize their coherence. We argue that a chart, its accompanying image, and the caption that links them form a multimodal unit, and that the inferential work required to read it varies systematically. To capture this variation, we develop a typology of reasoning gaps, R1 through R5, that characterizes how chart, image, and text jointly convey a scientific claim, and the interpretive work this demands of the reader. Some pairs restate the same data, while in other pairs, charts are used to quantify a structure the image localizes, project image content onto an external variable, audit an image-based claim, or jointly construct a frame that neither panel can establish alone. The typology is anchored in the grounding theory of communication and was derived bottom-up, with a neuroscience expert, from a corpus of 79 traumatic brain injury papers and 32 chart-image pairs. Crucially, the levels provide a systematic mechanism for identifying where grounding succeeds or breaks down, rather than leaving it to subjective inference. We show this in a study in which a domain expert and three non-experts judge vision-language model (VLM) descriptions of 25 pairs: the level predicts where their judgments align and where they diverge, isolating the points at which contextual knowledge, not the figure, carries coherence. This typology thus offers figure designers a systematic way to balance text against chart-image pairs, bridging the expert-to-non-expert divide in reading a scientific takeaway.
Figures
Reference graph
Works this paper leans on
-
[1]
V . Bonnelle, R. Leech, K. M. Kinnunen, T. E. Ham, C. F. Beckmann, X. De Boissezon, R. J. Greenwood, and D. J. Sharp. Default mode net- work connectivity predicts sustained attention deficits after traumatic brain injury.Journal of Neuroscience, 31(38):13442–13451, 2011. 3
work page 2011
- [2]
-
[3]
H. H. Clark.Using Language. Cambridge University Press, 1996. 1, 2
work page 1996
-
[4]
H. H. Clark and S. E. Brennan. Grounding in communication. In L. B. Resnick, J. M. Levine, and S. D. Teasley, eds.,Perspectives on Socially Shared Cognition, pp. 127–149. American Psychological As- sociation, 1991. 1, 2
work page 1991
-
[5]
H. H. Clark and T. B. Carlson. Hearers and speech acts.Language, 58(2):332–373, 1982. doi: 10.2307/414102 2
- [6]
-
[7]
S. L. Franconeri, L. M. Padilla, P. Shah, J. M. Zacks, and J. Hull- man. The science of visual data communication: What works.Psy- chological Science in the Public Interest, 22(3):110–161, 2021. doi: 10.1177/15291006211051956 2
-
[8]
H. P. Grice. Logic and conversation. In P. Cole and J. L. Morgan, eds.,Syntax and Semantics, Vol. 3: Speech Acts, pp. 41–58. Academic Press, 1975. 2
work page 1975
- [9]
-
[10]
J. Hullman and N. Diakopoulos. Visualization rhetoric: Framing effects in narrative visualization.IEEE Transactions on Visualiza- tion and Computer Graphics, 17(12):2231–2240, 2011. doi: 10.1109/ TVCG.2011.255 2
work page 2011
-
[11]
M. S. Islam, R. Rahman, A. Masry, M. T. R. Laskar, M. T. Nayeem, and E. Hoque. Are large vision language models up to the chal- lenge of chart comprehension and reasoning? an extensive investi- gation into the capabilities and limitations of LVLMs.arXiv preprint arXiv:2406.00257, 2024. doi: 10.48550/arXiv.2406.00257 1
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2406.00257 2024
-
[12]
S. E. Kahou, V . Michalski, A. Atkinson, ´A. K´ad´ar, A. Trischler, and Y . Bengio. Figureqa: An annotated figure dataset for visual reasoning. arXiv preprint arXiv:1710.07300, 2017. 2
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[13]
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi. A diagram is worth a dozen images. InEuropean Con- ference on Computer Vision (ECCV), pp. 235–251. Springer, 2016. 1, 2
work page 2016
-
[14]
D. H. Kim, E. Hoque, J. Kim, and M. Agrawala. Facilitating document reading by linking text and tables. InProceedings of the 31st Annual ACM Symposium on User Interface Software and Technology (UIST), pp. 423–434, 2018. doi: 10.1145/3242587.3242617 2
-
[15]
S. M. Kosslyn. Understanding charts and graphs.Applied Cognitive Psychology, 3(3):185–225, 1989. doi: 10.1002/acp.2350030302 2
-
[16]
Z. Li, X. Yang, K. Choi, W. Zhu, R. Hsieh, H. Kim, J. H. Lim, S. Ji, B. Lee, X. Yan, L. R. Petzold, S. D. Wilson, W. Lim, and W. Y . Wang. MMSci: A dataset for graduate-level multi-discipline multimodal sci- entific understanding.arXiv preprint arXiv:2407.04903, 2024. doi: 10 .48550/arXiv.2407.04903 2
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[17]
A. Lundgard and A. Satyanarayan. Accessible visualization via nat- ural language descriptions: A four-level model of semantic con- tent.IEEE Transactions on Visualization and Computer Graphics, 28(1):1073–1083, 2021. doi: 10.1109/TVCG.2021.3114770 1, 4
- [18]
-
[19]
R. Raizman, I. Tavor, A. Biegon, S. Harnof, C. Hoffmann, G. Tsar- faty, E. Fruchter, L. Tatsa-Laur, M. Weiser, and A. Livny. Traumatic brain injury severity in a network perspective: A diffusion mri based connectome study.Scientific reports, 10(1):9121, 2020. 3
work page 2020
-
[20]
S. Sen, A. Nakarmi, X. Song, and A. Dasgupta. Pointers at uzh shared task 2026: Reasoning probes for argumentation mining in un resolu- tions. In13th Workshop on Argument Mining and Reasoning, p. 145,
work page 2026
- [21]
-
[22]
M. S. F. Sorg, L. Delano-Wood, M. N. Luc, D. M. Schiehser, K. L. Hanson, D. A. Nation, M. E. Lanni, A. J. Jak, K. Lu, M. Meloy, et al. White matter integrity in veterans with mild traumatic brain injury: associations with executive function and loss of consciousness.The Journal of head trauma rehabilitation, 29(1):21, 2014. 3
work page 2014
-
[23]
S. Vaidya and A. Dasgupta. Knowing what to look for: A fact-evidence reasoning framework for decoding communicative visu- alization. In2020 IEEE Visualization Conference (VIS), pp. 231–235. IEEE, 2020. 2
work page 2020
-
[24]
Z. Wang, M. Xia, L. He, H. Chen, Y . Liu, R. Zhu, K. Liang, X. Wu, H. Liu, S. Malladi, et al. Charxiv: Charting gaps in realistic chart understanding in multimodal llms. vol. 37, pp. 113569–113697, 2024. 2
work page 2024
-
[25]
Q. Zhi, A. Ottley, and R. Metoyer. Linking and layout: Exploring the integration of text and visualization in storytelling. 38(3):675–685,
-
[26]
J. Zong, C. Lee, A. Lundgard, J. Jang, D. Hajas, and A. Satya- narayan. Rich screen reader experiences for accessible data visual- ization. 41(3):15–27, 2022. 1 5
work page 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.