Pith. sign in

REVIEW 2 cited by

Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11229 v2 pith:B5LR5SEI submitted 2024-07-15 cs.CL cs.AIcs.CVcs.HCcs.LG

classification cs.CLcs.AIcs.CVcs.HCcs.LG
keywords chartmodelsquestioncurrentrobustnessvisualvlmsconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Chart question answering (CQA) is a crucial area of Visual Language Understanding. However, the robustness and consistency of current Visual Language Models (VLMs) in this field remain under-explored. This paper evaluates state-of-the-art VLMs on comprehensive datasets, developed specifically for this study, encompassing diverse question categories and chart formats. We investigate two key aspects: 1) the models' ability to handle varying levels of chart and question complexity, and 2) their robustness across different visual representations of the same underlying data. Our analysis reveals significant performance variations based on question and chart types, highlighting both strengths and weaknesses of current models. Additionally, we identify areas for improvement and propose future research directions to build more robust and reliable CQA systems. This study sheds light on the limitations of current models and paves the way for future advancements in the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

    cs.LG 2026-08 conditional novelty 6.0 of 10

    An edit-based reliability check for vision-language models can only miss errors that commute with the edit's answer transform, yielding a closed-form design rule for edit suites.

  2. ViStruct: Simulating Expert-Like Reasoning Through Task Decomposition and Visual Attention Cues

    cs.HC 2025-06 conditional novelty 5.0 of 10

    ViStruct automatically breaks visualization questions into ordered subtasks tied to highlighted chart regions, imitating expert analysis strategies for chart reading.

Pith tools