Pith. sign in

REVIEW 21 cited by

CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18521 v1 pith:LDYST4ZS submitted 2024-06-26 cs.CL cs.CV

classification cs.CLcs.CV
keywords chartquestionscharxivchartsmodelsunderstandingachieveselements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an over-optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions can deteriorate performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress. Project page and leaderboard: https://charxiv.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Exam for Active Observers

    cs.CV 2026-07 conditional novelty 7.0 of 10

    On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.

  2. ChartCap: Mitigating Hallucination of Dense Chart Captioning

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A new 565K-pair chart-caption dataset with schema-based dense captions and a reference-free visual consistency metric improves VLM captioning and reduces hallucination.

  3. Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers

    cs.CL 2025-07 conditional novelty 7.0 of 10

    MISS-QA, a new benchmark for information-seeking QA over schematic diagrams, shows the best open-source multimodal model at 61.6% accuracy versus 89.0% for human experts.

  4. SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification

    cs.CL 2025-06 conditional novelty 7.0 of 10

    SciVer is the first benchmark for multimodal scientific claim verification over full paper context, and current foundation models score about 16 points below expert humans.

  5. Chartography: A Benchmark for Professional Chart Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A 100-task expert-authored benchmark of professional charts resists saturation: the best frontier model scores 45.0%, with failures concentrated in visual perception and domain conventions.

  6. Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Replacing returned images with a fixed text placeholder in tool-augmented visual reasoning preserves benchmark accuracy, suggesting the tool-call text, not the returned pixels, carries the gain.

  7. Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    RLVR training on 64,000 procedurally generated Trace instances improves Qwen2.5-VL macro-average on 24 external visual reasoning benchmarks by 3.51 points at 3B and 4.06 points at 7B.

  8. In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

    cs.CL 2025-07 conditional novelty 6.0 of 10

    ChartScope, using a template-based synthetic data pipeline and dual-path reasoning training, outperforms prior chart-reading models on several advanced chart benchmarks.

  9. Multilingual Multimodal Software Developer for Code Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.

  10. Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approach

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A draft-and-repair agentic loop using GPT-4o-mini reduces text-to-chart execution errors to 4.5-4.6% on two benchmarks, suggesting execution is nearly solved and future work should focus on quality and accessibility.

  11. ChartLens: Fine-grained Visual Attribution in Charts

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ChartLens uses segmentation and set-of-marks prompting to attribute chart-based answers to specific visual elements, and the authors release a new benchmark for evaluating such attribution.

  12. ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

    cs.AI 2025-01 conditional novelty 6.0 of 10

    ChartCoder, a 7B multimodal LLM with a code-LLM backbone trained on 160k synthetic chart-code pairs, surpasses previous open-source models at converting chart images into executable plotting code.

  13. ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Visual editing of input images as a chain of thought improves multimodal LLM accuracy on structured image tasks by 3 to 12 points.

  14. Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new multimodal reasoning benchmark shows that state-of-the-art AI models lag human experts by more than 30 percentage points, with visual reasoning errors as the main bottleneck.

  15. Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new benchmark built from real scientific paper charts, including flowcharts and context-dependent questions, shows large multimodal models perform far below human level on chart understanding.

  16. AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A new large benchmark of AI-generated videos with human ratings across four dimensions, plus an LMM-based video quality assessor that outperforms prior metrics.

  17. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

    cs.AI 2026-07 conditional novelty 5.5 of 10

    A real-device-centric foundation GUI agent with hybrid GUI+CLI batched actions, AutoResearch data flywheel, online RL, and a proactive harness reaches SOTA mobile and competitive desktop/web scores.

  18. VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism

    cs.CV 2025-06 conditional novelty 5.0 of 10

    VReST combines Monte Carlo tree search with a self-reward signal inside a vision-language model to get higher accuracy than CoT, ToT, or voting baselines on MathVista, MathVision, and CharXiv, while spending several t...

  19. ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A human-scored benchmark shows that even GPT-4o averages below 4/5 correctness and all tested models struggle with scientific diagram prompts that combine spatial, numeric, and attribute requirements.

  20. Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    The preprint's abstract claims a sparse softmax variant that masks non-competitive classes and accelerates training, but the provided body contains an unrelated chart-captioning paper and none of the claimed method.

  21. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.

Pith tools