Pith. sign in

REVIEW 20 cited by

UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14761 v3 pith:6J7AOHSZ submitted 2023-05-24 cs.CL

classification cs.CL
keywords taskschartdatachartsmodelreasoningelementslanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Charts are very popular for analyzing data, visualizing key insights and answering complex reasoning questions about data. To facilitate chart-based data analysis using natural language, several downstream tasks have been introduced recently such as chart question answering and chart summarization. However, most of the methods that solve these tasks use pretraining on language or vision-language tasks that do not attempt to explicitly model the structure of the charts (e.g., how data is visually encoded and how chart elements are related to each other). To address this, we first build a large corpus of charts covering a wide variety of topics and visual styles. We then present UniChart, a pretrained model for chart comprehension and reasoning. UniChart encodes the relevant text, data, and visual elements of charts and then uses a chart-grounded text decoder to generate the expected output in natural language. We propose several chart-specific pretraining tasks that include: (i) low-level tasks to extract the visual elements (e.g., bars, lines) and data from charts, and (ii) high-level tasks to acquire chart understanding and reasoning skills. We find that pretraining the model on a large corpus with chart-specific low- and high-level tasks followed by finetuning on three down-streaming tasks results in state-of-the-art performance on three downstream tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A JEPA encoder finetuned on synthetic bar charts enables a lightweight decoder to recover bar values from chart images, but the method remains behind state-of-the-art supervised systems.

  2. FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    FinReportBench is a fine-grained, expert-grounded benchmark for institution-grade LLM financial report generation, and its skill-evolution method improves G1 and G2 scores across model families.

  3. Visual Programmability: A Guide for Code-as-Thought in Chart Understanding

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A vision-language model learns to dynamically switch between code-based and visual reasoning for chart questions, improving average accuracy by about one point over fixed strategies.

  4. FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark of real-world financial charts shows current vision-language models lag badly on questions that require reading values from chart axes.

  5. In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

    cs.CL 2025-07 conditional novelty 6.0 of 10

    ChartScope, using a template-based synthetic data pipeline and dual-path reasoning training, outperforms prior chart-reading models on several advanced chart benchmarks.

  6. CoMemo: LVLMs Need Image Context with Image Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CoMemo adds a cross-attention image-memory path and thumbnail-anchored position encoding to reduce visual neglect in long-context and multi-image LVLM tasks.

  7. VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VisTA uses GRPO reinforcement learning to train a vision-language agent to select external visual tools for a frozen reasoning model, improving accuracy on ChartQA, Geometry3K, BlindTest, and MathVerse.

  8. ChartLens: Fine-grained Visual Attribution in Charts

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ChartLens uses segmentation and set-of-marks prompting to attribute chart-based answers to specific visual elements, and the authors release a new benchmark for evaluating such attribution.

  9. Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023

    cs.CL 2025-01 conditional novelty 6.0 of 10

    In human evaluations by three professional editors, GPT-4V captions for scientific figures were preferred over author-written captions and over captions from challenge-winning models.

  10. Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Eagle2-9B matches or outperforms much larger vision-language models on many benchmarks through a carefully constructed post-training data strategy.

  11. ChartInsighter: An Approach for Mitigating Hallucination in Time-series Chart Summary Generation with A Benchmark Dataset

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A multi-agent LLM pipeline with external computation and self-consistency checking produces time-series chart summaries with fewer annotated hallucinations than GPT-4 or VL2NL on the authors' new benchmark.

  12. ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

    cs.AI 2025-01 conditional novelty 6.0 of 10

    ChartCoder, a 7B multimodal LLM with a code-LLM backbone trained on 160k synthetic chart-code pairs, surpasses previous open-source models at converting chart images into executable plotting code.

  13. ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Visual editing of input images as a chain of thought improves multimodal LLM accuracy on structured image tasks by 3 to 12 points.

  14. HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HoVLE is a monolithic VLM whose holistic embedding module maps images and text into one shared space, letting a frozen LLM reach near-compositional performance.

  15. PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PVC unifies image and video token compression in VLMs by repeating images as static videos and using causal temporal attention with adaptive compression, achieving strong benchmark results at 64 tokens per frame.

  16. Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning

    cs.CV 2024-12 reject novelty 6.0 of 10

    Geo-LLaVA combines retrieval-augmented fine-tuning with in-context learning, reporting 65.25% and 42.36% on selected subsets of GeoQA and the new GeoMath dataset.

  17. Chimera: Improving Generalist Model with Domain-Specific Experts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Chimera fuses frozen domain-expert encoders into a generalist multimodal LLM via routing and a 30% masking of general tokens, lifting InternVL2-8B from 61.6 to 64.9 on MathVista and from 31.3 to 32.4 on MathVerse.

  18. SigLIP-HD by Fine-to-Coarse Supervision

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Fine-to-coarse L1 supervision lets a standard-resolution SigLIP 2 encoder produce better visual tokens for MLLMs without higher-resolution inference.

  19. Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

    cs.CV 2025-07 conditional novelty 4.0 of 10

    On the SciVQA 2025 benchmark, an InternVL3 model with optimized prompts and chain-of-thought instructions reaches ROUGE-1 and ROUGE-L F1 of 0.740, and a figure-type-aware ensemble ranks 5th.

  20. CLaSP: Learning Concepts for Time-Series Signals from Natural Language Supervision

    cs.CL 2024-11 conditional novelty 4.0 of 10

    CLaSP trains separate encoders for signals and text with a contrastive loss so that a natural language query can retrieve matching time-series signals without any predefined synonym dictionary.

Pith tools