Pith. sign in

REVIEW 42 cited by

ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.16483 v1 pith:KNUN3WD7 submitted 2023-11-27 cs.CV cs.CL

classification cs.CVcs.CL
keywords chartdatachartllamadatasetgenerationmulti-modaladditionallydatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for specific domain data, particularly when it comes to interpreting chart figures. This is mainly due to the lack of relevant multi-modal instruction tuning datasets. In this article, we create a high-quality instruction-tuning dataset leveraging GPT-4. We develop a multi-step data generation process in which different steps are responsible for generating tabular data, creating chart figures, and designing instruction tuning data separately. Our method's flexibility enables us to generate diverse, high-quality instruction-tuning data consistently and efficiently while maintaining a low resource expenditure. Additionally, it allows us to incorporate a wider variety of chart and task types not yet featured in existing datasets. Next, we introduce ChartLlama, a multi-modal large language model that we've trained using our created dataset. ChartLlama outperforms all prior methods in ChartQA, Chart-to-text, and Chart-extraction evaluation benchmarks. Additionally, ChartLlama significantly improves upon the baseline in our specially compiled chart dataset, which includes new chart and task types. The results of ChartLlama confirm the value and huge potential of our proposed data generation method in enhancing chart comprehension.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 42 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    ChartArena unifies eight chart families across three real-world visual scenarios and two languages under a format-agnostic triple/graph evaluation protocol, revealing clear gaps among 26 MLLMs.

  2. SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

    cs.AI 2026-03 conditional novelty 7.0 of 10

    SciVisAgentBench provides 108 expert-crafted tasks and a mixed LLM-plus-deterministic evaluation pipeline for benchmarking AI agents that perform scientific visualization workflows.

  3. ChartCap: Mitigating Hallucination of Dense Chart Captioning

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A new 565K-pair chart-caption dataset with schema-based dense captions and a reference-free visual consistency metric improves VLM captioning and reduces hallucination.

  4. Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

    cs.CV 2025-02 conditional novelty 7.0 of 10

    SPARC selectively and progressively reinforces attention to relevant image tokens during decoding, improving both precision and recall in detailed image captioning compared to baselines and prior hallucination-mitigat...

  5. SketchAgent: Language-Driven Sequential Sketch Generation

    cs.CV 2024-11 conditional novelty 7.0 of 10

    SketchAgent uses a multimodal LLM prompted with a numbered-grid sketching language to generate, edit, and collaboratively draw sequential vector sketches without any training.

  6. LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning

    cs.CL 2026-08 reject novelty 6.0 of 10

    LongChart is a graph-consistent multi-chart VQA benchmark where 10 multimodal LLMs lose accuracy as question reasoning hops grow.

  7. VisCanvas: A Node-Based Interface for Exploratory Visualization Authoring with LLMs

    cs.HC 2026-07 conditional novelty 6.0 of 10

    A node-based interface for LLM chart authoring led 20 users to explore in more branched, tree-like patterns than a chat interface, without raising measured workload.

  8. MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Current multimodal LLMs can copy the look of multi-view dashboards but mostly fail to bind real data and implement cross-view interactions.

  9. Visual Programmability: A Guide for Code-as-Thought in Chart Understanding

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A vision-language model learns to dynamically switch between code-based and visual reasoning for chart questions, improving average accuracy by about one point over fixed strategies.

  10. SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.

  11. BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Replotting real-world charts into code-backed images, then fine-tuning with supervised learning and GRPO reinforcement learning, produces chart QA models that beat prior chart-specific models on several benchmarks.

  12. FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark of real-world financial charts shows current vision-language models lag badly on questions that require reading values from chart axes.

  13. In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding

    cs.CL 2025-07 conditional novelty 6.0 of 10

    ChartScope, using a template-based synthetic data pipeline and dual-path reasoning training, outperforms prior chart-reading models on several advanced chart benchmarks.

  14. Multilingual Multimodal Software Developer for Code Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.

  15. VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 7-billion-parameter multimodal model fine-tuned on 2,500 expert critiques of data visualizations matches or beats much larger models at identifying visualization defects.

  16. Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approach

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A draft-and-repair agentic loop using GPT-4o-mini reduces text-to-chart execution errors to 4.5-4.6% on two benchmarks, suggesting execution is nearly solved and future work should focus on quality and accessibility.

  17. ChartLens: Fine-grained Visual Attribution in Charts

    cs.CL 2025-05 conditional novelty 6.0 of 10

    ChartLens uses segmentation and set-of-marks prompting to attribute chart-based answers to specific visual elements, and the authors release a new benchmark for evaluating such attribution.

  18. Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Multimodal LLMs underperform humans at directly rating charts' experiential impact, but they are substantially better at pairwise comparisons, especially when the human ratings differ clearly.

  19. ChartInsighter: An Approach for Mitigating Hallucination in Time-series Chart Summary Generation with A Benchmark Dataset

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A multi-agent LLM pipeline with external computation and self-consistency checking produces time-series chart summaries with fewer annotated hallucinations than GPT-4 or VL2NL on the authors' new benchmark.

  20. ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation

    cs.AI 2025-01 conditional novelty 6.0 of 10

    ChartCoder, a 7B multimodal LLM with a code-LLM backbone trained on 160k synthetic chart-code pairs, surpasses previous open-source models at converting chart images into executable plotting code.

  21. ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Visual editing of input images as a chain of thought improves multimodal LLM accuracy on structured image tasks by 3 to 12 points.

  22. Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A new multimodal reasoning benchmark shows that state-of-the-art AI models lag human experts by more than 30 percentage points, with visual reasoning errors as the main bottleneck.

  23. MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MMFactory automatically generates and benchmarks a pool of reusable programmatic vision-language solutions from a few examples, letting users pick one that fits their accuracy and speed constraints.

  24. SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A synthetic figure generation pipeline and a one-million-image dataset that improve pre-training for figure question answering.

  25. Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new benchmark built from real scientific paper charts, including flowcharts and context-dependent questions, shows large multimodal models perform far below human level on chart understanding.

  26. Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Aggregating a VLM's attention over visual tokens and mapping it back to image patches produces saliency maps that, in a 13-sample deletion test on ChartGemma, appear to localize the visual evidence behind chart answers.

  27. ChatImage: Navigating Long-Form LLM Answers through Interactive Images

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ChatImage renders LLM answers as images, then uses visual grounding to place clickable hotspots on rendered regions for interactive follow-up.

  28. VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    Merging a coding LLM into a vision-language model via task vectors yields an open-source multimodal coder that reaches near-GPT-4o performance on the authors' new benchmark.

  29. CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A fine-tuned Llama 3 model with a trained document critic and program-based reasoning beats GPT-4o and other baselines on a new carbon footprint QA benchmark.

  30. Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Injecting self-verified bounding boxes into Chain-of-Thought data improves few-shot adaptation of multimodal LLMs on charts, tables, receipts, and reports.

  31. ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ChartMind is a new bilingual chart QA benchmark, and ChartLLM's structured context extraction yields higher scores than three existing prompting paradigms in the paper's evaluations.

  32. CHAOS: Chart Analysis with Outlier Samples

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A chart perturbation robustness benchmark with five textual and ten visual distortion types, three human-calibrated severity levels, and evaluations of 13 MLLMs on ChartQA and chart summarization.

  33. PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents

    cs.IR 2025-01 conditional novelty 5.0 of 10

    PlotEdit, a self-reflective multi-agent LLM pipeline, claims state-of-the-art results for natural-language chart editing on ChartCraft across style, layout, format, and data edits.

  34. ChartAdapter: Large Vision-Language Model for Chart Summarization

    cs.MM 2024-12 reject novelty 5.0 of 10

    ChartAdapter, a learnable-query cross-modal adapter, reportedly beats prior chart-summarization models on Chart-to-Text, but its evaluation may be contaminated by training on the test benchmark's source data.

  35. Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant

    cs.LG 2025-08 unverdicted novelty 4.0 of 10

    The preprint's abstract claims a sparse softmax variant that masks non-competitive classes and accelerates training, but the provided body contains an unrelated chart-captioning paper and none of the claimed method.

  36. Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

    cs.CV 2025-07 conditional novelty 4.0 of 10

    On the SciVQA 2025 benchmark, an InternVL3 model with optimized prompts and chain-of-thought instructions reaches ROUGE-1 and ROUGE-L F1 of 0.740, and a figure-type-aware ensemble ranks 5th.

  37. Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models

    cs.CL 2025-07 accept novelty 4.0 of 10

    An ensemble of InternVL3-78B and Pixtral-Large with retrievable few-shot examples and a confidence threshold beats individual models on scientific figure question answering.

  38. Evaluating LLMs for Visualization Generation and Understanding

    cs.HC 2025-06 conditional novelty 4.0 of 10

    GPT-4o led on most chart generation and understanding tests in this sample, but all four models struggled with complex charts, dotted lines, and close bar lengths.

  39. ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding

    cs.CV 2025-05 conditional novelty 4.0 of 10

    ChartSketcher has a multimodal LLM sketch intermediate reasoning steps directly on chart images and feed those sketches back as visual feedback, improving chart QA accuracy over its base model.

  40. Graph-Based Multimodal Contrastive Learning for Chart Question Answering

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A multimodal scene graph with contrastive learning, injected as a decoder soft prompt, gives small ChartQA gains, but ablations show the visual graph alone often matches or beats it.

  41. Text2Insight: Transform natural language text into insights seamlessly using multi-model architecture

    cs.AI 2024-12 reject novelty 3.0 of 10

    Text2Insight combines an LLM text-to-SQL step with a rule-based chart predictor and BERT-based question answering and prediction, but its end-to-end performance claims rest on circular or missing evaluation.

  42. Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era

    cs.CV 2024-11 unverdicted novelty 3.0 of 10

    A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.

Pith tools