REVIEW 42 cited by
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for specific domain data, particularly when it comes to interpreting chart figures. This is mainly due to the lack of relevant multi-modal instruction tuning datasets. In this article, we create a high-quality instruction-tuning dataset leveraging GPT-4. We develop a multi-step data generation process in which different steps are responsible for generating tabular data, creating chart figures, and designing instruction tuning data separately. Our method's flexibility enables us to generate diverse, high-quality instruction-tuning data consistently and efficiently while maintaining a low resource expenditure. Additionally, it allows us to incorporate a wider variety of chart and task types not yet featured in existing datasets. Next, we introduce ChartLlama, a multi-modal large language model that we've trained using our created dataset. ChartLlama outperforms all prior methods in ChartQA, Chart-to-text, and Chart-extraction evaluation benchmarks. Additionally, ChartLlama significantly improves upon the baseline in our specially compiled chart dataset, which includes new chart and task types. The results of ChartLlama confirm the value and huge potential of our proposed data generation method in enhancing chart comprehension.
Forward citations
Cited by 42 Pith papers
-
ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
ChartArena unifies eight chart families across three real-world visual scenarios and two languages under a format-agnostic triple/graph evaluation protocol, revealing clear gaps among 26 MLLMs.
-
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
SciVisAgentBench provides 108 expert-crafted tasks and a mixed LLM-plus-deterministic evaluation pipeline for benchmarking AI agents that perform scientific visualization workflows.
-
ChartCap: Mitigating Hallucination of Dense Chart Captioning
A new 565K-pair chart-caption dataset with schema-based dense captions and a reference-free visual consistency metric improves VLM captioning and reduces hallucination.
-
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
SPARC selectively and progressively reinforces attention to relevant image tokens during decoding, improving both precision and recall in detailed image captioning compared to baselines and prior hallucination-mitigat...
-
SketchAgent: Language-Driven Sequential Sketch Generation
SketchAgent uses a multimodal LLM prompted with a numbered-grid sketching language to generate, edit, and collaboratively draw sequential vector sketches without any training.
-
LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning
LongChart is a graph-consistent multi-chart VQA benchmark where 10 multimodal LLMs lose accuracy as question reasoning hops grow.
-
VisCanvas: A Node-Based Interface for Exploratory Visualization Authoring with LLMs
A node-based interface for LLM chart authoring led 20 users to explore in more branched, tree-like patterns than a chat interface, without raising measured workload.
-
MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction
Current multimodal LLMs can copy the look of multi-view dashboards but mostly fail to bind real data and implement cross-view interactions.
-
Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
A vision-language model learns to dynamically switch between code-based and visual reasoning for chart questions, improving average accuracy by about one point over fixed strategies.
-
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.
-
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
Replotting real-world charts into code-backed images, then fine-tuning with supervised learning and GRPO reinforcement learning, produces chart QA models that beat prior chart-specific models on several benchmarks.
-
FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models
A new benchmark of real-world financial charts shows current vision-language models lag badly on questions that require reading values from chart axes.
-
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
ChartScope, using a template-based synthetic data pipeline and dual-path reasoning training, outperforms prior chart-reading models on several advanced chart benchmarks.
-
Multilingual Multimodal Software Developer for Code Generation
A 7B vision-language model trained on synthetic diagram-to-code data outperforms several larger open-weight models on a new 10-language UML/flowchart code-generation benchmark.
-
VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
A 7-billion-parameter multimodal model fine-tuned on 2,500 expert critiques of data visualizations matches or beats much larger models at identifying visualization defects.
-
Does It Run and Is That Enough? Revisiting Text-to-Chart Generation with a Multi-Agent Approach
A draft-and-repair agentic loop using GPT-4o-mini reduces text-to-chart execution errors to 4.5-4.6% on two benchmarks, suggesting execution is nearly solved and future work should focus on quality and accessibility.
-
ChartLens: Fine-grained Visual Attribution in Charts
ChartLens uses segmentation and set-of-marks prompting to attribute chart-based answers to specific visual elements, and the authors release a new benchmark for evaluating such attribution.
-
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
Multimodal LLMs underperform humans at directly rating charts' experiential impact, but they are substantially better at pairwise comparisons, especially when the human ratings differ clearly.
-
ChartInsighter: An Approach for Mitigating Hallucination in Time-series Chart Summary Generation with A Benchmark Dataset
A multi-agent LLM pipeline with external computation and self-consistency checking produces time-series chart summaries with fewer annotated hallucinations than GPT-4 or VL2NL on the authors' new benchmark.
-
ChartCoder: Advancing Multimodal Large Language Model for Chart-to-Code Generation
ChartCoder, a 7B multimodal LLM with a code-LLM backbone trained on 160k synthetic chart-code pairs, surpasses previous open-source models at converting chart images into executable plotting code.
-
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Visual editing of input images as a chain of thought improves multimodal LLM accuracy on structured image tasks by 3 to 12 points.
-
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
A new multimodal reasoning benchmark shows that state-of-the-art AI models lag human experts by more than 30 percentage points, with visual reasoning errors as the main bottleneck.
-
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
MMFactory automatically generates and benchmarks a pool of reusable programmatic vision-language solutions from a few examples, letting users pick one that fits their accuracy and speed constraints.
-
SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images
A synthetic figure generation pipeline and a one-million-image dataset that improve pre-training for figure question answering.
-
Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature
A new benchmark built from real scientific paper charts, including flowcharts and context-dependent questions, shows large multimodal models perform far below human level on chart understanding.
-
Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs
Aggregating a VLM's attention over visual tokens and mapping it back to image patches produces saliency maps that, in a 13-sample deletion test on ChartGemma, appear to localize the visual evidence behind chart answers.
-
ChatImage: Navigating Long-Form LLM Answers through Interactive Images
ChatImage renders LLM answers as images, then uses visual grounding to place clickable hotspots on rendered regions for interactive follow-up.
-
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
Merging a coding LLM into a vision-language model via task vectors yields an open-source multimodal coder that reaches near-GPT-4o performance on the authors' new benchmark.
-
CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
A fine-tuned Llama 3 model with a trained document critic and program-based reasoning beats GPT-4o and other baselines on a new carbon footprint QA benchmark.
-
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
Injecting self-verified bounding boxes into Chain-of-Thought data improves few-shot adaptation of multimodal LLMs on charts, tables, receipts, and reports.
-
ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering
ChartMind is a new bilingual chart QA benchmark, and ChartLLM's structured context extraction yields higher scores than three existing prompting paradigms in the paper's evaluations.
-
CHAOS: Chart Analysis with Outlier Samples
A chart perturbation robustness benchmark with five textual and ten visual distortion types, three human-calibrated severity levels, and evaluations of 13 MLLMs on ChartQA and chart summarization.
-
PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
PlotEdit, a self-reflective multi-agent LLM pipeline, claims state-of-the-art results for natural-language chart editing on ChartCraft across style, layout, format, and data edits.
-
ChartAdapter: Large Vision-Language Model for Chart Summarization
ChartAdapter, a learnable-query cross-modal adapter, reportedly beats prior chart-summarization models on Chart-to-Text, but its evaluation may be contaminated by training on the test benchmark's source data.
-
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
The preprint's abstract claims a sparse softmax variant that masks non-competitive classes and accelerates training, but the provided body contains an unrelated chart-captioning paper and none of the claimed method.
-
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
On the SciVQA 2025 benchmark, an InternVL3 model with optimized prompts and chain-of-thought instructions reaches ROUGE-1 and ROUGE-L F1 of 0.740, and a figure-type-aware ensemble ranks 5th.
-
Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models
An ensemble of InternVL3-78B and Pixtral-Large with retrievable few-shot examples and a confidence threshold beats individual models on scientific figure question answering.
-
Evaluating LLMs for Visualization Generation and Understanding
GPT-4o led on most chart generation and understanding tests in this sample, but all four models struggled with complex charts, dotted lines, and close bar lengths.
-
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
ChartSketcher has a multimodal LLM sketch intermediate reasoning steps directly on chart images and feed those sketches back as visual feedback, improving chart QA accuracy over its base model.
-
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
A multimodal scene graph with contrastive learning, injected as a decoder soft prompt, gives small ChartQA gains, but ablations show the visual graph alone often matches or beats it.
-
Text2Insight: Transform natural language text into insights seamlessly using multi-model architecture
Text2Insight combines an LLM text-to-SQL step with a rule-based chart predictor and BERT-based question answering and prediction, but its end-to-end performance claims rest on circular or missing evaluation.
-
Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era
A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.
Discussion (0). Continue with ORCID to comment.