Pith. sign in

REVIEW 2 cited by

Visualization Literacy of Multimodal Large Language Models: A Comparative Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10996 v1 pith:BOSNESAF submitted 2024-06-24 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords visualizationmllmslanguageliteracylargemodelsmultimodalbeen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The recent introduction of multimodal large language models (MLLMs) combine the inherent power of large language models (LLMs) with the renewed capabilities to reason about the multimodal context. The potential usage scenarios for MLLMs significantly outpace their text-only counterparts. Many recent works in visualization have demonstrated MLLMs' capability to understand and interpret visualization results and explain the content of the visualization to users in natural language. In the machine learning community, the general vision capabilities of MLLMs have been evaluated and tested through various visual understanding benchmarks. However, the ability of MLLMs to accomplish specific visualization tasks based on visual perception has not been properly explored and evaluated, particularly, from a visualization-centric perspective. In this work, we aim to fill the gap by utilizing the concept of visualization literacy to evaluate MLLMs. We assess MLLMs' performance over two popular visualization literacy evaluation datasets (VLAT and mini-VLAT). Under the framework of visualization literacy, we develop a general setup to compare different multimodal large language models (e.g., GPT4-o, Claude 3 Opus, Gemini 1.5 Pro) as well as against existing human baselines. Our study demonstrates MLLMs' competitive performance in visualization literacy, where they outperform humans in certain tasks such as identifying correlations, clusters, and hierarchical structures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Animating Petascale Time-varying Data on Commodity Hardware with LLM-assisted Scripting

    cs.AI 2026-03 conditional novelty 7.0 of 10

    An LLM-assisted, keyframe-based animation framework streams cloud-hosted petascale datasets to commodity hardware and generates 3D scientific animations from natural-language requests.

  2. Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study

    cs.HC 2025-05 conditional novelty 6.0 of 10

    LLM agents can partially reproduce human ratings in visualization studies, and their agreement is highest for basic conclusions that expert evaluators predict with high confidence.

Pith tools