Introduces CVLAT and VFRI to disentangle visual vs factual correctness in 15 LVLMs, classifies models by reliance sign, compares to human baseline, and tests prompt interventions.
Visualization literacy of multimodal large language models: a comparative study
3 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
3
Pith papers citing it
1
external citations · external index
representative citing papers
MLLMs given the same instructions as human participants achieve expert-level performance on perceiving stress in network visualizations and rely on similar visual proxies.
An LLM-assisted, keyframe-based animation framework streams cloud-hosted petascale datasets to commodity hardware and generates 3D scientific animations from natural-language requests.
citing papers explorer
-
Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy
Introduces CVLAT and VFRI to disentangle visual vs factual correctness in 15 LVLMs, classifies models by reliance sign, compares to human baseline, and tests prompt interventions.