REVIEW 4 major objections 3 minor 110 references
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
T0 review · 4 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read SciFigBench shows a vision-language model can score highest on figure description quality yet fabricate answers for blurred content in 96% of cases.
desk verdict Solid new benchmark and A-R-I framework for VLM uncertainty behavior, but the headline admittance/fabrication reversal needs human validation of the judge labels. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the A-R-I framework applied through controlled visual stress tests. Its key operational distinction is between admittance blur and inductance blur: a selectively blurred element is either unrecoverable from the remaining context (an admittance probe, where a reliable model should say it cannot tell) or inferable from surrounding cues such as axis scales or legend colors (an inductance probe, where a reliable model should infer the value). Around this distinction, resistance probes with misleading captions, non-existent elements, false numerical anchors, and unanswerable questions quantify how a model handles context that contradicts the figure. This A-R-I trio gives each model a behavioral profile that is independent of its MQM and reasoning scores.
What would settle it
Take model responses to the 228 selectively blurred admittance figures and have human annotators independently label each response as admitting or fabricating; if the human admit rates for the two leading models converge or invert relative to the automated judge, the central behavioral gap is an artifact of the judge.
Extended reading notes
Core claim
The central discovery is that behavioral reliability under uncertainty is empirically separable from perception and reasoning accuracy. Using selectively blurred chart labels that are either unrecoverable from context (admittance probes) or inferable from remaining cues (inductance probes), plus caption-bias and false-premise probes, the paper finds that GPT-5.2, despite the highest MQM description quality (91.6) and strong reasoning accuracy (78.4%), states a specific answer for a blurred element 96% of the time and acknowledges the visual limitation only 8% of the time under direct questioning. Gemini 3.1 Pro, with MQM 90.2 and reasoning accuracy 81.0%, admits uncertainty in 71% of these cases and achieves the highest resistance score (0.91). Population-level correlations between quality and behavior are high (Spearman rho 0.83 to 0.95 across dimensions), but the reversal at the top of the leaderboard shows that high perception and reasoning scores do not guarantee reliable behavior; a benchmark reporting only MQM would rank these two models as near-equivalent while missing opposite behavioral profiles.
Load-bearing premise
The admittance and fabrication percentages rest on a single automated judge's binary labels, and the paper's human validation covers only the MQM description-quality scores, so if that judge labels wording differently across models the headline divergence could be a judge artifact rather than a behavioral difference.
Editorial extensions
If this is right
- Models with nearly equal MQM and reasoning scores can be assigned opposite behavioral profiles, so accuracy-based leaderboards cannot stand in for deployment-safety evaluation.
- The A-R-I dimensions are empirically separable: a model that resists false premises can still fabricate when evidence is missing, so both dimensions need to be measured independently.
- Selecting a model for scientific-figure processing based only on description quality could embed silent fabrication into summarization, comparison, or question-answering pipelines.
- Presupposition-embedded false premises (for example, asking about a benchmark line that does not exist) are the hardest resistance probes across all models, more so than explicitly wrong numerical values.
- Caption-bias resistance does not track model quality monotonically, suggesting that how much a model trusts provided context is shaped by instruction-following behavior rather than raw capability alone.
Reading between the lines
- The same A-R-I split could serve as a cheap pre-deployment screen: run admittance blurs on a small set of figures and measure the admit rate before trusting a model with scientific documents.
- The distinction between unrecoverable and inferable evidence should transfer to medical imaging or document processing, where a model that says 'I cannot tell' is often safer than one that guesses a plausible value.
- Because the admittance gap is assigned by a single automated judge and the paper's human validation covered MQM description quality rather than the binary admittance and fabrication labels, a human re-rating study of the blurred-label responses would directly test whether the 8% versus 71% gap is a model property or a judge artifact.
- The 'must-answer' bias finding points to a causal follow-up: train an open-weight model on examples that reward explicit uncertainty for unrecoverable elements, then measure whether this behavior transfers to unseen chart types and languages.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SciFigBench, a diagnostic benchmark for vision-language model (VLM) understanding of scientific figures, and the Admittance–Resistance–Inductance (A-R-I) framework for evaluating behavioral reliability under uncertainty. The benchmark contains 250 arXiv figures with expert descriptions, 1,000 human-reviewed reasoning questions, and over 34,000 evaluation setups including image transformations, caption-bias probes, false-premise resistance probes, and selective-blur admittance/inductance probes. Eight VLMs are evaluated on perception (MQM description quality), reasoning accuracy, and behavior. The central empirical claim is that models with similar perception and reasoning accuracy can be behavioral opposites under uncertainty: GPT-5.2 achieves the highest description quality (MQM 91.6) but is reported to hallucinate unreadable content in 96% of admittance-blur cases, whereas Gemini 3.1 Pro, with comparable accuracy, admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). The authors argue that accuracy-based benchmarks therefore rank behaviorally opposite models as near-equivalents, and that behavioral reliability must be evaluated separately.
Significance. If the behavioral measurements are valid, this is a substantial contribution. The paper's strengths are considerable: a large human annotation effort (600+ hours) with reported inter-annotator agreement; externally grounded benchmark construction from arXiv figures; a clear A-R-I decomposition that is conceptually useful; deterministic inference settings; bootstrap CIs; split-half reliability; cross-judge ablations for reasoning; and a probe-designer independence check. The framework addresses a real gap in chart and scientific-figure evaluation, and the finding that accuracy scores alone do not predict uncertainty acknowledgment would be practically important for deploying VLMs in scientific workflows. The appendices are unusually transparent, with full prompt and rubric inventories, which supports reproducibility.
major comments (4)
- [§4.4, Appendix B.7, Appendix F (S2/S3)] The headline admittance, fabrication, and resistance numbers are produced entirely by a GPT-4o judge applying the S2/S3 rubrics, yet Appendix C.1 validates the judge only for MQM description scoring (Krippendorff's α = 0.91, error-type F1s). No human agreement or error analysis is reported for the binary admits/fabricates/correct labels, nor for the 1.0/0.5/0.0 resistance scoring. This is load-bearing because the central reversal—GPT-5.2 hallucinates 96% vs. Gemini admits 71%—is a statement about these unvalidated binary labels. The paper's Limitations section overstates the case by citing the MQM human agreement as validation for 'automated evaluation' broadly. Please report human-annotator agreement for the S2/S3 rubrics (or at minimum a representative error analysis) and for the resistance scoring rubric before these figures are used to support the paper's main conclusion.
- [Appendix F (S2), §4.4] The S2 rubric defines 'fabricates' as 'provided a specific answer' with no confidence or hedging requirement, so a response such as 'The label is unreadable, but it might be X' is simultaneously scored as 'admits' and 'fabricates.' The abstract and §4.4 nonetheless describe GPT-5.2 as a 'confident fabricator' and state it 'hallucinates unreadable content in 96% of cases.' Confidence is never measured by the rubric. Please report the joint distribution of the admits and fabricates labels (admits∧fabricates, admits∧¬fabricates, ¬admits∧fabricates, ¬admits∧¬fabricates) for all models, and re-express the headline as 'fabrication without admission' or add a third label such as 'confident fabrication' if that is the intended construct. Without this, the 96% figure conflates genuinely confident hallucinations with hedged guesses and overstates the behavioral gap.
- [Table 4, §4.4] The behavioral comparison is asymmetric: Gemini's fabrication rate on the admittance probes is never reported, and neither is the overall admittance-by-fabrication contingency for either model. Only GPT-5.2's fabrication rate (96%) and Gemini's admittance rate (71%) are given. To support the claim of 'behaviorally opposite' profiles, the full contingency for both models is required. If Gemini also fabricates a large fraction of the time while admitting, the apparent opposition may reflect a difference in hedging style rather than in fabrication proclivity. This is especially important given that Gemini uses a 16k max-token setting and the judge may be sensitive to response length or hedging phrasing; the required table would let a reader evaluate whether the divergence is a model-level behavioral difference or a judge artifact.
- [Appendix C.1, Table 9] The paper's own validation data in Table 9 show that GPT-4o as an MQM judge has very low recall for Hallucinated Content (recall = 0.07, F1 = 0.12). This under-detection of hallucinated content in the description-quality channel is a further reason to require direct validation of the behavioral fabrication labels: the same judge model that under-detects hallucination in MQM scoring is being asked to classify fabrication in the S2/S3 rubrics. The manuscript currently does not demonstrate that the judge can reliably discriminate fabricated from non-fabricated blur-region content, which is the exact discrimination on which the headline claim rests. Please provide a direct human evaluation of the S2/S3 labels, ideally stratified by model and by admits/fabricates combination.
minor comments (3)
- [§4.4, Table 4] The column header 'Capability (%) Resistance Admittance (%) Inductance (%)' is ambiguous because the Resistance columns are scores on a 0–1 scale while the others are percentages; the caption explains this, but the table would be clearer if the header rows visually separated the four blocks, as in the subcaptions.
- [Appendix E.3] Please state explicitly in §4.4 or in the limitations that Gemini 3.1 Pro was run with max_tokens = 16,000 while all other models used 2,048, because this output-length asymmetry is a plausible confound for the judge-based admittance measurement and should be addressed in the analysis.
- [§3.2, MQM formula] The MQM formula 'max(0, 100 − P×100/(N×5))' assumes every checklist item can incur a maximum Major penalty of 5.0; since the checklist items carry Major/Minor severity tags, it would be helpful to state explicitly that N×5 is an upper bound and that the score is accordingly a normalized rather than exact percentage.
Circularity Check
No significant circularity: SciFigBench's central findings are externally grounded measurements, not derivations from their inputs.
full rationale
The paper's derivation chain is a benchmark-measurement chain, not an analytic derivation. The inputs are externally grounded: 250 figures from 187 arXiv papers, expert descriptions produced by two annotators with 94% agreement and a third adjudicator, 1,000 human-reviewed reasoning questions, selective-blur targets confirmed by human annotators, and model outputs generated at temperature 0. The headline behavioral quantities are empirical measurements: GPT-5.2's MQM of 91.6 versus Gemini 3.1 Pro's 90.2, and active admittance of 8% versus 71%, are read off model outputs through the S2/S3 rubrics; they are not implied by the MQM formula max(0,100−P×100/(N×5)) or by any fitted parameter. Nothing in the definition of MQM or the A-R-I rubric forces the observed dissociation between perception quality and behavioral reliability, so no self-definitional or fitted-input-as-prediction pattern is present. The authors' self-citations (Eger et al. 2025, Greisinger and Eger 2026, Mukherjee et al. 2025, Zhang et al. 2025) appear only as general context or related-work comparisons; none supplies a load-bearing premise, a uniqueness theorem, or an ansatz that is imported as external support. The genuine limitation flagged in the paper is a validity concern rather than circularity: Appendix C.1 validates the GPT-4o MQM judge against humans (Krippendorff alpha = 0.91) but reports that the judge under-detects Hallucinated Content (recall = 0.07), and the admittance/fabrication binary labels are not separately human-validated. That affects confidence in the headline behavioral numbers, but it is a measurement-validity issue; the numbers are not constructed so as to equal the paper's claims by definition. The benchmark is self-contained against external data and human annotations, and the central divergence is an empirical contingency rather than a tautology.
Assumptions & free parameters
free parameters (5)
- MQM penalty weights =
Accuracy and Completeness Major=5.0, Minor=2.0; Clarity Major=2.5, Minor=1.0
- MQM numerical tolerance =
Percentages +/-3 percentage points; axis ranges +/-10% of stated value
- Resistance scoring rubric =
1.0 resists, 0.5 hedges, 0.0 accepts
- Blur parameters =
grey blend 0.7, Gaussian kernel size 75
- Caption-bias false-claim count =
2-3 false claims per caption; 70/30 correct versus poisoned split
assumptions (4)
- domain assumption The GPT-4o judge's behavioral classifications (admits, fabricates, correct) are reliable proxies for human judgment of model outputs on admittance, resistance, and inductance probes.
- domain assumption Human-confirmed selective-blur targets are genuinely unrecoverable (admittance) or inferable (inductance) as classified.
- domain assumption Temperature-0 API outputs from commercial models are stable and representative; provider-side drift does not change relative rankings.
- domain assumption The 250 arXiv figures (bar, line, pie, mostly NLP and ML) are representative enough to support conclusions about scientific figure understanding in general.
invented entities (1)
-
A-R-I framework dimensions (Admittance, Resistance, Inductance)
Cite this review
Pith. "Pith review of How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures." pith.science (2026). https://pith.science/paper/L4NCZO3Q
@misc{pith2026260813267,
author = {Pith},
title = {Pith review of: How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures},
year = {2026},
howpublished = {\url{https://pith.science/paper/L4NCZO3Q}},
note = {Machine review of arXiv:2608.13267}
}
read the original abstract
Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures with high-quality human annotations across three evaluation aspects, totaling 600+ hours of annotation effort. We further extend these figures via image transformations, reasoning questions, resistance probes, caption-bias probes, and confirmed selective-blur targets, producing over 34,000 evaluation setups for stress testing. We further propose the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist misleading context, and infer cautiously from partial information. Our results reveal substantial behavioral differences among models. GPT-5.2 achieves the highest description quality (MQM 91.6) with strong reasoning accuracy (78.4%), yet hallucinates unreadable content in 96% of cases, whereas Gemini 3.1 Pro, a comparably capable model (MQM 90.2, reasoning 81.0%), admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). These findings show that high perception and reasoning accuracy alone do not guarantee behavioral reliability, a dimension critical for deployment in scientific workflows.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[2]
Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, and others . 2025. https://doi.org/10.48550/arXiv.2511.21631 Qwen3-VL technical report . arXiv preprint arXiv:2511.21631
-
[4]
Qingxing Cao, Junhao Cheng, Xiaodan Liang, and Liang Lin. 2024. https://doi.org/10.18653/v1/2024.acl-long.658 Visdiahalbench: A visual dialogue benchmark for diagnosing hallucination in large vision-language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, volume 1, pages 12161--12176
-
[5]
Steffi Chern, Zhulin Hu, Yuqing Yang, Ethan Chern, Yuan Guo, Jiahe Jin, Binjie Wang, and Pengfei Liu. 2024. https://arxiv.org/abs/2406.13261 BeHonest : Benchmarking H onesty in L arge L anguage M odels . arXiv preprint arXiv:2406.13261
arXiv 2024
-
[6]
Steffen Eger, Yong Cao, Jennifer D'Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, Chenghua Lin, Nafise Sadat Moosavi, Wei Zhao, and Tristan Miller. 2025. https://doi.org/10.48550/arXiv.2502.05151 Transforming science with large language models: A survey on AI -assisted scientific discover...
-
[7]
Agarwal, Joanna Lin, Anson Zhou, Sonnet Xu, Vasiliki Bikia, Roxana Daneshjou, and Sanmi Koyejo
Aaron Fanous, Jacob Goldberg, Ank A. Agarwal, Joanna Lin, Anson Zhou, Sonnet Xu, Vasiliki Bikia, Roxana Daneshjou, and Sanmi Koyejo. 2025. https://doi.org/10.1609/aies.v8i1.36598 SycEval : Evaluating LLM sycophancy . In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 8 of AIES '25, pages 893--900
-
[8]
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021. https://doi.org/10.1162/tacl_a_00437 Experts, errors, and context: A large-scale study of human evaluation for machine translation . Transactions of the Association for Computational Linguistics, 9:1460--1474
-
[9]
Christian Greisinger and Steffen Eger. 2026. https://openreview.net/forum?id=rJv2byEWA3 TikZilla : Scaling text-to- TikZ with high-quality data and reinforcement learning . In Proceedings of the 14th International Conference on Learning Representations
2026
-
[10]
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. 2024. https://doi.org/10.1109/CVPR52733.2024.01363 HallusionBench : An advanced diagnostic suite for entangled language hallucination and visual illusion in L arge V ision- L anguage M odels . In P...
arXiv 2024
Show all 110 references
-
[11]
Ming Hu, Chenglong Ma, Wei Li, Wanghan Xu, Jiamin Wu, Jucheng Hu, Tianbin Li, Guohang Zhuang, Jiaqi Liu, Yingzhou Lu, and others . 2025. https://doi.org/10.48550/arXiv.2508.21148 A survey of scientific large language models: From data foundations to agent frontiers . arXiv pre...
2025 doi
-
[12]
Kung-Hsiang Huang, Mingyang Zhou, Hou Pong Chan, Yi Fung, Zhenhailong Wang, Lingyu Zhang, Shih-Fu Chang, and Heng Ji. 2024. https://doi.org/10.18653/v1/2024.findings-acl.41 Do LVLMs understand charts? analyzing and correcting factual errors in chart captioning . In Findings of...
2024 doi
-
[14]
Klaus Krippendorff. 2011. Computing Krippendorff's alpha-reliability. Departmental Papers (ASC)
2011
-
[15]
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.20 Evaluating object hallucination in large vision-language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Pr...
2023 doi
-
[16]
Elizabeth F Loftus. 1975. Leading questions and the eyewitness report. Cognitive Psychology, 7(4):560--572
1975
-
[17]
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014. https://doi.org/10.5565/rev/tradumatica.77 Multidimensional quality metrics ( MQM ): A framework for declaring and describing translation quality metrics . Tradum \`a tica , (12):455--463
2014 doi
-
[18]
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2024. MathVista : Evaluating mathematical reasoning of foundation models in visual contexts. In Proceedings of the International Conferen...
2024
-
[19]
Ridwan Mahbub, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mizanur Rahman, Mir Tafseer Nayeem, and Enamul Hoque. 2025. https://doi.org/10.1109/VIS60296.2025.00006 The perils of chart deception: How misleading visualizations affect vision-language models . In 2025 IEEE Visu...
2025
-
[20]
Ahmed Masry, Mohammed Saidul Islam, Mahir Ahmed, Aayush Bajaj, Firoz Kabir, Aaryaman Kartha, Md Tahmid Rahman Laskar, Mizanur Rahman, Shadikur Rahman, Mehrad Shahmohammadi, Megh Thakkar, Md Rizwan Parvez, Enamul Hoque, and Shafiq Joty. 2025. https://doi.org/10.18653/v1/2025.fi...
2025 doi
-
[21]
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. https://doi.org/10.18653/v1/2022.findings-acl.177 ChartQA : A benchmark for question answering about charts with visual and logical reasoning . In Findings of the Association for Computational Lingui...
2022 doi
-
[22]
Meta AI . 2025. The llama 4 herd: The beginning of a new era of natively multimodal AI innovation. https://ai.meta.com/blog/llama-4-multimodal-intelligence/. Blog post; accessed 2026-05-13
2025
- [23]
-
[24]
Kushin Mukherjee, Donghao Ren, Dominik Moritz, and Yannick Assogba. 2025. https://doi.org/10.1109/TVCG.2025.3634249 EncQA : Benchmarking vision-language models on visual encodings for charts . IEEE Transactions on Visualization and Computer Graphics, pages 648--658
2025
- [25]
- [26]
-
[27]
Jonathan Roberts, Kai Han, Neil Houlsby, and Samuel Albanie. 2024. SciFIBench : Benchmarking large multimodal models for scientific figure interpretation. In Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track
2024
-
[28]
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. https://doi.org/10.18653/v1/D18-1437 Object hallucination in image captioning . In Proceedings of the 2018 conference on Empirical Methods in Natural Language Processing (EMNLP), pages 403...
2018 doi
-
[29]
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and...
2024
-
[30]
selective prediction
Tejas Srinivasan, Jack Hessel, Tanmay Gupta, Bill Yuchen Lin, Yejin Choi, Jesse Thomason, and Khyathi Chandu. 2024. https://doi.org/10.18653/v1/2024.findings-acl.767 Selective “selective prediction”: Reducing unnecessary abstention in vision-language reasoning . In Findings of...
2024 doi
-
[31]
Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake, Wenxuan Ding, Fangcong Yin, Prasann Singhal, Manya Wadhwa, Zeyu Liu, Zayne Sprague, Ramya Namuduri, Bodun Hu, Juan Rodriguez, Puyuan Peng, and Greg Durrett. 2025. ChartMuseum : Testing visual reasoning capabilities of large vision-...
2025
-
[32]
Gemini 3.1 pro model card
Gemini Team. Gemini 3.1 pro model card. 2026. URL https://storage.googleapis. com/deepmind-media/Mod el-Cards/Gemini-3-1-Pro-Model-Card.pdf
2026
-
[34]
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.330 Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language M...
2023 doi
-
[35]
Jonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, and Iryna Gurevych. 2026. https://doi.org/10.18653/v1/2026.acl-long.377 Protecting multimodal large language models against misleading visualizations . In Proceedings of the 64th Annual Meeting of the Association for Co...
2026 doi
-
[36]
Amos Tversky and Daniel Kahneman. 1974. https://doi.org/10.1126/science.185.4157.1124 Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty. Science, 185(4157):1124--1131
1974
-
[37]
Xingqi Wang, Yiming Cui, Xin Yao, Shijin Wang, Guoping Hu, and Xiaoyu Qin. 2025. https://doi.org/10.48550/arXiv.2509.17481 ChartHal : A fine-grained framework evaluating hallucination of large vision language models in chart understanding . arXiv preprint arXiv:2509.17481
2025 doi
-
[38]
Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Haotian Liu, Sadhika Malladi, Alexis Chevalier, Sanjeev Arora, and Danqi Chen. 2024. https://doi.org/10.52202/079017-3609 Charxiv: Charting gaps in realistic chart understanding in m...
2024 doi
- [39]
-
[40]
Bingbing Wen, Jihan Yao, Shangbin Feng, Chenjun Xu, Yulia Tsvetkov, Bill Howe, and Lucy Lu Wang. 2025. https://doi.org/10.1162/tacl_a_00754 Know your limits: A survey of abstention in large language models . Transactions of the Association for Computational Linguistics, 13:529--556
2025 doi
- [41]
-
[42]
Liang Yan, Xu Jiang, Jian Ma, Yuhang Liu, Tian Bian, Qichao Wang, Abhishek Basu, Yu Rong, Tingyang Xu, Pengcheng Wu, Le Song, Imran Razzak, Junchi Yan, Zengfeng Huang, and Yutong Xie. 2025. https://openreview.net/forum?id=HSz1Kr5BeC A comprehensive survey of multimodal LLM s f...
2025
-
[43]
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, and 3 others. 2024. MMMU : A massive multi...
2024
-
[44]
Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai, Jonas Belouadi, Fahimeh Moafian, and Zhixue Zhao. 2025. https://openreview.net/forum?id=ugyqNEOjoU ScImage : How good are multimodal large language models at scientific text-to-image generation? In International Conference ...
2025
-
[45]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. https://doi.org/10.52202/075280-2020 Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena...
2023 doi
-
[46]
Zifeng Zhu, Mengzhao Jia, Zhihan Zhang, Lang Li, and Meng Jiang. 2025. https://doi.org/10.18653/v1/2025.naacl-long.566 MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems . In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of th...
2025 doi
-
[47]
A Comprehensive Survey of Multimodal
Liang Yan and Xu Jiang and Jian Ma and Yuhang Liu and Tian Bian and Qichao Wang and Abhishek Basu and Yu Rong and Tingyang Xu and Pengcheng Wu and Le Song and Imran Razzak and Junchi Yan and Zengfeng Huang and Yutong Xie , booktitle=. A Comprehensive Survey of Multimodal. 2025 , url=
2025
-
[48]
A Comprehensive Survey of Multimodal
Yan, Liang and Jiang, Xu and Ma, Jian and Liu, Yuhang and Bian, Tian and Wang, Qichao and Basu, Abhishek and Rong, Yu and Xu, Tingyang and Wu, Pengcheng and Song, Le and Razzak, Imran and Yan, Junchi and Huang, Zenfeng and Xie, Yutong , booktitle=. A Comprehensive Survey of Multimodal
-
[49]
2025 , url=
Zhang, Leixin and Eger, Steffen and Cheng, Yinjie and Zhai, Weihe and Belouadi, Jonas and Moafian, Fahimeh and Zhao, Zhixue , booktitle=. 2025 , url=
2025
-
[50]
2026 , url =
Greisinger, Christian and Eger, Steffen , booktitle =. 2026 , url =
2026
-
[51]
Transforming Science with Large Language Models: A Survey on
Eger, Steffen and Cao, Yong and D'Souza, Jennifer and Geiger, Andreas and Greisinger, Christian and Gross, Stephanie and Hou, Yufang and Krenn, Brigitte and Lauscher, Anne and Li, Yizhi and Lin, Chenghua and Moosavi, Nafise Sadat and Zhao, Wei and Miller, Tristan , journal =. ...
2025
-
[52]
arXiv preprint arXiv:2508.21148 , year =
A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers , author =. arXiv preprint arXiv:2508.21148 , year =. 2508.21148 , archivePrefix =
-
[53]
arXiv preprint arXiv:2410.21276 , year=
GPT-4o System Card , author=. arXiv preprint arXiv:2410.21276 , year=
-
[54]
2026 , author=
Gemini 3.1 pro model card. 2026 , author=
2026
-
[55]
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs , url =
Abdelrahman Abouelenin and Atabak Ashfaq and Adam Atkinson and Hany Awadalla and Nguyen Bach and Jianmin Bao and Alon Benhaim and Martin Cai and Vishrav Chaudhary and Congcong Chen and Dong Chen and Dongdong Chen and Jun. Phi-4-Mini Technical Report: Compact yet Powerful Multi...
-
[56]
2025 , howpublished=
The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal. 2025 , howpublished=
2025
-
[57]
arXiv preprint arXiv:2511.21631 , year=
Bai, Shuai and Cai, Yuxuan and Chen, Ruizhe and Chen, Keqin and Chen, Xionghui and Cheng, Zesen and Deng, Lianghao and Ding, Wei and Gao, Chang and Ge, Chunjiang and. arXiv preprint arXiv:2511.21631 , year=
-
[58]
arXiv preprint arXiv:2503.19786 , year=
Gemma 3 technical report , author=. arXiv preprint arXiv:2503.19786 , year=
-
[59]
2025 , howpublished=
Introducing Mistral 3 , author=. 2025 , howpublished=
2025
-
[60]
Cognitive Psychology , volume=
Leading questions and the eyewitness report , author=. Cognitive Psychology , volume=. 1975 , publisher=
1975
-
[61]
, author=
Judgment under Uncertainty: Heuristics and Biases: Biases in judgments reveal some heuristics of thinking under uncertainty. , author=. Science , volume=. 1974 , publisher=
1974
-
[62]
Syntax and Semantics , volume=
Logic and conversation , author=. Syntax and Semantics , volume=
-
[63]
Proceedings of the International Conference on Learning Representations (ICLR) , year=
Towards Understanding Sycophancy in Language Models , author=. Proceedings of the International Conference on Learning Representations (ICLR) , year=
-
[64]
Proceedings of ICLR Workshop , year=
Kahou, Samira Ebrahimi and Michalski, Vincent and Atkinson, Adam and K. Proceedings of ICLR Workshop , year=
-
[65]
Methani, Nitesh and Ganguly, Pritha and Khapra, Mitesh M and Kumar, Pratyush , booktitle=
-
[66]
2022 , doi =
Masry, Ahmed and Long, Do Xuan and Tan, Jia Qing and Joty, Shafiq and Hoque, Enamul , booktitle =. 2022 , doi =
2022
-
[67]
2023 , doi=
Xu, Zhengzhuo and Du, Sinan and Qi, Yiyan and Xu, Chengjin and Yuan, Chun and Guo, Jian , journal=. 2023 , doi=
2023
-
[68]
Li, Shengzhi and Tajbakhsh, Nima , booktitle=
-
[69]
Proceedings of the 2018 conference on Empirical Methods in Natural Language Processing (EMNLP) , publisher=
Object Hallucination in Image Captioning , author=. Proceedings of the 2018 conference on Empirical Methods in Natural Language Processing (EMNLP) , publisher=. 2018 , doi=
2018
-
[70]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=
Li, Yifan and Du, Yifan and Zhou, Kun and Wang, Jinpeng and Zhao, Wayne Xin and Wen, Ji-Rong , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year=
2023
-
[71]
2024 , doi=
Guan, Tianrui and Liu, Fuxiao and Wu, Xiyang and Xian, Ruiqi and Li, Zongxia and Liu, Xiaoyu and Wang, Xijun and Chen, Lichang and Huang, Furong and Yacoob, Yaser and Manocha, Dinesh and Zhou, Tianyi , booktitle=. 2024 , doi=
2024
-
[72]
Aligning Large Multimodal Models with Factually Augmented
Sun, Zhiqing and Shen, Sheng and Cao, Shengcao and Liu, Haotian and Li, Chunyuan and Shen, Yikang and Gan, Chuang and Gui, Liang-Yan and Wang, Yu-Xiong and Yang, Yiming and others , booktitle=. Aligning Large Multimodal Models with Factually Augmented
-
[73]
arXiv preprint arXiv:2404.18930 , year =
Hallucination of Multimodal Large Language Models: A Survey , author =. arXiv preprint arXiv:2404.18930 , year =. 2404.18930 , archivePrefix =
-
[74]
Multidimensional Quality Metrics (
Lommel, Arle and Uszkoreit, Hans and Burchardt, Aljoscha , journal=. Multidimensional Quality Metrics (. 2014 , doi=
2014
-
[75]
Transactions of the Association for Computational Linguistics , volume =
Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation , author =. Transactions of the Association for Computational Linguistics , volume =. 2021 , doi =
2021
-
[76]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , journal=. 2023 , doi=
2023
-
[77]
2025 , month = feb, howpublished =
Use Lens to Search Your Screen While You Browse on. 2025 , month = feb, howpublished =
2025
-
[78]
2025 , month = jul, howpublished =
2025
-
[79]
2026 , howpublished =
Morgan Stanley Uses. 2026 , howpublished =
2026
-
[80]
2026 , month = mar, howpublished =
Using Research on. 2026 , month = mar, howpublished =
2026
-
[81]
2024 , month = dec, howpublished =
Try Deep Research and Our New Experimental Model in. 2024 , month = dec, howpublished =
2024
-
[82]
, title =
Beede, Emma and Baylor, Elizabeth and Hersch, Fred and Iurchenko, Anna and Wilcox, Lauren and Ruamviboonsuk, Paisan and Vardoulakis, Laura M. , title =. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , pages =. 2020 , doi =
2020
-
[83]
2024 , month = nov, howpublished =
Information Request. 2024 , month = nov, howpublished =
2024
-
[84]
Advances in Neural Information Processing Systems , volume=
Charxiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs , author=. Advances in Neural Information Processing Systems , volume=. 2024 , doi =
2024
-
[85]
Roberts, Jonathan and Han, Kai and Houlsby, Neil and Albanie, Samuel , booktitle=
-
[86]
Tang, Liyan and Kim, Grace and Zhao, Xinyu and Lake, Thom and Ding, Wenxuan and Yin, Fangcong and Singhal, Prasann and Wadhwa, Manya and Liu, Zeyu and Sprague, Zayne and Namuduri, Ramya and Hu, Bodun and Rodriguez, Juan and Peng, Puyuan and Durrett, Greg , journal=
-
[87]
Lu, Pan and Bansal, Hritik and Xia, Tony and Liu, Jiacheng and Li, Chunyuan and Hajishirzi, Hannaneh and Cheng, Hao and Chang, Kai-Wei and Galley, Michel and Gao, Jianfeng , booktitle=
-
[88]
Yue, Xiang and Ni, Yuansheng and Zhang, Kai and Zheng, Tianyu and Liu, Ruoqi and Zhang, Ge and Stevens, Samuel and Jiang, Dongfu and Ren, Weiming and Sun, Yuxuan and Wei, Cong and Yu, Botao and Yuan, Ruibin and Sun, Renliang and Yin, Ming and Zheng, Boyuan and Yang, Zhenzhu an...
-
[89]
2025 , doi =
Masry, Ahmed and Islam, Mohammed Saidul and Ahmed, Mahir and Bajaj, Aayush and Kabir, Firoz and Kartha, Aaryaman and Laskar, Md Tahmid Rahman and Rahman, Mizanur and Rahman, Shadikur and Shahmohammadi, Mehrad and Thakkar, Megh and Parvez, Md Rizwan and Hoque, Enamul and Joty, ...
2025
-
[90]
2025 , publisher=
Mukherjee, Kushin and Ren, Donghao and Moritz, Dominik and Assogba, Yannick , journal=. 2025 , publisher=
2025
-
[91]
2025 , doi=
Zhu, Zifeng and Jia, Mengzhao and Zhang, Zhihan and Li, Lang and Jiang, Meng , booktitle=. 2025 , doi=
2025
-
[92]
Huang, Kung-Hsiang and Zhou, Mingyang and Chan, Hou Pong and Fung, Yi and Wang, Zhenhailong and Zhang, Lingyu and Chang, Shih-Fu and Ji, Heng , booktitle =. Do. 2024 , pages =
2024
-
[93]
2025 , doi =
Wang, Xingqi and Cui, Yiming and Yao, Xin and Wang, Shijin and Hu, Guoping and Qin, Xiaoyu , journal =. 2025 , doi =
2025
-
[94]
arXiv preprint arXiv:2505.17235 , year =
Moured, Omar and Chen, Yufan and Liu, Ruiping and Rei. arXiv preprint arXiv:2505.17235 , year =
-
[95]
Losing the Plot: How
Shin, Philip Wootaek and Sampson, Jack and Narayanan, Vijaykrishnan and Marquez, Andres and Halappanavar, Mahantesh , journal =. Losing the Plot: How
-
[96]
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , volume=
Protecting Multimodal Large Language Models Against Misleading Visualizations , author=. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , volume=. 2026 , doi =
2026
-
[97]
2025 IEEE Visualization and Visual Analytics (VIS) , pages=
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models , author=. 2025 IEEE Visualization and Visual Analytics (VIS) , pages=. 2025 , organization=
2025
-
[98]
2024 , eprint =
Chern, Steffi and Hu, Zhulin and Yang, Yuqing and Chern, Ethan and Guo, Yuan and Jin, Jiahe and Wang, Binjie and Liu, Pengfei , journal =. 2024 , eprint =
2024
-
[99]
Transactions of the Association for Computational Linguistics , volume=
Know Your Limits: A survey of Abstention in Large Language Models , author=. Transactions of the Association for Computational Linguistics , volume=. 2025 , publisher=
2025
-
[100]
arXiv preprint arXiv:2411.04368 , year =
Measuring Short-Form Factuality in Large Language Models , author =. arXiv preprint arXiv:2411.04368 , year =
-
[101]
Findings of the Association for Computational Linguistics: ACL 2024 , pages=
Srinivasan, Tejas and Hessel, Jack and Gupta, Tanmay and Lin, Bill Yuchen and Choi, Yejin and Thomason, Jesse and Chandu, Khyathi , title =. Findings of the Association for Computational Linguistics: ACL 2024 , pages=. 2024 , doi =
2024
-
[102]
and Lin, Joanna and Zhou, Anson and Xu, Sonnet and Bikia, Vasiliki and Daneshjou, Roxana and Koyejo, Sanmi , booktitle =
Fanous, Aaron and Goldberg, Jacob and Agarwal, Ank A. and Lin, Joanna and Zhou, Anson and Xu, Sonnet and Bikia, Vasiliki and Daneshjou, Roxana and Koyejo, Sanmi , booktitle =. 2025 , doi =
2025
-
[103]
arXiv preprint arXiv:2207.05221 , year=
Language Models (Mostly) Know What They Know , author=. arXiv preprint arXiv:2207.05221 , year=. 2207.05221 , archivePrefix =
-
[104]
Computing
Krippendorff, Klaus , journal=. Computing. 2011 , publisher=
2011
-
[105]
2023 , doi =
Tian, Katherine and Mitchell, Eric and Zhou, Allan and Sharma, Archit and Rafailov, Rafael and Yao, Huaxiu and Finn, Chelsea and Manning, Christopher D , booktitle=. 2023 , doi =
2023
-
[106]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , volume=
Visdiahalbench: A Visual Dialogue Benchmark for Diagnosing Hallucination in Large Vision-Language Models , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , volume=. 2024 , doi=
2024
-
[107]
Aho and Jeffrey D
Alfred V. Aho and Jeffrey D. Ullman , title =. 1972
1972
-
[108]
Publications Manual , year = "1983", publisher =
1983
-
[109]
Chandra and Dexter C
Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243
1981
-
[110]
Scalable training of
Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of
-
[111]
Dan Gusfield , title =. 1997
1997
-
[112]
Tetreault , title =
Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =
2015
-
[113]
A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =
Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =
-
[114]
Proceedings of Translating and the Computer 36 , year =
Multidimensional Quality Metrics (MQM): A Framework for Declaring and Describing Translation Quality Metrics , author =. Proceedings of Translating and the Computer 36 , year =
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.