REVIEW 5 major objections 5 minor 54 references
Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation
T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read AI vision models cannot reliably count nodes and edges in rendered knowledge graphs or spot wrong ones, a 26-model benchmark finds.
desk verdict A useful new benchmark for MLLM evaluation on graph-structured visual inputs, but the 'abstractive visual understanding' claim is confounded by the paper's own modality ablation and missing human baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the M3STR benchmark and its 'multi-modal map.' To build a map, the pipeline samples a subgraph from the FB15K-237 multi-modal knowledge graph, applies a task-specific modification—none for counting, an injected anomalous entity or relation for detection, a masked entity or relation with five options for completion—and renders the result through GraphViz, so each node shows a text label and an image. A task-specific prompt restricts the output to a number, Yes/No, or a choice letter, and the overall score averages seven subtasks (entity/relation counting, entity/relation/mix detection, entity/relation completion). Random-choice baselines are reported for every task, and the detection task uses a 1:1 positive/negative ratio.
What would settle it
Render the same M3STR subgraphs in layouts that differ only in legibility—larger fonts, wider node spacing, fewer edge crossings, no entity images—and keep the graph structure fixed. If near-random counting and detection scores rise substantially, the deficit is perceptual rather than abstractive; if they stay near random, the paper's claim of an abstractive-understanding gap is supported.
Extended reading notes
Core claim
The central claim is that current MLLMs do not convert a rendered knowledge-graph image into its abstract relational structure. The evidence from M3STR is that many models score below 30 percent on entity counting, almost all models sit at or below random-choice levels on anomaly detection, and confusion matrices show systematic answer bias rather than guessing. The modality ablation sharpens the point: removing entity images or entity texts from the input often improves accuracy, and feeding the same knowledge as plain text beats the visual map for several models. The authors interpret this as small models suffering cognitive overload when the image carries many visual details, so visual processing of structured knowledge lags behind purely textual reasoning.
Load-bearing premise
The load-bearing premise is that a rendered image of a knowledge-graph subgraph presents the relational topology to an MLLM as legibly as the graph does to a human, so low counting and detection scores measure abstractive understanding rather than difficulty reading dense, cluttered images; the paper's own finding that text-only inputs score higher directly weakens this premise.
Editorial extensions
If this is right
- M3STR tracks a capability that existing natural-image, chart, and math benchmarks do not surface, so a high leaderboard rank on those benchmarks does not imply structured abstract visual reasoning.
- Half the evaluated models cannot reliably count the nodes or edges in a rendered knowledge graph, indicating a basic bottleneck in turning visual layout into relational structure.
- Anomaly detection near the random baseline means models will accept factually wrong entities or relations in a visual knowledge graph, which matters for any application that displays structured knowledge to users.
- Because text-only knowledge-graph input often beats the visual map, improving MLLMs on this benchmark likely requires better visual encoding of structured layouts rather than additional world knowledge.
- Scaling within a model family helps: larger Qwen2 and Qwen2.5 models improve substantially on M3STR, the best overall model is Qwen2.5-VL-72B, and several API models rank lower.
Reading between the lines
- The paper's design mixes low-level legibility with abstraction, since a dense graph rendering can defeat a vision encoder before relational reasoning starts; a controlled rendering study would separate the two.
- The high multiple-choice completion scores, especially with text input, could partly reflect memorization of FB15K-237, a widely used dataset; the authors raise this for text inputs, and the same concern extends to any visual answer that depends on OCR of node labels.
- If clutter is the real obstacle, a topology-only version of M3STR with entity images removed should raise counting and detection accuracy modestly; if scores stay near random, the failure is in abstract relational reasoning.
- The systematic default-to-one-answer bias in detection suggests models carry a strong prior for 'no anomaly'; prompt calibration or negative-example instruction on graph renderings is a natural next experiment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a new evaluation paradigm for multimodal large language models (MLLMs), focusing on 'abstractive visual understanding of structured knowledge.' It introduces M3STR, a benchmark built from FB15K-237 multi-modal knowledge graphs, rendered as images via GraphViz. The benchmark contains three tasks (count, detection, completion) with entity/relation subtasks. The authors evaluate 26 MLLMs and report generally low accuracy, concluding that current MLLMs lack robust capabilities for abstract structural understanding of visual knowledge representations. Code and data are released. The paper also includes a modality-contribution analysis suggesting that removing entity images or substituting text-only KG inputs often improves performance.
Significance. If the central claim were fully supported, this would identify a genuine capability gap in MLLMs that existing benchmarks do not capture, and the released benchmark could be a useful resource. The paper is methodologically transparent in releasing code and data, and the scale of the evaluation (26 models) is commendable. However, the interpretation is currently weakened by several confounds, most notably the lack of a human baseline, the inconsistency of the random-choice baselines in the detection task, and the acknowledged possibility of pretraining memorization of FB15K-237. These issues do not invalidate the benchmark as a resource, but they do undermine the strength of the paper's central claim as stated.
major comments (5)
- [Section 4.3, Figure 4] The modality ablation weakens the central claim that MLLMs lack abstract structural understanding of visual knowledge representations. The paper reports that Qwen2.5-VL-7B improves from 44.81% to 71.12% on entity counting when entity images are removed, and that text-only KG inputs substantially boost completion accuracy for several models. If the bottleneck were abstract reasoning over the relational topology, text-only inputs should not be systematically easier than visual renderings. These results suggest that the GraphViz rendering itself, not the abstract structure, may be the limiting factor. Since the paper does not report a human baseline or any oracle that parses the image (e.g., extracting nodes and edges with a graph parser), the low MLLM scores conflate low-level visual parsing failures with deficits in abstractive understanding. The authors should either add a human evaluation on the same rendered images or an oracle baseline to calibrate the legibility of the images, and then revisit the claim in Section 4.3.
- [Table 1] The random-choice baselines reported for Task 2 detection are internally inconsistent with the task formulation. The prompt in Figure 3 asks for a binary Yes/No answer, and Section 3.4.2 states that the positive/negative ratio is controlled to be 1:1. A uniform random guess on a balanced binary task should yield 50% accuracy, yet the table lists 66.66% for entity detection and 83.66% for relation detection. If these numbers are correct, the task cannot be a balanced binary classification; if they are typos, then the claim in Section 4.2 (Perspective 3) that models perform at near-random levels on detection is miscalibrated. For example, LLaVA-1.5-7B scores 66.64% on entity detection, which would be substantially above a true 50% random baseline. Please correct the baselines and re-assess the corresponding conclusions.
- [Section 4.3] The completion task is confounded by possible pretraining memorization of FB15K-237. The authors themselves acknowledge that 'FB15K-237, as well-known open-source data, may have been included in the pre-trained corpus.' This directly undermines the interpretation that high accuracy on text-only KG inputs for Task 3 reflects abstractive reasoning; it may instead reflect memorization of facts from the pretraining corpus. To support the benchmark's validity as a measure of abstractive visual understanding, the authors should add a control that separates memorization from reasoning. Possible fixes include evaluating on a KG that is unlikely to be in pretraining data, using a temporal or random split that removes entities seen during pretraining, or testing on synthetic triples that are not part of any public KG.
- [Section 3.4.2] The negative-sampling procedure for the detection task relies on the assumption that replacing an entity or relation with a random one creates a factually incorrect subgraph. This is not generally valid under the open-world assumption: a triple that is absent from FB15K-237 may still be plausible or true in the real world. For example, replacing the tail entity of a 'president_of' relation with another person could result in a plausible triple that is not in the KG. Without a validation step to ensure that the replacement is genuinely factually incorrect (e.g., through human annotation or a closed-world subset), the ground-truth labels for the detection task may be unreliable, which would affect the interpretation of all detection results.
- [Section 4.2 vs. Section 4.3] There is an inconsistency in the construct being claimed. In Section 4.2 (Perspective 3), the authors attribute low counting accuracy to 'a fundamental deficit in basic visual perception,' while the abstract and title frame the contribution as a deficit in 'abstractive visual understanding.' These are different capacities: a model might fail to perceive the rendered graph clearly but still be able to reason abstractly over a legible representation, or vice versa. The paper should either separate these constructs explicitly by adding perceptual control tasks (e.g., counting text labels or nodes without relational reasoning) or temper the central claim to what the data actually support.
minor comments (5)
- [Throughout] There are several typos and inconsistencies: 'undetstanding' in the abstract and Figure 1, 'Scalind law' in Section 4.2, 'Geimini' for Gemini in Table 1, 'distrbution' in Figure 2, and the inconsistent spelling 'M3Str' vs. 'M3STR'. Please proofread carefully.
- [Section 3.4.3] For reproducibility, please specify the exact GraphViz parameters used (layout engine, node size, label font size, image resolution, overlap handling). These details affect the legibility of the rendered images and are essential for others to reproduce the benchmark.
- [Figure 4] The bar chart is difficult to read in grayscale and the legend for 'w/ Textual KG' is not clearly distinguished. Please use distinct patterns or colors and increase the figure resolution.
- [Section 4.3] The phrase '159% performance gain' is misleading. If Qwen2.5-VL-7B improves from 44.81% to 71.12% when entity images are removed, that is a relative gain of approximately 59%, not 159%. Please use unambiguous terminology (e.g., 'an increase of 26.31 percentage points' or 'a 1.59x relative improvement').
- [Table 1] The random-choice baseline for Task 1 assumes a uniform distribution over the possible counts. Please verify that the count distribution is indeed uniform, or compute the baseline from the empirical distribution of entity/relation counts in the benchmark.
Circularity Check
No significant circularity: M3STR outcomes are externally measured, the central claim is an interpretive generalization, and self-citations are background only.
full rationale
The paper's construction pipeline is explicit: subgraph sampling (Eq. 4), task-specific modification (Eq. 5), and GraphViz visual translation (Eq. 6) generate the image-text-answer triples (I,Q,A) before any model is run. Ground-truth labels are defined by the Modifier operations (counting; replacing an entity or relation with one outside the subgraph; masking with distractors sampled from the MMKG), so accuracy is computed against labels that do not depend on model outputs, and no fitted parameter is renamed as a prediction. The Section 4.3 sentence 'Current MLLMs lack robust capabilities for abstract structural understanding of visual knowledge representations' is an interpretation of the observed low accuracies, not a quantity forced by a definition that equates the construct with the scores; there is no equation-level reduction of the conclusion to the benchmark construction. Self-citations ([5], [45]-[47]) appear in background sections for MMKG, structural prompting, and KGC and are not used as a uniqueness theorem or as the sole support for the central claim. The modality ablation in Figure 4, where text-only KG inputs beat visual MMKG inputs and removing entity images often helps, plus the paper's own admission that FB15K-237 'may have been included in the pre-trained corpus', is a legitimate construct-validity concern about whether the low scores isolate abstractive understanding rather than perception or memorization; however, a validity threat is not a circular derivation. No quoted step equates the claimed prediction with its input by construction, so the honest finding is no significant circularity.
Assumptions & free parameters
free parameters (3)
- subgraph target size K =
not reported per task
- positive/negative ratio in detection task =
1:1
- replacement source for negative sampling =
another entity not in current subgraph
assumptions (3)
- domain assumption FB15K-237, as a multi-modal KG, contains correct factual triples with entity images and texts
- domain assumption GraphViz rendering preserves the subgraph structure and renders entity images and texts legibly
- ad hoc to paper A random replacement of an entity or relation creates a factually incorrect subgraph detectable from commonsense or KG knowledge
Cite this review
Pith. "Pith review of Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation." pith.science (2026). https://pith.science/paper/4YHWWJKY
@misc{pith2026250601293,
author = {Pith},
title = {Pith review of: Abstractive Visual Understanding of Multi-modal Structured Knowledge: A New Perspective for MLLM Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/4YHWWJKY}},
note = {Machine review of arXiv:2506.01293}
}
read the original abstract
Multi-modal large language models (MLLMs) incorporate heterogeneous modalities into LLMs, enabling a comprehensive understanding of diverse scenarios and objects. Despite the proliferation of evaluation benchmarks and leaderboards for MLLMs, they predominantly overlook the critical capacity of MLLMs to comprehend world knowledge with structured abstractions that appear in visual form. To address this gap, we propose a novel evaluation paradigm and devise M3STR, an innovative benchmark grounded in the Multi-Modal Map for STRuctured understanding. This benchmark leverages multi-modal knowledge graphs to synthesize images encapsulating subgraph architectures enriched with multi-modal entities. M3STR necessitates that MLLMs not only recognize the multi-modal entities within the visual inputs but also decipher intricate relational topologies among them. We delineate the benchmark's statistical profiles and automated construction pipeline, accompanied by an extensive empirical analysis of 26 state-of-the-art MLLMs. Our findings reveal persistent deficiencies in processing abstractive visual information with structured knowledge, thereby charting a pivotal trajectory for advancing MLLMs' holistic reasoning capacities. Our code and data are released at https://github.com/zjukg/M3STR
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Marah I Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harki- rat S. Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Mar- tin Cai, Caio César Teodoro Mendes, Weizhu Chen, Vishrav Chaudhary, Parul Chopra, Allie Del Giorno, Gustavo de Rosa, Matthew Dixon, Ro...
arXiv 2024
-
[2]
Bollacker, Colin Evans, Praveen K
Kurt D. Bollacker, Colin Evans, Praveen K. Paritosh, Tim Sturge, and Jamie Taylor
-
[3]
Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Ok- sana Yakhnenko. 2013. Translating Embeddings for Modeling Multi-relational Data. In NIPS. 2787–2795
work page 2013
-
[4]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yimin Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi Wang, Wenqi Shao, Pei Chu, Zhongying Tu, Tong He, Zhiyong Wu, Huipeng Deng, Jia...
arXiv 2024
-
[5]
Pan, Ningyu Zhang, and Huajun Chen
Zhuo Chen, Yichi Zhang, Yin Fang, Yuxia Geng, Lingbing Guo, Xiang Chen, Qian Li, Wen Zhang, Jiaoyan Chen, Yushan Zhu, Jiaqi Li, Xiaoze Liu, Jeff Z. Pan, Ningyu Zhang, and Huajun Chen. 2024. Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey. CoRR abs/2402.05391 (2024)
arXiv 2024
-
[6]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023. In- structBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. In NeurIPS
work page 2023
-
[7]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv:2010.11929 [cs.CV] https://arxiv.org/abs/2010.11929
arXiv 2021
-
[8]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
arXiv 2024
Show all 54 references
-
[9]
Chaoyou Fu, Yifan Zhang, Shukang Yin, Bo Li, Xinyu Fang, Sirui Zhao, Haodong Duan, Xing Sun, Ziwei Liu, Liang Wang, Caifeng Shan, and Ran He. 2024. MME- Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs. CoRR abs/2411.15296 (2024)
2024 arXiv
-
[10]
Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, and Xiangyu Yue. 2025. OneLLM: One Framework to Align All Modalities with Language. arXiv:2312.03700 [cs.CV] https://arxiv.org/ abs/2312.03700
2025 arXiv
-
[11]
Zheqi He, Xinya Wu, Pengfei Zhou, Richeng Xuan, Guang Liu, Xi Yang, Qiannan Zhu, and Hua Huang. 2024. CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning. In IJCAI. ijcai.org, 830–838
2024
-
[12]
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, Yi Ren, Yuexian Zou, Zhou Zhao, and Shinji Watanabe. 2024. AudioGPT: Understanding and Generat- ing Speech, Music, Sound, and Talking Head. InAA...
2024
-
[13]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling Laws for Neural Language Models. CoRR abs/2001.08361 (2020)
2020 arXiv
-
[14]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Mem- ory Management for Large Language Model Serving with PagedAttention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating...
2023
-
[15]
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. 2018. A dataset of clinically generated visual questions and answers about radiology images. Scientific data 5, 1 (2018), 1–10. Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Trovato et al
2018
-
[16]
Ke Liang, Lingyuan Meng, Meng Liu, Yue Liu, Wenxuan Tu, Siwei Wang, Sihang Zhou, Xinwang Liu, Fuchun Sun, and Kunlun He. 2024. A survey of knowl- edge graph reasoning on graph types: Static, dynamic, and multi-modal. IEEE Transactions on Pattern Analysis and Machine Intelligen...
2024
-
[17]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023. Improved Baselines with Visual Instruction Tuning
2023
-
[18]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruc- tion Tuning. In NeurIPS
2023
-
[19]
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin
-
[20]
Rosenblum
Ye Liu, Hui Li, Alberto García-Durán, Mathias Niepert, Daniel Oñoro-Rubio, and David S. Rosenblum. 2019. MMKG: Multi-modal Knowledge Graphs. In ESWC (Lecture Notes in Computer Science, Vol. 11503) . Springer, 459–474
2019
-
[21]
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. 2024. DeepSeek-VL: Towards Real-World Vision-Language Understanding. arXiv:2403.05525 [cs.AI]
2024 arXiv
-
[22]
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2024. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. In ICLR. OpenReview.net
2024
-
[23]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In NeurIPS
2022
-
[24]
Yougang Lyu, Lingyong Yan, Shuaiqiang Wang, Haibo Shi, Dawei Yin, Pengjie Ren, Zhumin Chen, Maarten de Rijke, and Zhaochun Ren. 2024. KnowTuning: Knowledge-aware Fine-tuning for Large Language Models. In EMNLP. Associa- tion for Computational Linguistics, 14535–14556
2024
-
[25]
Joty, and Enamul Hoque
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq R. Joty, and Enamul Hoque
-
[26]
OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)
2023 arXiv
-
[27]
OpenAI. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] https://arxiv.org/ abs/2410.21276
2024 arXiv
-
[28]
Haojie Pan, Yuzhou Zhang, Zepeng Zhai, Ruiji Fu, Ming Liu, Yangqiu Song, Zhongyuan Wang, and Bing Qin. 2022. Kuaipedia: a Large-scale Multi-modal Short-video Encyclopedia. CoRR abs/2211.00732 (2022)
2022 arXiv
-
[29]
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu
-
[30]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.000...
2021 arXiv
-
[31]
Shezheng Song, Xiaopeng Li, and Shasha Li. 2023. How to Bridge the Gap between Modalities: A Comprehensive Survey on Multimodal Large Language Model. CoRR abs/2311.07594 (2023)
2023 arXiv
- [32]
-
[33]
IEEE Trans
Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Trans. Knowl. Data Eng. 36, 7 (2024), 3580–3599
2024
-
[34]
Chawla, and Panpan Xu
Yijun Tian, Huan Song, Zichen Wang, Haozhu Wang, Ziqing Hu, Fang Wang, Nitesh V. Chawla, and Panpan Xu. 2024. Graph Neural Prompting with Large Language Models. In AAAI. AAAI Press, 19080–19088
2024
-
[35]
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. 2024. Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs. In CVPR. IEEE, 9568–9578
2024
-
[36]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-VL: Enhancing Vision-Language Mode...
2024 arXiv
-
[37]
Qwen Team. 2025. Qwen2.5-VL. https://qwenlm.github.io/blog/qwen2.5-vl/
2025
-
[38]
Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Haotian Liu, Sadhika Malladi, Alexis Chevalier, Sanjeev Arora, and Danqi Chen. 2024. CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs. In NeurIPS
2024
-
[39]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[40]
Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, Zhenda Xie, Yu Wu, Kai Hu, Jiawei Wang, Yaofeng Sun, Yukun Li, Yishi Piao, Kang Guan, Aixin Liu, Xin Xie, Yuxiang You, Kai Dong, Xingkai Yu, Haowei Zhang,...
2024 arXiv
-
[41]
Xin Wang, Benyuan Meng, Hong Chen, Yuan Meng, Ke Lv, and Wenwu Zhu
-
[42]
Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, and Wenqi Shao. 2024. MMT-Bench: A...
2024
-
[43]
Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2024. MMM...
2024
-
[44]
Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, Peng Jin, Wenqi Zhang, Fan Wang, Lidong Bing, and Deli Zhao. 2025. VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understandi...
2025 arXiv
-
[45]
Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Shaokai Chen, Mengshu Sun, Binbin Hu, Zhiqiang Zhang, Lei Liang, Wen Zhang, and Huajun Chen. 2025. Have We Designed Generalizable Structural Knowledge Promptings? Systematic Evaluation and Rethinking. CoRR abs/2501.00244 (2025)
2025 arXiv
-
[46]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2...
2024 arXiv
-
[47]
Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Wen Zhang, and Huajun Chen. 2024. Making Large Language Models Perform Better in Knowledge Graph Completion. In ACM Multimedia. ACM, 233–242
2024
-
[48]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv:2304.10592 [cs.CV] https://arxiv.org/abs/2304.10592
2023 arXiv
-
[49]
Xun Zhu, Zheng Zhang, Xi Chen, Yiming Shi, Miao Li, and Ji Wu. 2025. Connector- S: A Survey of Connectors in Multi-modal Large Language Models. CoRR abs/2502.11453 (2025)
2025 arXiv
-
[51]
Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, and Huajun Chen. 2025. Tokenization, Fusion, and Augmentation: To- wards Fine-grained Multi-modal Entity Representation. In AAAI. AAAI Press, 13322–13330
2025
-
[2008]
In SIGMOD Conference
Freebase: a collaboratively created graph database for structuring human knowledge. In SIGMOD Conference. ACM, 1247–1250
-
[2022]
In ACL (Findings)
ChartQA: A Benchmark for Question Answering about Charts with Vi- sual and Logical Reasoning. In ACL (Findings). Association for Computational Linguistics, 2263–2279
-
[2023]
In ACM Multimedia
TIVA-KG: A Multimodal Knowledge Graph with Text, Image, Video and Audio. In ACM Multimedia. ACM, 2391–2399
-
[2024]
In ECCV (6) (Lecture Notes in Computer Science, Vol
MMBench: Is Your Multi-modal Model an All-Around Player?. In ECCV (6) (Lecture Notes in Computer Science, Vol. 15064) . Springer, 216–233
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.