REVIEW 3 major objections 2 minor 1 cited by
SEAM benchmark finds vision models consistently lag language on semantically equivalent tasks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection SEAM is a well-motivated benchmark, but its core premise—semantic equivalence across notation systems—is asserted, not validated, so the headline finding is provisional. the 3 major comments →
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
SEAM's central claim is that contemporary vision-language models are not modality-agnostic: when the same problem is posed in text and in an equivalent visual form, models answer the text version correctly more often, and their two answers frequently disagree. The benchmark avoids the usual confound of OCR-style image-text pairs by using distinct notation systems for each modality, so the information content is intended to be identical rather than one being a picture of the other. The paper reports that across 21 models this vision-lags-language pattern is systematic, and error analysis identifies tokenization failures in domain notation for the textual side and visual perception failures th
What carries the argument
The load-bearing object is the SEAM item pair: a problem expressed in a standardized textual-symbolic notation (for example, an equation or symbolic expression) and in a matching visual-spatial notation (for example, a geometry or diagrammatic form). Because the two notations are different in kind, any accuracy difference between them cannot be explained by one being a photocopy of the other; the pair is designed to hold the semantic content fixed while changing only the modality, so the gap measures reasoning rather than information asymmetry.
Load-bearing premise
The textual and visual versions of each problem must carry exactly the same information, so any accuracy gap between them is due to modality, not to one version being easier.
What would settle it
Have a panel of human experts solve the same SEAM items in both forms; if their accuracy is also lower on the visual form by a similar margin, the two notations are not informationally equivalent and the reported model gap would not prove a modality reasoning deficit.
If this is right
- If the SEAM result holds, current VLMs are not modality-agnostic, and their visual reasoning is weaker than their textual reasoning on content-equivalent problems.
- SEAM gives the field a controlled testbed for measuring modality-agnostic reasoning, so future models can be compared on the same semantically equivalent pairs.
- The error analysis points to concrete improvement targets: fixing domain-notation tokenization for text and reducing visual hallucinations for images.
- Robustness to visual transformations implies the modality gap is tied to semantic content and reasoning, not to how the image happens to be rendered.
Where Pith is reading between the lines
- If human performance on the same pairs reproduces the vision-lags-language gap, the equivalence assumption would be suspect and the benchmark would be measuring task difficulty rather than model bias.
- Extending SEAM to other standardized notation pairs, such as musical scores or chemical structures, would test whether the modality imbalance generalizes beyond the four domains studied.
- If the gap persists after improving visual encoders, the bottleneck may be cross-modal alignment rather than perception, implying that training should pair symbolic and visual formats explicitly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SEAM, a benchmark of paired visual and textual inputs that are intended to be semantically equivalent across four domains with standardized notation. It evaluates 21 contemporary vision-language models, reporting a systematic modality imbalance (vision lagging language), low cross-modal agreement, and robustness to visual transformations. The central claim is that SEAM enables a controlled, semantically equivalent comparison of textual-symbolic versus visual-spatial reasoning, isolating modality effects from content differences.
Significance. If the equivalence claim is validated, SEAM addresses a well-known confound in VLM evaluation—the asymmetric information between text and images—and would be a useful tool for measuring modality-agnostic reasoning. The breadth across 21 models and the attention to visual transformations are strengths. However, the evidence for the central premise is not reported in the abstract; the benchmark's value and the interpretation of the empirical results depend on an independent demonstration of semantic equivalence.
major comments (3)
- [Abstract, 'semantically equivalent inputs'] The central assumption of information equivalence is asserted but not demonstrated. The abstract contrasts SEAM with OCR-based pairing and claims the inputs are semantically equivalent, yet no construction protocol, human rating, inversion check, or information-theoretic measure is reported. If the visual notation carries extra spatial cues (e.g., implicit geometric relations) or the textual notation is more compact or less ambiguous, the observed 'vision frequently lags language' could be a task-difficulty artifact. The manuscript must describe how equivalence was established independently of model performance and provide validation evidence, or substantially qualify the interpretation of the modality gap.
- [Abstract, 'vision frequently lags language' and 'cross-modal agreement relatively low'] These aggregate claims are based on 21 models, but no effect sizes, confidence intervals, or paired significance tests are reported. Since the same models are evaluated on matched content, paired statistical comparisons are straightforward and should be given. The manuscript should also specify the evaluation metric (accuracy, calibration, etc.), how ties are handled, and how the 'largely robust to visual transformations' claim is supported, including the list of transformations and statistical controls.
- [Abstract, error analysis] The error analysis attributes the modality gap to two factors: textual perception failures from tokenization and visual perception failures inducing hallucinations. No coding methodology, inter-rater reliability, or concrete examples are provided. Without a systematic and falsifiable error taxonomy, this explanation cannot be assessed. The full manuscript should describe the error annotation process, operational definitions, and example items.
minor comments (2)
- [Abstract, 'OCR-based image-text pairing'] This term is not defined. Briefly explain what OCR-based pairing means and how SEAM's notation-based pairing differs from it, especially for readers not familiar with existing VLM benchmarks.
- [General] For a benchmark paper, consider explicitly stating a plan to release the benchmark instances, construction code, and evaluation scripts to facilitate independent validation of the equivalence claim.
Circularity Check
No significant circularity; the paper reports an empirical benchmark result rather than deriving its conclusion from its own assumptions.
full rationale
The paper's primary claim is an empirical measurement: across 21 models, vision frequently lags language on paired inputs that the authors constructed to be semantically equivalent. This is not a derivation from the definition of the benchmark; it is an observed outcome, and nothing in the abstract indicates that the observed modality gap is forced by construction or by a fitted parameter renamed as a prediction. The load-bearing premise of semantic equivalence is asserted rather than independently verified, but an unvalidated assumption is not circularity under the stated criteria. No equations, fitted parameters, or self-citation chains are present in the available evidence to show that any result reduces to its own inputs. The 'semantically equivalent' characterization is a design assumption, not a result derived from the measurements; even if it were false, the appropriate critique would be construct validity or task-difficulty confounding, not circular reasoning. Because the available manuscript text contains no specific reduction of a claimed derivation to its inputs, the circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption Paired textual and visual notations are semantically equivalent.
Cite this review
Pith. "Pith review of SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models." pith.science (2026). https://pith.science/paper/AC2I3VWD
@misc{pith2026250818179,
author = {Pith},
title = {Pith review of: SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/AC2I3VWD}},
note = {Machine review of arXiv:2508.18179}
}
read the original abstract
Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences and asymmetric information. We introduce SEAM, a benchmark that pairs semantically equivalent inputs across four domains that have existing standardized textual and visual notations. By employing distinct notation systems across modalities, in contrast to OCR-based image-text pairing, SEAM provides a rigorous comparative assessment of the textual-symbolic and visual-spatial reasoning capabilities of VLMs. Across 21 contemporary models, we observe systematic modality imbalance: vision frequently lags language in overall performance, despite the problems containing semantically equivalent information, and cross-modal agreement is relatively low. Our error analysis reveals two main drivers: textual perception failures from tokenization in domain notation and visual perception failures that induce hallucinations. We also show that our results are largely robust to visual transformations. SEAM establishes a controlled, semantically equivalent setting for measuring and improving modality-agnostic reasoning.
Forward citations
Cited by 1 Pith paper
-
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
TokenSwap measures and mitigates the MLLM modality gap: swapping textual concepts for matched images lowers accuracy by 4-47% across 42 models, and training with such swaps reduces the gap.
Reference graph
Works this paper leans on
-
[1]
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, et al. Pixtral 12b. arXiv preprint arXiv:2410.07073, 2024
Pith/arXiv arXiv 2024
-
[2]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35: 0 23716--23736, 2022
2022
-
[3]
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic. The claude 3 model family: Opus, sonnet, haiku, 2024. URL https://assets.anthropic.com/m/61e7d27f8c8f5919/original/Claude-3-Model-Card.pdf
2024
-
[4]
Claude 3.7 sonnet system card, 2025 a
Anthropic. Claude 3.7 sonnet system card, 2025 a . URL https://anthropic.com/claude-3-7-sonnet-system-card
2025
-
[5]
System card: Claude opus 4 & claude sonnet 4, 2025 b
Anthropic. System card: Claude opus 4 & claude sonnet 4, 2025 b . URL https://www-cdn.anthropic.com/4263b940cabb546aa0e3283f35b686f4f3b2ff47.pdf
2025
-
[6]
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pp.\ 2425--2433, 2015
2015
-
[7]
Openflamingo: An open-source framework for training large autoregressive vision-language models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al. Openflamingo: An open-source framework for training large autoregressive vision-language models. arXiv preprint arXiv:2308.01390, 2023
Pith/arXiv arXiv 2023
-
[8]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023
Pith/arXiv arXiv 2023
-
[9]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025
Pith/arXiv arXiv 2025
-
[10]
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint arXiv:2310.09478, 2023
Pith/arXiv arXiv 2023
-
[11]
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pp.\ 104--120. Springer, 2020
2020
-
[12]
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024 a
Pith/arXiv arXiv 2024
-
[13]
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Science China Information Sciences, 67 0 (12): 0 220101, 2024 b
2024
-
[14]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 24185--24198, 2024 c
2024
-
[15]
Chess.com, 2025
Chess.com . Chess.com, 2025. URL https://www.chess.com/. Accessed: March 2025
2025
-
[16]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025
Pith/arXiv arXiv 2025
-
[17]
Opencompass: A universal evaluation platform for foundation models
OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023
2023
-
[18]
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, and Huaxiu Yao. Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges. arXiv preprint arXiv:2311.03287, 2023
Pith/arXiv arXiv 2023
-
[19]
music21: A toolkit for computer-aided musicology and symbolic music analysis
Michael Scott Cuthbert and Christopher Ariza. music21: A toolkit for computer-aided musicology and symbolic music analysis. In 11th International Society for Music Information Retrieval Conference, pp.\ 637--642, 2010
2010
-
[20]
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023. URL https://arxiv.org/abs/2305.06500
Pith/arXiv arXiv 2023
-
[21]
Gemini: a family of highly capable multimodal models
Google DeepMind. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023
Pith/arXiv arXiv 2023
-
[22]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Google DeepMind. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024 a
Pith/arXiv arXiv 2024
-
[23]
Gemma: Open models based on gemini research and technology
Google DeepMind. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024 b
Pith/arXiv arXiv 2024
-
[24]
Gemma 2: Improving open language models at a practical size
Google DeepMind. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024 c
Pith/arXiv arXiv 2024
-
[25]
Introducing gemini 2.0: our new ai model for the agentic era, 2025 a
Google DeepMind. Introducing gemini 2.0: our new ai model for the agentic era, 2025 a . URL https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/
2025
-
[26]
Gemma 3 technical report, 2025 b
Google DeepMind. Gemma 3 technical report, 2025 b . URL https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf
2025
-
[27]
Portable game notation specification and implementation guide, 1994
Steven J Edwards. Portable game notation specification and implementation guide, 1994. Chess programming standard
1994
-
[28]
Pmr: Prototypical modal rebalance for multimodal learning
Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang, and Song Guo. Pmr: Prototypical modal rebalance for multimodal learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 20029--20038, 2023
2023
-
[29]
Layoutgpt: Compositional visual planning and generation with large language models
Weixi Feng, Wanrong Zhu, Tsu-jui Fu, Varun Jampani, Arjun Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang. Layoutgpt: Compositional visual planning and generation with large language models. Advances in Neural Information Processing Systems, 36: 0 18225--18250, 2023
2023
-
[30]
python-chess: A chess library for python, 2025
Niklas Fiekas. python-chess: A chess library for python, 2025. URL https://python-chess.readthedocs.io/. Accessed: March 2025
2025
-
[31]
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large language models can see but not perceive. In European Conference on Computer Vision, pp.\ 148--166. Springer, 2024
2024
-
[32]
Llama-adapter v2: Parameter-efficient visual instruction model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010, 2023
Pith/arXiv arXiv 2023
-
[33]
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 6904--6913, 2017
2017
-
[34]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024
Pith/arXiv arXiv 2024
-
[35]
What can large language models do in chemistry? a comprehensive benchmark on eight tasks
Taicheng Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh Chawla, Olaf Wiest, Xiangliang Zhang, et al. What can large language models do in chemistry? a comprehensive benchmark on eight tasks. Advances in Neural Information Processing Systems, 36: 0 59662--59688, 2023
2023
-
[36]
Exploring network structure, dynamics, and function using networkx
Aric Hagberg, Daniel Schult, and Pieter Swart. Exploring network structure, dynamics, and function using networkx. In Proceedings of the 7th Python in Science Conference, pp.\ 11--15, 2008
2008
-
[37]
Chartllama: A multimodal llm for chart understanding and generation
Yucheng Han, Chi Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang. Chartllama: A multimodal llm for chart understanding and generation. arXiv preprint arXiv:2311.16483, 2023
Pith/arXiv arXiv 2023
-
[38]
Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark
Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Yu Cheng. Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark. arXiv preprint arXiv:2501.05444, 2025
Pith/arXiv arXiv 2025
-
[39]
Deciphering cross-modal alignment in large vision-language models with modality integration rate
Qidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. Deciphering cross-modal alignment in large vision-language models with modality integration rate. arXiv preprint arXiv:2410.07167, 2024
Pith/arXiv arXiv 2024
-
[40]
Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably)
Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang, and Longbo Huang. Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably). In International conference on machine learning, pp.\ 9226--9259. PMLR, 2022
2022
-
[41]
Mantis: Interleaved multi-image instruction tuning
Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. Mantis: Interleaved multi-image instruction tuning. arXiv preprint arXiv:2405.01483, 2024
Pith/arXiv arXiv 2024
-
[42]
Spin: Sparsifying and integrating internal neurons in large language models for text classification
Difan Jiao, Yilun Liu, Zhenwei Tang, Daniel Matter, J \"u rgen Pfeffer, and Ashton Anderson. Spin: Sparsifying and integrating internal neurons in large language models for text classification. In Findings of the Association for Computational Linguistics ACL 2024, pp.\ 4666--4682, 2024
2024
-
[43]
Pubchem 2025 update
Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. Pubchem 2025 update. Nucleic Acids Research, 53 0 (D1): 0 D1516--D1525, 2025
2025
-
[44]
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 0 22199--22213, 2022
2022
-
[45]
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009
2009
-
[46]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023
2023
-
[47]
Snap: A general-purpose network analysis and graph-mining library
Jure Leskovec and Rok Sosi c . Snap: A general-purpose network analysis and graph-mining library. ACM Transactions on Intelligent Systems and Technology (TIST), 8 0 (1): 0 1--20, 2016
2016
-
[48]
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024 a
Pith/arXiv arXiv 2024
-
[49]
Seed-bench: Benchmarking multimodal large language models
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. Seed-bench: Benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 13299--13308, 2024 b
2024
-
[50]
Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024 c
Pith/arXiv arXiv 2024
-
[51]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning, pp.\ 12888--12900. PMLR, 2022
2022
-
[52]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp.\ 19730--19742. PMLR, 2023
2023
-
[53]
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XXX 16, pp.\ 121--137. Springer, 2020
2020
-
[54]
Lichess evaluation database, 2025 a
Lichess Team . Lichess evaluation database, 2025 a . URL https://database.lichess.org/#evals. Accessed: March 2025
2025
-
[55]
Lichess: Free online chess, 2025 b
Lichess Team . Lichess: Free online chess, 2025 b . URL https://lichess.org/. Accessed: March 2025
2025
-
[56]
Lichess puzzle database, 2025 c
Lichess Team . Lichess puzzle database, 2025 c . URL https://database.lichess.org/#puzzles. Accessed: March 2025
2025
-
[57]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \'a r, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer vision--ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, pp.\ 740--755. Springer, 2014
2014
-
[58]
Fuxiao Liu, Tianrui Guan, Zongxia Li, Lichang Chen, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models. arXiv preprint arXiv:2310.14566, 2 0 (3): 0 9, 2023 a
Pith/arXiv arXiv 2023
-
[59]
Improved baselines with visual instruction tuning, 2023 b
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023 b
2023
-
[60]
Visual instruction tuning, 2023 c
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning, 2023 c
2023
-
[61]
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a . URL https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
-
[62]
Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233. Springer, 2024 b
2024
-
[63]
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019
2019
-
[64]
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023
Pith/arXiv arXiv 2023
-
[65]
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pp.\ 3195--3204, 2019
2019
-
[66]
Gpt-4v(ision) system card, 2023
OpenAI. Gpt-4v(ision) system card, 2023. URL https://cdn.openai.com/papers/GPTV_System_Card.pdf
work page 2023
-
[67]
OpenAI. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024 a
Pith/arXiv arXiv 2024
-
[68]
Gpt-4o mini: Advancing cost-efficient intelligence, 2024 b
OpenAI. Gpt-4o mini: Advancing cost-efficient intelligence, 2024 b . URL https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/
work page 2024
-
[69]
Openai o1 system card, December 2024 c
OpenAI. Openai o1 system card, December 2024 c . URL https://cdn.openai.com/o1-system-card-20241205.pdf
work page 2024
-
[70]
OpenAI. Gpt-4.5 system card, 2025 a . URL https://cdn.openai.com/gpt-4-5-system-card-2272025.pdf
work page 2025
-
[71]
OpenAI. Gpt-5 system card, 2025 b . URL https://cdn.openai.com/pdf/8124a3ce-ab78-4f06-96eb-49ea29ffb52f/gpt5-system-card-aug7.pdf
work page 2025
-
[72]
Openai o3-mini system card, February 2025 c
OpenAI. Openai o3-mini system card, February 2025 c . URL https://cdn.openai.com/o3-mini-system-card-feb10.pdf
work page 2025
-
[73]
Openai o3 and o4-mini system card, 2025 d
OpenAI. Openai o3 and o4-mini system card, 2025 d . URL https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf
work page 2025
-
[74]
Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment
Rohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov, and Louis-Philippe Morency. Cross-modal attention congruence regularization for vision-language relation alignment. arXiv preprint arXiv:2212.10549, 2022
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[75]
Simon Park, Abhishek Panigrahi, Yun Cheng, Dingli Yu, Anirudh Goyal, and Sanjeev Arora. Generalizing from simple to hard visual reasoning: Can we mitigate modality imbalance in vlms? arXiv preprint arXiv:2501.02669, 2025
Pith/arXiv arXiv 2025
-
[76]
Balanced multimodal learning via on-the-fly gradient modulation
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. Balanced multimodal learning via on-the-fly gradient modulation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 8238--8247, 2022
work page 2022
-
[77]
Rdkit: Open-source cheminformatics, 2025
RDKit Development Team . Rdkit: Open-source cheminformatics, 2025. URL https://www.rdkit.org/. Accessed: March 2025
work page 2025
-
[78]
The music encoding initiative (mei)
Perry Roland. The music encoding initiative (mei). In Proceedings of the First International Conference on Musical Applications Using XML, volume 1060, pp.\ 55--59. Citeseer, 2002
work page 2002
-
[79]
Music abc notation with music theory dataset, 2025
Seeker38. Music abc notation with music theory dataset, 2025. URL https://huggingface.co/datasets/Seeker38/music_abc_notation_with_music_theory. Accessed: March 2025
work page 2025
-
[80]
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019
Pith/arXiv arXiv 1908
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.