Pith. sign in

REVIEW 3 major objections 2 minor 1 cited by

SEAM benchmark finds vision models consistently lag language on semantically equivalent tasks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

SEAM measures VLM reasoning consistency across modalities using paired semantically equivalent textual and visual notations, and finds systematic vision-language imbalance.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection SEAM is a well-motivated benchmark, but its core premise—semantic equivalence across notation systems—is asserted, not validated, so the headline finding is provisional. the 3 major comments →

arxiv 2508.18179 v1 pith:AC2I3VWD submitted 2025-08-25 cs.AI cs.CV

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models

classification cs.AI cs.CV
keywords vision-language modelsmodality imbalancebenchmarksemantic equivalencecross-modal reasoningvisual-spatial reasoningtextual-symbolic reasoninghallucination
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces SEAM, a benchmark built from pairs of problems that are meant to contain exactly the same information, one written in a standard symbolic notation and the other in a visual-spatial notation such as a diagram. Using these matched pairs, the paper measures whether vision-language models reason as well from pictures as from text. Across 21 contemporary models it finds a systematic gap: accuracy is higher on the textual form, and a given model often gives different answers to the two forms of the same problem. The authors trace the gap to perception failures in both modalities and show it survives visual transformations, so it is not a rendering artifact.

Core claim

SEAM's central claim is that contemporary vision-language models are not modality-agnostic: when the same problem is posed in text and in an equivalent visual form, models answer the text version correctly more often, and their two answers frequently disagree. The benchmark avoids the usual confound of OCR-style image-text pairs by using distinct notation systems for each modality, so the information content is intended to be identical rather than one being a picture of the other. The paper reports that across 21 models this vision-lags-language pattern is systematic, and error analysis identifies tokenization failures in domain notation for the textual side and visual perception failures th

What carries the argument

The load-bearing object is the SEAM item pair: a problem expressed in a standardized textual-symbolic notation (for example, an equation or symbolic expression) and in a matching visual-spatial notation (for example, a geometry or diagrammatic form). Because the two notations are different in kind, any accuracy difference between them cannot be explained by one being a photocopy of the other; the pair is designed to hold the semantic content fixed while changing only the modality, so the gap measures reasoning rather than information asymmetry.

Load-bearing premise

The textual and visual versions of each problem must carry exactly the same information, so any accuracy gap between them is due to modality, not to one version being easier.

What would settle it

Have a panel of human experts solve the same SEAM items in both forms; if their accuracy is also lower on the visual form by a similar margin, the two notations are not informationally equivalent and the reported model gap would not prove a modality reasoning deficit.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the SEAM result holds, current VLMs are not modality-agnostic, and their visual reasoning is weaker than their textual reasoning on content-equivalent problems.
  • SEAM gives the field a controlled testbed for measuring modality-agnostic reasoning, so future models can be compared on the same semantically equivalent pairs.
  • The error analysis points to concrete improvement targets: fixing domain-notation tokenization for text and reducing visual hallucinations for images.
  • Robustness to visual transformations implies the modality gap is tied to semantic content and reasoning, not to how the image happens to be rendered.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If human performance on the same pairs reproduces the vision-lags-language gap, the equivalence assumption would be suspect and the benchmark would be measuring task difficulty rather than model bias.
  • Extending SEAM to other standardized notation pairs, such as musical scores or chemical structures, would test whether the modality imbalance generalizes beyond the four domains studied.
  • If the gap persists after improving visual encoders, the bottleneck may be cross-modal alignment rather than perception, implying that training should pair symbolic and visual formats explicitly.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper introduces SEAM, a benchmark of paired visual and textual inputs that are intended to be semantically equivalent across four domains with standardized notation. It evaluates 21 contemporary vision-language models, reporting a systematic modality imbalance (vision lagging language), low cross-modal agreement, and robustness to visual transformations. The central claim is that SEAM enables a controlled, semantically equivalent comparison of textual-symbolic versus visual-spatial reasoning, isolating modality effects from content differences.

Significance. If the equivalence claim is validated, SEAM addresses a well-known confound in VLM evaluation—the asymmetric information between text and images—and would be a useful tool for measuring modality-agnostic reasoning. The breadth across 21 models and the attention to visual transformations are strengths. However, the evidence for the central premise is not reported in the abstract; the benchmark's value and the interpretation of the empirical results depend on an independent demonstration of semantic equivalence.

major comments (3)
  1. [Abstract, 'semantically equivalent inputs'] The central assumption of information equivalence is asserted but not demonstrated. The abstract contrasts SEAM with OCR-based pairing and claims the inputs are semantically equivalent, yet no construction protocol, human rating, inversion check, or information-theoretic measure is reported. If the visual notation carries extra spatial cues (e.g., implicit geometric relations) or the textual notation is more compact or less ambiguous, the observed 'vision frequently lags language' could be a task-difficulty artifact. The manuscript must describe how equivalence was established independently of model performance and provide validation evidence, or substantially qualify the interpretation of the modality gap.
  2. [Abstract, 'vision frequently lags language' and 'cross-modal agreement relatively low'] These aggregate claims are based on 21 models, but no effect sizes, confidence intervals, or paired significance tests are reported. Since the same models are evaluated on matched content, paired statistical comparisons are straightforward and should be given. The manuscript should also specify the evaluation metric (accuracy, calibration, etc.), how ties are handled, and how the 'largely robust to visual transformations' claim is supported, including the list of transformations and statistical controls.
  3. [Abstract, error analysis] The error analysis attributes the modality gap to two factors: textual perception failures from tokenization and visual perception failures inducing hallucinations. No coding methodology, inter-rater reliability, or concrete examples are provided. Without a systematic and falsifiable error taxonomy, this explanation cannot be assessed. The full manuscript should describe the error annotation process, operational definitions, and example items.
minor comments (2)
  1. [Abstract, 'OCR-based image-text pairing'] This term is not defined. Briefly explain what OCR-based pairing means and how SEAM's notation-based pairing differs from it, especially for readers not familiar with existing VLM benchmarks.
  2. [General] For a benchmark paper, consider explicitly stating a plan to release the benchmark instances, construction code, and evaluation scripts to facilitate independent validation of the equivalence claim.

Circularity Check

0 steps flagged

No significant circularity; the paper reports an empirical benchmark result rather than deriving its conclusion from its own assumptions.

full rationale

The paper's primary claim is an empirical measurement: across 21 models, vision frequently lags language on paired inputs that the authors constructed to be semantically equivalent. This is not a derivation from the definition of the benchmark; it is an observed outcome, and nothing in the abstract indicates that the observed modality gap is forced by construction or by a fitted parameter renamed as a prediction. The load-bearing premise of semantic equivalence is asserted rather than independently verified, but an unvalidated assumption is not circularity under the stated criteria. No equations, fitted parameters, or self-citation chains are present in the available evidence to show that any result reduces to its own inputs. The 'semantically equivalent' characterization is a design assumption, not a result derived from the measurements; even if it were false, the appropriate critique would be construct validity or task-difficulty confounding, not circular reasoning. Because the available manuscript text contains no specific reduction of a claimed derivation to its inputs, the circularity score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central design assumption is the semantic equivalence of paired inputs. No free parameters or invented entities are disclosed in the abstract.

axioms (1)
  • domain assumption Paired textual and visual notations are semantically equivalent.
    The benchmark design relies on the assumption that the same information is conveyed exactly in both modalities. This is stated in the abstract: 'pairs semantically equivalent inputs across four domains that have existing standardized textual and visual notations.'

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models." pith.science (2026). https://pith.science/paper/AC2I3VWD

@misc{pith2026250818179,
  author       = {Pith},
  title        = {Pith review of: SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AC2I3VWD}},
  note         = {Machine review of arXiv:2508.18179}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences and asymmetric information. We introduce SEAM, a benchmark that pairs semantically equivalent inputs across four domains that have existing standardized textual and visual notations. By employing distinct notation systems across modalities, in contrast to OCR-based image-text pairing, SEAM provides a rigorous comparative assessment of the textual-symbolic and visual-spatial reasoning capabilities of VLMs. Across 21 contemporary models, we observe systematic modality imbalance: vision frequently lags language in overall performance, despite the problems containing semantically equivalent information, and cross-modal agreement is relatively low. Our error analysis reveals two main drivers: textual perception failures from tokenization in domain notation and visual perception failures that induce hallucinations. We also show that our results are largely robust to visual transformations. SEAM establishes a controlled, semantically equivalent setting for measuring and improving modality-agnostic reasoning.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

    cs.CL 2026-05 conditional novelty 7.0

    TokenSwap measures and mitigates the MLLM modality gap: swapping textual concepts for matched images lowers accuracy by 4-47% across 42 models, and training with such swaps reduces the gap.

Reference graph

Works this paper leans on

121 extracted references · 27 canonical work pages · cited by 1 Pith paper · 2 internal anchors

  1. [1]

    Pixtral 12b

    Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, et al. Pixtral 12b. arXiv preprint arXiv:2410.07073, 2024

  2. [2]

    Flamingo: a visual language model for few-shot learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35: 0 23716--23736, 2022

  3. [3]

    The claude 3 model family: Opus, sonnet, haiku, 2024

    Anthropic. The claude 3 model family: Opus, sonnet, haiku, 2024. URL https://assets.anthropic.com/m/61e7d27f8c8f5919/original/Claude-3-Model-Card.pdf

  4. [4]

    Claude 3.7 sonnet system card, 2025 a

    Anthropic. Claude 3.7 sonnet system card, 2025 a . URL https://anthropic.com/claude-3-7-sonnet-system-card

  5. [5]

    System card: Claude opus 4 & claude sonnet 4, 2025 b

    Anthropic. System card: Claude opus 4 & claude sonnet 4, 2025 b . URL https://www-cdn.anthropic.com/4263b940cabb546aa0e3283f35b686f4f3b2ff47.pdf

  6. [6]

    Vqa: Visual question answering

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pp.\ 2425--2433, 2015

  7. [7]

    Openflamingo: An open-source framework for training large autoregressive vision-language models

    Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al. Openflamingo: An open-source framework for training large autoregressive vision-language models. arXiv preprint arXiv:2308.01390, 2023

  8. [8]

    Qwen technical report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023

  9. [9]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025

  10. [10]

    Minigpt-v2: large language model as a unified interface for vision-language multi-task learning

    Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint arXiv:2310.09478, 2023

  11. [11]

    Uniter: Universal image-text representation learning

    Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pp.\ 104--120. Springer, 2020

  12. [12]

    Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024 a

  13. [13]

    How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

    Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Science China Information Sciences, 67 0 (12): 0 220101, 2024 b

  14. [14]

    Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 24185--24198, 2024 c

  15. [15]

    Chess.com, 2025

    Chess.com . Chess.com, 2025. URL https://www.chess.com/. Accessed: March 2025

  16. [16]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025

  17. [17]

    Opencompass: A universal evaluation platform for foundation models

    OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023

  18. [18]

    Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges

    Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, and Huaxiu Yao. Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges. arXiv preprint arXiv:2311.03287, 2023

  19. [19]

    music21: A toolkit for computer-aided musicology and symbolic music analysis

    Michael Scott Cuthbert and Christopher Ariza. music21: A toolkit for computer-aided musicology and symbolic music analysis. In 11th International Society for Music Information Retrieval Conference, pp.\ 637--642, 2010

  20. [20]

    Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023. URL https://arxiv.org/abs/2305.06500

  21. [21]

    Gemini: a family of highly capable multimodal models

    Google DeepMind. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023

  22. [22]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Google DeepMind. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024 a

  23. [23]

    Gemma: Open models based on gemini research and technology

    Google DeepMind. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024 b

  24. [24]

    Gemma 2: Improving open language models at a practical size

    Google DeepMind. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024 c

  25. [25]

    Introducing gemini 2.0: our new ai model for the agentic era, 2025 a

    Google DeepMind. Introducing gemini 2.0: our new ai model for the agentic era, 2025 a . URL https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/

  26. [26]

    Gemma 3 technical report, 2025 b

    Google DeepMind. Gemma 3 technical report, 2025 b . URL https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf

  27. [27]

    Portable game notation specification and implementation guide, 1994

    Steven J Edwards. Portable game notation specification and implementation guide, 1994. Chess programming standard

  28. [28]

    Pmr: Prototypical modal rebalance for multimodal learning

    Yunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang, and Song Guo. Pmr: Prototypical modal rebalance for multimodal learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 20029--20038, 2023

  29. [29]

    Layoutgpt: Compositional visual planning and generation with large language models

    Weixi Feng, Wanrong Zhu, Tsu-jui Fu, Varun Jampani, Arjun Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang. Layoutgpt: Compositional visual planning and generation with large language models. Advances in Neural Information Processing Systems, 36: 0 18225--18250, 2023

  30. [30]

    python-chess: A chess library for python, 2025

    Niklas Fiekas. python-chess: A chess library for python, 2025. URL https://python-chess.readthedocs.io/. Accessed: March 2025

  31. [31]

    Blink: Multimodal large language models can see but not perceive

    Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large language models can see but not perceive. In European Conference on Computer Vision, pp.\ 148--166. Springer, 2024

  32. [32]

    Llama-adapter v2: Parameter-efficient visual instruction model

    Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010, 2023

  33. [33]

    Making the v in vqa matter: Elevating the role of image understanding in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp.\ 6904--6913, 2017

  34. [34]

    The llama 3 herd of models

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024

  35. [35]

    What can large language models do in chemistry? a comprehensive benchmark on eight tasks

    Taicheng Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh Chawla, Olaf Wiest, Xiangliang Zhang, et al. What can large language models do in chemistry? a comprehensive benchmark on eight tasks. Advances in Neural Information Processing Systems, 36: 0 59662--59688, 2023

  36. [36]

    Exploring network structure, dynamics, and function using networkx

    Aric Hagberg, Daniel Schult, and Pieter Swart. Exploring network structure, dynamics, and function using networkx. In Proceedings of the 7th Python in Science Conference, pp.\ 11--15, 2008

  37. [37]

    Chartllama: A multimodal llm for chart understanding and generation

    Yucheng Han, Chi Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang. Chartllama: A multimodal llm for chart understanding and generation. arXiv preprint arXiv:2311.16483, 2023

  38. [38]

    Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark

    Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Yu Cheng. Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark. arXiv preprint arXiv:2501.05444, 2025

  39. [39]

    Deciphering cross-modal alignment in large vision-language models with modality integration rate

    Qidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. Deciphering cross-modal alignment in large vision-language models with modality integration rate. arXiv preprint arXiv:2410.07167, 2024

  40. [40]

    Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably)

    Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang, and Longbo Huang. Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably). In International conference on machine learning, pp.\ 9226--9259. PMLR, 2022

  41. [41]

    Mantis: Interleaved multi-image instruction tuning

    Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. Mantis: Interleaved multi-image instruction tuning. arXiv preprint arXiv:2405.01483, 2024

  42. [42]

    Spin: Sparsifying and integrating internal neurons in large language models for text classification

    Difan Jiao, Yilun Liu, Zhenwei Tang, Daniel Matter, J \"u rgen Pfeffer, and Ashton Anderson. Spin: Sparsifying and integrating internal neurons in large language models for text classification. In Findings of the Association for Computational Linguistics ACL 2024, pp.\ 4666--4682, 2024

  43. [43]

    Pubchem 2025 update

    Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, et al. Pubchem 2025 update. Nucleic Acids Research, 53 0 (D1): 0 D1516--D1525, 2025

  44. [44]

    Large language models are zero-shot reasoners

    Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35: 0 22199--22213, 2022

  45. [45]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009

  46. [46]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023

  47. [47]

    Snap: A general-purpose network analysis and graph-mining library

    Jure Leskovec and Rok Sosi c . Snap: A general-purpose network analysis and graph-mining library. ACM Transactions on Intelligent Systems and Technology (TIST), 8 0 (1): 0 1--20, 2016

  48. [48]

    Llava-onevision: Easy visual task transfer

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024 a

  49. [49]

    Seed-bench: Benchmarking multimodal large language models

    Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. Seed-bench: Benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.\ 13299--13308, 2024 b

  50. [50]

    Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models

    Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024 c

  51. [51]

    Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning, pp.\ 12888--12900. PMLR, 2022

  52. [52]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp.\ 19730--19742. PMLR, 2023

  53. [53]

    Oscar: Object-semantics aligned pre-training for vision-language tasks

    Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XXX 16, pp.\ 121--137. Springer, 2020

  54. [54]

    Lichess evaluation database, 2025 a

    Lichess Team . Lichess evaluation database, 2025 a . URL https://database.lichess.org/#evals. Accessed: March 2025

  55. [55]

    Lichess: Free online chess, 2025 b

    Lichess Team . Lichess: Free online chess, 2025 b . URL https://lichess.org/. Accessed: March 2025

  56. [56]

    Lichess puzzle database, 2025 c

    Lichess Team . Lichess puzzle database, 2025 c . URL https://database.lichess.org/#puzzles. Accessed: March 2025

  57. [57]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \'a r, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer vision--ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, pp.\ 740--755. Springer, 2014

  58. [58]

    Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models

    Fuxiao Liu, Tianrui Guan, Zongxia Li, Lichang Chen, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models. arXiv preprint arXiv:2310.14566, 2 0 (3): 0 9, 2023 a

  59. [59]

    Improved baselines with visual instruction tuning, 2023 b

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023 b

  60. [60]

    Visual instruction tuning, 2023 c

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning, 2023 c

  61. [61]

    Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a . URL https://llava-vl.github.io/blog/2024-01-30-llava-next/

  62. [62]

    Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233

    Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pp.\ 216--233. Springer, 2024 b

  63. [63]

    Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019

  64. [64]

    Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

    Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023

  65. [65]

    Ok-vqa: A visual question answering benchmark requiring external knowledge

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pp.\ 3195--3204, 2019

  66. [66]

    Gpt-4v(ision) system card, 2023

    OpenAI. Gpt-4v(ision) system card, 2023. URL https://cdn.openai.com/papers/GPTV_System_Card.pdf

  67. [67]

    Gpt-4o system card

    OpenAI. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024 a

  68. [68]

    Gpt-4o mini: Advancing cost-efficient intelligence, 2024 b

    OpenAI. Gpt-4o mini: Advancing cost-efficient intelligence, 2024 b . URL https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/

  69. [69]

    Openai o1 system card, December 2024 c

    OpenAI. Openai o1 system card, December 2024 c . URL https://cdn.openai.com/o1-system-card-20241205.pdf

  70. [70]

    Gpt-4.5 system card, 2025 a

    OpenAI. Gpt-4.5 system card, 2025 a . URL https://cdn.openai.com/gpt-4-5-system-card-2272025.pdf

  71. [71]

    Gpt-5 system card, 2025 b

    OpenAI. Gpt-5 system card, 2025 b . URL https://cdn.openai.com/pdf/8124a3ce-ab78-4f06-96eb-49ea29ffb52f/gpt5-system-card-aug7.pdf

  72. [72]

    Openai o3-mini system card, February 2025 c

    OpenAI. Openai o3-mini system card, February 2025 c . URL https://cdn.openai.com/o3-mini-system-card-feb10.pdf

  73. [73]

    Openai o3 and o4-mini system card, 2025 d

    OpenAI. Openai o3 and o4-mini system card, 2025 d . URL https://cdn.openai.com/pdf/2221c875-02dc-4789-800b-e7758f3722c1/o3-and-o4-mini-system-card.pdf

  74. [74]

    Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment

    Rohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov, and Louis-Philippe Morency. Cross-modal attention congruence regularization for vision-language relation alignment. arXiv preprint arXiv:2212.10549, 2022

  75. [75]

    Generalizing from simple to hard visual reasoning: Can we mitigate modality imbalance in vlms? arXiv preprint arXiv:2501.02669, 2025

    Simon Park, Abhishek Panigrahi, Yun Cheng, Dingli Yu, Anirudh Goyal, and Sanjeev Arora. Generalizing from simple to hard visual reasoning: Can we mitigate modality imbalance in vlms? arXiv preprint arXiv:2501.02669, 2025

  76. [76]

    Balanced multimodal learning via on-the-fly gradient modulation

    Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. Balanced multimodal learning via on-the-fly gradient modulation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp.\ 8238--8247, 2022

  77. [77]

    Rdkit: Open-source cheminformatics, 2025

    RDKit Development Team . Rdkit: Open-source cheminformatics, 2025. URL https://www.rdkit.org/. Accessed: March 2025

  78. [78]

    The music encoding initiative (mei)

    Perry Roland. The music encoding initiative (mei). In Proceedings of the First International Conference on Musical Applications Using XML, volume 1060, pp.\ 55--59. Citeseer, 2002

  79. [79]

    Music abc notation with music theory dataset, 2025

    Seeker38. Music abc notation with music theory dataset, 2025. URL https://huggingface.co/datasets/Seeker38/music_abc_notation_with_music_theory. Accessed: March 2025

  80. [80]

    Lxmert: Learning cross-modality encoder representations from transformers

    Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.