Pith. sign in

REVIEW 4 major objections 4 minor 205 references

Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read General-purpose multimodal models fall short in specialized fields, so this survey organizes the domain-specific benchmarks that measure that gap into a single eight-discipline taxonomy.

desk verdict A workmanlike survey that will be handy as a pointer resource, but its central coverage claim is not checkable as written and several table entries are off-scope. read the letter →

arxiv 2506.12958 v2 pith:HWYET6G2 submitted 2025-06-15 cs.LG

classification cs.LG
keywords multimodallargelanguagemodelsdomain-specificbenchmarksbenchmarktaxonomyLLMevaluationsurveyartificialgeneralintelligencedomainadaptation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is trying to establish that general-purpose multimodal large language models (MLLMs) are not enough: measuring and advancing progress in specialized fields requires domain-specific benchmarks. It surveys the benchmark literature across the eight disciplines listed in the methodology—engineering, science, technology, mathematics, humanities, finance, healthcare, and language understanding—and organizes the results into a taxonomy of domains, sub-domains, and application areas. The paper's contribution to a reader is a consolidated map of which benchmarks exist, what input modalities they use, and where current models succeed or fail. A sympathetic reading is that it makes the 'last mile problem' concrete by assembling repeated evidence that strong general models underperform when faced with specialized data and reasoning.

What carries the argument

The organizing device is the eight-branch domain hierarchy (Figure 1): disciplines → domains → sub-domains → application areas. Each branch is paired with a consolidation table that records the benchmark's scale, task type, input modality, models evaluated, and performance. That hierarchy carries the survey's argument: it converts scattered benchmark papers into a structured picture of where the MLLM evaluation ecosystem is dense, where it is thin, and where model failure is systematic.

What would settle it

A systematic replication that runs the paper's own search terms against the same databases and screens for domain-specific MLLM benchmarks could settle the coverage claim: finding a substantial discipline or an established benchmark family that fits none of the eight taxonomy branches would refute comprehensiveness. A cheaper check is row-level: every benchmark named in the domain tables should be traceable to a paper that actually introduces or evaluates it, and any miscategorized row would weaken the resource's reliability.

Watch

Extended reading notes

Core claim

The central claim is that the field has reached the point where general benchmarks no longer tell the full story: MLLMs need domain-specific benchmarks to expose and guide their specialized capabilities. The paper's positive contribution is a taxonomy, announced as seven disciplines in the abstract but presented as eight in the methodology and Figure 1, with per-domain tables consolidating each benchmark's scale, task type, input modality, models evaluated, and reported performance. Across the domains the assembled evidence shows a recurring pattern: frontier models handle high-level reasoning and familiar text well, but stumble on fine-grained perception, specialized data formats, and domain-specific reasoning—low accuracy on financial question answering, weak geospatial localization, pathology understanding far below human experts, and poor materials property prediction. The survey frames these failures not as isolated results but as a systematic gap that domain-specific benchmarking is meant to close.

Load-bearing premise

The load-bearing premise is that the search strategy described in Section 2 captured the full landscape of domain-specific MLLM benchmarks, so that the taxonomy and summary tables are representative and complete; the paper's own inconsistency between seven disciplines (abstract) and eight (methodology) shows those coverage boundaries are not clearly pinned down.

Editorial extensions

If this is right

  • If the taxonomy is accurate, a researcher entering an unfamiliar domain can locate the relevant benchmarks and their reported baselines in one place, lowering the cost of designing a new evaluation.
  • The cross-domain pattern of failures—text shortcuts in pathology, poor counting and localization in geospatial data, low accuracy on financial QA—implies that benchmark design should include controls that isolate each input modality.
  • The survey's evidence supports making future benchmarks 'living' and multimodal, and scoring robustness, efficiency, and safety as well as accuracy.
  • Domain-specific benchmark results argue for domain-adapted fine-tuning and hybrid pipelines (LLM plus retrieval, symbolic solvers, or verification tools) rather than relying on a generalist model alone.
  • If the map is complete, it gives a concrete way to track whether MLLM progress toward generally capable AI is actually broadening across disciplines, the paper's stated long-term aim.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: its own tables suggest a testable hypothesis that MLLM performance on a benchmark depends less on model size than on how closely the task's input format and vocabulary match the model's training distribution.
  • The abstract's seven-discipline count versus the methodology's eight suggests the taxonomy's boundaries were still shifting; a natural extension is a living, community-maintained registry that updates the map as new benchmarks appear.
  • The recurring 'text shortcut' failure—models answering from language cues instead of analyzing images—implies that future benchmarks should include diagnostic variants that remove one modality at a time, so scores cannot be gamed by linguistic priors.
  • The survey's cited fine-tuning results imply that domain benchmarks can serve as training signal, not just evaluation tools; adversarially noisy or domain-specific examples measurably improve robustness and accuracy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper is a survey of domain-specific benchmarks for multimodal large language models (MLLMs). It proposes a taxonomy of disciplines—Engineering, Science, Technology, Mathematics, Humanities, Finance, Healthcare, and Language Understanding—and, for each, provides summary tables with scale, task type, input modality, models, performance, and key focus, along with narrative discussion of trends and limitations. The stated contribution is a comprehensive, accessible resource that maps the domain-specific evaluation landscape and highlights where current MLLMs succeed or fail in specialized fields.

Significance. If the survey's scope and categorization were reliable, it would be a useful entry point for researchers seeking domain-specific evaluation resources. The paper's strength is its breadth: it organizes a large number of benchmarks and studies across eight disciplines and identifies recurring gaps (e.g., in medical image reasoning and in modality-dependent performance). However, the value of the survey depends entirely on the reproducibility of its selection process and the accuracy of its categorizations; the current inconsistencies substantially weaken the central claim of comprehensiveness. The paper does not provide machine-checked proofs or reproducible code; its contribution is the structured synthesis itself.

major comments (4)
  1. [§1 (Abstract) vs §2] The abstract states that the paper introduces 'a taxonomy of seven key disciplines,' while Section 2 explicitly lists eight disciplines and Figure 1 displays eight. This is a direct internal contradiction about the scope of the survey. Since the paper's central claim is that it provides a comprehensive taxonomy, the intended number of disciplines must be clarified; as written, a reader cannot tell whether one section is extraneous or the abstract is wrong.
  2. [§2, Search Strategy and Scope] The search strategy paragraph names databases and general keyword categories but provides no screening criteria, deduplication procedure, date range, number of records retrieved, number of records screened, or exclusion log. The paper claims to be a 'systematic review' and concludes that its resource is comprehensive, but without these details the coverage is not checkable or reproducible. The claim that the review 'examines eight key disciplines' and that the tables consolidate 'relevant benchmarks and survey papers' is therefore not falsifiable; the authors should add a principled inclusion/exclusion protocol and a flow diagram or equivalent transparency measure.
  3. [Tables 1, 2, and 8] Several entries in the summary tables are not domain-specific benchmarks, which contradicts the paper's stated scope. Table 1 places BIG-bench under Software Engineering/Knowledge Graphs & Semantic Systems, but BIG-bench is a general-purpose benchmark spanning 204 tasks across many areas. Table 2 includes WeatherBench 2, a weather forecasting benchmark that is not an LLM or MLLM evaluation benchmark. Table 8 lists ToT (a prompting method), LongLLaVA, LLaVA-OneVision, KOSMOS-1, KOSMOS-2, and ChatGLM (models or architectures) as if they were benchmarks. These inclusions make the consolidated resource unreliable and undermine the central claim that the paper catalogs domain-specific benchmarks for evaluating MLLMs.
  4. [§1.2, paragraph 3] The manuscript states: "OpenAI's GPT-4o (referred to as 'o3') at 69.1%." This is a factual error: GPT-4o and o3 are different models, and the parenthetical misattributes a reported score. In a survey whose purpose is to summarize model performance on benchmarks, such a mislabeling is a load-bearing accuracy problem because readers may rely on the reported comparisons for model selection. The sentence should be corrected and the source of the 69.1% figure should be cited.
minor comments (4)
  1. [Figure 1] The label 'Chain & Crypto' should read 'Blockchain & Cryptocurrency' to match the terminology used in Section 5.3 and Table 3.
  2. [Throughout] The model name 'LLaVA' is repeatedly typeset as 'LLaV A' with an erroneous space; please correct this globally.
  3. [§6.1 vs §8.2.1] The benchmark cited as [104] is called 'KnowledgeMath' in Table 4 and Section 6.3 but 'FinanceMath' in Table 6; the text in §6.3 notes the alternative name, but the tables should use one canonical name or explicitly cross-reference the alias.
  4. [Table 6] The row for FinanceBench lists 'GPT-4+retriever' with performance '19% correct'; the surrounding text (Section 8.2.3) says GPT-4 provided correct responses to only 19% of questions, but the table does not indicate whether this is the best result among the evaluated models or a representative one; please clarify.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey organizes external benchmarks and derives no predictions from fitted inputs or self-citation chains.

full rationale

This paper is a survey and taxonomy of domain-specific MLLM benchmarks. It contains no equations, fitted parameters, or derived quantitative predictions whose outputs could reduce to its inputs by construction. The central contribution is an organizing framework: the authors select eight disciplines, assign existing benchmark papers to sub-domains, and summarize their characteristics. The taxonomy is a descriptive categorization of external work, not a result derived from those same papers, so the self-definitional and fitted-input patterns do not apply. There is no load-bearing appeal to a prior uniqueness theorem by the same authors, and no ansatz is smuggled in via citation; the benchmark descriptions are taken from the cited papers themselves, which are independent external sources. The paper's coverage claim is supported by a stated search strategy over standard databases, and while that strategy is high-level, the absence of screening detail is a reproducibility or rigor concern, not a circularity concern. Internal inconsistencies such as the abstract stating 'seven key disciplines' while Section 2 and Figure 1 list eight, and table entries that are models or methods rather than benchmarks (e.g., ToT, LongLLaVA, KOSMOS-1, ChatGLM) or general-purpose rather than domain-specific (e.g., BIG-bench under Software Engineering), are scope and consistency defects. They do not make the survey's claims equivalent to its inputs by definition. Under the hard rule requiring a specific reduction to label circularity, no such reduction can be quoted from this paper. The honest finding is therefore no significant circularity, with a score of 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper is a survey, so the central claim depends on domain assumptions about completeness and classification rather than on fitted parameters or invented entities. No equations or predictions are made.

assumptions (2)
  • domain assumption The eight discipline taxonomy is complete and non-overlapping.
    The entire survey structure and all summary tables rest on this classification, introduced in Section 2 and Figure 1, but no external criterion justifies the choice of eight versus seven disciplines.
  • domain assumption The literature search and inclusion criteria capture all or most relevant domain-specific benchmarks.
    Section 2 describes queries to IEEE Xplore, ACM, Scopus, and arXiv with general keywords, but gives no reproducible protocol or count of screened papers, so representativeness is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Domain Specific Benchmarks for Evaluating Multimodal Large Language Models." pith.science (2026). https://pith.science/paper/HWYET6G2

@misc{pith2026250612958,
  author       = {Pith},
  title        = {Pith review of: Domain Specific Benchmarks for Evaluating Multimodal Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HWYET6G2}},
  note         = {Machine review of arXiv:2506.12958}
}
read the original abstract

Large language models (LLMs) are increasingly being deployed across disciplines due to their advanced reasoning and problem solving capabilities. To measure their effectiveness, various benchmarks have been developed that measure aspects of LLM reasoning, comprehension, and problem-solving. While several surveys address LLM evaluation and benchmarks, a domain-specific analysis remains underexplored in the literature. This paper introduces a taxonomy of seven key disciplines, encompassing various domains and application areas where LLMs are extensively utilized. Additionally, we provide a comprehensive review of LLM benchmarks and survey papers within each domain, highlighting the unique capabilities of LLMs and the challenges faced in their application. Finally, we compile and categorize these benchmarks by domain to create an accessible resource for researchers, aiming to pave the way for advancements toward artificial general intelligence (AGI)

Figures

Figures reproduced from arXiv: 2506.12958 by the authors.

Figure 1
Figure 1. The domain hierarchy of the eight disciplines covered in this article. The hierarchy is based on the domains and sub-domains of each discipline, and the [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

205 extracted references · 11 canonical work pages

  1. [1]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al., Gpt-4 technical report, arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Milli- can, et al., Gemini: a family of highly capable multimodal models, arXiv preprint arXiv:2312.11805 (2023)

  3. [3]

    Bommasani, D

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosse- lut, E. Brunskill, et al., On the opportunities and risks of foundation models, arXiv preprint arXiv:2108.07258 (2021)

  4. [4]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Ka- dian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al., The llama 3 herd of models, arXiv preprint arXiv:2407.21783 (2024)

  5. [5]

    Abdin, J

    M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kauffmann, et al., Phi-4 technical report, arXiv preprint arXiv:2412.08905 (2024)

  6. [6]

    Islam, A

    P. Islam, A. Kannappan, D. Kiela, R. Qian, N. Scherrer, B. Vidgen, Financebench: A new benchmark for financial question answering, arXiv preprint (2023). arXiv:2311. 11944

  7. [7]

    Zheng, Y

    D. Zheng, Y . Wang, E. Shi, X. Liu, Y . Ma, H. Zhang, Z. Zheng, Top general performance=top domain perfor- mance? domaincodebench: A multi-domain code genera- tion benchmark (2025).arXiv:2412.18573. URLhttps://arxiv.org/abs/2412.18573

  8. [8]

    J. Li, Y . Zhu, Z. Xu, J. Gu, M. Zhu, X. Liu, N. Liu, Y . Peng, F. Feng, J. Tang, Mmro: Are multimodal llms eligible as the brain for in-home robotics?, arXiv preprint arXiv:2406.19693 (2024)

Show all 205 references
  1. [9]

    A. C. Doris, D. Grandi, R. Tomich, M. F. Alam, M. Ataei, H. Cheong, F. Ahmed, Designqa: A multimodal bench- mark for evaluating large language models’ understanding of engineering documentation, Journal of Computing and Information Science in Engineering 25 (2) (2024) 021009. ...

  2. [10]

    Kernan Freire, C

    S. Kernan Freire, C. Wang, M. Foosherian, S. Wellsandt, S. Ruiz-Arenas, E. Niforatos, Knowledge sharing in man- ufacturing using llm-powered tools: user study and model benchmarking, Frontiers in Artificial intelligence 7 (2024) 1293084

  3. [11]

    X. Liu, S. Yang, X. Dong, H. Rong, B. Fu, Manu-eval: A chinese language understanding benchmark for manu- facturing industry, in: China Conference on Knowledge Graph and Semantic Computing, Springer, 2024, pp. 309– 317

  4. [12]

    Eslaminia, A

    A. Eslaminia, A. Jackson, B. Tian, A. Stern, H. Gor- don, R. Malhotra, K. Nahrstedt, C. Shao, Fdm-bench: A comprehensive benchmark for evaluating large language models in additive manufacturing tasks, arXiv preprint arXiv:2412.09819 (2024)

  5. [13]

    Fakih, R

    M. Fakih, R. Dharmaji, Y . Moghaddas, G. Quiros, O. Ogundare, M. A. Al Faruque, Llm4plc: Harnessing large language models for verifiable programming of plcs in industrial control systems, in: Proceedings of the 46th International Conference on Software Engineering: Software En...

  6. [14]

    Tizaoui, R

    T. Tizaoui, R. Tan, Towards a benchmark dataset for large language models in the context of process automation, Digital Chemical Engineering (2024) 100186

  7. [15]

    Y . Xia, J. Zhang, N. Jazdi, M. Weyrich, Incorporating large language models into production systems for en- hanced task automation and flexibility, arXiv preprint arXiv:2407.08550 (2024)

  8. [16]

    Ogundare, S

    O. Ogundare, S. Madasu, N. Wiggins, Industrial engi- neering with large language models: A case study of chatgpt’s performance on oil & gas problems, in: 2023 11th International Conference on Control, Mechatronics and Automation (ICCMA), IEEE, 2023, pp. 458–461

  9. [17]

    S. A. Rahman, S. Chawla, M. Yaqot, B. Menezes, Lever- aging large language models for supply chain manage- ment optimization: A case study, in: International Confer- ence on Innovative Intelligent Industrial Production and Logistics, Springer, 2024, pp. 175–197

  10. [18]

    B. Li, K. Mellou, B. Zhang, J. Pathuri, I. Menache, Large language models for supply chain optimization, arXiv preprint arXiv:2307.03875 (2023)

  11. [19]

    Raman, A

    R. Raman, A. Sreenivasan, M. Suresh, A. Gunasekaran, P. Nedungadi, Ai-driven education: a comparative study on chatgpt and bard in supply chain management contexts, Cogent Business & Management 11 (1) (2024) 2412742. 32

  12. [20]

    Meyer, J

    L.-P. Meyer, J. Frey, K. Junghanns, F. Brei, K. Bulert, S. Gründer-Fahrer, M. Martin, Developing a scalable benchmark for assessing large language models in knowl- edge graph engineering (2023).arXiv:2308.16622. URLhttps://arxiv.org/abs/2308.16622

  13. [21]

    bench authors, Beyond the imitation game: Quantify- ing and extrapolating the capabilities of language models, Transactions on Machine Learning Research (2023)

    B. bench authors, Beyond the imitation game: Quantify- ing and extrapolating the capabilities of language models, Transactions on Machine Learning Research (2023). URL https://openreview.net/forum?id= uyTL5Bvosj

  14. [22]

    Azanza, B

    M. Azanza, B. P. Lamancha, E. Pizarro, Tracking the moving target: A framework for continuous evaluation of llm test generation in industry (2025). arXiv:2504. 18985. URLhttps://arxiv.org/abs/2504.18985

  15. [23]

    N. Shah, Z. Genc, D. Araci, Stackeval: Benchmarking llms in coding assistance, Advances in Neural Informa- tion Processing Systems 37 (2024) 36976–36994

  16. [24]

    Zhang, C

    Q. Zhang, C. Fang, Y . Xie, Y . Zhang, Y . Yang, W. Sun, S. Yu, Z. Chen, A survey on large language models for software engineering (2024).arXiv:2312.15223. URLhttps://arxiv.org/abs/2312.15223

  17. [25]

    R. Bell, R. Longshore, R. Madachy, Introducing syseng- bench: A novel benchmark for assessing large language models in systems engineering, Tech. rep., Acquisition Research Program (2024)

  18. [26]

    Y . Hu, Y . Goktas, D. D. Yellamati, C. De Tassigny, The use and misuse of pre-trained generative large language models in reliability engineering, in: 2024 Annual Reli- ability and Maintainability Symposium (RAMS), IEEE, 2024, pp. 1–7

  19. [27]

    Vendrow, E

    J. Vendrow, E. Vendrow, S. Beery, A. Madry, Do large lan- guage model benchmarks test reliability?, arXiv preprint arXiv:2502.03461 (2025)

  20. [28]

    Y . Liu, Y . Yao, J.-F. Ton, X. Zhang, R. Guo, H. Cheng, Y . Klochkov, M. F. Taufiq, H. Li, Trustworthy llms: a sur- vey and guideline for evaluating large language models’ alignment (2024).arXiv:2308.05374. URLhttps://arxiv.org/abs/2308.05374

  21. [29]

    J. A. Irvin, E. R. Liu, J. C. Chen, I. Dormoy, J. Kim, S. Khanna, Z. Zheng, S. Ermon, TEOChat: A Large Vision-Language Assistant for Temporal Earth Observa- tion Data, _eprint: 2410.06234 (2024). URLhttps://arxiv.org/abs/2410.06234

  22. [30]

    Xiong, F

    Z. Xiong, F. Zhang, Y . Wang, Y . Shi, X. X. Zhu, Earth- Nets: Empowering artificial intelligence for Earth obser- vation, IEEE Geoscience and Remote Sensing Magazine (2024) 2–36doi:10.1109/MGRS.2024.3466998

  23. [31]

    Zhang, S

    C. Zhang, S. Wang, Good at captioning, bad at counting: Benchmarking gpt-4v on earth observation data, arXiv preprint arXiv:2401.17600 (2024)

  24. [32]

    W. Li, D. Yao, R. Zhao, W. Chen, Z. Xu, C. Luo, C. Gong, Q. Jing, H. Tan, J. Bi, STBench: Assessing the Ability of Large Language Models in Spatio-Temporal Analysis, arXiv preprint arXiv:2406.19065 (2024)

  25. [33]

    Sapkota, R

    R. Sapkota, R. Qureshi, S. Z. Hassan, J. Shutske, M. Shoman, M. Sajjad, F. A. Dharejo, A. Paudel, J. Li, Z. Meng, others, Multi-modal LLMs in agriculture: A comprehensive review, Authorea PreprintsPublisher: Au- thorea (2024)

  26. [34]

    M. S. Danish, M. A. Munir, S. R. A. Shah, K. Kuck- reja, F. S. Khan, P. Fraccaro, A. Lacoste, S. Khan, GEOBench-VLM: Benchmarking Vision-Language Mod- els for Geospatial Tasks, _eprint: 2411.19325 (2024). URLhttps://arxiv.org/abs/2411.19325

  27. [35]

    Roberts, T

    J. Roberts, T. Lüddecke, R. Sheikh, K. Han, S. Albanie, Charting new territories: Exploring the geographic and geospatial capabilities of multimodal llms, in: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 554–563

  28. [36]

    X. Liu, Z. Lian, RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mix- ture of Experts, _eprint: 2412.05679 (2024). URLhttps://arxiv.org/abs/2412.05679

  29. [37]

    C. Lin, H. Lyu, X. Xu, J. Luo, INS-MMBench: A Com- prehensive Benchmark for Evaluating LVLMs’ Perfor- mance in Insurance, arXiv preprint arXiv:2406.09105 (2024)

  30. [38]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, J. Steinhardt, Measuring massive multitask lan- guage understanding, arXiv preprint arXiv:2009.03300 (2020)

  31. [39]

    P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, A. Kalyan, Learn to explain: Multi- modal reasoning via thought chains for science question answering, Advances in Neural Information Processing Systems 35 (2022) 2507–2521

  32. [40]

    D. Fu, R. Guo, G. Khalighinejad, O. Liu, B. Dhingra, D. Yogatama, R. Jia, W. Neiswanger, Isobench: Bench- marking multimodal foundation models on isomorphic representations, arXiv preprint arXiv:2404.01266 (2024)

  33. [41]

    Jiang, Z

    Z. Jiang, Z. Yang, J. Chen, Z. Du, W. Wang, B. Xu, J. Tang, Visscience: An extensive benchmark for evalu- ating k12 educational multi-modal scientific reasoning, arXiv preprint arXiv:2409.13730 (2024)

  34. [42]

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y . Pang, J. Dirani, J. Michael, S. R. Bowman, Gpqa: A graduate- level google-proof q&a benchmark, in: First Conference on Language Modeling, 2024. 33

  35. [43]

    Anand, J

    A. Anand, J. Kapuriya, A. Singh, J. Saraf, N. Lal, A. Verma, R. Gupta, R. Shah, Mm-phyqa: Multimodal physics question-answering with multi-image cot prompt- ing, in: Pacific-Asia Conference on Knowledge Discovery and Data Mining, Springer, 2024, pp. 53–64

  36. [44]

    Mirza, N

    A. Mirza, N. Alampara, S. Kunchapu, M. Ríos-García, B. Emoekabu, A. Krishnan, T. Gupta, M. Schilling- Wilhelmi, M. Okereke, A. Aneesh, et al., Are large language models superhuman chemists?, arXiv preprint arXiv:2404.01475 (2024)

  37. [45]

    S. Zhu, X. Liu, G. Khalighinejad, Chemqa: a multimodal question-and-answering dataset on chemistry reasoning, https://huggingface.co/datasets/shangzhu/ ChemQA(2024)

  38. [46]

    T. Guo, B. Nan, Z. Liang, Z. Guo, N. Chawla, O. Wiest, X. Zhang, et al., What can large language models do in chemistry? a comprehensive benchmark on eight tasks, Advances in Neural Information Processing Systems 36 (2023) 59662–59688

  39. [47]

    B. Yu, F. N. Baker, Z. Chen, X. Ning, H. Sun, Llasmol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tun- ing dataset, arXiv preprint arXiv:2402.09391 (2024)

  40. [48]

    M. Zaki, N. Krishnan, et al., Mascqa: A question answering dataset for investigating materials science knowledge of large language models, arXiv preprint arXiv:2308.09115 (2023)

  41. [49]

    A. N. Rubungo, K. Li, J. Hattrick-Simpers, A. B. Di- eng, Llm4mat-bench: benchmarking large language mod- els for materials property prediction, arXiv preprint arXiv:2411.00177 (2024)

  42. [50]

    H. Cao, Y . Shao, Z. Liu, Z. Liu, X. Tang, Y . Yao, Y . Li, Presto: progressive pretraining enhances synthetic chem- istry outcomes, arXiv preprint arXiv:2406.13193 (2024)

  43. [51]

    X. Liu, Y . Guo, H. Li, J. Liu, S. Huang, B. Ke, J. Lv, Drugllm: Open large language model for few-shot molecule generation, arXiv preprint arXiv:2405.06690 (2024)

  44. [52]

    S. Rasp, S. Hoyer, A. Merose, I. Langmore, P. Battaglia, T. Russell, A. Sanchez-Gonzalez, V . Yang, R. Carver, S. Agrawal, et al., Weatherbench 2: A benchmark for the next generation of data-driven global weather models, Journal of Advances in Modeling Earth Systems 16 (6) (20...

  45. [53]

    J. Chen, P. Zhou, Y . Hua, D. Chong, M. Cao, Y . Li, Z. Yuan, B. Zhu, J. Liang, Vision-language models meet meteorology: Developing models for extreme weather events detection with heatmaps, arXiv preprint arXiv:2406.09838 (2024)

  46. [54]

    C. Ma, Z. Hua, A. Anderson-Frey, V . Iyer, X. Liu, L. Qin, Weatherqa: Can multimodal language models reason about severe weather?, arXiv preprint arXiv:2406.11217 (2024)

  47. [55]

    H. Li, Z. Wang, J. Wang, A. K. H. Lau, H. Qu, Cllmate: A multimodal llm for weather and climate events fore- casting, arXiv preprint arXiv:2409.19058 (2024)

  48. [56]

    Y . Sun, C. Wang, Y . Peng, Unleashing the potential of large language model: Zero-shot vqa for flood disaster scenario, in: Proceedings of the 4th International Confer- ence on Artificial Intelligence and Computer Engineering, 2023, pp. 368–373

  49. [57]

    Rawat, Disasterqa: A benchmark for assessing the performance of llms in disaster response, arXiv preprint arXiv:2410.20707 (2024)

    R. Rawat, Disasterqa: A benchmark for assessing the performance of llms in disaster response, arXiv preprint arXiv:2410.20707 (2024)

  50. [58]

    Z. B. Patel, Y . Bachwana, N. Sharma, S. Gut- tikunda, N. Batra, Vayubuddy: an llm-powered chat- bot to democratize air quality insights, arXiv preprint arXiv:2411.12760 (2024)

  51. [59]

    Pafilis, S

    E. Pafilis, S. P. Frankild, L. Fanini, S. Faulwetter, C. Pavloudi, A. Vasileiadou, C. Arvanitidis, L. J. Jensen, The species and organisms resources for fast and accurate identification of taxonomic names in text, PloS one 8 (6) (2013) e65390

  52. [60]

    Abdelmageed, F

    N. Abdelmageed, F. Löffler, L. Feddoul, A. Algergawy, S. Samuel, J. Gaikwad, A. Kazem, B. König-Ries, Bio- divnere: Gold standard corpora for named entity recog- nition and relation extraction in the biodiversity domain, Biodiversity Data Journal 10 (2022) e89481

  53. [61]

    A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y . Zhao, X. Du, M. R. G. Madani, et al., Are we done with mmlu?, arXiv preprint arXiv:2406.04127 (2024)

  54. [62]

    A. M. Richard, R. Huang, S. Waidyanatha, P. Shinn, B. J. Collins, I. Thillainadarajah, C. M. Grulke, A. J. Williams, R. R. Lougee, R. S. Judson, et al., The tox21 10k com- pound library: collaborative chemistry advancing toxi- cology, Chemical Research in Toxicology 34 (2) (20...

  55. [63]

    S. Kim, P. A. Thiessen, E. E. Bolton, J. Chen, G. Fu, A. Gindulyte, L. Han, J. He, S. He, B. A. Shoemaker, et al., Pubchem substance and compound databases, Nu- cleic acids research 44 (D1) (2016) D1202–D1213

  56. [64]

    W. Jin, C. Coley, R. Barzilay, T. Jaakkola, Predicting or- ganic reaction outcomes with weisfeiler-lehman network, Advances in neural information processing systems 30 (2017)

  57. [65]

    Edwards, C

    C. Edwards, C. Zhai, H. Ji, Text2mol: Cross-modal molecule retrieval with natural language queries, in: Pro- ceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 595–607. 34

  58. [66]

    J. J. Irwin, K. G. Tang, J. Young, C. Dandarchuluun, B. R. Wong, M. Khurelbaatar, Y . S. Moroz, J. Mayfield, R. A. Sayle, Zinc20—a free ultralarge-scale chemical database for ligand discovery, Journal of chemical information and modeling 60 (12) (2020) 6065–6073

  59. [67]

    Davies, M

    M. Davies, M. Nowotka, G. Papadatos, N. Dedman, A. Gaulton, F. Atkinson, L. Bellis, J. P. Overington, Chembl web services: streamlining access to drug dis- covery data and utilities, Nucleic acids research 43 (W1) (2015) W612–W620

  60. [68]

    S. Rasp, P. D. Dueben, S. Scher, J. A. Weyn, S. Mouatadid, N. Thuerey, Weatherbench: a benchmark data set for data- driven weather forecasting, Journal of Advances in Mod- eling Earth Systems 12 (11) (2020) e2020MS002203

  61. [69]

    S. Li, W. Yang, P. Zhang, X. Xiao, D. Cao, Y . Qin, X. Zhang, Y . Zhao, P. Bogdan, Climatellm: Efficient weather forecasting via frequency-aware large language models, arXiv preprint arXiv:2502.11059 (2025)

  62. [70]

    Sachdeva, N

    E. Sachdeva, N. Agarwal, S. Chundi, S. Roelofs, J. Li, M. Kochenderfer, C. Choi, B. Dariush, Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2024, pp. 7513–7522

  63. [71]

    X. Ding, J. Han, H. Xu, X. Liang, W. Zhang, X. Li, Holistic autonomous driving understanding by bird’s-eye- view injected multi-modal large models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13668–13677

  64. [72]

    T. Qian, J. Chen, L. Zhuo, Y . Jiao, Y .-G. Jiang, NuScenes-QA: A Multi-Modal Visual Question An- swering Benchmark for Autonomous Driving Scenario, Proceedings of the AAAI Conference on Artificial Intelligence 38 (5) (2024) 4542–4550, number: 5. doi:10.1609/aaai.v38i5.28253. ...

  65. [73]

    S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, others, Cambrian-1: A fully open, vision-centric exploration of multimodal llms, arXiv preprint arXiv:2406.16860 (2024)

  66. [74]

    Q. Zhou, S. Chen, Y . Wang, H. Xu, W. Du, H. Zhang, Y . Du, J. B. Tenenbaum, C. Gan, HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments, ArxivUniv Massachusetts Amherst Peking Univ MIT MIT (2024)

  67. [75]

    Z. Yang, X. Jia, H. Li, J. Yan, LLM4Drive: A Survey of Large Language Models for Autonomous Driving, Arx- ivOpenDriveLab (2024)

  68. [76]

    Q. Kong, Y . Kawana, R. Saini, A. Kumar, J. Pan, T. Gu, Y . Ozao, B. Opra, Y . Sato, N. Kobori, WTS: A Pedestrian- Centric Traffic Video Dataset for Fine-Grained Spatial- Temporal Understanding, V ol. 15134, 2025, pp. 1–18, woven Toyota. doi:10.1007/978-3-031-73116-7_ 1

  69. [77]

    Malla, C

    S. Malla, C. Choi, I. Dwivedi, J. H. Choi, J. Li, Drama: Joint risk localization and captioning in driving, in: Pro- ceedings of the IEEE/CVF winter conference on applica- tions of computer vision, 2023, pp. 1043–1052

  70. [78]

    J. Yang, S. Gao, Y . Qiu, L. Chen, T. Li, B. Dai, K. Chitta, P. Wu, J. Zeng, P. Luo, J. Zhang, A. Geiger, Y . Qiao, H. Li, Generalized Predictive Model for Autonomous Driving, 2024, pp. 14662–14672. URL https://openaccess.thecvf.com/content/ CVPR2024/html/Yang_Generalized_Pred...

  71. [79]

    M. Nie, R. Peng, C. Wang, X. Cai, J. Han, H. Xu, L. Zhang, Reason2Drive: Towards Interpretable and Chain-Based Reasoning for Autonomous Driving, in: A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sat- tler, G. Varol (Eds.), Computer Vision – ECCV 2024, Springer Nature Swi...

  72. [80]

    C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, H. Li, DriveLM: Driving with Graph Visual Question Answering, in: A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sat- tler, G. Varol (Eds.), Computer Vision – ECCV 2024, Springer Nat...

  73. [81]

    D. Wu, W. Han, T. Wang, Y . Liu, X. Zhang, J. Shen, Language Prompt for Autonomous Driving, arXiv:2309.04379 [cs] (Sep. 2023). doi:10.48550/ arXiv.2309.04379. URLhttp://arxiv.org/abs/2309.04379

  74. [82]

    Z. Xiao, Q. Wang, H. Pearce, S. Chen, Logic meets magic: Llms cracking smart contract vulnerabilities (2025).arXiv:2501.07058. URLhttps://arxiv.org/abs/2501.07058

  75. [83]

    Z. Wei, J. Sun, Z. Zhang, X. Zhang, M. Li, Z. Hou, Llm- smartaudit: Advanced smart contract vulnerability detec- tion (2024).arXiv:2410.09381. URLhttps://arxiv.org/abs/2410.09381

  76. [84]

    Zhang, K

    L. Zhang, K. Li, K. Sun, D. Wu, Y . Liu, H. Tian, Y . Liu, Acfix: Guiding llms with mined common rbac practices for context-aware repair of access control vulnerabilities in smart contracts (2024).arXiv:2403.06838. URLhttps://arxiv.org/abs/2403.06838 35

  77. [85]

    K. I. Roumeliotis, N. D. Tselikas, D. K. Nasiopoulos, Llms and nlp models in cryptocurrency sentiment anal- ysis: A comparative classification study, Big Data and Cognitive Computing 8 (6) (2024) 63. doi:10.3390/ bdcc8060063

  78. [86]

    Makri, G

    E. Makri, G. Palaiokrassas, S. Bouraga, A. Polychroni- adou, L. Tassiulas, Ethereum price prediction employing large language models for short-term and few-shot fore- casting (2025).arXiv:2503.23190. URLhttps://arxiv.org/abs/2503.23190

  79. [87]

    Q. Wang, Y . Gao, Z. Tang, B. Luo, N. Chen, B. He, Exploring llm cryptocurrency trading through fact- subjectivity aware reasoning (2025). arXiv:2410. 12464. URLhttps://arxiv.org/abs/2410.12464

  80. [88]

    Z. He, Z. Li, S. Yang, H. Ye, A. Qiao, X. Zhang, X. Luo, T. Chen, Large language models for blockchain secu- rity: A systematic literature review (2025). arXiv: 2403.14280. URLhttps://arxiv.org/abs/2403.14280

  81. [89]

    Trozze, T

    A. Trozze, T. Davies, B. Kleinberg, Large language models in cryptocurrency securities cases: Can a gpt model meaningfully assist lawyers? (2024). arXiv: 2308.06032. URLhttps://arxiv.org/abs/2308.06032

  82. [90]

    C. Cui, Y . Ma, X. Cao, W. Ye, Y . Zhou, K. Liang, J. Chen, J. Lu, Z. Yang, K.-D. Liao, et al., A survey on multi- modal large language models for autonomous driving, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 958–979

  83. [91]

    Y . Shi, K. Jiang, J. Li, Z. Qian, J. Wen, M. Yang, K. Wang, D. Yang, Grid-centric traffic scenario percep- tion for autonomous driving: A comprehensive review, IEEE Transactions on Neural Networks and Learning Systems (2024)

  84. [92]

    P. Tong, E. Brown, P. Wu, S. Woo, A. J. V . IYER, S. C. Akula, S. Yang, J. Yang, M. Middepogu, Z. Wang, et al., Cambrian-1: A fully open, vision-centric exploration of multimodal llms, Advances in Neural Information Pro- cessing Systems 37 (2024) 87310–87356

  85. [93]

    J. Yang, S. Gao, Y . Qiu, L. Chen, T. Li, B. Dai, K. Chitta, P. Wu, J. Zeng, P. Luo, et al., Generalized predictive model for autonomous driving, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 14662–14672

  86. [94]

    R. S. Shah, K. Chawla, D. Eidnani, A. Shah, W. Du, S. Chava, N. Raman, C. Smiley, J. Chen, D. Yang, When flue meets flang: Benchmarks and large pre-trained lan- guage model for financial domain, arXiv preprint (2022). arXiv:2211.00083

  87. [95]

    W. Guan, J. Cao, S. Qian, J. Gao, C. Ouyang, Logllm: Log-based anomaly detection using large language mod- els (2025).arXiv:2411.08561. URLhttps://arxiv.org/abs/2411.08561

  88. [96]

    Sinha, A

    R. Sinha, A. Elhafsi, C. Agia, M. Foutter, E. Schmerling, M. Pavone, Real-time anomaly detection and reactive planning with large language models (2024). arXiv: 2407.08735. URLhttps://arxiv.org/abs/2407.08735

  89. [97]

    Zhang, D

    R. Zhang, D. Jiang, Y . Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K.-W. Chang, Y . Qiao, et al., Mathverse: Does your multi-modal llm truly see the diagrams in vi- sual math problems?, in: European Conference on Com- puter Vision, Springer, 2024, pp. 169–186

  90. [98]

    J. Fan, S. Martinson, E. Y . Wang, K. Hausknecht, J. Brenner, D. Liu, N. Peng, C. Wang, M. Brenner, HARDMATH: A benchmark dataset for challenging problems in applied mathematics, in: The 4th Workshop on Mathematical Reasoning and AI at NeurIPS’24, 2024. URL https://openreview....

  91. [99]

    H. Liu, Y . Zhang, Y . Luo, A. C.-C. Yao, Augmenting Math Word Problems via Iterative Question Composing, ArxivShanghai Qizhi Inst Beijing Univ Posts & Telecom- mun (2024)

  92. [100]

    Liang, D

    Z. Liang, D. Yu, W. Yu, W. Yao, Z. Zhang, X. Zhang, D. Yu, Mathchat: Benchmarking mathematical reasoning and instruction following in multi-turn interactions, arXiv preprint arXiv:2405.19444 (2024)

  93. [101]

    Zhang, L

    Z. Zhang, L. Xu, Z. Jiang, H. Hao, R. Wang, Multiple- choice questions are efficient and robust llm evaluators. 2024d, URL https://doi. org/10.48550/arXiv 2405

  94. [102]

    Anantheswaran, H

    U. Anantheswaran, H. Gupta, K. Scaria, S. Verma, C. Baral, S. Mishra, Cutting through the noise: Boosting llm performance on math word problems, arXiv preprint arXiv:2406.15444 (2024)

  95. [103]

    K. Yang, A. Swope, A. Gu, R. Chalamala, P. Song, S. Yu, S. Godil, R. J. Prenger, A. Anandkumar, Leandojo: Theo- rem proving with retrieval-augmented language models, Advances in Neural Information Processing Systems 36 (2023) 21573–21612

  96. [104]

    Y . Zhao, H. Liu, Y . Long, R. Zhang, C. Zhao, A. Cohan, Financemath: Knowledge-intensive math reasoning in finance domains, in: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 2024, pp. 12841–12858

  97. [105]

    Pezeshkpour, E

    P. Pezeshkpour, E. Hruschka, Large language models sensitivity to the order of options in multiple-choice questions, in: K. Duh, H. Gomez, S. Bethard (Eds.), Findings of the Association for Computational Linguis- tics: NAACL 2024, Association for Computational 36 Linguistics, ...

  98. [106]

    LaBelle, Monte carlo tree search applications to neural theorem proving, Ph.D

    E. LaBelle, Monte carlo tree search applications to neural theorem proving, Ph.D. thesis, Massachusetts Institute of Technology (2024)

  99. [107]

    Z. Yuan, K. Wang, S. Zhu, Y . Yuan, J. Zhou, Y . Zhu, W. Wei, Finllms: A framework for financial reasoning dataset generation with large language models, IEEE Transactions on Big Data (2024)

  100. [108]

    Y . Jin, M. Choi, G. Verma, J. Wang, S. Kumar, Mm- soc: Benchmarking multimodal large language models in social media platforms, arXiv preprint arXiv:2402.14154 (2024)

  101. [109]

    Y . Chen, S. Yan, Q. Guo, J. Jia, Z. Li, Y . Xiao, Hotv- com: Generating buzzworthy comments for videos, in: L.-W. Ku, A. Martins, V . Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 2024, pp. 2198–2224. doi: 10.18653/v1...

  102. [110]

    Y . Chen, S. Yan, Z. Zhu, Z. Li, Y . Xiao, Xmecap: Meme caption generation with sub-image adaptability, in: Pro- ceedings of the 32nd ACM International Conference on Multimedia, MM ’24, Association for Computing Ma- chinery, New York, NY , USA, 2024, pp. 3352–3361. doi:10.1145...

  103. [111]

    Shahriar, R

    S. Shahriar, R. Dara, Priv-iq: A benchmark and com- parative evaluation of large multimodal models on pri- vacy competencies, AI 6 (2) (2025) 29. doi:10.3390/ ai6020029

  104. [112]

    I. O. Gallegos, et al., Bias and fairness in large language models: A survey, Comput. Linguist. 50 (3) (2024) 1097– 1179.doi:10.1162/coli_a_00524

  105. [113]

    Liu, et al., Culturevlm: Characterizing and improving cultural understanding of vision-language models for over 100 countries, arXiv (Jan

    S. Liu, et al., Culturevlm: Characterizing and improving cultural understanding of vision-language models for over 100 countries, arXiv (Jan. 2025). arXiv:arXiv:2501. 01282,doi:10.48550/arXiv.2501.01282

  106. [114]

    Ghaboura, et al., Time travel: A comprehensive bench- mark to evaluate lmms on historical and cultural arti- facts, arXiv (Feb

    S. Ghaboura, et al., Time travel: A comprehensive bench- mark to evaluate lmms on historical and cultural arti- facts, arXiv (Feb. 2025). arXiv:arXiv:2502.14865, doi:10.48550/arXiv.2502.14865

  107. [115]

    Hu, et al., Emobench-m: Benchmarking emotional intelligence for multimodal large language models, arXiv (Feb

    H. Hu, et al., Emobench-m: Benchmarking emotional intelligence for multimodal large language models, arXiv (Feb. 2025). arXiv:arXiv:2502.04424, doi:10. 48550/arXiv.2502.04424

  108. [116]

    Y . Chen, S. Yan, S. Liu, Y . Li, Y . Xiao, Emotionqueen: A benchmark for evaluating empathy of large language mod- els, in: L.-W. Ku, A. Martins, V . Srikumar (Eds.), Find- ings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 2024, pp. 2149–2176...

  109. [117]

    Yang, et al., Editworld: Simulating world dynamics for instruction-following image editing, CoRRAccessed: Feb

    L. Yang, et al., Editworld: Simulating world dynamics for instruction-following image editing, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= aLGe5a823O

  110. [118]

    He, et al., Llms meet multimodal generation and editing: A survey, CoRRAccessed: Feb

    Y . He, et al., Llms meet multimodal generation and editing: A survey, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= sYfBSrHed5

  111. [119]

    J. Cao, Y . Liu, Y . Shi, K. Ding, L. Jin, Wenmind: A comprehensive benchmark for evaluating large language models in chinese classical literature and language arts, Adv. Neural Inf. Process. Syst. 37 (2025) 51358–51410

  112. [120]

    Li, et al., The music maestro or the musically chal- lenged, a massive music evaluation benchmark for large language models, in: L.-W

    J. Li, et al., The music maestro or the musically chal- lenged, a massive music evaluation benchmark for large language models, in: L.-W. Ku, A. Martins, V . Srikumar (Eds.), Findings of the Association for Computational Lin- guistics: ACL 2024, Bangkok, Thailand, 2024, pp. 32...

  113. [121]

    B. Weck, I. Manco, E. Benetos, E. Quinton, G. Fazekas, D. Bogdanov, Muchomusic: Evaluating music un- derstanding in multimodal audio-language models, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= ViQaFx6bmj

  114. [122]

    Zhou, et al., Can llms ’reason’ in music? an evaluation of llms’ capability of music understanding and generation, CoRRAccessed: Feb

    Z. Zhou, et al., Can llms ’reason’ in music? an evaluation of llms’ capability of music understanding and generation, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= 39KGGrrCoj

  115. [123]

    Hachmeier, R

    S. Hachmeier, R. Jäschke, A benchmark and robustness study of in-context-learning with large language models in music entity detection, in: O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schock- aert (Eds.), Proceedings of the 31st International Conferen...

  116. [124]

    Z. Wang, et al., Muchin: a chinese colloquial descrip- tion benchmark for evaluating language models in the field of music, in: Proceedings of the Thirty-Third In- ternational Joint Conference on Artificial Intelligence, IJCAI ’24, Jeju, Korea, 2024, pp. 7771–7779. doi: 10.249...

  117. [125]

    Zhao, et al., Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space, arXiv (Feb

    Y . Zhao, et al., Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space, arXiv (Feb. 2025). arXiv:arXiv:2502.12532, doi: 10.48550/arXiv.2502.12532. 37

  118. [126]

    Zhang, et al., Transportationgames: Benchmarking transportation knowledge of (multimodal) large language models, CoRRAccessed: Feb

    X. Zhang, et al., Transportationgames: Benchmarking transportation knowledge of (multimodal) large language models, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= S2LbBjEs7x

  119. [127]

    T. Nie, J. Sun, W. Ma, Exploring the roles of large language models in reshaping transportation sys- tems: A survey, framework, and roadmap, arXiv (Mar. 2025). arXiv:arXiv:2503.21411, doi:10.48550/ arXiv.2503.21411

  120. [128]

    Zhang, J

    W. Zhang, J. Han, Z. Xu, H. Ni, H. Liu, H. Xiong, Urban foundation models: A survey, in: Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, Association for Computing Machinery, New York, NY , USA, 2024, pp. 6633–6643. doi:10.1145/363...

  121. [129]

    Zheng, et al., Urbanplanbench: A comprehensive assessment of urban planning abilities in large language modelsAccessed: Feb

    Y . Zheng, et al., Urbanplanbench: A comprehensive assessment of urban planning abilities in large language modelsAccessed: Feb. 25, 2025 (Oct. 2024). URL https://openreview.net/forum?id= Dl5JaX7zoN

  122. [130]

    Feng, et al., Citybench: Evaluating the capabilities of large language model as world model, CoRRAccessed: Feb

    J. Feng, et al., Citybench: Evaluating the capabilities of large language model as world model, CoRRAccessed: Feb. 25, 2025 (Jan. 2024). URL https://openreview.net/forum?id= TCDS3eYNtD

  123. [131]

    J. Ji, Y . Chen, M. Jin, W. Xu, W. Hua, Y . Zhang, Moralbench: Moral evaluation of llms, CoRRAccessed: Feb. 27, 2025 (Jan. 2024). URL https://openreview.net/forum?id= uZWqTso8oK

  124. [132]

    B. Yan, J. Zhang, Z. Chen, S. Shan, X. Chen, M3oralbench: A multimodal moral benchmark for lvlms, arXiv (Dec. 2024). arXiv:arXiv:2412.20718, doi: 10.48550/arXiv.2412.20718

  125. [133]

    G. F. G. Marraffini, A. Cotton, N. F. Hsueh, A. Frid- man, J. Wisznia, L. D. Corro, The greatest good bench- mark: Measuring llms’ alignment with utilitarian moral dilemmas, in: Y . Al-Onaizan, M. Bansal, Y .-N. Chen (Eds.), Proceedings of the 2024 Conference on Empir- ical Me...

  126. [134]

    Yao, et al., Value compass leaderboard: A platform for fundamental and validated evaluation of llms values, arXiv (Jan

    J. Yao, et al., Value compass leaderboard: A platform for fundamental and validated evaluation of llms values, arXiv (Jan. 2025). arXiv:arXiv:2501.07071, doi: 10.48550/arXiv.2501.07071

  127. [135]

    Y . Li, Y . Huang, Y . Lin, S. Wu, Y . Wan, L. Sun, I think, therefore i am: Benchmarking awareness of large language models using awarebench, arXiv (Feb. 2024). arXiv:arXiv:2401.17882, doi:10.48550/ arXiv.2401.17882

  128. [136]

    Y . Yang, Y . Xu, C. Huang, J. Jurgensen, H. Hu, Interideas: An llm and expert-enhanced dataset for philosophical intertextualityAccessed: Feb. 27, 2025 (Oct. 2024). URL https://openreview.net/forum?id= cA8iQJFioL

  129. [137]

    Trepczy´nski, Religion, theology, and philosophical skills of llm–powered chatbots, Disput

    M. Trepczy´nski, Religion, theology, and philosophical skills of llm–powered chatbots, Disput. Philos. Int. J. Phi- los. Relig. 25 (1) (2023) 19–36

  130. [138]

    Deng, et al., Deconstructing the ethics of large lan- guage models from long-standing issues to new-emerging dilemmas, CoRRAccessed: Feb

    C. Deng, et al., Deconstructing the ethics of large lan- guage models from long-standing issues to new-emerging dilemmas, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= fa1CtyfDrd

  131. [139]

    Wang, et al., Piecing it all together: Verifying multi-hop multimodal claims, in: O

    H. Wang, et al., Piecing it all together: Verifying multi-hop multimodal claims, in: O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, S. Schock- aert (Eds.), Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, UAE, 2025, ...

  132. [140]

    Jin, et al., Agentreview: Exploring peer review dy- namics with llm agents, in: Y

    Y . Jin, et al., Agentreview: Exploring peer review dy- namics with llm agents, in: Y . Al-Onaizan, M. Bansal, Y .-N. Chen (Eds.), Proceedings of the 2024 Confer- ence on Empirical Methods in Natural Language Pro- cessing, Miami, Florida, USA, 2024, pp. 1208–1226. doi:10.18653...

  133. [141]

    Song, et al., Mosabench: Multi-object sentiment analy- sis benchmark for evaluating multimodal large language models understanding of complex image, arXiv (Nov

    S. Song, et al., Mosabench: Multi-object sentiment analy- sis benchmark for evaluating multimodal large language models understanding of complex image, arXiv (Nov. 2024). arXiv:arXiv:2412.00060, doi:10.48550/ arXiv.2412.00060

  134. [142]

    M. F. Adilazuarda, et al., Towards measuring and mod- eling ’culture’ in llms: A survey, in: Y . Al-Onaizan, M. Bansal, Y .-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, 2024, pp. 15763– 15784.doi:1...

  135. [143]

    Y . Chen, Y . Xiao, Recent advancement of emotion cognition in large language models, CoRRAccessed: Feb. 24, 2025 (Jan. 2024). URL https://openreview.net/forum?id= BMOLRz7ko6

  136. [144]

    Chakrabarty, P

    T. Chakrabarty, P. Laban, D. Agarwal, S. Muresan, C.- S. Wu, Art or artifice? large language models and the false promise of creativity, in: Proceedings of the 2024 CHI Conference on Human Factors in Comput- ing Systems, CHI ’24, Association for Computing Ma- chinery, New York...

  137. [145]

    C. Y .-l. Cheng, S. A. Hale, Beyond english: Evaluating automated measurement of moral foundations in non- english discourse with a chinese case study, arXiv (Feb. 2025). arXiv:arXiv:2502.02451, doi:10.48550/ arXiv.2502.02451

  138. [146]

    Bulla, S

    L. Bulla, S. De Giorgis, M. Mongiovì, A. Gangemi, Large language models meet moral values: A comprehen- sive assessment of moral abilities, Comput. Hum. Behav. Rep. 17 (2025) 100609. doi:10.1016/j.chbr.2025. 100609

  139. [147]

    Q. Xie, W. Han, X. Zhang, Y . Lai, M. Peng, A. Lopez- Lira, J. Huang, Pixiu: A large language model, instruction data and evaluation benchmark for finance, arXiv preprint (2023).arXiv:2306.05443

  140. [148]

    Z. Chen, W. Chen, C. Smiley, S. Shah, I. Borova, D. Lang- don, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, W. Y . Wang, Finqa: A dataset of numerical reason- ing over financial data, arXiv preprint (2021). arXiv: 2109.00122

  141. [149]

    Reddy, R

    V . Reddy, R. Koncel-Kedziorski, V . D. Lai, M. Krumdick, C. Lovering, C. Tanner, Docfinqa: A long-context finan- cial reasoning dataset, arXiv preprint (2024). arXiv: 2401.06915

  142. [150]

    Z. Chen, S. Li, C. Smiley, Z. Ma, S. Shah, W. Y . Wang, Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering, arXiv preprint (2022).arXiv:2210.03849

  143. [151]

    Webersinke, M

    N. Webersinke, M. Kraus, J. A. Bingler, M. Leippold, Climatebert: A pretrained language model for climate- related text, arXiv preprint (2021). arXiv:2110.12010

  144. [152]

    Sharma, T

    S. Sharma, T. Nayak, A. Bose, A. K. Meena, K. Dasgupta, N. Ganguly, P. Goyal, Finred: A dataset for relation ex- traction in financial domain, in: Companion Proceedings of the Web Conference 2022, 2022, pp. 595–597

  145. [153]

    X. Wu, J. Liu, H. Su, Z. Lin, Y . Qi, C. Xu, J. Su, J. Zhong, F. Wang, S. Wang, F. Hua, Golden touchstone: A comprehensive bilingual benchmark for evaluating fi- nancial large language models, arXiv preprint (2024). arXiv:2411.06272

  146. [154]

    Subrahmanyam, Behavioural finance: A review and synthesis, European Financial Management 14 (1) (2008) 12–29

    A. Subrahmanyam, Behavioural finance: A review and synthesis, European Financial Management 14 (1) (2008) 12–29

  147. [155]

    Rubbaniy, A

    G. Rubbaniy, A. A. Khalid, K. Syriopoulos, E. Polyzos, Dynamic returns connectedness: Portfolio hedging impli- cations during the COVID-19 pandemic and the Russia– Ukraine war, Journal of Futures Markets 44 (10) (2024) 1613–1639

  148. [156]

    S. Li, H. Hoque, J. Liu, Investor sentiment and firm capital structure, Journal of Corporate Finance 80 (2023) 102426

  149. [157]

    Karadima, H

    M. Karadima, H. Louri, Economic policy uncertainty and non-performing loans: The moderating role of bank con- centration, Finance Research Letters 38 (2021) 101458

  150. [158]

    Hodbod, S

    A. Hodbod, S. Huber, K. Vasilev, Sectoral risk-weights and macroprudential policy, Journal of Banking and Fi- nance 112 (2020) 105336

  151. [159]

    Danielsson, K

    J. Danielsson, K. R. James, M. Valenzuela, I. Zer, Model risk of risk models, Journal of Financial Stability 23 (2016) 79–91

  152. [160]

    Oehler, M

    A. Oehler, M. Horn, Does chatgpt provide better advice than robo-advisors?, Finance Research Letters 60 (2024) 104898

  153. [161]

    Dowling, B

    M. Dowling, B. Lucey, Chatgpt for (finance) research: The bananarama conjecture, Finance Research Letters 53 (2023) 103662

  154. [162]

    Polyzos, A

    E. Polyzos, A. Fotiadis, T.-C. Huan, The asymmetric im- pact of Twitter sentiment and emotions: Impulse response analysis on European tourism firms using micro-data, Tourism Management 104 (2024) 104909

  155. [163]

    Kalamara, A

    E. Kalamara, A. Turrell, C. Redl, G. Kapetanios, S. Ka- padia, Making text count: economic forecasting using newspaper text, Journal of Applied Econometrics 37 (5) (2022) 896–919

  156. [164]

    G. P. Herrera, M. Constantino, J.-J. Su, A. Naranpanawa, Renewable energy stocks forecast using twitter investor sentiment and deep learning, Energy Economics 114 (2022) 106285

  157. [165]

    Dicks, P

    D. Dicks, P. Fulghieri, Uncertainty, investor sentiment, and innovation, The Review of Financial Studies 34 (3) (2021) 1236–1279

  158. [166]

    X.-Y . Liu, G. Wang, H. Yang, D. Zha, Fingpt: Democ- ratizing internet-scale data for financial large language models, arXiv preprint arXiv:2307.10485 (2023)

  159. [167]

    S. Wu, O. Irsoy, S. Lu, V . Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, G. Mann, Bloomberggpt: A large language model for finance, arXiv preprint (2023).arXiv:2303.17564

  160. [168]

    Y . Liu, N. Bu, Z. Li, Y . Zhang, Z. Zhao, At-fingpt: Fi- nancial risk prediction via an audio-text large language model, Finance Research Letters (2025) 106967

  161. [169]

    E. Polyzos, Inflation and the war in Ukraine: Evidence using impulse response functions on economic indicators and twitter sentiment, Research in International Business and Finance 66 (2023) 102044

  162. [170]

    S. Long, B. Lucey, Y . Xie, L. Yarovaya, I just like the stock, The role of Reddit sentiment in the GameStop share rally. Financial Review 58 (1) (2023) 19–37. 39

  163. [171]

    Lyócsa, E

    Š. Lyócsa, E. Baumöhl, T. Výrost, Yolo trading: Rid- ing with the herd during the gamestop episode, Finance Research Letters 46 (2022) 102359

  164. [172]

    Polyzos, A

    E. Polyzos, A. Samitas, I. Kampouris, Quantifying mar- ket efficiency: Information dissemination through social media, Available at SSRN 4082899 (2022)

  165. [173]

    Vasileiou, Does the short squeeze lead to market ab- normality and antileverage effect? evidence from the gamestop case, Journal of Economic Studies 49 (8) (2022) 1360–1373

    E. Vasileiou, Does the short squeeze lead to market ab- normality and antileverage effect? evidence from the gamestop case, Journal of Economic Studies 49 (8) (2022) 1360–1373

  166. [174]

    A. H. Huang, H. Wang, Y . Yang, Finbert: A large lan- guage model for extracting information from financial text, Contemporary Accounting Research 40 (2) (2023) 806–841

  167. [175]

    Loughran, B

    T. Loughran, B. McDonald, When is a liability not a liability? textual analysis, dictionaries, and 10-ks, The Journal of Finance 66 (1) (2011) 35–65

  168. [176]

    Ni¸ toi, M

    M. Ni¸ toi, M. M. Pochea, ¸ S. C. Radu, Unveiling the senti- ment behind central bank narratives: A novel deep learn- ing index, Journal of Behavioral and Experimental Fi- nance 38 (2023) 100809. doi:10.1016/j.jbef.2023. 100809

  169. [177]

    Schimanski, A

    T. Schimanski, A. Reding, N. Reding, J. A. Bingler, M. Kraus, M. Leippold, Bridging the gap in esg mea- surement: Using nlp to quantify environmental, social, and governance communication, Finance Research Let- ters 61 (2024) 104979

  170. [178]

    Polyzos, A

    E. Polyzos, A. Samitas, M.-S. Katsaiti, Who is unhappy for Brexit? a machine-learning, agent-based study on financial instability, International Review of Financial Analysis 72 (2020) 101590

  171. [179]

    Polyzos, K

    E. Polyzos, K. Abdulrahman, A. Christopoulos, Good management or good finances? An agent-based study on the causes of bank failure, Banks & Bank Systems 13, Iss. 3 (2018) 95–105

  172. [180]

    Stevens Institute of Technology, Applying large language models to financial decision- making, https://www.stevens.edu/news/ applying-large-language-models-to-financial-decision-making , [Accessed 3 March 2025] (2023)

  173. [181]

    J. Ye, G. Wang, Y . Li, Z. Deng, W. Li, T. Li, H. Duan, Z. Huang, Y . Su, B. Wang, et al., Gmai-mmbench: A com- prehensive multimodal evaluation benchmark towards general medical ai, Advances in Neural Information Pro- cessing Systems 37 (2024) 94327–94427

  174. [182]

    J. Liu, W. Wang, Y . Su, J. Huan, W. Chen, Y . Zhang, C.-Y . Li, K.-J. Chang, X. Xin, L. Shen, et al., A spectrum evalu- ation benchmark for medical multi-modal large language models, arXiv preprint arXiv:2402.11217 (2024)

  175. [183]

    K. Keat, R. Venkatesh, Y . Huang, R. Kumar, S. Tuteja, K. Sangkuhl, B. Li, L. Gong, M. Whirl-Carrillo, T. E. Klein, et al., Pgxqa: A resource for evaluating llm perfor- mance for pharmacogenomic qa tasks, in: Biocomputing 2025: Proceedings of the Pacific Symposium, World Sci- ...

  176. [184]

    H. Liu, H. Wang, Genotex: A benchmark for eval- uating llm-based exploration of gene expression data in alignment with bioinformaticians, arXiv preprint arXiv:2406.15341 (2024)

  177. [185]

    F. Bai, Y . Du, T. Huang, M. Q.-H. Meng, B. Zhao, M3d: Advancing 3d medical image analysis with multi-modal large language models, arXiv preprint arXiv:2404.00578 (2024)

  178. [186]

    P. R. A. S. Bassi, M. C. Yavuz, K. Wang, X. Chen, W. Li, S. Decherchi, A. Cavalli, Y . Yang, A. Yuille, Z. Zhou, Radgpt: Constructing 3d image-text tumor datasets (2025).arXiv:2501.04678. URLhttps://arxiv.org/abs/2501.04678

  179. [187]

    S. Mo, P. P. Liang, MultiMed: Massively Multimodal and Multitask Medical Understanding, arXiv preprint arXiv:2408.12682 (2024)

  180. [188]

    M. S. Sepehri, Z. Fabian, M. Soltanolkotabi, M. Soltanolkotabi, MediConfusion: Can you trust your AI radiologist? Probing the reliability of mul- timodal medical foundation models, arXiv preprint arXiv:2409.15477 (2024)

  181. [189]

    Neehal, B

    N. Neehal, B. Wang, S. Debopadhaya, S. Dan, K. Muruge- san, V . Anand, K. P. Bennett, Ctbench: A comprehensive benchmark for evaluating language model capabilities in clinical trial design, arXiv preprint arXiv:2406.17888 (2024)

  182. [190]

    Khandekar, Q

    N. Khandekar, Q. Jin, G. Xiong, S. Dunn, S. Apple- baum, Z. Anwar, M. Sarfo-Gyamfi, C. Safranek, A. An- war, A. Zhang, et al., Medcalc-bench: Evaluating large language models for medical calculations, Advances in Neural Information Processing Systems 37 (2024) 84730– 84745

  183. [191]

    Jiang, P

    J. Jiang, P. Chen, J. Wang, D. He, Z. Wei, L. Hong, L. Zong, S. Wang, Q. Yu, Z. Ma, et al., Benchmarking large language models on multiple tasks in bioinformat- ics nlp with prompting, arXiv preprint arXiv:2503.04013 (2025)

  184. [192]

    L. Chen, X. Han, S. Lin, H. Mai, H. Ran, Trimedlm: Advancing three-dimensional medical image analysis with multi-modal llm, in: 2024 IEEE International Con- ference on Bioinformatics and Biomedicine (BIBM), 2024, pp. 4505–4512. doi:10.1109/BIBM62325.2024. 10822809. 40

  185. [193]

    Lozano, J

    A. Lozano, J. Nirschl, J. Burgess, S. R. Gupte, Y . Zhang, A. Unell, S. Yeung, Micro-bench: A microscopy bench- mark for vision-language understanding, Advances in Neural Information Processing Systems 37 (2024) 30670– 30685

  186. [194]

    Y . Sun, H. Wu, C. Zhu, S. Zheng, Q. Chen, K. Zhang, Y . Zhang, D. Wan, X. Lan, M. Zheng, et al., Pathmmu: A massive multimodal expert-level benchmark for un- derstanding and reasoning in pathology, in: European Conference on Computer Vision, Springer, 2024, pp. 56– 73

  187. [195]

    Burgess, J

    J. Burgess, J. J. Nirschl, L. Bravo-Sánchez, A. Lozano, S. R. Gupte, J. G. Galaz-Montoya, Y . Zhang, Y . Su, D. Bhowmik, Z. Coman, S. M. Hasan, A. Johannesson, W. D. Leineweber, M. G. Nair, R. Yarlagadda, C. Zuraski, W. Chiu, S. Cohen, J. N. Hansen, M. D. Leonetti, C. Liu, E. ...

  188. [196]

    Y . Chen, G. Wang, Y . Ji, Y . Li, J. Ye, T. Li, M. Hu, R. Yu, Y . Qiao, J. He, Slidechat: A large vision-language assistant for whole-slide pathology image understanding (2025).arXiv:2410.11761. URLhttps://arxiv.org/abs/2410.11761

  189. [197]

    Lozano, J

    A. Lozano, J. Nirschl, J. Burgess, S. R. Gupte, Y . Zhang, A. Unell, S. Yeung-Levy, µ-bench: A vision-language benchmark for microscopy understanding (2024). arXiv: 2407.01791. URLhttps://arxiv.org/abs/2407.01791

  190. [198]

    X. Wang, D. Song, S. Chen, C. Zhang, B. Wang, LongLLaV A: Scaling Multi-modal LLMs to 1000 Im- ages Efficiently via a Hybrid Architecture, arXiv preprint arXiv:2409.02889 (2024)

  191. [199]

    B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y . Li, Z. Liu, et al., Llava- onevision: Easy visual task transfer, arXiv preprint arXiv:2408.03326 (2024)

  192. [200]

    Huang, L

    S. Huang, L. Dong, W. Wang, Y . Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, Q. Liu, K. Ag- garwal, Z. Chi, J. Bjorck, V . Chaudhary, S. Som, X. Song, F. Wei, Language is not all you need: Aligning perception with language models (2023).arXiv:2302.14045. UR...

  193. [201]

    Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, F. Wei, Kosmos-2: Grounding multimodal large language models to the world (2023).arXiv:2306.14824. URLhttps://arxiv.org/abs/2306.14824

  194. [202]

    T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, et al., Chatglm: A family of large language models from glm-130b to glm-4 all tools, arXiv preprint arXiv:2406.12793 (2024)

  195. [203]

    Y . Wang, Y . Liu, F. Yu, C. Huang, K. Li, Z. Wan, W. Che, H. Chen, Cvlue: A new benchmark dataset for chinese vision-language understanding evaluation, in: Proceed- ings of the AAAI Conference on Artificial Intelligence, V ol. 39, 2025, pp. 8196–8204

  196. [204]

    S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y . Cao, K. Narasimhan, Tree of thoughts: Deliberate problem solving with large language models, Advances in Neural Information Processing Systems 36 (2024)

  197. [205]

    M. A. Arshad, T. Z. Jubery, T. Roy, R. Nassiri, A. K. Singh, A. Singh, C. Hegde, B. Ganapathysubramanian, A. Balu, A. Krishnamurthy, et al., Leveraging vision lan- guage models for specialized agricultural tasks, in: 2025 IEEE/CVF Winter Conference on Applications of Com- pute...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.