REVIEW 3 major objections 3 minor 2 cited by
External tools can overcome the three persistent weaknesses of multimodal LLMs, this survey argues.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection A useful but possibly over-promotional survey of tool-augmented MLLMs; the four-part taxonomy is sound, but the 'comprehensive' claim needs a check on coverage and negative results. the 3 major comments →
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that augmenting MLLMs with external tools is a promising and increasingly necessary strategy to overcome fundamental limitations that remain despite scaling. It asserts that the problem is not merely model size but the quality of multimodal data, the difficulty of complex downstream tasks, and the lack of comprehensive evaluation protocols. By surveying the literature along four dimensions, the paper argues that tools can generate and annotate high-quality multimodal data, decompose and solve challenging tasks via expert APIs, and enable richer evaluation of MLLM outputs. The paper positions tool-augmentation not as an optional add-on but as a transformative path
What carries the argument
The organizing device is a four-dimensional taxonomy: (1) tools for data acquisition and annotation, (2) tools for task performance, (3) tools for evaluation, (4) current limitations and future directions. This taxonomy carries the survey's argument by converting the scattered tool-augmentation literature into a systematic case that external tools address each of the MLLM failure modes in turn.
Load-bearing premise
The claim rests on the assumption that the surveyed tool-augmentation methods are representative of the field and that their reported benefits are genuine — not artifacts of benchmark selection or cherry-picked tasks.
What would settle it
A meta-analysis of published tool-augmented MLLM results showing no average improvement over non-tool baselines on held-out multimodal benchmarks (e.g., MMBench, MM-Vet, or similar), or showing that tool chains produce cascading errors that outweigh their gains.
If this is right
- Tool-augmentation will likely become a standard component of MLLM architectures rather than an afterthought.
- Tool-assisted data pipelines could lower the cost of high-quality multimodal datasets, making MLLM training more accessible.
- Evaluation benchmarks will need to incorporate tool use and API interaction to assess real-world MLLM competence.
- Future MLLMs will be judged by their ability to select, call, and integrate external tools, not just by parametric knowledge.
- Tool-mediated evaluation may expose failure modes that static benchmarks currently hide.
Where Pith is reading between the lines
- The survey's four-dimensional frame implies a fifth open question it does not develop: how to make tool selection itself reliable and safe, since a model that delegates poorly inherits the tool's errors.
- Extending the logic, tool-augmented MLLMs could become testbeds for broader questions about reasoning, planning, and delegation in AI systems.
- One testable extension: measure whether tool-augmented models outperform non-tool baselines on standard benchmarks like MMBench or MM-Vet when tools are available.
- The framing suggests a convergence: as tools become more powerful, the bottleneck shifts from parametric knowledge to orchestration and error handling.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a survey of methods that augment Multimodal Large Language Models (MLLMs) with external tools such as APIs, expert models, and knowledge bases. It organizes the literature along four dimensions: (1) using tools to acquire and annotate high-quality multimodal data, (2) using tools to improve performance on complex downstream tasks, (3) using tools to enable comprehensive evaluation, and (4) limitations and future directions. The central claim, stated in the abstract, is that external tools offer a promising strategy to overcome the limitations of MLLMs in data quality, task performance, and evaluation. The paper also points to a publicly available GitHub repository as a companion project page.
Significance. If the survey is genuinely comprehensive and balanced, it would be a valuable resource for a fast-moving area, providing a shared taxonomy and a structured view of a fragmented literature. The public GitHub repository is a useful community artifact and will likely be used by many researchers. The four-dimensional organization is a sensible and pedagogically useful framing. However, the strength of the contribution depends entirely on the representativeness and critical appraisal of the selected literature, which cannot be assessed from the abstract alone.
major comments (3)
- [Abstract] The abstract's central claim, that external tools are a 'promising strategy' to overcome MLLM limitations, is presented as a synthesis of the surveyed literature. The support for this claim depends on the selection of papers being representative and including balanced evidence. The abstract gives no indication of a systematic search strategy, inclusion/exclusion criteria, or a timeline of coverage. Without such methodological transparency, the 'comprehensive' descriptor is unsupported. This is load-bearing because the paper's conclusion is only as strong as the completeness of the survey. Please add a section describing the literature search protocol, and preferably a table of excluded/non-representative topics.
- [Abstract / Project page (GitHub)] The project page is a curated GitHub repository maintained by the authors. If the survey's coverage is drawn from this repository, the selection is potentially self-referential: the authors curate the list and then use it to substantiate a 'comprehensive' survey. This is not an accusation of misconduct, but a methodological concern. The manuscript should state how the repository was constructed independently (e.g., database queries, multiple reviewers, pre-registered criteria) and how the survey's scope differs from the repository's. Otherwise, the reader cannot rule out survivorship bias in the cited works.
- [Abstract, dimension (4)] The paper's stated goal to 'underscore the transformative potential' of external tools creates a risk of confirmation bias. A survey that is genuinely useful must engage with negative results, emerging failures, and the costs of tool augmentation (e.g., error propagation, dependency on external API reliability, added latency, and computational overhead). The abstract mentions 'current limitations and future directions' but does not signal that counterevidence to the 'promising strategy' thesis will be examined. Please make it explicit that the survey will include a critical assessment of cases where tools do not help or introduce new issues, not just a collection of successes.
minor comments (3)
- [Abstract] The term 'external tools' is defined by examples (APIs, expert models, knowledge bases), but the boundary is unclear. For instance, retrieval-augmented generation is a form of tool use, yet it is not listed. A precise definition or a short taxonomy would improve the abstract.
- [Abstract] GPT-4V is given as an example of a successful MLLM. It may be outdated by the time of publication; consider citing a few representative open-source MLLMs as well (e.g., LLaVA) to avoid the appearance of a closed-set perspective.
- [Project page] The GitHub URL appears as a raw link. Please provide it as a proper reference or footnote, and consider adding a version/DOI to enable citation.
Circularity Check
No circularity; survey synthesizes literature without deriving claims from its own premises.
full rationale
This paper is a survey. Its central claim — that external tools offer a promising strategy to overcome MLLM data quality, task performance, and evaluation challenges — is a synthesis of the surveyed literature rather than a derivational result. No equation, fitted parameter, or predictive model is introduced, so there is no step in which an output is equivalent to an input by construction. The organizing dimensions in the abstract (data acquisition, task performance, evaluation, future directions) map onto the stated challenges, but this is a framing choice, not a circular derivation. The only self-referential element is the project page link (github.com/Lackel/Awesome-Tools-for-MLLMs), which the authors maintain; while this creates a potential selection-framing concern, it does not constitute circular reasoning because the survey's conclusions are not justified by that repository alone. The skeptic's concern about selection bias and omission of negative results is a validity/coverage concern, not a circularity concern under the specified criteria. No quoteable circular step exists in the available text. Therefore score 0.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption Multimodal large language models have meaningful limitations in data quality, task performance, and evaluation.
- ad hoc to paper External tools can be meaningfully categorized into the four dimensions used in the paper.
Cite this review
Pith. "Pith review of Empowering Multimodal LLMs with External Tools: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/UQNYXUYJ
@misc{pith2026250810955,
author = {Pith},
title = {Pith review of: Empowering Multimodal LLMs with External Tools: A Comprehensive Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/UQNYXUYJ}},
note = {Machine review of arXiv:2508.10955}
}
read the original abstract
By integrating the perception capabilities of multimodal encoders with the generative power of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs), exemplified by GPT-4V, have achieved great success in various multimodal tasks, pointing toward a promising pathway to artificial general intelligence. Despite this progress, the limited quality of multimodal data, poor performance on many complex downstream tasks, and inadequate evaluation protocols continue to hinder the reliability and broader applicability of MLLMs across diverse domains. Inspired by the human ability to leverage external tools for enhanced reasoning and problem-solving, augmenting MLLMs with external tools (e.g., APIs, expert models, and knowledge bases) offers a promising strategy to overcome these challenges. In this paper, we present a comprehensive survey on leveraging external tools to enhance MLLM performance. Our discussion is structured along four key dimensions about external tools: (1) how they can facilitate the acquisition and annotation of high-quality multimodal data; (2) how they can assist in improving MLLM performance on challenging downstream tasks; (3) how they enable comprehensive and accurate evaluation of MLLMs; (4) the current limitations and future directions of tool-augmented MLLMs. Through this survey, we aim to underscore the transformative potential of external tools in advancing MLLM capabilities, offering a forward-looking perspective on their development and applications. The project page of this paper is publicly available athttps://github.com/Lackel/Awesome-Tools-for-MLLMs.
Forward citations
Cited by 2 Pith papers
-
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning
MathVis-Fine proposes a dataset with fine-grained visual annotations and dependency ratings plus a progressive two-stage training paradigm to align visual supervision with sample-specific necessity in multimodal mathe...
-
Concurrency without Model Changes: Future-based Asynchronous Function Calling for LLMs
AsyncFC decouples LLM decoding from function execution via symbolic futures, enabling overlap and parallelism to reduce end-to-end latency on function-calling benchmarks while preserving accuracy.
Reference graph
Works this paper leans on
-
[1]
11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEauthordefaulttextstyle \@IEEEcompsocnotconfonly \@IEEEauthor...
- [2]
-
[3]
H. Touvron et al. Llama: Open and efficient foundation language models. arXiv:2302.13971 , 2023
Pith/arXiv arXiv 2023
-
[4]
Chiang et al
W.-L. Chiang et al. Vicuna: An open-source chatbot impressing gpt-4 with 90\
- [5]
- [6]
-
[7]
An et al
W. An et al. Generalized category discovery with large language models in the loop. In ACL Findings , 2024
2024
-
[8]
Kasneci et al
E. Kasneci et al. Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual differences , 2023
2023
-
[9]
W. An et al. Knowledge acquisition disentanglement for knowledge-based visual question answering with large language models. arXiv:2407.15346 , 2024
Pith/arXiv arXiv 2024
-
[10]
W. Dai et al. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv:2306.04387 , 2023
Pith/arXiv arXiv 2023
- [11]
-
[12]
J. Zhu et al. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv:2504.10479 , 2025
Pith/arXiv arXiv 2025
- [13]
-
[14]
Z. Wu et al. Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv:2412.10302 , 2024
Pith/arXiv arXiv 2024
-
[15]
Yin et al
S. Yin et al. A survey on multimodal large language models. National Science Review , November 2024
2024
-
[16]
Lin et al
H. Lin et al. Schedule your edit: A simple yet effective diffusion noise schedule for image editing. NeurIPS , 2024
2024
-
[17]
Z. Bai et al. Hallucination of multimodal large language models: A survey. arXiv:2404.18930 , 2024
Pith/arXiv arXiv 2024
-
[18]
Zhang et al
X. Zhang et al. From redundancy to relevance: Enhancing explainability in multimodal large language models. arXiv e-prints , pp. arXiv--2406, 2024
2024
-
[19]
S. Sun et al. A review of multimodal explainable artificial intelligence: Past, present and future. arXiv preprint arXiv:2412.14056 , 2024
Pith/arXiv arXiv 2024
-
[20]
Y. Wang et al. Multimodal chain-of-thought reasoning: A comprehensive survey. arXiv:2503.12605 , 2025
Pith/arXiv arXiv 2025
-
[21]
M. Hu et al. Advancing medical imaging with language models: A journey from n-grams to chatgpt. arXiv:2304.04920 , 2023
Pith/arXiv arXiv 2023
-
[22]
L. Chen et al. Driving with llms: Fusing object-level vector modality for explainable autonomous driving. arXiv:2310.01957 , 2023
Pith/arXiv arXiv 2023
- [23]
-
[24]
Radford et al
A. Radford et al. Learning transferable visual models from natural language supervision. In ICML , 2021
2021
-
[25]
Liu et al
Z. Liu et al. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV , 2021
2021
-
[26]
Elizalde et al
B. Elizalde et al. Clap learning audio concepts from natural language supervision. In ICASSP , 2023
2023
-
[27]
H. Lauren c on et al. What matters when building vision-language models? arXiv:2405.02246 , 2024
Pith/arXiv arXiv 2024
- [28]
-
[29]
Qin et al
Y. Qin et al. Tool learning with foundation models. ACM Computing Surveys , 2024
2024
-
[30]
G. Team et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv:2403.05530 , 2024
Pith/arXiv arXiv 2024
-
[31]
W. Huang et al. Visual hallucinations of multi-modal large language models. arXiv:2402.14683 , 2024
Pith/arXiv arXiv 2024
-
[32]
T. Han et al. The instinctive bias: Spurious images lead to illusion in mllms. arXiv:2402.03757 , 2024
Pith/arXiv arXiv 2024
-
[33]
Tong et al
S. Tong et al. Eyes wide shut? exploring the visual shortcomings of multimodal llms. In CVPR , 2024
2024
-
[34]
Sharma et al
P. Sharma et al. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL , 2018
2018
-
[35]
Chambers et al
C. Chambers et al. Flumejava: easy, efficient data-parallel pipelines. ACM Sigplan Notices , 2010
2010
-
[36]
X. Chen et al. Pali: A jointly-scaled multilingual language-image model. arXiv:2209.06794 , 2022
Pith/arXiv arXiv 2022
-
[37]
Schuhmann et al
C. Schuhmann et al. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS , 2022
2022
-
[38]
S. Y. Gadre et al. Datacomp: In search of the next generation of multimodal datasets. NeurIPS , 2023
2023
-
[39]
Jia et al
C. Jia et al. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML , 2021
2021
-
[40]
Zhu et al
W. Zhu et al. Multimodal c4: An open, billion-scale corpus of images interleaved with text. NeurIPS , 2023
2023
-
[41]
Gu et al
J. Gu et al. Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark. NeurIPS , 2022
2022
-
[42]
Srinivasan et al
K. Srinivasan et al. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In ACM SIGIR conference on research and development in information retrieval , 2021
2021
-
[43]
Changpinyo et al
S. Changpinyo et al. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR , 2021
2021
-
[44]
K. Desai et al. Redcaps: Web-curated image-text data created by the people, for the people. arXiv:2111.11431 , 2021
Pith/arXiv arXiv 2021
-
[45]
L. Li et al. Multimodal arxiv: A dataset for improving scientific comprehension of large vision-language models. arXiv:2403.00231 , 2024
Pith/arXiv arXiv 2024
-
[46]
Goyal et al
Y. Goyal et al. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In CVPR , 2017
2017
-
[47]
J. Wu et al. Ai challenger: A large-scale dataset for going deeper in image understanding. arXiv:1711.06475 , 2017
Pith/arXiv arXiv 2017
-
[48]
Ordonez et al
V. Ordonez et al. Im2text: Describing images using 1 million captioned photographs. NeurIPS , 2011
2011
-
[49]
H. Qiu et al. Longhalqa: Long-context hallucination evaluation for multimodal large language models. arXiv:2410.09962 , 2024
Pith/arXiv arXiv 2024
-
[50]
X. Wu et al. Autohallusion: Automatic generation of hallucination benchmarks for vision-language models. arXiv:2406.10900 , 2024
Pith/arXiv arXiv 2024
-
[51]
Kaul et al
P. Kaul et al. Throne: An object-based hallucination benchmark for the free-form generations of large vision-language models. In CVPR , 2024
2024
-
[52]
J. Liu et al. Phd: A chatgpt-prompted visual hallucination evaluation dataset. arXiv:2403.11116 , 2024
Pith/arXiv arXiv 2024
-
[53]
Gao et al
Y. Gao et al. Aigcs confuse ai too: Investigating and explaining synthetic image-induced hallucinations in large vision-language models. In ACMMM , 2024
2024
-
[54]
Y. Qian et al. How easy is it to fool your multimodal llms? an empirical analysis on deceptive prompts. arXiv:2402.13220 , 2024
Pith/arXiv arXiv 2024
-
[55]
A. Ben-Kish et al. Mitigating open-vocabulary caption hallucinations. arXiv:2312.03631 , 2023
Pith/arXiv arXiv 2023
-
[56]
Wang et al
L. Wang et al. Mitigating fine-grained hallucination by fine-tuning large vision-language models with caption rewrites. In International Conference on Multimedia Modeling , 2024
2024
-
[57]
H. Lovenia et al. Negative object presence evaluation (nope) to measure object hallucination in vision-language models. arXiv:2310.05338 , 2023
Pith/arXiv arXiv 2023
-
[58]
F. Liu et al. Mitigating hallucination in large multi-modal models via robust instruction tuning. arXiv:2306.14565 , 2023
Pith/arXiv arXiv 2023
-
[59]
H. Hu et al. Ciem: Contrastive instruction evaluation method for better instruction tuning. arXiv:2309.02301 , 2023
Pith/arXiv arXiv 2023
-
[60]
W. Zhong et al. Investigating and mitigating the multimodal hallucination snowballing in large vision-language models. arXiv:2407.00569 , 2024
Pith/arXiv arXiv 2024
-
[61]
L. Li et al. Silkie: Preference distillation for large visual language models. arXiv:2312.10665 , 2023
Pith/arXiv arXiv 2023
-
[62]
Kafle et al
K. Kafle et al. Dvqa: Understanding data visualizations via question answering. In CVPR , 2018
2018
-
[63]
Nie et al
J. Nie et al. Mmrel: A relation understanding dataset and benchmark in the mllm era. arXiv e-prints , pp. arXiv--2406, 2024
2024
-
[64]
Y. Liu et al. Investigating and mitigating object hallucinations in pretrained vision-language (clip) models. arXiv:2410.03176 , 2024
Pith/arXiv arXiv 2024
-
[65]
Ziyang et al
M. Ziyang et al. Vga: Vision gui assistant-minimizing hallucinations through image-centric fine-tuning. In EMNLP Findings , 2024
2024
-
[66]
S. Petryk et al. Aloha: A new measure for hallucination in captioning models. arXiv:2404.02904 , 2024
Pith/arXiv arXiv 2024
-
[67]
Jiang et al
C. Jiang et al. Hal-eval: A universal and fine-grained hallucination evaluation framework for large vision language models. In ACMMM , 2024
2024
-
[68]
X. Chen et al. Unified hallucination detection for multimodal large language models. arXiv:2402.03190 , 2024
Pith/arXiv arXiv 2024
-
[69]
Villa et al
A. Villa et al. Behind the magic, merlim: Multi-modal evaluation benchmark for large image-language models. In CVPR , 2025
2025
-
[70]
Z. Chen et al. Mitigating hallucination in visual language models with visual supervision. arXiv:2311.16479 , 2023
Pith/arXiv arXiv 2023
-
[71]
B. Zhai et al. Halle-control: controlling object hallucination in large multimodal models. arXiv:2310.01779 , 2023
Pith/arXiv arXiv 2023
-
[72]
J. Wang et al. Evaluation and analysis of hallucination in large vision-language models. arXiv:2308.15126 , 2023
Pith/arXiv arXiv 2023
-
[73]
Y. Li et al. Evaluating object hallucination in large vision-language models. arXiv:2305.10355 , 2023
Pith/arXiv arXiv 2023
-
[74]
A. Yan et al. List items one by one: A new data source and learning paradigm for multimodal llms. arXiv:2404.16375 , 2024
Pith/arXiv arXiv 2024
-
[75]
Gunjal et al
A. Gunjal et al. Detecting and preventing hallucinations in large vision language models. In AAAI , 2024
2024
-
[76]
A. Seth et al. Towards a systematic evaluation of hallucinations in large-vision language models. arXiv:2412.20622 , 2024
Pith/arXiv arXiv 2024
-
[77]
Y. Wang et al. Videohallucer: Evaluating intrinsic and extrinsic hallucinations in large video-language models. arXiv:2406.16338 , 2024
Pith/arXiv arXiv 2024
-
[78]
H. S. Shahgir et al. Illusionvqa: A challenging optical illusion dataset for vision language models. arXiv:2403.15952 , 2024
Pith/arXiv arXiv 2024
-
[79]
S. Ghosh et al. Visual description grounding reduces hallucinations and boosts reasoning in lvlms. arXiv:2405.15683 , 2024
Pith/arXiv arXiv 2024
-
[80]
Li et al
C. Li et al. Vidhalluc: Evaluating temporal hallucinations in multimodal large language models for video understanding. In CVPR , 2025
2025
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.