Pith. sign in

Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that this interpretation can be premature: a tool-call trace alone does not show whether the tool supplied answer-critical information. We study two representative ``thinking with images'' agents, Thyme and DeepEyesV2, across real-world understanding, OCR, chart understanding, and mathematical reasoning. Each agent is compared with its Tool-Free counterpart and with a Pure-Text Reasoner trained from the same source pool without tool-calling trajectories. Tool access yields little consistent aggregate improvement, does not reliably reduce generated-token cost, and leaves only a small tool-only solved set: 93% of DeepEyesV2's tool-solved problems and 96% of Thyme's are also solved by at least one non-tool setting. Mechanism ablations further show that the full tool-use loop does not consistently outperform either the tool-call format or the returned execution result alone. In the settings we study, the analyzed agents appear to learn tool-calling patterns more reliably than tool-contributed capabilities, suggesting that evaluation should distinguish tool availability from whether tools actually expand what agents can solve.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Self-Evolving Code-with-Image Reasoning

cs.CV · 2026-08-11 · conditional · novelty 6.0

A frozen vision-language model more than doubles its accuracy on precise visual-computation questions by writing Python code and refining plain-text skills from its own failed executions.

citing papers explorer

Showing 1 of 1 citing paper.

  • Self-Evolving Code-with-Image Reasoning cs.CV · 2026-08-11 · conditional · none · ref 8 · internal anchor

    A frozen vision-language model more than doubles its accuracy on precise visual-computation questions by writing Python code and refining plain-text skills from its own failed executions.