Pith. sign in

REVIEW 5 cited by

Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15606 v2 pith:4LPIUM4K submitted 2024-12-20 cs.AI cs.CV

classification cs.AIcs.CV
keywords datamulti-modalagenttool-usageunderlineusagevlmscontroller
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The advancement of large language models (LLMs) prompts the development of multi-modal agents, which are used as a controller to call external tools, providing a feasible way to solve practical tasks. In this paper, we propose a multi-modal agent tuning method that automatically generates multi-modal tool-usage data and tunes a vision-language model (VLM) as the controller for powerful tool-usage reasoning. To preserve the data quality, we prompt the GPT-4o mini model to generate queries, files, and trajectories, followed by query-file and trajectory verifiers. Based on the data synthesis pipeline, we collect the MM-Traj dataset that contains 20K tasks with trajectories of tool usage. Then, we develop the T3-Agent via \underline{T}rajectory \underline{T}uning on VLMs for \underline{T}ool usage using MM-Traj. Evaluations on the GTA and GAIA benchmarks show that the T3-Agent consistently achieves improvements on two popular VLMs: MiniCPM-V-8.5B and {Qwen2-VL-7B}, which outperforms untrained VLMs by $20\%$, showing the effectiveness of the proposed data synthesis pipeline, leading to high-quality data for tool-usage capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A three-stage teacher-hypothesize, student-verify, then-train pipeline lets a post-trained tool-using agent adopt new visual tools without retraining from scratch.

  2. Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A modular 8B agent with episodic visual memory and RL-trained retrieval reaches 91.4% cross-turn image recall over 20 turns, outperforming 32B all-context baselines with ~1.8× lower latency.

  3. A Multimodal Foundation Model of Spatial Transcriptomics and Histology for Biological Discovery and Clinical Prediction

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    A hierarchical multimodal foundation model (STORM) maps H&E morphology to spatial gene expression and improves immunotherapy and prognosis prediction across 7,245 patients.

  4. Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

    cs.AI 2025-10 unverdicted novelty 6.0 of 10

    TRACE uses an evidence bank to score tool-augmented LLM agents on efficiency, hallucination, and adaptivity without ground-truth trajectories.

  5. Allen: Rethinking MAS Design through Step-Level Policy Autonomy

    cs.MA 2025-08 reject novelty 4.0 of 10

    Allen is a step-level policy-autonomy architecture for multi-agent systems that lets agents piece together their own execution plans, but it lacks any empirical validation.

Pith tools