REVIEW 8 cited by
What Are Tools Anyway? A Survey from the Language Model Perspective
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Language models (LMs) are powerful yet mostly for text generation tasks. Tools have substantially enhanced their performance for tasks that require complex skills. However, many works adopt the term "tool" in different ways, raising the question: What is a tool anyway? Subsequently, where and how do tools help LMs? In this survey, we provide a unified definition of tools as external programs used by LMs, and perform a systematic review of LM tooling scenarios and approaches. Grounded on this review, we empirically study the efficiency of various tooling methods by measuring their required compute and performance gains on various benchmarks, and highlight some challenges and potential future research in the field.
Forward citations
Cited by 8 Pith papers
-
PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents
A multi-agent repair framework that samples multiple edit locations and iteratively reflects on patch attempts reaches 76.0% Pass@1 on SWE-bench-Verified, up to a 7.8% relative gain over SWE-agent.
-
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
A hybrid rule-plus-VLM detector for mobile GUI agents, tested on a new 204-trajectory Android benchmark, reports 10-30% gains over baselines.
-
Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning
A 7B open-source model fine-tuned on GPT-4-verified tool-use trajectories plus reflection data reports state-of-the-art pass rates on StableToolBench and high error-correction rates on a new benchmark.
-
RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving
RepoMaster, a repository-aware code agent, lifts the task pass rate from 40.7% to 62.9% and cuts token use by about 95% versus OpenHands on the new GitTaskBench benchmark.
-
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Rule-based RL with a binary format-and-tool-match reward produces tool-calling LLMs that outperform GPT-4o on BFCL, API-Bank, and ACEBench, and can rival or beat SFT-then-RL under equal data budgets.
-
A Compute-Matched Re-Evaluation of TroVE on MATH
After matching computational budget, TroVE's toolbox mechanism yields only a marginal, statistically non-significant 1% accuracy gain over a plain sampling baseline on MATH.
-
OMuleT: Orchestrating Multiple Tools for Practicable Conversational Recommendation
A fixed-policy multi-tool harness with over ten generic retrieval and lookup tools improves the relevance, novelty, and diversity of LLM recommendations for real Roblox user requests compared to LLM prompting alone.
-
CyberMentor: AI Powered Learning Tool Platform to Address Diverse Student Needs in Cybersecurity Education
The paper introduces CyberMentor, an open-source LLM-based tutoring platform for cybersecurity education, and reports automated evaluation scores that are not independently validated.
Discussion (0). Continue with ORCID to comment.