REVIEW 8 cited by
Tool Documentation Enables Zero-Shot Tool-Usage with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Today, large language models (LLMs) are taught to use new tools by providing a few demonstrations of the tool's usage. Unfortunately, demonstrations are hard to acquire, and can result in undesirable biased usage if the wrong demonstration is chosen. Even in the rare scenario that demonstrations are readily available, there is no principled selection protocol to determine how many and which ones to provide. As tasks grow more complex, the selection search grows combinatorially and invariably becomes intractable. Our work provides an alternative to demonstrations: tool documentation. We advocate the use of tool documentation, descriptions for the individual tool usage, over demonstrations. We substantiate our claim through three main empirical findings on 6 tasks across both vision and language modalities. First, on existing benchmarks, zero-shot prompts with only tool documentation are sufficient for eliciting proper tool usage, achieving performance on par with few-shot prompts. Second, on a newly collected realistic tool-use dataset with hundreds of available tool APIs, we show that tool documentation is significantly more valuable than demonstrations, with zero-shot documentation significantly outperforming few-shot without documentation. Third, we highlight the benefits of tool documentations by tackling image generation and video tracking using just-released unseen state-of-the-art models as tools. Finally, we highlight the possibility of using tool documentation to automatically enable new applications: by using nothing more than the documentation of GroundingDino, Stable Diffusion, XMem, and SAM, LLMs can re-invent the functionalities of the just-released Grounded-SAM and Track Anything models.
Forward citations
Cited by 8 Pith papers
-
ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory
Provider-side tool memory graphs, expanded by execution-verified frontier probing and queried by adaptive traversal, improve and transfer tool use across agents and environments.
-
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
Most MCP tool descriptions (97.1%) contain quality smells, and augmenting them improves agent success by a median of 5.85 percentage points at a 67.46% increase in execution steps.
-
Extended Version: It Should Be Easy but... New Users Experiences and Challenges with Secret Management Tools
New users of secret management tools can store and read secrets fairly easily, but injecting them into an app is much harder because official documentation omits relevant examples and flag explanations, driving users ...
-
ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution
ASPERA generates a benchmark of 250 executable assistant tasks and finds that LLMs, even with full API documentation, solve only 10 to 80 percent of them.
-
CODEMENV: Benchmarking Large Language Models on Code Migration
CODEMENV provides 922 examples and three tasks for evaluating LLMs on cross-version code migration, finding models are much better at migrating old code to new versions (up to 43.84% pass@1) than the reverse.
-
S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering
Structured scene-event memory with anchor-sensitive retrieval beats vanilla RAG and several recent memory baselines on long-horizon interactive QA while using far fewer evidence tokens.
-
ReQuestNet: A Foundational Learning model for Channel Estimation
A single neural network model is claimed to estimate 5G channels across varying resource blocks, MIMO layers, and precoding patterns, beating a statistics-aware MMSE baseline by up to 10 dB.
-
Augmented Vision-Language Models: A Systematic Review
A structured taxonomy of inference-time augmentation techniques that connect vision-language models to external symbolic systems, tools, and knowledge sources.
Discussion (0). Sign in to comment.