Pith. sign in

REVIEW 1 cited by

Look Before You Leap: Towards Decision-Aware and Generalizable Tool-Usage for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.16696 v3 pith:INBY3EHI submitted 2024-02-26 cs.CL

classification cs.CL
keywords llmstool-usagetoolscapabilitiesdatasetsdecision-awaredeerdiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tool-augmented large language models (LLMs) are attracting widespread attention when accessing up-to-date knowledge and alleviating hallucination issues. Nowadays, advanced closed-source LLMs (e.g., ChatGPT) have demonstrated surprising tool-usage capabilities through prompting and in-context learning techniques. To empower the capabilities of open-source LLMs (e.g., LLaMA) in manipulating tools, current efforts focus on either template-driven or token-triggered tool-usage. However, the former hampers LLMs' flexibility to address diverse user's queries due to constrained tool interactions, while the latter limits the generalizability when engaging with new tools, since tool-usage learning is based on task- and tool-specific datasets. To alleviate these concerns, in this paper, we propose a decision-aware and generalizable tool-usage framework (DEER). Specifically, we first construct the tool-usage samples with multiple decision branches via an automatic generation pipeline, thereby inspiring the decision-making awareness of LLMs under diverse scenarios. Meanwhile, we propose a novel tool sampling strategy to enhance the generalizability of LLMs over unseen tools. Extensive experiments demonstrate that our proposed DEER is effective and significantly outperforms baselines across various datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A contrastively trained retriever that selects behaviorally consistent call/no-call demonstrations raises H2A direct-response rate by 8.5 points and ToolDEER no-search accuracy by 4.2 points on average, without fine-t...

Pith tools