REVIEW 14 cited by
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLM) have achieved remarkable performance on various NLP tasks and are augmented by tools for broader applications. Yet, how to evaluate and analyze the tool-utilization capability of LLMs is still under-explored. In contrast to previous works that evaluate models holistically, we comprehensively decompose the tool utilization into multiple sub-processes, including instruction following, planning, reasoning, retrieval, understanding, and review. Based on that, we further introduce T-Eval to evaluate the tool utilization capability step by step. T-Eval disentangles the tool utilization evaluation into several sub-domains along model capabilities, facilitating the inner understanding of both holistic and isolated competency of LLMs. We conduct extensive experiments on T-Eval and in-depth analysis of various LLMs. T-Eval not only exhibits consistency with the outcome-oriented evaluation but also provides a more fine-grained analysis of the capabilities of LLMs, providing a new perspective in LLM evaluation on tool-utilization ability. The benchmark will be available at https://github.com/open-compass/T-Eval.
Forward citations
Cited by 14 Pith papers
-
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
A new benchmark and metric show that large language models still struggle to call tools when the needed details are scattered across multi-party, multi-round group dialogues.
-
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
WorkSurface-Bench measures surface routing separately from answer correctness and finds near-perfect routing still leaves 25–44% answer errors across four LLM backbones.
-
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis
A large synthetic instruction corpus with guidelines, preference rules, and format variants improves LLM performance on five NLU benchmarks by an average of 3.1%.
-
Reducing Tool Hallucination via Reliability Alignment
A reliability alignment framework, Relign, that expands the LLM tool-use action space with indecisive actions reduces tool hallucination rates and improves task success on the new RelyToolBench benchmark.
-
Qwen-Audio-VAE Technical Report
A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.
-
Teaching a Language Model to Speak the Language of Tools
LoRA fine-tuning of BgGPT models on a bilingual Bulgarian function-calling dataset yields large gains on a self-built 120-case benchmark while keeping knowledge benchmarks stable.
-
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.
-
Implementing Long Text Style Transfer with LLMs through Dual-Layered Sentence and Paragraph Structure Extraction and Mapping
A dual-layered sentence and paragraph template method for zero-shot long-text style transfer, with a reported average gain of 0.20 over direct prompting but limited statistical and external support.
-
Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant
Auto-SLURP is a new end-to-end benchmark for LLM multi-agent personal assistants, and the best current framework still fails a majority of user requests.
-
A Framework for Testing and Adapting REST APIs as LLM Tools
A framework that converts REST APIs into LLM-callable tools, generates data-aware natural-language test cases, and reports an error taxonomy from over 2,400 executions.
-
TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use
TL-Training, a task-feature-based training framework, lets a 7B CodeLLaMA-2 model reach competitive tool-use performance using only 1,217 training trajectories.
-
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need
The paper claims that proxy tasks selected by cross-model performance correlation and small-model variance ratios can predict LLM tool-use capability rankings at early training stages.
-
Evaluation and Benchmarking of LLM Agents: A Survey
A review that proposes a two-dimensional taxonomy for evaluating LLM agents and highlights enterprise-specific evaluation gaps.
-
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
ACBench tests compressed LLMs on agentic tasks and finds 4-bit quantization keeps tool use and workflow generation strong while hurting real-world application performance.
Discussion (0). Continue with ORCID to comment.