Pith. sign in

REVIEW 14 cited by

T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.14033 v3 pith:CH7J2YGB submitted 2023-12-21 cs.CL

classification cs.CL
keywords t-evalllmssteptoolutilizationcapabilityevaluateevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLM) have achieved remarkable performance on various NLP tasks and are augmented by tools for broader applications. Yet, how to evaluate and analyze the tool-utilization capability of LLMs is still under-explored. In contrast to previous works that evaluate models holistically, we comprehensively decompose the tool utilization into multiple sub-processes, including instruction following, planning, reasoning, retrieval, understanding, and review. Based on that, we further introduce T-Eval to evaluate the tool utilization capability step by step. T-Eval disentangles the tool utilization evaluation into several sub-domains along model capabilities, facilitating the inner understanding of both holistic and isolated competency of LLMs. We conduct extensive experiments on T-Eval and in-depth analysis of various LLMs. T-Eval not only exhibits consistency with the outcome-oriented evaluation but also provides a more fine-grained analysis of the capabilities of LLMs, providing a new perspective in LLM evaluation on tool-utilization ability. The benchmark will be available at https://github.com/open-compass/T-Eval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A new benchmark and metric show that large language models still struggle to call tools when the needed details are scattered across multi-party, multi-round group dialogues.

  2. WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

    cs.CL 2026-07 conditional novelty 6.0 of 10

    WorkSurface-Bench measures surface routing separately from answer correctness and finds near-perfect routing still leaves 25–44% answer errors across four LLM backbones.

  3. Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A large synthetic instruction corpus with guidelines, preference rules, and format variants improves LLM performance on five NLU benchmarks by an average of 3.1%.

  4. Reducing Tool Hallucination via Reliability Alignment

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A reliability alignment framework, Relign, that expands the LLM tool-use action space with indecisive actions reduces tool hallucination rates and improves task success on the new RelyToolBench benchmark.

  5. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  6. Teaching a Language Model to Speak the Language of Tools

    cs.IR 2025-06 conditional novelty 5.0 of 10

    LoRA fine-tuning of BgGPT models on a bilingual Bulgarian function-calling dataset yields large gains on a self-built 120-case benchmark while keeping knowledge benchmarks stable.

  7. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  8. Implementing Long Text Style Transfer with LLMs through Dual-Layered Sentence and Paragraph Structure Extraction and Mapping

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A dual-layered sentence and paragraph template method for zero-shot long-text style transfer, with a reported average gain of 0.20 over direct prompting but limited statistical and external support.

  9. Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Auto-SLURP is a new end-to-end benchmark for LLM multi-agent personal assistants, and the best current framework still fails a majority of user requests.

  10. A Framework for Testing and Adapting REST APIs as LLM Tools

    cs.SE 2025-04 conditional novelty 5.0 of 10

    A framework that converts REST APIs into LLM-callable tools, generates data-aware natural-language test cases, and reports an error taxonomy from over 2,400 executions.

  11. TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use

    cs.CL 2024-12 conditional novelty 5.0 of 10

    TL-Training, a task-feature-based training framework, lets a 7B CodeLLaMA-2 model reach competitive tool-use performance using only 1,217 training trajectories.

  12. Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need

    cs.CL 2024-12 reject novelty 5.0 of 10

    The paper claims that proxy tasks selected by cross-model performance correlation and small-model variance ratios can predict LLM tool-use capability rankings at early training stages.

  13. Evaluation and Benchmarking of LLM Agents: A Survey

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A review that proposes a two-dimensional taxonomy for evaluating LLM agents and highlights enterprise-specific evaluation gaps.

  14. Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression

    cs.LG 2025-05 conditional novelty 4.0 of 10

    ACBench tests compressed LLMs on agentic tasks and finds 4-bit quantization keeps tool use and workflow generation strong while hurting real-world application performance.

Pith tools