Pith. sign in

REVIEW 11 cited by

ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00741 v3 pith:ZSSPAYVH submitted 2024-01-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords toollearningllmsscenariostooleyescapabilitiestoolsalignment
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Existing evaluations of tool learning primarily focus on validating the alignment of selected tools for large language models (LLMs) with expected outcomes. However, these approaches rely on a limited set of scenarios where answers can be pre-determined, diverging from genuine needs. Furthermore, a sole emphasis on outcomes disregards the complex capabilities required for LLMs to effectively use tools. To tackle this issue, we propose ToolEyes, a fine-grained system tailored for the evaluation of the LLMs' tool learning capabilities in authentic scenarios. The system meticulously examines seven real-world scenarios, analyzing five dimensions crucial to LLMs in tool learning: format alignment, intent comprehension, behavior planning, tool selection, and answer organization. Additionally, ToolEyes incorporates a tool library boasting approximately 600 tools, serving as an intermediary between LLMs and the physical world. Evaluations involving ten LLMs across three categories reveal a preference for specific scenarios and limited cognitive abilities in tool learning. Intriguingly, expanding the model size even exacerbates the hindrance to tool learning. The code and data are available at https://github.com/Junjie-Ye/ToolEyes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A new benchmark and metric show that large language models still struggle to call tools when the needed details are scattered across multi-party, multi-round group dialogues.

  2. Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A new bilingual benchmark with per-question human accuracy and common mistakes shows current multimodal AI models still underperform humans on reasoning.

  3. Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems

    cs.SE 2025-07 conditional novelty 6.0 of 10

    LLM tool agents fail at parameter filling in five recurring ways; perturbing tool documents and user queries drives most failures, and invented parameter names are tied to the model rather than the input.

  4. RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving

    cs.SE 2025-05 conditional novelty 6.0 of 10

    RepoMaster, a repository-aware code agent, lifts the task pass rate from 40.7% to 62.9% and cuts token use by about 95% versus OpenHands on the new GitTaskBench benchmark.

  5. When2Call: When (not) to Call Tools

    cs.CL 2025-04 conditional novelty 6.0 of 10

    When2Call measures when language models should call tools versus ask questions or refuse, and shows that RPO training substantially improves this decision-making.

  6. VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    On VisuLogic's 1,000 vision-centric puzzles, the best multimodal models reach 28.1% accuracy versus a 24.9% random baseline and 51.4% human accuracy, and an RL baseline lifts accuracy by up to 5.6 points.

  7. Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training

    cs.CL 2025-02 reject novelty 6.0 of 10

    Continual pre-training on a large agent-focused corpus of API docs and tool trajectories improves an 8B LLM's function-calling and planning, but the reported generalization is undermined by benchmark data appearing in...

  8. CallNavi, A Challenge and Empirical Study on LLM Function Calling and Routing

    cs.SE 2025-01 conditional novelty 6.0 of 10

    CallNavi is a new benchmark for LLM function calling with unfiltered, nested, multi-step API tasks; GPT-4o leads the leaderboard and a two-step routing pipeline improves fine-tuned models.

  9. LegalAgentBench: Evaluating LLM Agents in Legal Domain

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new Chinese legal-domain benchmark with 17 real-world corpora, 37 tools, 300 human-verified tasks, and a fine-grained evaluation metric shows GPT-4o leads with 79% success under ReAct.

  10. CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

    cs.SE 2025-06 conditional novelty 5.0 of 10

    CRITICTOOL, a benchmark of 2,740 tool-calling error scenarios built from BFCL and T-Eval with GPT-4o-based error injection, finds most LLMs rarely recover from tool-use errors, with GPT-4o best at 69.01 overall and to...

  11. Auto-SLURP: A Benchmark Dataset for Evaluating Multi-Agent Frameworks in Smart Personal Assistant

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Auto-SLURP is a new end-to-end benchmark for LLM multi-agent personal assistants, and the best current framework still fails a majority of user requests.

Pith tools