Pith. sign in

Paper Citation Record · LEDGER

MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2310.03128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.03128 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:48.010835Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:18:59.715072Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bfdcb510-4117-4c31-bdb4-3897e0a75dd1 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.521093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:5ac69ae693f4014f695adb21350cd361e8b83854ca4f178a01df5a51e3673d4b

Observation 15ec5b6d-4d56-46f8-a061-989c9f6bc2bc · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.181459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:ed2702984e58b2b253ea5025d1d164c7add43ac525ad2a4d940643f86024a38b

Observation 55495220-d2ad-4deb-8d59-2736b19284fe · inbound

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains cites this paper.

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:19:00.864649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T03:19:00.831153Z digest=sha256:d22385b806d699f892af832a9cfd7ff08839d352a39719f7bf87294ca4d3dd4e

Observation 3e0fd3d7-5ddd-4c39-9843-d7f7858f7ca4 · inbound

FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs cites this paper.

FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:38:20.876271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:38:20.876271Z digest=sha256:58226311213352fd093697dea38458a8a1dc1d9a542c233086c666a3bbc9bfbf

Observation 4c483699-7693-414e-bf15-0b1d645a95c1 · inbound

LegalAgentBench: Evaluating LLM Agents in Legal Domain cites this paper.

LegalAgentBench: Evaluating LLM Agents in Legal Domain MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T05:43:30.989852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:43:30.989852Z digest=sha256:01f02d236c4d0807c5b47345171171dfdbb54223f4b23b3c54762901a10ea9d2

Observation 34424258-f287-4c29-95a8-71d578960cf5 · inbound

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond cites this paper.

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T20:11:11.989677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:11:11.989677Z digest=sha256:95be6ee78f9086ebf3eea50fa1de380134f29239865a7c757562347276e9dd4b

Observation f8e8b2de-7d13-4cb2-974c-66d4b10d318c · inbound

When2Call: When (not) to Call Tools cites this paper.

When2Call: When (not) to Call Tools MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:48.010835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:48.010835Z digest=sha256:c58d64a533c624b7920a31d07f9ff7bd05a391d200d243db11e988d34e91c35c

Observation c18c7bb9-1a8c-4176-9085-053352f1ff88 · inbound

Prompt Injection Attack to Tool Selection in LLM Agents cites this paper.

Prompt Injection Attack to Tool Selection in LLM Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:08:28.977781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T17:08:28.933831Z digest=sha256:0b3bc04125fe6c5caae1d048e5d90dc007910dee610bf187746dd2f63ea71ec6

Observation 4f54f3af-95b8-43f1-8d21-d802d29b36c8 · inbound

ToolSpectrum : Towards Personalized Tool Utilization for Large Language Models cites this paper.

ToolSpectrum : Towards Personalized Tool Utilization for Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T20:23:48.870809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:23:48.870809Z digest=sha256:18ca7c182675e22ff83959414f1c5f8a83058fa69c3a85bee54f7d972fb907b8

Observation 9e16af64-e263-45dc-b0d4-27ecf863a911 · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.470715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:bcdc0bcd72530c82563e044703ccee0f6662124b40d17b6c36183e26181467f5

Observation b59b1997-5a74-45d7-8a59-a4d8437cf4cc · inbound

The Curious Language Model: Strategic Test-Time Information Acquisition cites this paper.

The Curious Language Model: Strategic Test-Time Information Acquisition MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:45.870083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:45.870083Z digest=sha256:0b063350d4c61fa7fc7bec298a60ac7418c4e7e49af8998750670badf3eb28da

Observation 20756517-6596-4bec-9837-d139ff5bdd0b · inbound

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues cites this paper.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.861022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.861022Z digest=sha256:6b12eeade2d2de0447b025aef59d52b6f97db923b315c5f549c7387447594384

Observation 6ac1df20-9361-4c43-82c9-f4898d7e2d0e · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:47.549550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:47.549550Z digest=sha256:d89034be5efa1acfa1b635f0496f73e7f7c6991cf44a7fd48a5524feec28c0ff

Observation 337c765f-ae53-452a-8937-39caf61300d2 · inbound

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models cites this paper.

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:43.754386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:43.754386Z digest=sha256:c2fa14e87f7f78c06fce95da83c37ce4532e25f54acb1d49f1395b5034f35d44

Observation 0b1ca4a9-c1c7-4390-8c31-1e095790d3e2 · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:57.024887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:57.024887Z digest=sha256:d557a926170d16fad7c66970efcdbd142a946a9f457c155b02c337636697accf

Observation 22550fcd-f160-4d2f-bbd9-269ad5bc33d1 · inbound

Towards Compute-Optimal Many-Shot In-Context Learning cites this paper.

Towards Compute-Optimal Many-Shot In-Context Learning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:38.656905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:38.656905Z digest=sha256:1d052633ab10c67a6ad700af4212f76d51da02dfce34d539a29f3de3f9b163d6

Observation 51babeb4-227a-4516-95b0-2f094cd24194 · inbound

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis cites this paper.

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:40:51.463887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T00:37:11.945418Z digest=sha256:7d85d86934645f0695471f4789e16f6bb1d68dfb8722150aeacf63ccbff82577

Observation 1576480d-e593-4139-8f1a-353f60d92f52 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.609937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.609937Z digest=sha256:83e7161c7c1165094f59fa16a318db70b7c100d481a42cf64c02d665d3ce3a12

Observation 67a0b2f9-4580-4f30-9e34-68085d9089bd · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.335428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.335428Z digest=sha256:16a969ff58cc3486c302e85d3e6da3f2de733f4a4eb9b98dd4929d0851eb3ad4

Observation 915671c3-e5cc-4ec2-9679-5d17a1a45d50 · inbound

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever cites this paper.

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:16.475327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:46:16.475327Z digest=sha256:d7f9ee5676625614ce0f1e24dbac5ac9d0c7b22fc1b43263a39ef74de89c506a

Observation a03ba34e-3ea1-4c44-a50e-849e2df3cda7 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.792208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T11:02:55.529271Z digest=sha256:d617f4dccd3ced14b21cd6089239ba22aa8ce3a9567d52b094c2881c2dac71e6

Observation 0e9d9af4-cf15-48a8-824e-b99be0ce07c2 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:12.592641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:12.592641Z digest=sha256:a1b3c59ab8ed04e258e3b161319939588a4fecea95182266b399e9c975efefde

Observation 523d99e1-82e0-4b96-84b8-3501b8c110a9 · inbound

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception cites this paper.

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:50:52.539483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T03:46:03.228969Z digest=sha256:8f7ddfb43b22aaf50b68eabd1f09411ba648a0c91e42bcaa204c39decb699977

Observation 0700dca1-1a6d-46f8-8cd4-7a0dabd4525f · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 261

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:18:20.250588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:5cedf612402ab702cdb66b37c499a8af1083c25fd9e535f441d6fe650a33be88

Observation ad97d443-4774-437f-b967-c37f48d6ce05 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:36.280577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:36.280577Z digest=sha256:f80acb41c1ea4375c25f12b348d686279a8da9c6e331f160cbffae4bb59bc11f

Observation 59ec2ede-71a8-4033-86a9-01bfcebf3c9f · inbound

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls cites this paper.

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:46:33.268448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:46:33.268448Z digest=sha256:e870c4e9e5dc4c962b17ccb71df73aa23ef120513fba381fe6ee752425c8f690

Observation 4b853e52-1d2d-404e-83a7-6d6775680c58 · inbound

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents cites this paper.

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T19:59:59.030585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T19:59:59.030585Z digest=sha256:e7251e6067a26d0c284c64d4600fb3a0548b964a83aad3e20247e150c06f39ca

Observation 4f2bdf27-338b-487a-8c00-d32a047c8fd4 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:10.333342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:35378f1c729c45b828e9c851f277896f08498af5cbb2167843bfe03991e78b6d

Observation 4a5476e6-df85-4c89-9155-3d38d16c87f8 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:33.837524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T19:04:53.639217Z digest=sha256:ed4c1acf3cb1874bf9e2c61eba630176ba790ce7c171b27ed49a9f1cc0e85115

Observation d00cbb03-51f7-491b-9c00-d4a541116888 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:12.424114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T00:25:46.507401Z digest=sha256:c8a13b8b56f5a18b7b52e240f19efdfcc1f2519faf3a9a0f4d880ad458b1f757

Observation f87b296a-6676-48f1-9d59-9f30edbfe7a2 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T05:24:44.710671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:24:44.710671Z digest=sha256:f8da84f70e18868948cb706462b82995662b5797acf4d0016948dbd368c260c0

Observation 7709fb7c-2ad0-4ef3-886e-0853e6383a01 · inbound

Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI cites this paper.

Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:37.672983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T19:06:36.952203Z digest=sha256:de063e7e4c2be6dc054f0f646800847fcff546a6a9a9fd47f0990521e6a0058e

Observation 684b6e09-e7cb-4db7-9017-2bb70f976afd · inbound

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation cites this paper.

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:42.967662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T16:19:39.392516Z digest=sha256:9cea7c0e6e02c611caa416d7700d03063ba6c7f36c502e69664cccbb755d62fc

Observation a86b0777-73cc-4fea-906b-4dfb1a7f260f · inbound

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning cites this paper.

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:51.785672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T01:29:47.384341Z digest=sha256:a9ef6fc41f2e0f96fa7caf6d7471a7d4b1d350c5f6203a2ba14c9f32917239ea

Observation b053461d-5763-4d1d-8f70-7b549a050ee9 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:35:04.476298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T05:31:54.252816Z digest=sha256:c80f5ab58ea0aa0c00841728f105de08dfa66083d917a49629e57de32e0bf76c

Observation 3e545d30-a3b9-44e7-b915-75cc50001943 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:53:43.552641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T20:52:20.975459Z digest=sha256:4bf09370b2792538ffd065daf2f2ad39cb3de1bbc388d9df9ad4d3efeec51095

Observation d473db4c-3e82-4cd3-9767-bcb8063c5b44 · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:57.998107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-15T03:07:38.232966Z digest=sha256:ffa9e5d5e4382ad80f09754cf66809f4ec4af8c67ca1697d1a8cb60f68f999a1

Observation 749d96c9-e863-41ee-9a88-6548c8b301dd · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 165

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:52:39.943407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-19T16:51:13.491389Z digest=sha256:52f3148f0c955d3dbc0385bc42630609bea577235f2918f8ca555951de414f99

Observation b36e908b-d4c6-4335-a441-73f48c1b2f6c · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.401757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:17094ab125701357ecda5ef1c34c9a0552e9c9222b52670e00b2be42df826edf

Observation b7301589-2913-4ce2-9e68-39c437d74e07 · inbound

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning cites this paper.

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.732455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T00:25:59.532059Z digest=sha256:16f92b714d85ac38051201afe52da23bd77b536e5a26be8716057f8fbb061607

Observation df94b234-2959-4db7-a73e-2469787000e9 · inbound

Capability Self-Assessment: Teaching LLMs to Know Their Limits cites this paper.

Capability Self-Assessment: Teaching LLMs to Know Their Limits MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.643921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:25:50.579196Z digest=sha256:ebaad1ce2836895937bca86c56ced04543b5d6f21348538b38c23aa3432c35e2

Observation 0cbf862b-fdcf-46a9-adff-cdc6cc2f28a2 · inbound

NTILC: Neural Tool Invocation via Learned Compression cites this paper.

NTILC: Neural Tool Invocation via Learned Compression MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:04.773446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T00:05:36.600389Z digest=sha256:c8cbf14680a76062c3d5734722398ff6e08c0c2887be8f9fdae50a76f401787f

Observation 7f6a9ba2-b68d-4f41-8fe9-02eb8f016c6e · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.155289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:34:23.382705Z digest=sha256:ee211fdd7ac9863ef1ba6b6fedbb07a0557ce6e823a2d77e609cb54b4331b99c

Observation c152ec48-6f8b-4634-82ed-6532c4cc7091 · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T14:59:04.152474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:59:04.152474Z digest=sha256:eb722a608671099c83a10c4ff020a98b9fe66a30f56e348a13692d7323789daa

Observation b490de2a-c3a7-4ef9-8b45-7932e70decb2 · inbound

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose cites this paper.

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:18:59.716659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T00:37:50.570757Z digest=sha256:9e1a3982e7484ff156c627a68d3b7155c251cfd95344d7ce9d857d6fd1675265

Observation cf78261e-76a9-4938-8b66-59b2e68a263f · inbound

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing cites this paper.

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T01:35:09.657698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:35:09.657698Z digest=sha256:a414a8abbe00ed1080754ca32a4c6b132160aacb92330244089738414f4f7980

Observation f5bd2c1c-a64a-4a8f-bc81-35180af2a342 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:29.777251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:29.777251Z digest=sha256:1561c2648253827e837eb44b29a3cbfb4753e7ae9c09dbec56f02c772563eb94