Pith. sign in

Paper Citation Record · LEDGER

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 76 inbound Pith citation observations for arXiv:2304.08244.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.08244 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T20:51:41.198767Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 76 of 76 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:02:39.336321Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch17

External citation measurements

12
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7e8c62eb-e041-47e0-8dcf-1a388158b592 · outbound

This paper cites Advances in neural information processing systems, 33:1877–1901.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Advances in neural information processing systems, 33:1877–1901

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:51:41.274960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:b3413939af3337d509a91bffa13bf643f69f9a0a1ea9c0781f4595bf7dd3a547

Observation 30f8b389-5a5a-42b2-89b7-5f88fa2c3b2f · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.214006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:95551ffec217229ace2201f49f74e3d7175ddc72ee2fd92648707d6635ef8648

Observation eb4fc248-f863-459b-bfd1-6b4f69cc96e1 · outbound

This paper cites Large Language Models as Tool Makers.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Large Language Models as Tool Makers

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.217174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:9fa04ec8785d2c9a500e6e58d3457dd97a7d06782ee18736cd383deab11e9df5

Observation ef3ea066-0f13-48cd-8186-9dc0bbfc0b41 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Evaluating Large Language Models Trained on Code

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.220201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:537f0086dfb2559b9bfed95823cb21f13e3ca82b03e83d30594c7a653a81aabe

Observation b5af33ed-2587-4617-8fdd-de8a36f0ca8d · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.223381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:1d25b9c56b8dd3222bbc00a5ca0261e035272bbf32c9d15499e421ea0530c2f9

Observation 5156ba47-0c6d-492d-a45b-2e621718da54 · outbound

This paper cites Atlas: Few-shot Learning with Retrieval Augmented Language Models.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Atlas: Few-shot Learning with Retrieval Augmented Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:48:43.695710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:6393372416b63c8ac826d0cafadd8a32382033898e98b4f563a1d3ec3f23a73d

Observation da151480-a189-4469-b097-b10dbc9d831b · outbound

This paper cites TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.229535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:9a8cdb2fb29af0031542b98e75ee5f3c275a3b77f7b7d472b8d553461c8e41bd

Observation 8299eb5f-16fb-4597-acc7-f89f102dd58a · outbound

This paper cites Augmented Language Models: a Survey.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Augmented Language Models: a Survey

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:41:28.390682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:c38ca4e537191fd486d81cb728963f1f2c0ed16a3271663506ad93ac608059e5

Observation 250c8fff-585b-4489-a77f-d19bf8baa484 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs WebGPT: Browser-assisted question-answering with human feedback

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.235370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:c0a6f456742116f2527bbb2a921090abd1926d448970d5782acdc639787a7be5

Observation e7b489bd-d3ff-47e7-924e-5d54054b8650 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ART: Automatic multi-step reasoning and tool-use for large language models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:03:06.497846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:5046b3e64609ad6a3914ec1849a4cf141f8ac62e10a4457ec79a21566c19ed11

Observation a5fb5a6a-5f0e-43ed-806d-e70ad41e22b0 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Gorilla: Large Language Model Connected with Massive APIs

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.241440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:79680d302824ddee9637f2e04a7bfff848fdd876738f008bc4601ccbf057b1b5

Observation 978e2a16-ee6e-4f7b-b841-2adbd758b011 · outbound

This paper cites CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.244375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:2256404cb3e549bdfdf2935af2a33546a63df46ca245c7d1d858923259b87577

Observation 70825f6d-7bd5-4910-acb7-0d9e37811786 · outbound

This paper cites Making Language Models Better Tool Learners with Execution Feedback.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Making Language Models Better Tool Learners with Execution Feedback

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.247322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:cdf110cc211b9e3d85e2e8b85affec51e765bfcd0fbc2264a3764978f4589ee7

Observation 156bed37-3106-45fe-a0c1-bbf1decb5782 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.250457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:d1fa5f5ea14cdec4a2d1bd46dd579272bf41d728ea4e424ff90e534eec50ac0b

Observation 40d25fbf-7ebd-45be-9e3e-a9900ab27456 · outbound

This paper cites Preference Ranking Optimization for Human Alignment.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Preference Ranking Optimization for Human Alignment

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.253387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:e36b816eeccce9b38af8d9f481a1501ca8197ba458fcecb32e780584b0463ceb

Observation 6a311f75-1698-456f-8832-8f53030f1dbd · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:03:48.638764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:f530c3c46c4fda8369aa91e2e1f43c6aad3cd81add708d769f48c5fbbbf0329d

Observation ba776283-7bd2-4bbc-bf83-38c71dffbe25 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.259217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:6840ed99088799f80b8c8596499c594f975dc0091143bdcc79dd5eb7cd0434c3

Observation bfbdf7ce-c539-4cbb-bd42-0af8fbe72369 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.261979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:7a37c02e704f5f5cb9a2e1153e1c56e5637958759c7bb6260de01c7030249a4b

Observation f235feca-3722-451c-9108-83b41a289ca4 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ReAct: Synergizing Reasoning and Acting in Language Models

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:51:41.264870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:5442fada622991d440669f76f0e7879a74b74f46e326ddd0037c67acba73df19

Observation a1cb8f7d-c6b6-4807-910f-338c384ceca0 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:50:01.041344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:931527d314d710708ff5fc354a22cd21ca5281f9b708de26c4072b6b48175858

Observation ce24e540-d6f4-49c7-8ed1-be0308682643 · outbound

This paper cites A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.270956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:d856dc8efdfc24751887274bc4a25660cb4ea5dc6a69f293b390ef86d8a3bbd3

Observation 44c9d07d-b1b4-4e14-9e1f-d4cfcb900f9b · outbound

This paper cites ToolQA: A Dataset for LLM Question Answering with External Tools.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ToolQA: A Dataset for LLM Question Answering with External Tools

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.210477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:836c245198bd17d58cbc1e76950c7f9508d40a4e1c421d046accecbcc7238a2f

Observation bf9fc090-dc9b-459a-84de-acffc165fabd · outbound

This paper cites name": "ToolSearcher.

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs name": "ToolSearcher

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:51:41.273003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:51:41.198767Z digest=sha256:46a8d39f51a696ba191cf8943a85dd1b8e5657eb5bd699de688e64f19d412d69

Pith citing papers

Observation cea748f9-ea52-4557-aae7-92268217187b · inbound

Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory cites this paper.

Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:37:20.889942Z digest=sha256:70e6e3e0fdab367dc97aff57e3d7e64fafb3607b9661de46c94409d450d37164

Observation 3254d231-6c20-4e95-8005-f8cd33db1e87 · inbound

ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases cites this paper.

ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:03:48.465611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:03:48.426204Z digest=sha256:35a73ff3ddddf9500583c37225286eeec736816b9ba238ceb2d272b3a4be00c5

Observation afbfab5c-6e56-4b0c-a33a-bf472e40bc3c · inbound

Mind2Web: Towards a Generalist Agent for the Web cites this paper.

Mind2Web: Towards a Generalist Agent for the Web API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:05:15.992207Z digest=sha256:e68b54b0ac1f28f0bd892cc60f7d1c1eebff0e448138b0c72bd6e7151a6206c7

Observation e1c98ac3-f231-4fb0-a5f2-f65336e93c80 · inbound

A Survey on Large Language Model based Autonomous Agents cites this paper.

A Survey on Large Language Model based Autonomous Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:03:00.340349Z digest=sha256:b1ab40be375febfdc22de2bd325a0c1469857238bad08f3eb70f9dc8b0e1d294

Observation 648ad77e-fc10-466a-8385-0b591ab0782c · inbound

GAIA: a benchmark for General AI Assistants cites this paper.

GAIA: a benchmark for General AI Assistants API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T15:46:03.247029Z digest=sha256:16579f5b73bccf41ae03a7ff67614b1cf39a2efd40bea8906bc1894da48fa38e

Observation 93c65495-1427-4f0a-b1cb-4e78ee6135a6 · inbound

Learning to Ask: When LLM Agents Meet Unclear Instruction cites this paper.

Learning to Ask: When LLM Agents Meet Unclear Instruction API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:13:28.115525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T21:08:42.276002Z digest=sha256:19b7db1d2113588471a4ad6b05a06afbbbc07378ab54be06731e151822e9dbdf

Observation 165a7397-528a-4460-b29d-b9c72399118a · inbound

Prompt Injection Attack to Tool Selection in LLM Agents cites this paper.

Prompt Injection Attack to Tool Selection in LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:08:29.053456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T17:08:28.933831Z digest=sha256:c9f5dd0df1044acc76295aaf3398301dc434b4a91d50d31a8083c1a19262facb

Observation dbee5958-f76d-4e3a-ae8c-9eb264c3da18 · inbound

NaviAgent: Bilevel Planning on Tool Navigation Graph for Large-Scale Orchestration cites this paper.

NaviAgent: Bilevel Planning on Tool Navigation Graph for Large-Scale Orchestration API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:13:01.834735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T08:12:10.542542Z digest=sha256:6a4d21a717af9c6cdd16983ef0c606d6f9988cce9c6cd2f559a08eca372bfd2e

Observation a4d59512-941d-4138-8afd-e64b0a39d3d4 · inbound

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues cites this paper.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.336321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.336321Z digest=sha256:803c44e14334e36800ee82a64f1cfe8791fd5eb2da13b18bdff8352b330f75ef

Observation 5e8bd2b4-b095-4eaa-96f2-ef35ab2f289f · inbound

AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks cites this paper.

AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:54.130776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:55:54.130776Z digest=sha256:fab72a70382374bcff430c0fd6ba6d434e39978ddf1bc87e6f90b8d9867fc1c7

Observation e53ccb40-cafc-4294-95b5-09604d98476e · inbound

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models cites this paper.

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:43.768837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:43.768837Z digest=sha256:88202253ebe0520f708ccbb7c4debc960835b5c1373b9e6a8381238aa0a2f8b4

Observation c0c88fd4-d6ba-4d6b-90b2-40d720b4d7b9 · inbound

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering cites this paper.

DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T17:11:33.897035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:11:33.897035Z digest=sha256:6fcbe1a385284c830788e61279c1840a389f1ba4c6b3203e5d79a9e900406f14

Observation cc626abe-0d63-4856-9a82-f1567f756248 · inbound

GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts cites this paper.

GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:47.117662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:47.117662Z digest=sha256:25134f7488d0d50b1d471d584c2bee9ec5de0b342b13948e0a1c7d2117931837

Observation 03992d48-a956-461c-9576-8456f8801208 · inbound

FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain cites this paper.

FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:59:37.812022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:59:37.812022Z digest=sha256:9d0a0eafe0f44aeff2c734de472a9b2aa9ce0c55cca4333fd1c4cdf8053834fe

Observation 29a34b5f-864e-42dd-b212-e88a4fef4cd2 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:eb1463a8bc446fd44cb0aa9605d432ea7983250c2ec5b11bd2b13311d8c49543

Observation bf497c9a-95ca-4e9a-b2db-90c07482a652 · inbound

MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations cites this paper.

MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:53:45.371850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:53:45.371850Z digest=sha256:47fbc0d92ef513795d690c613bd6c9747102aa907ad04d4d03925cc72186229d

Observation 4f5134ab-59be-4e29-8c6c-d1a72f2f2df9 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.642196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.642196Z digest=sha256:a893daf8fba4f26827e237660a6d614dd10b92664b6ed05f6b6258967ced9d8f

Observation 3701e0ee-3fa4-42a1-a9e1-f2222427131e · inbound

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench cites this paper.

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:00.677620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:00.677620Z digest=sha256:32dc3b3102d117aa274b1b7e057fa01934a2ce9a4a68a5b45899171ec4228c9a

Observation 879e3b24-c647-48e2-8516-817f99aa7375 · inbound

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions cites this paper.

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T15:01:31.337055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T15:00:51.162221Z digest=sha256:12e4fd5d5d419296dfad824a381b6501b8b12349d783f5162c05482be9af0b97

Observation 779cdcf4-e961-4759-9f23-5a1a1158c2f8 · inbound

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling cites this paper.

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T06:30:59.574221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T06:30:39.858246Z digest=sha256:d06fe14b18a7877516cfbaf67c885dd93c12c6d275779df1add3961728617ab5

Observation a8325615-c5d3-4968-82b9-7a5427425ec7 · inbound

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception cites this paper.

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:50:52.576462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:46:03.228969Z digest=sha256:4f8de474841c6344288913de896b43e659958203d15cb1b73982822b375af8b0

Observation a2666c32-226d-4844-8154-8abf2f7506e3 · inbound

Asking LLMs to Verify First is Almost Free Lunch cites this paper.

Asking LLMs to Verify First is Almost Free Lunch API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T21:03:18.066668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T21:03:18.066668Z digest=sha256:e42dbd8bbcf566667544758d3c38a4eaf818ce2d2aa6a3e6a183ac3e91d6a078

Observation 847dd99e-6431-4773-ac71-84f4d28dc959 · inbound

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction cites this paper.

AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:40:44.011544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:38:37.818623Z digest=sha256:05e870a9dc6493ea1e5224e8b1e3aa7efaf4fb6a1d46f3b285751dda5da1d703

Observation a4d50127-be3c-47ff-98a4-5d172e8a2203 · inbound

RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation cites this paper.

RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:09:51.803474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:09:51.803474Z digest=sha256:1d5906d4ca960d10e5e5173541fd76912e76ea5f6ebbd41917f1ac09a2afebab

Observation ab1047b2-a35d-471f-8661-f38d8ed2d744 · inbound

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation cites this paper.

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:07:11.528454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T03:04:17.755968Z digest=sha256:97d1966c50d9163db619f65175f95e886d2de19a1c3c1fc9ae04d9a366aeaa9b

Observation 636fa33c-ccd1-465e-82ac-8d3600073afe · inbound

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints cites this paper.

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T12:00:04.121440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:32c27295e5d7c0964f23e903885c2e9d9f569961659d40002630b255fc24992e

Observation f50ff977-8629-4a1f-a475-43247594d8a6 · inbound

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use cites this paper.

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T05:54:53.390032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:54:53.390032Z digest=sha256:6cd82bf44e6f27bfa7e7da311d05e5e57f8ea2d9649778ae8ac87c14fd462743

Observation d10a8c0c-04cd-4db8-9426-48fa64c28d0f · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:21:23.204138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:5c03048102e9e25df33241875f0db50503dbe23e4e1ac13a8baf66f0f2e48245

Observation ca9df927-aee1-41ec-bee8-957e6d3e6a7a · inbound

MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security cites this paper.

MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:10:50.283791Z digest=sha256:ac664549610d8d2eb2715566941227598d857e035a8c761f095300eaadf4ac97

Observation 0061d9e7-b0f0-4fd1-81d2-ae2e51da7d0c · inbound

SAGE: A Service Agent Graph-guided Evaluation Benchmark cites this paper.

SAGE: A Service Agent Graph-guided Evaluation Benchmark API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:41:23.956104Z digest=sha256:40c334f090f5f3c20b26771c4e969cbc10657186bf002df6ede6f5ef750c8f2f

Observation 2a58aa53-804e-484f-9e7b-1ff3a98a0502 · inbound

A Periodic Space of Distributed Computing: Vision & Framework cites this paper.

A Periodic Space of Distributed Computing: Vision & Framework API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:55:44.933926Z digest=sha256:92a69c6c073ccc9c3f9ba9d279f7a14c7554353c2ca43eb0f4911953a05b493b

Observation 9e87fac9-b753-4da4-8aab-a8a075835d1d · inbound

English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training cites this paper.

English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:08:32.517020Z digest=sha256:389d9844bccac7df4c0d5a0ac8b188c86fd115ead57169a3f924d7b4ba1b8808

Observation 9cfb58f7-7bbf-4682-bdba-238f0af61c55 · inbound

GraSP: Graph-Structured Skill Compositions for LLM Agents cites this paper.

GraSP: Graph-Structured Skill Compositions for LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:16:48.625528Z digest=sha256:7a3669caababa8a3da10364865280c643ba659f0b7b7b77fa6879a35d681c506

Observation 5c724b56-4960-4e53-b244-a3d4218aadd9 · inbound

PARM: Pipeline-Adapted Reward Model cites this paper.

PARM: Pipeline-Adapted Reward Model API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:15:26.015817Z digest=sha256:8a125c18903c12e9052b1e54b69d3ea36162c055e3147a218c03b1f953c310f5

Observation 3ebb8397-8721-4bde-a128-e3f0b428e1dc · inbound

Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language cites this paper.

Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:56:53.055513Z digest=sha256:78ebdf349c3f1b5afedb5758f527bc718a5bcebc9047d6b26e4378c842d7d3ea

Observation 93c2ddaa-9498-4fb2-b04f-ec25359d1462 · inbound

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models cites this paper.

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T00:50:27.257888Z digest=sha256:3f34a8d9118a6d6ee9010e95de601634f9dc825fd8c4815b2d45647eb27a682b

Observation 239173b7-1c36-4516-91c7-d301940f608e · inbound

Quantifying Divergence in Inter-LLM Communication Through API Retrieval and Ranking cites this paper.

Quantifying Divergence in Inter-LLM Communication Through API Retrieval and Ranking API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:59:56.133502Z digest=sha256:950686c8526e0f3568cb1ee6d6f9eb65ada838d30c5352ac5baa96cedbb08995

Observation 759fa504-1109-4f66-8091-e3684506dbd9 · inbound

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows cites this paper.

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T05:52:13.100867Z digest=sha256:f8f010cab37021f0dc4e1ece23dc063cf903dcb3086a5f74a071f134d05193ae

Observation fd25fa72-8feb-46a5-b24e-4d9c62576748 · inbound

Tool Calling is Linearly Readable and Steerable in Language Models cites this paper.

Tool Calling is Linearly Readable and Steerable in Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T03:09:11.013914Z digest=sha256:1a969d11d2e377caacb0ce576bd99ac17db6ebc61c1135809f75c32341429514

Observation 54975448-a0e9-48ef-93e9-4394a9ad97bb · inbound

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning cites this paper.

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T02:47:31.830410Z digest=sha256:0538c15bf88b8ba134db7ff47bfb340b81312e69c70c434ee48b54c0de4faf5e

Observation 0b97cd65-7bea-4e58-a6b0-e3668a8e6abf · inbound

Trajectory Supervision for Continual Tool-Use Learning in LLMs cites this paper.

Trajectory Supervision for Continual Tool-Use Learning in LLMs API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:20:20.779419Z digest=sha256:4228f36fb2be0363bbe63677e0c98e935f1969ee47bd491e48f7f95b97bdc7f1

Observation 2209d1b2-d772-49b3-8ae0-03fa309fffde · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:51:41.275571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T05:31:54.252816Z digest=sha256:19ac64c734b14c1fab09242f027e1315badc6c4316f556dbc54b6b4c61316235

Observation 5b7dde3c-31f3-4ae2-8ac3-39aae584ba88 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:53:43.544451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T20:52:20.975459Z digest=sha256:fc6ea5726e6eb2cd9eea38949438d632daba37ca8dd8823a50430823c17daecf

Observation 6c47fd8d-faee-444e-8d3f-da9f3b66d61f · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:13:37.648430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:bed29a98c27077a1acdc8389ac8220c7ca1722113df1b1f3923512603f8b3879

Observation 90dc69b7-7f6e-4b53-aff8-4fe5b1b3a533 · inbound

Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs cites this paper.

Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T22:32:49.692323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T22:30:21.756864Z digest=sha256:6bd27324957e724037f2b5eb0f5d5bd4816d982432a258faf900fc4a7de09134

Observation c3c71dd9-7383-498e-8b7b-17c10c0a095a · inbound

Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning cites this paper.

Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:15:00.685356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:08:57.644677Z digest=sha256:aa4ff6113150199f68932116cff7968fe71e2b7e990f51088a691207b7842e59

Observation d230f240-1a2f-4e96-811f-2bde4f1dacd8 · inbound

An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments cites this paper.

An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T09:58:10.845623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T09:58:05.349147Z digest=sha256:26bd066d0f291e0d0601c1a8fa9f73d00ff892547d6aa9a0461633e41e2463ba

Observation f9a0d735-9023-45df-ba13-1427bfd3b862 · inbound

Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents cites this paper.

Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:50:20.462079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:47:01.153482Z digest=sha256:36d85e465cadc0a3560b3f83580f3b047879a245a0710d4c0f98003f41615aa9

Observation 462c8bb5-f17f-4410-b4ef-16686b16daa5 · inbound

Testing Agentic Workflows with Structural Coverage Criteria cites this paper.

Testing Agentic Workflows with Structural Coverage Criteria API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.597917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T16:17:28.497424Z digest=sha256:a86d4d0ce4c8ce40780c46ea84fc2b847008a39c4f085f5d0af57c12a8ad08bd

Observation 1d3c429f-c7ef-4adb-97e0-a6f671ca604c · inbound

Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation cites this paper.

Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:16:11.834422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:25:51.439330Z digest=sha256:be3f81b7aa4f5b771b6a226e307e38cf975f5cac3d0104d41cd9023d16dccf30

Observation 919b48d2-47cd-4d29-91fc-a3a499af0a07 · inbound

CLI-Anything: Towards Agent-Native Computer Use cites this paper.

CLI-Anything: Towards Agent-Native Computer Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T05:16:39.588710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T08:20:03.855573Z digest=sha256:c0f217bc4604cd0ad48eb57872f18433c645dbcfdbe4c5a5f1581e014e29a19c

Observation f603fe7d-7fbd-48ac-811a-c1f17e964cf5 · inbound

NTILC: Neural Tool Invocation via Learned Compression cites this paper.

NTILC: Neural Tool Invocation via Learned Compression API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:07:04.802789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T00:05:36.600389Z digest=sha256:12f2b463daa7a118b7206429cdedb0af9af05c9ba6f90e3349caac0460caac73

Observation 20807d09-fb00-4ec9-b8cc-6b686d58eee2 · inbound

Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents cites this paper.

Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:17:18.709187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:32:06.237095Z digest=sha256:8d32ed410fcb23268f84b897819ef71b19a0208ebbf422150cf26b9d69eb9d3c

Observation 8d9934ab-f58c-49e5-bc0a-366ad3d090b8 · inbound

What makes a harness a harness: necessary and sufficient conditions for an agent harness cites this paper.

What makes a harness a harness: necessary and sufficient conditions for an agent harness API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-27T19:31:10.865229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:15:57.372858Z digest=sha256:92e59b0943245cb8d82be695d46a133a3773a65f5b23106c01ba855686a1b461

Observation b138915d-492a-4264-ab6d-6a0b7eadd3e1 · inbound

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task cites this paper.

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:37:56.932757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:55:01.205573Z digest=sha256:faf921942ac11b32a4ed837e4c88eb988b236d425b2322e3e0d78205e9e52f66

Observation c879f485-c2c6-4d80-8a1b-9e02d4ce5885 · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T14:58:33.242010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:e3ca5f3db579923e69f243c1591e2759378f473196ec50c578aac8fa025f5c52

Observation 36a52fef-6599-4810-afd2-29e4bad1929e · inbound

MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision cites this paper.

MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:48:45.994355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:43:37.671325Z digest=sha256:34d4d8310d662e1048a8ad1921d5068d9e28668e859bee4536b7e5622e3c76d3

Observation 15941e6f-3eb8-4382-9ea3-2ff3da0c28cb · inbound

OpenRath: Session-Centered Runtime State for Agent Systems cites this paper.

OpenRath: Session-Centered Runtime State for Agent Systems API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T01:29:23.012616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:18:48.243426Z digest=sha256:2933602c1ac62f9c3698a6a888d81e91bf6345700048f865c4e3274f67d26ec4

Observation 81ef9cf6-3008-4e06-8613-639d8eca84c9 · inbound

PhoneBuddy: Training Open Models for Agentic Phone Use cites this paper.

PhoneBuddy: Training Open Models for Agentic Phone Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:59:46.469016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:15:49.428124Z digest=sha256:9cc2fe491ae82f300c89459a48b2d180ca2b684c26c89a7d1a132a49c75bcb56

Observation 479681bb-ff1f-4019-9b38-40818e53bcda · inbound

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages cites this paper.

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T10:14:36.201722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T10:12:45.090257Z digest=sha256:cae23d7c7cb028a764a0f2eadf602556c02e8e3f98f2f8655aa7ed5e7c45bcc2

Observation a8d4f677-7147-4af1-88ec-f9d958d93e3f · inbound

A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization cites this paper.

A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 175

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T12:05:43.578203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T02:32:19.425550Z digest=sha256:ecede6df6876918a2e27948417a6f8e9ba491818d18d21c13fe9e4deeb50e3d2

Observation ba1ca775-7917-43df-b323-7d70bdee1b3b · inbound

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation cites this paper.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:f1b056022a12302a20e76b711b9587fbfd0e2b0496256f995103a8b16760d2fd

Observation d16bf914-0b5c-4ffd-9d89-4c313ed8f701 · inbound

GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning cites this paper.

GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T05:59:23.238387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:59:23.238387Z digest=sha256:d5e3f0971d931513d9908dd58a9b7c97075ee2b22a59af11bde11fb1b52dfa27

Observation 491e6537-7e74-40ae-8b6f-40011538cbab · inbound

Mach-Mind-4-Flash Technical Report cites this paper.

Mach-Mind-4-Flash Technical Report API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T03:29:34.486347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:29:34.486347Z digest=sha256:dadfc0caf6a3436939615194bcab574d310420b496908d7cffa594e8273b4885

Observation 1db43143-b52d-451f-9bf6-74d1f40a6b1d · inbound

Ceci n'est pas une pipe: AI systems as semantic abstractions cites this paper.

Ceci n'est pas une pipe: AI systems as semantic abstractions API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-13T02:40:34.279112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:40:34.279112Z digest=sha256:186621ad9f2b76ba2e6370329b70991bbe60188beda0bf7d8dfb21a72c582cf8

Observation 9ea1bd16-bafd-4775-a936-3167f14c6c39 · inbound

LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans cites this paper.

LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T08:37:17.973015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:37:17.973015Z digest=sha256:1ab8b1a675bf85681d1f4c853d98697b2a9f65134cb8937263a3822e09565046

Observation 0d4518c6-0ccc-4314-a513-d460d619016c · inbound

LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization cites this paper.

LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T21:29:41.718592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:29:41.718592Z digest=sha256:7df1eb67f7c5cfd11363fc70fcb4f703259cb94a2d66be9352d848b2fd35f170

Observation f4806c61-d969-4397-857a-c0651b09f023 · inbound

ProEvent: An Event-centric Benchmark for Proactive Agents cites this paper.

ProEvent: An Event-centric Benchmark for Proactive Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T17:18:05.283498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:18:05.283498Z digest=sha256:81f5ae94ea975afebc4f4a28ed41a200a2aad0fd7f420dffec5f969c64aa53db

Observation 3e48715e-bf87-4070-8042-f448aa32867b · inbound

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains cites this paper.

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:25.830028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:25.830028Z digest=sha256:264f5b16471950da46b349bc1c97c44833124a61842a7ae5b19831e8cdfa6775

Observation abd8a518-7536-49d2-bcca-e703f9fe52ed · inbound

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions cites this paper.

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 27

Resolution
malformed identifier
no resolver link, observed 2026-08-01T15:03:42.730768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:03:42.730768Z digest=sha256:0602f4de943601bbb61cfb1bac62fc8e0def6a153300973c534dc4ec6f2d65b7

Observation e233c450-2dbf-48af-ba42-cade0ca1db5a · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:37.495334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:37.495334Z digest=sha256:3e8e429dfc285168b163592dd1bd60272d7680e04c1144a3977e0f03ca2a2055

Observation d0dda371-ab98-4916-b787-30621babd5a5 · inbound

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows cites this paper.

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T03:35:43.031171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:35:43.031171Z digest=sha256:091ef117fe62cbfadbc70df760a9306bd7bf5a203e6e107b5f2f30bffa82d139

Observation ceb3b68c-bc60-4a3a-9cb4-8172151e9d01 · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:14.153015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:14.153015Z digest=sha256:fb2fa4bc0264da301dfa568e67c03872bb2078c87c73331e1f818618ab224731

Observation f53554aa-2c8a-4243-a619-d863f7379586 · inbound

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing cites this paper.

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T01:35:09.670186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:35:09.670186Z digest=sha256:5abb476981951e3d2d7a6cfbce72ca733a96a49f33a50b1474557b06f05c58e5

Observation 18030774-6878-4337-9381-1420c505e1b7 · inbound

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents cites this paper.

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:17:55.456005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:17:55.456005Z digest=sha256:8f778a37857aa4d6fbb8aa77ecd786a019097f1bdcc94478126564bede7392aa

Observation df2f28b3-d29a-4b28-a0b7-389c682f1a60 · inbound

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models cites this paper.

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:38.571002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:38.571002Z digest=sha256:e9589525a0b8371ae6339f5ca7632a45c7a6d946d471b73c6d038220b3a4bdb3