Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T20:51:41.198767Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 76 inbound Pith citation observations for arXiv:2304.08244.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T20:51:41.198767Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:02:39.336321Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
23 of 23 outbound references displayed
External citation measurements
12
pith, observed 2026-08-05T02:28:24.338817Z
Observation 7e8c62eb-e041-47e0-8dcf-1a388158b592 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Advances in neural information processing systems, 33:1877–1901
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30f8b389-5a5a-42b2-89b7-5f88fa2c3b2f · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb4fc248-f863-459b-bfd1-6b4f69cc96e1 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Large Language Models as Tool Makers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef3ea066-0f13-48cd-8186-9dc0bbfc0b41 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5af33ed-2587-4617-8fdd-de8a36f0ca8d · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5156ba47-0c6d-492d-a45b-2e621718da54 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Atlas: Few-shot Learning with Retrieval Augmented Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da151480-a189-4469-b097-b10dbc9d831b · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8299eb5f-16fb-4597-acc7-f89f102dd58a · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Augmented Language Models: a Survey
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 250c8fff-585b-4489-a77f-d19bf8baa484 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs WebGPT: Browser-assisted question-answering with human feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7b489bd-d3ff-47e7-924e-5d54054b8650 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ART: Automatic multi-step reasoning and tool-use for large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5fb5a6a-5f0e-43ed-806d-e70ad41e22b0 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Gorilla: Large Language Model Connected with Massive APIs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 978e2a16-ee6e-4f7b-b841-2adbd758b011 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70825f6d-7bd5-4910-acb7-0d9e37811786 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Making Language Models Better Tool Learners with Execution Feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 156bed37-3106-45fe-a0c1-bbf1decb5782 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40d25fbf-7ebd-45be-9e3e-a9900ab27456 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Preference Ranking Optimization for Human Alignment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a311f75-1698-456f-8832-8f53030f1dbd · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba776283-7bd2-4bbc-bf83-38c71dffbe25 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfbdf7ce-c539-4cbb-bd42-0af8fbe72369 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f235feca-3722-451c-9108-83b41a289ca4 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ReAct: Synergizing Reasoning and Acting in Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1cb8f7d-c6b6-4807-910f-338c384ceca0 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce24e540-d6f4-49c7-8ed1-be0308682643 · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44c9d07d-b1b4-4e14-9e1f-d4cfcb900f9b · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs ToolQA: A Dataset for LLM Question Answering with External Tools
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf9fc090-dc9b-459a-84de-acffc165fabd · outbound
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs name": "ToolSearcher
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cea748f9-ea52-4557-aae7-92268217187b · inbound
Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3254d231-6c20-4e95-8005-f8cd33db1e87 · inbound
ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afbfab5c-6e56-4b0c-a33a-bf472e40bc3c · inbound
Mind2Web: Towards a Generalist Agent for the Web API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1c98ac3-f231-4fb0-a5f2-f65336e93c80 · inbound
A Survey on Large Language Model based Autonomous Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 648ad77e-fc10-466a-8385-0b591ab0782c · inbound
GAIA: a benchmark for General AI Assistants API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93c65495-1427-4f0a-b1cb-4e78ee6135a6 · inbound
Learning to Ask: When LLM Agents Meet Unclear Instruction API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 165a7397-528a-4460-b29d-b9c72399118a · inbound
Prompt Injection Attack to Tool Selection in LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbee5958-f76d-4e3a-ae8c-9eb264c3da18 · inbound
NaviAgent: Bilevel Planning on Tool Navigation Graph for Large-Scale Orchestration API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4d59512-941d-4138-8afd-e64b0a39d3d4 · inbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e8bd2b4-b095-4eaa-96f2-ef35ab2f289f · inbound
AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e53ccb40-cafc-4294-95b5-09604d98476e · inbound
MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c88fd4-d6ba-4d6b-90b2-40d720b4d7b9 · inbound
DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc626abe-0d63-4856-9a82-f1567f756248 · inbound
GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03992d48-a956-461c-9576-8456f8801208 · inbound
FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a34b5f-864e-42dd-b212-e88a4fef4cd2 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf497c9a-95ca-4e9a-b2db-90c07482a652 · inbound
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5134ab-59be-4e29-8c6c-d1a72f2f2df9 · inbound
Evaluation and Benchmarking of LLM Agents: A Survey API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3701e0ee-3fa4-42a1-a9e1-f2222427131e · inbound
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879e3b24-c647-48e2-8516-817f99aa7375 · inbound
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 779cdcf4-e961-4759-9f23-5a1a1158c2f8 · inbound
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8325615-c5d3-4968-82b9-7a5427425ec7 · inbound
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2666c32-226d-4844-8154-8abf2f7506e3 · inbound
Asking LLMs to Verify First is Almost Free Lunch API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 847dd99e-6431-4773-ac71-84f4d28dc959 · inbound
AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4d50127-be3c-47ff-98a4-5d172e8a2203 · inbound
RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab1047b2-a35d-471f-8661-f38d8ed2d744 · inbound
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 636fa33c-ccd1-465e-82ac-8d3600073afe · inbound
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f50ff977-8629-4a1f-a475-43247594d8a6 · inbound
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10a8c0c-04cd-4db8-9426-48fa64c28d0f · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca9df927-aee1-41ec-bee8-957e6d3e6a7a · inbound
MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0061d9e7-b0f0-4fd1-81d2-ae2e51da7d0c · inbound
SAGE: A Service Agent Graph-guided Evaluation Benchmark API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a58aa53-804e-484f-9e7b-1ff3a98a0502 · inbound
A Periodic Space of Distributed Computing: Vision & Framework API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e87fac9-b753-4da4-8aab-a8a075835d1d · inbound
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cfb58f7-7bbf-4682-bdba-238f0af61c55 · inbound
GraSP: Graph-Structured Skill Compositions for LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c724b56-4960-4e53-b244-a3d4218aadd9 · inbound
PARM: Pipeline-Adapted Reward Model API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ebb8397-8721-4bde-a128-e3f0b428e1dc · inbound
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93c2ddaa-9498-4fb2-b04f-ec25359d1462 · inbound
Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 239173b7-1c36-4516-91c7-d301940f608e · inbound
Quantifying Divergence in Inter-LLM Communication Through API Retrieval and Ranking API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 759fa504-1109-4f66-8091-e3684506dbd9 · inbound
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd25fa72-8feb-46a5-b24e-4d9c62576748 · inbound
Tool Calling is Linearly Readable and Steerable in Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54975448-a0e9-48ef-93e9-4394a9ad97bb · inbound
TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b97cd65-7bea-4e58-a6b0-e3668a8e6abf · inbound
Trajectory Supervision for Continual Tool-Use Learning in LLMs API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2209d1b2-d772-49b3-8ae0-03fa309fffde · inbound
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b7dde3c-31f3-4ae2-8ac3-39aae584ba88 · inbound
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c47fd8d-faee-444e-8d3f-da9f3b66d61f · inbound
The Scaling Laws of Skills in LLM Agent Systems API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90dc69b7-7f6e-4b53-aff8-4fe5b1b3a533 · inbound
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3c71dd9-7383-498e-8b7b-17c10c0a095a · inbound
Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuning API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d230f240-1a2f-4e96-811f-2bde4f1dacd8 · inbound
An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9a0d735-9023-45df-ba13-1427bfd3b862 · inbound
Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 462c8bb5-f17f-4410-b4ef-16686b16daa5 · inbound
Testing Agentic Workflows with Structural Coverage Criteria API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d3c429f-c7ef-4adb-97e0-a6f671ca604c · inbound
Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 919b48d2-47cd-4d29-91fc-a3a499af0a07 · inbound
CLI-Anything: Towards Agent-Native Computer Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f603fe7d-7fbd-48ac-811a-c1f17e964cf5 · inbound
NTILC: Neural Tool Invocation via Learned Compression API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20807d09-fb00-4ec9-b8cc-6b686d58eee2 · inbound
Contract2Tool: Learning Preconditions and Effects for Reliable Tool-Augmented LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d9934ab-f58c-49e5-bc0a-366ad3d090b8 · inbound
What makes a harness a harness: necessary and sufficient conditions for an agent harness API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b138915d-492a-4264-ab6d-6a0b7eadd3e1 · inbound
Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c879f485-c2c6-4d80-8a1b-9e02d4ce5885 · inbound
SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36a52fef-6599-4810-afd2-29e4bad1929e · inbound
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15941e6f-3eb8-4382-9ea3-2ff3da0c28cb · inbound
OpenRath: Session-Centered Runtime State for Agent Systems API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81ef9cf6-3008-4e06-8613-639d8eca84c9 · inbound
PhoneBuddy: Training Open Models for Agentic Phone Use API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 479681bb-ff1f-4019-9b38-40818e53bcda · inbound
SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a8d4f677-7147-4af1-88ec-f9d958d93e3f · inbound
A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba1ca775-7917-43df-b323-7d70bdee1b3b · inbound
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d16bf914-0b5c-4ffd-9d89-4c313ed8f701 · inbound
GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491e6537-7e74-40ae-8b6f-40011538cbab · inbound
Mach-Mind-4-Flash Technical Report API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db43143-b52d-451f-9bf6-74d1f40a6b1d · inbound
Ceci n'est pas une pipe: AI systems as semantic abstractions API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea1bd16-bafd-4775-a936-3167f14c6c39 · inbound
LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4518c6-0ccc-4314-a513-d460d619016c · inbound
LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4806c61-d969-4397-857a-c0651b09f023 · inbound
ProEvent: An Event-centric Benchmark for Proactive Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e48715e-bf87-4070-8042-f448aa32867b · inbound
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd8a518-7536-49d2-bcca-e703f9fe52ed · inbound
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e233c450-2dbf-48af-ba42-cade0ca1db5a · inbound
CRAFT: Learn the Schema, Execute the Plan API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0dda371-ab98-4916-b787-30621babd5a5 · inbound
SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb3b68c-bc60-4a3a-9cb4-8172151e9d01 · inbound
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f53554aa-2c8a-4243-a619-d863f7379586 · inbound
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18030774-6878-4337-9381-1420c505e1b7 · inbound
Execution-First Synthetic Tool-Use Trace Generation for LLM Agents API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df2f28b3-d29a-4b28-a0b7-389c682f1a60 · inbound
Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.