Pith. sign in

Paper Citation Record · LEDGER

ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2408.04682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04682 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:42:31.292282Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ad1fd50-ca3d-45cd-9304-6f5c1a5f1f4f · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:35:51.215921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:f8b0a5a57c589a6246885631e0695475b63090b482fa7f3079da7ce39a44e0f2

Observation a301ce5c-2bd5-4879-8d8f-971613a711b6 · inbound

Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents cites this paper.

Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-12T20:36:02.129282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:36:02.129282Z digest=sha256:e87d15d06241113d9a7d18b4d5bff721e7e2357bde2d0e07717c161a6664af85

Observation 7b358280-6ffe-4e51-98d5-6d9caa0fc493 · inbound

CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning cites this paper.

CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:06.916173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:06.916173Z digest=sha256:d8909d00b0a11ef72a607e407d30b9faa08c97daa9d6fa6b1f6bd23020088fab

Observation ec77af0a-d3e9-43b4-a0e5-ba071b0f6dfe · inbound

Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications cites this paper.

Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:47:14.563824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:47:14.563824Z digest=sha256:6a751c889a6630bcc03225f4f88548558a165a0072916f9660f1e2dbc000b663

Observation 9602fd78-0b64-4afa-9e40-bb39925730e3 · inbound

Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions cites this paper.

Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:02:40.376710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T09:02:40.294491Z digest=sha256:10a0a54423929090dd3a6965a9bafe3601dddb60056adf62405b993e75122806

Observation dedd5286-1cc7-4070-b1ca-b34d47eb14cd · inbound

MARFT: Multi-Agent Reinforcement Fine-Tuning cites this paper.

MARFT: Multi-Agent Reinforcement Fine-Tuning ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.292282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:42:31.292282Z digest=sha256:616b444da316044b84b960783efe793d62aaedb0dab8bd16ac19932f42436460

Observation 2b9a74e7-ea75-4326-ad4f-18ce6946824e · inbound

When2Call: When (not) to Call Tools cites this paper.

When2Call: When (not) to Call Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:48.025171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:48.025171Z digest=sha256:3fa1bd683b174f04bfbe7bc13612b03f235eabc27dc859e94e4d43d19d5ff5cb

Observation 62e9db9a-bf21-434f-b710-3500d82941f0 · inbound

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction cites this paper.

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:35.085270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:05:35.085270Z digest=sha256:b84526e29719d6c4b180502a987c777a2c5e917e2b5c2bbc74eae7c71ea1a89e

Observation 1873e3ea-ba02-4f5c-a9c5-b8504ac0da62 · inbound

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges cites this paper.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.209839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.209839Z digest=sha256:a15af24dc8b67d28c4fd6534ccfafaf3da4ce371df8887f195e029a0edb7743f

Observation dbd88b35-1cd6-4f4d-a67a-46f92a7aa064 · inbound

Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services cites this paper.

Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:34.819687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:34.819687Z digest=sha256:6f3e68cdb38e5bb8c629d58317eb1331c0518be0bd2766213630c3042b0df3db

Observation 55ff60d1-758d-4466-8142-ea7b84b5ab8c · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:01.284289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:01.284289Z digest=sha256:10d85b0fcd27dd7aaa5dd390f4d07b3b843af00ef5f423246ee2a28913ae7cc8

Observation cd43e93d-4be5-4004-9d58-f5834c6c6f73 · inbound

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions cites this paper.

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:38.571897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:36:38.571897Z digest=sha256:d5cd4f75c23aed8c7dbf3aca644d9af731f93e73036fff15d39f0007846b0224

Observation f533e183-3092-407c-839d-ffecd6c4733d · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:52:17.432112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:4c8df34d1d2b5c19349cd69e17e33a2dc7e477695345f51b10bd1caa59592e1f

Observation 546c32d9-a530-4dc0-9c0d-1bde942dbe87 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.628526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.628526Z digest=sha256:a79b63977825746a1eee8c2f733e8768166d867b530cca26dad64eed20af118d

Observation 90fb7440-f45f-4f79-b44e-d8638061cfa0 · inbound

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities cites this paper.

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:54.497019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:54.497019Z digest=sha256:89d46f4bbbf4023e571c600ac3656a50cdd00b9504ee24f86746e6e16be1c751

Observation 4e5ac1d4-53b1-4a8f-881b-74a20ced8816 · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:47.975577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:47.975577Z digest=sha256:8034185edc847b1f1c5bfcd163dbd9f2cc776dc1c3e30dad54e1bd36d98d8ce1

Observation 9c5b7e91-0ef3-44f7-ba03-612cb3886353 · inbound

Apple Intelligence Foundation Language Models: Tech Report 2025 cites this paper.

Apple Intelligence Foundation Language Models: Tech Report 2025 ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:58.622125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:58.622125Z digest=sha256:43537e4b249c2aa1332e624ca54d4f5063790f3b063ebe76aa6828b972d6c2b8

Observation 2065dfca-4776-4ec4-b6fd-34ad20d7434b · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:57.418255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:57.418255Z digest=sha256:4726566ec63900dd46115792f1a5450320aec67e277ceb40c8ff420c722e9c41

Observation 17b85ac1-5fe6-4332-a284-bd791f769b3d · inbound

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution cites this paper.

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 4976

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:30.254612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:30.254612Z digest=sha256:806093fc4c012dcc7fd33d4369f9b053b1d536259a87da249a211d68505b8ce2

Observation 8985ac9c-af6b-44df-94c6-5f026bf80eae · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.486170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.486170Z digest=sha256:0cdb068127755e9335902626067d51e624f09f433c1b3137333dc46d9b2fb4e1

Observation 384f4a7a-5d99-4c71-ad96-646df71eaa06 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.733725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:3a8ce07a134f5548b1636fe61c712f3711ffe6d4cdc6fcc5eef0ba21bb65626d

Observation 32adb8cd-f031-49ee-82e3-dc87d140ae16 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.408309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.408309Z digest=sha256:5342bfb4b408ebb4bd671f2514385f5af84a1d0d8b04158ace5d544fed3a1ee9

Observation c8254a3f-bd8b-4506-bc88-76b597e87b29 · inbound

Hell or High Water: Evaluating Agentic Recovery from External Failures cites this paper.

Hell or High Water: Evaluating Agentic Recovery from External Failures ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:35:04.652799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:35:04.652799Z digest=sha256:2f188a62f19ee964520f394cc840f59617fffb28ad57366f81c1791285861988

Observation 96b63d17-0dd0-4f62-96de-03c8dc107cf8 · inbound

PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback cites this paper.

PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:28.102393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:57:28.102393Z digest=sha256:5fb0e25cf01cb552eab13b7003ff32b659ab3ae056f2ba6a07b0bd7b0b33eafa

Observation dcde5d88-8512-4bb9-bea9-bbb2d01f6287 · inbound

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench cites this paper.

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:00.685400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:00.685400Z digest=sha256:0fa0ff5829f61d45731435d92c5d2c43cf1fe95b4bd14f46e282262621e56a54

Observation 2acc08f8-df26-4463-b47c-75c18ec09782 · inbound

COMPASS: Benchmarking Constrained Optimization in LLM Agents cites this paper.

COMPASS: Benchmarking Constrained Optimization in LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:46:07.880209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T08:45:11.334594Z digest=sha256:6aba6cf7cbb4d190e8efe4a3e98a34526634ea38fd66d28b43fc926efa7fba5b

Observation cdcb1055-214e-49e5-9335-e7073b300fa1 · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:37.341452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:37.341452Z digest=sha256:530c4e9132ab6c01176feaa3737b36894cfd98e0736788da22ab5ed81cb58667

Observation cbc00d68-b687-49dd-93f7-3e8128ef1c4f · inbound

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents cites this paper.

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:31.274272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:31.274272Z digest=sha256:117aaefca4ba2c073d57c4660773002dc23eff1d72554bd0b818c1447720c001

Observation 7671de91-1088-487d-ae66-8bcdfc4e3373 · inbound

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls cites this paper.

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:46:33.312636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:46:33.312636Z digest=sha256:c2229bb7e6d3bcceab54fc78c0b6cc5ab885535f400c1c129b3a8c8cb000506e

Observation 1e99cad8-ad5f-4498-af27-6b33117f8164 · inbound

Mind the Sim2Real Gap in User Simulation for Agentic Tasks cites this paper.

Mind the Sim2Real Gap in User Simulation for Agentic Tasks ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:29.394088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:51:29.394088Z digest=sha256:ad368aad0ae67abed1868482f18fc22022fa14985e43728672bad7bba9ebf6b4

Observation 303b2907-089e-4a5c-b292-3012610353e9 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.276299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:bf8c1ef03a0d3bbbd9a92a91a5f63065f808d790dd26c47c7ae4e0e0341254e5

Observation f512d45f-3c96-4cc1-8cb9-a12791cb8bc4 · inbound

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills cites this paper.

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T09:29:51.204678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:29:51.204678Z digest=sha256:a42086568a7ae1126a15768314177968e142147b8d63eb44ca9f7a376cdcc07b

Observation b64fd55d-3560-4fae-b973-1d1382b4c779 · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:02.948205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:23:26.128263Z digest=sha256:cadb8c699e4778f9af8100a02d7685c4206c7547d38098252f1eee506ff5e016

Observation cb9d7391-7d34-42db-a282-a7f59a02b766 · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T21:35:10.173253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:35:10.173253Z digest=sha256:6d9ce5047ce4ff00ff835258a70e3bc56097b6468125e2950b03b9337f5a924e

Observation a9d309d2-bd28-4d6d-b504-e79a0a1ae000 · inbound

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation cites this paper.

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:25.973584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T14:01:23.894966Z digest=sha256:1ff1923352bbedc013025a1dfab1391742f0defebcbddbffe08e2330b699ac27

Observation 78643d9a-d7a8-4f1b-9013-46192bceb7f3 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.975780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:d9532eb49b093c4e8edb7bacc6d69a3fc57777f3fbe3901bca873ad3676a5190

Observation bf0365b3-03c4-4603-84ae-bfa4a8c932e4 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.057004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:81c035e299c8498e46114db465250c13908c1257d8bb96818345f788df525755

Observation 77ad442d-6a89-4507-8670-4154703a5604 · inbound

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions cites this paper.

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:13:44.721016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T17:10:02.936605Z digest=sha256:28d38187a60a9fb78c05f99b8a5d7dc481fc0b7365eabf0c355649782d047c33

Observation 5df495e5-804e-45c2-8469-186133156986 · inbound

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents cites this paper.

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:46:23.684356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T14:03:02.703499Z digest=sha256:a9eff3741422876270bfa2560588438e40957a3c30550319d2d4add57210048f

Observation 5fa44723-d6a5-4bca-a7d3-fb6e66293b35 · inbound

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose cites this paper.

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:59.725697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T00:37:50.570757Z digest=sha256:bbe71f47e78e342a28b173382dfcd8a91f100a2c93b7912156cb8741e2f635f6

Observation 91ffc82a-47cf-4552-9eec-f94d765e5ebb · inbound

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows cites this paper.

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:55.533094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T20:18:07.134598Z digest=sha256:1e3b18e66fea32cb2afeb46081b10074fc3e55c8552fd9488560f6f6f9ec4769

Observation 29be73f0-dce0-478d-bf47-45b0d02832d4 · inbound

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use cites this paper.

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:36:53.348031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:36:53.348031Z digest=sha256:6c7cd9d2fd193aef1a6e51142da09d26de4bc15dc03b0a9ef1c1992deab5e32e

Observation 63af01f3-1fd4-4326-a9cd-47ddd31efa8f · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:06.175548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:06.175548Z digest=sha256:27aaba18d58e7a797428d0cfa50dde24c01e012295a58fc925878453b679113b

Observation 5eea9802-d3ed-463f-8f49-be91be9f9510 · inbound

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools cites this paper.

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:39.396840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:39.396840Z digest=sha256:d435ceb6ef4bed7a636ebb7180f38a37a7847c3190611780057235bdd2e31d4d