Pith. sign in

Paper Citation Record · LEDGER

ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2408.04682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04682 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:34.819687Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ad1fd50-ca3d-45cd-9304-6f5c1a5f1f4f · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:35:51.215921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:4391c14794928897be31345e2760db8125bdd369361a724a7a5918f795be1737

Observation 9602fd78-0b64-4afa-9e40-bb39925730e3 · inbound

Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions cites this paper.

Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:02:40.376710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T09:02:40.294491Z digest=sha256:a312322491f337277d317b2d23c09a09916283c6f39847d1f2b4ee05190f8145

Observation dbd88b35-1cd6-4f4d-a67a-46f92a7aa064 · inbound

Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services cites this paper.

Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:34.819687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:34.819687Z digest=sha256:290559baf387c6cefbeaea182ad1d171a7dbf77c6ac7e59bfa05fb510f288636

Observation 55ff60d1-758d-4466-8142-ea7b84b5ab8c · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:01.284289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:01.284289Z digest=sha256:d442ae11a4d5b5ee61691ea79b2b1d94b427568cdbe66257fc74c9d8ede7166a

Observation cd43e93d-4be5-4004-9d58-f5834c6c6f73 · inbound

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions cites this paper.

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:38.571897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:36:38.571897Z digest=sha256:c9cfe2dd2a813dafd4eaa561a82af2b8ab60c2649016248de1d99f60a6948a85

Observation f533e183-3092-407c-839d-ffecd6c4733d · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:52:17.432112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:04696143cf41625be804c062a9fd3bd6ea40dee4c80f7f0d36bde7a05b82acc9

Observation 546c32d9-a530-4dc0-9c0d-1bde942dbe87 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.628526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.628526Z digest=sha256:58ac80406b309a61eb493ab2ecebccdde82e29bb44cd0c4bec31dec7f521ed48

Observation 90fb7440-f45f-4f79-b44e-d8638061cfa0 · inbound

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities cites this paper.

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:54.497019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:54.497019Z digest=sha256:0bdc1faf78a8b5476b50f3c68bd9f982ff5591ba36afcc62e15a3f08944cfc12

Observation 4e5ac1d4-53b1-4a8f-881b-74a20ced8816 · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:47.975577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:47.975577Z digest=sha256:8d997361dfb9c7d9904bf11064a53d12b8f35ad5c964b826ffd26159105618d5

Observation 9c5b7e91-0ef3-44f7-ba03-612cb3886353 · inbound

Apple Intelligence Foundation Language Models: Tech Report 2025 cites this paper.

Apple Intelligence Foundation Language Models: Tech Report 2025 ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:58.622125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:58.622125Z digest=sha256:919e94d1015eae83733764cff639c341cf26d396d5ed49aa489123dca7f720f3

Observation 2065dfca-4776-4ec4-b6fd-34ad20d7434b · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:57.418255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:57.418255Z digest=sha256:13c2fd0da0e687cccf93f3d8659b4b10e5f96cb7c2a1da14a6ba322515e9e92e

Observation 17b85ac1-5fe6-4332-a284-bd791f769b3d · inbound

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution cites this paper.

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 4976

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:30.254612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:30.254612Z digest=sha256:e36bf10089198a5ca9c01f4e51d179708ab19a9200b3417727d39bf923278a19

Observation 8985ac9c-af6b-44df-94c6-5f026bf80eae · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.486170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.486170Z digest=sha256:98c6b13838621bea9b9e9e8799edeac48298616fb132c347569c3ae53226d673

Observation 384f4a7a-5d99-4c71-ad96-646df71eaa06 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.733725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:f5f12f43a5b7384e6f43e649f365c5c8b17a32496510eed3311b60880839a932

Observation 32adb8cd-f031-49ee-82e3-dc87d140ae16 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.408309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.408309Z digest=sha256:71a9500445dea639238bd22aae47bd38202caf87f757e922ff9330dc8b0f1f4b

Observation 96b63d17-0dd0-4f62-96de-03c8dc107cf8 · inbound

PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback cites this paper.

PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:28.102393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:57:28.102393Z digest=sha256:2a9dd355383eefbb0b8c48a6ab6a77ea3c30bb5ed976ef7ee342d92365cd271b

Observation dcde5d88-8512-4bb9-bea9-bbb2d01f6287 · inbound

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench cites this paper.

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:00.685400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:00.685400Z digest=sha256:15917f16d5a0c7ce373221c6dd0110f27e5625a1e111c62a2e7a47e302667feb

Observation 2acc08f8-df26-4463-b47c-75c18ec09782 · inbound

COMPASS: Benchmarking Constrained Optimization in LLM Agents cites this paper.

COMPASS: Benchmarking Constrained Optimization in LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:46:07.880209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T08:45:11.334594Z digest=sha256:549259ac16d3015285a2822338b18bd018a79530c2757bb98389b67440c5c779

Observation cdcb1055-214e-49e5-9335-e7073b300fa1 · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:37.341452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:37.341452Z digest=sha256:08e37c98fccf08273cc8ecb47d48f7badd2e2b8c9160d3b4eb961808476d70a8

Observation cbc00d68-b687-49dd-93f7-3e8128ef1c4f · inbound

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents cites this paper.

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:31.274272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:31.274272Z digest=sha256:3bb94c3581e469ab62d565531a5687c46686685642199e9be96a93dc564d45bd

Observation 7671de91-1088-487d-ae66-8bcdfc4e3373 · inbound

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls cites this paper.

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:46:33.312636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:46:33.312636Z digest=sha256:ce2b838c479ca21f7fd01f86951b5ea7c710337d2581f0e002f7faeb13f0c606

Observation 1e99cad8-ad5f-4498-af27-6b33117f8164 · inbound

Mind the Sim2Real Gap in User Simulation for Agentic Tasks cites this paper.

Mind the Sim2Real Gap in User Simulation for Agentic Tasks ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:29.394088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:51:29.394088Z digest=sha256:fcd54003be722f66f24bedc56cc1c3fdc584a1243726c3ef532a0adf4b1fb641

Observation 303b2907-089e-4a5c-b292-3012610353e9 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.276299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:7c74a008baf46167fe41043cde996846c093c5a9929eb3e9126f992c9ff22994

Observation f512d45f-3c96-4cc1-8cb9-a12791cb8bc4 · inbound

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills cites this paper.

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T09:29:51.204678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:29:51.204678Z digest=sha256:397b4653e6589c0af952214a16c3716c95f9f712aa305266ac38d98f921df057

Observation b64fd55d-3560-4fae-b973-1d1382b4c779 · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:02.948205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:23:26.128263Z digest=sha256:fe065a31b27a52e41afab20647ea6288576a48287dacabef8853a90086129da9

Observation cb9d7391-7d34-42db-a282-a7f59a02b766 · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T21:35:10.173253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:35:10.173253Z digest=sha256:5ffbb7cd7b330451a5d5664ed3932ba997afe20aa31b6ae993e1db41103f31b3

Observation a9d309d2-bd28-4d6d-b504-e79a0a1ae000 · inbound

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation cites this paper.

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:25.973584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T14:01:23.894966Z digest=sha256:7c9ea8cb16c4c7e59c49e9a7ba5ab84d3c47cc4ea6acf2ab94cea58aeda65c1d

Observation 78643d9a-d7a8-4f1b-9013-46192bceb7f3 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.975780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:3583d8208d123a14f8e50a8e9e774f882353ca4aa6bdec6c47406a7e66788c0c

Observation bf0365b3-03c4-4603-84ae-bfa4a8c932e4 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.057004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:9b7e5eea309cb4232593834bf30ff08c4989fcee6eff682c22ce3c3b418079df

Observation 77ad442d-6a89-4507-8670-4154703a5604 · inbound

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions cites this paper.

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:13:44.721016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T17:10:02.936605Z digest=sha256:6e61abca87052acaca3fa24fa394020935206806fabbfcc524b13e505580439c

Observation 5df495e5-804e-45c2-8469-186133156986 · inbound

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents cites this paper.

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:46:23.684356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:03:02.703499Z digest=sha256:d4b0497b687abc8717ac27af89ce35f90c9a4d45a7632f0befce7d992175e845

Observation 5fa44723-d6a5-4bca-a7d3-fb6e66293b35 · inbound

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose cites this paper.

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:59.725697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T00:37:50.570757Z digest=sha256:ed38df3c630a45b6aaa8564a9b587298ee72d04d03b0676f1bf489ea6660ed46

Observation 91ffc82a-47cf-4552-9eec-f94d765e5ebb · inbound

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows cites this paper.

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:55.533094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:18:07.134598Z digest=sha256:819f8c10fc85c52a8a31e94ac8b426ea1e98136ee450cac4d3f0bb39acf7969a

Observation 29be73f0-dce0-478d-bf47-45b0d02832d4 · inbound

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use cites this paper.

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:36:53.348031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:36:53.348031Z digest=sha256:594ba7dc5288ba367925efa74db6f5be7bbc3397f501fc6556b4384c5a97c66d

Observation 63af01f3-1fd4-4326-a9cd-47ddd31efa8f · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:06.175548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:06.175548Z digest=sha256:a03267ef62e5f5a535ce238113e7c2866fb4c4be0eec745f448af67c567f5384

Observation 5eea9802-d3ed-463f-8f49-be91be9f9510 · inbound

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools cites this paper.

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:39.396840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:39.396840Z digest=sha256:ef08dba799839b70a3ac1ab8a7031e6dc1e0de6393a3f06a4441155d9194b601