Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2408.04682.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:34.819687Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 0ad1fd50-ca3d-45cd-9304-6f5c1a5f1f4f · inbound
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9602fd78-0b64-4afa-9e40-bb39925730e3 · inbound
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbd88b35-1cd6-4f4d-a67a-46f92a7aa064 · inbound
Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ff60d1-758d-4466-8142-ea7b84b5ab8c · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd43e93d-4be5-4004-9d58-f5834c6c6f73 · inbound
CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f533e183-3092-407c-839d-ffecd6c4733d · inbound
$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 546c32d9-a530-4dc0-9c0d-1bde942dbe87 · inbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 129
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fb7440-f45f-4f79-b44e-d8638061cfa0 · inbound
Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 216
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e5ac1d4-53b1-4a8f-881b-74a20ced8816 · inbound
Teaching a Language Model to Speak the Language of Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5b7e91-0ef3-44f7-ba03-612cb3886353 · inbound
Apple Intelligence Foundation Language Models: Tech Report 2025 ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2065dfca-4776-4ec4-b6fd-34ad20d7434b · inbound
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b85ac1-5fe6-4332-a284-bd791f769b3d · inbound
ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 4976
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8985ac9c-af6b-44df-94c6-5f026bf80eae · inbound
Agent Identity Evals: Measuring Agentic Identity ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 384f4a7a-5d99-4c71-ad96-646df71eaa06 · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32adb8cd-f031-49ee-82e3-dc87d140ae16 · inbound
UserBench: An Interactive Gym Environment for User-Centric Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b63d17-0dd0-4f62-96de-03c8dc107cf8 · inbound
PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcde5d88-8512-4bb9-bea9-bbb2d01f6287 · inbound
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acc08f8-df26-4463-b47c-75c18ec09782 · inbound
COMPASS: Benchmarking Constrained Optimization in LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdcb1055-214e-49e5-9335-e7073b300fa1 · inbound
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc00d68-b687-49dd-93f7-3e8128ef1c4f · inbound
One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7671de91-1088-487d-ae66-8bcdfc4e3373 · inbound
Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e99cad8-ad5f-4498-af27-6b33117f8164 · inbound
Mind the Sim2Real Gap in User Simulation for Agentic Tasks ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 303b2907-089e-4a5c-b292-3012610353e9 · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f512d45f-3c96-4cc1-8cb9-a12791cb8bc4 · inbound
Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64fd55d-3560-4fae-b973-1d1382b4c779 · inbound
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb9d7391-7d34-42db-a282-a7f59a02b766 · inbound
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d309d2-bd28-4d6d-b504-e79a0a1ae000 · inbound
Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78643d9a-d7a8-4f1b-9013-46192bceb7f3 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf0365b3-03c4-4603-84ae-bfa4a8c932e4 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77ad442d-6a89-4507-8670-4154703a5604 · inbound
VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5df495e5-804e-45c2-8469-186133156986 · inbound
Designing for Doubt: The Case for Informed Abstention in Autonomous Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fa44723-d6a5-4bca-a7d3-fb6e66293b35 · inbound
Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91ffc82a-47cf-4552-9eec-f94d765e5ebb · inbound
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29be73f0-dce0-478d-bf47-45b0d02832d4 · inbound
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63af01f3-1fd4-4326-a9cd-47ddd31efa8f · inbound
Quo Vadis, World Modeling? ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eea9802-d3ed-463f-8f49-be91be9f9510 · inbound
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.