Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:02:41.100115Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.22853.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:02:41.100115Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 67af86c5-852a-4eab-9f36-fc04dcc5593f · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb98c640-b9ef-4d10-b882-d412b61048b5 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Phi-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1263e47b-b410-4cc9-862a-b4c46cd4ebdd · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ace2d624-af41-4043-99fc-7f8c256c5689 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f59721-6407-409e-955d-616765d4661a · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee157638-d7a2-4999-a40a-7d55726c6ca0 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Genie: A Generator of Natural Language Semantic Parsers for Virtual Assistant Commands
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6512ffbd-b7dd-481e-88e9-49d88c10eb49 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa1891e-46d6-412f-9683-9fdaeac463ce · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TinyAgent: Function Calling at the Edge
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7820be4-1746-493c-bf47-d3385e69a83e · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolTalk: Evaluating Tool-Usage in a Conversational Setting
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a22051-a2c3-4fc7-9b9b-e797ebfe2b16 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cbef17-dec1-4b55-b70c-7db2756b1bc5 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140e6b49-c3e1-4f90-99d9-53875517b35e · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c732c0-1389-4afa-8e90-b22d0b70dc3b · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Intelligent Virtual Assistants with LLM-based Process Automation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20756517-6596-4bec-9837-d139ff5bdd0b · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b10d36-0521-4cdb-9bc0-99c1e690b5a6 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1890a7c7-d408-4896-9b3f-b6f43833978d · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Mistral 7B
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9920ca20-b5c2-400f-a140-a0c3904cd7f0 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abb7d8da-005f-4b43-89b6-c9ed2ea968b6 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce720a0-115d-4a8e-88d7-e34a0c75fab4 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues SEAL: Suite for Evaluating API-use of LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ac36575-d17c-4731-8683-e28c29e9355e · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4d59512-941d-4138-8afd-e64b0a39d3d4 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f8e3684-0fda-440d-884d-98e4ea610dc6 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolACE: Winning the Points of LLM Function Calling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657ffe74-1b1a-4030-b055-d8fcbe58a610 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a800961-81c6-4fc3-9866-d9d3e3ff30c3 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Manning and Hinrich Schütze
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0c6393b-284a-4b21-9170-17b734c0e26e · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPT-4 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eaa2706-e189-4aef-9f14-8d5dd678067d · outbound
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ecc55d3b-2e3a-4235-9a5c-e8ae82c90c84 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Gorilla: Large Language Model Connected with Massive APIs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966f52ad-265e-4f2f-9823-e73d8ad708d8 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Foundation Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8eefb7e-6129-43c4-9966-2846c830af1d · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd8a6f6-34b6-4430-8e74-939421a62030 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Large Language Models: A Survey
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c161d1-5d57-4b99-b867-84df2740b0c3 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Qwen2.5 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194f2059-bec4-4236-aa48-eaba2aac480e · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ad492a3-1457-44b2-90e9-cdee7a3a7618 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa50049-4b1e-41f4-801b-d5d2a67165f4 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7584eeff-89c0-452b-a71c-7ac4129b577f · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ddb17d27-535d-41e0-9462-2883dd5cd10f · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598b52fb-ac15-4bd5-a9b2-6d0320e92621 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab839f1b-689b-47f7-a359-3fff8177c7b7 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TaskBench: Benchmarking Large Language Models for Task Automation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da63ca7c-fa85-4fa8-903e-c1c44a098e9e · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c1a3027-16e2-4f09-8e63-345342b65479 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Language Models are Few-Shot Learners
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfc23798-d9ac-4d89-8eed-a0dc20e9b9e5 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c97b5fb-b823-42e4-a498-5d6d718bb907 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues LLaMA: Open and Efficient Foundation Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f80d99-8a0d-437b-bc30-9babfb099cdb · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65dc56bb-2c0b-442d-b3db-66e711f76d3d · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8e6265f-9550-45e8-9f25-2c6b399afcd5 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5909157e-e86c-44ea-9604-ca43eb156c05 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223e56c2-0253-48b2-8fd1-41e74a31a625 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97428b19-a228-4f2d-af5d-9c1cf2c353fc · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5321c14a-8d40-4db5-a641-3352980a350c · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues On the Tool Manipulation Capability of Open-source Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa42965a-feaf-4205-9eb2-6acc7788c794 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14228d02-f6a2-4b96-bd90-60acba2b90c2 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23c8b628-30fe-47e5-a5ad-8bcf7a1084c4 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Schweitzer, and Alison Wood Brooks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 381dc53e-59e5-4fd9-b493-50ebe7b64b8d · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues A Survey on Multi-Turn Interaction Capabilities of Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e5d6cd4-2439-4fcc-9d6c-a6c20b2f61bf · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolQA: A Dataset for LLM Question Answering with External Tools
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d363ed-31e3-499e-9178-35297422add4 · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues online" 'onlinestring :=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c7cba1e-c6ed-4806-8dfb-91e3dfc4747b · outbound
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues write newline
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.