Pith. sign in

Paper Citation Record · LEDGER

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues

As of 7 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.22853.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22853 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:02:41.100115Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact5
  • verified fuzzy3
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67af86c5-852a-4eab-9f36-fc04dcc5593f · outbound

This paper cites Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.650566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.650566Z digest=sha256:8ae03a4609e2ac25674412427d06fd9d74c75ef11651a3d1bd3ba932273bc699

Observation fb98c640-b9ef-4d10-b882-d412b61048b5 · outbound

This paper cites Phi-4 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Phi-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.731201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.731201Z digest=sha256:1618f72dff44ae361808cd5b1808a3804c21b4fd551940a69c9ecbaa0255b692

Observation 1263e47b-b410-4cc9-862a-b4c46cd4ebdd · outbound

This paper cites Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.849489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.849489Z digest=sha256:01b933e9a06cd4b75fea29483e2a754448adf7aa04b473bb13c66bdad8376c78

Observation ace2d624-af41-4043-99fc-7f8c256c5689 · outbound

This paper cites API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:37.949602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:37.949602Z digest=sha256:3be43100cd8c67f293159803a3153e216809d646dcacd14c15cdfd69af1b2baf

Observation 07f59721-6407-409e-955d-616765d4661a · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:45.244990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.066276Z digest=sha256:0844b4c853a9304a3ed0ff007bf4d27a57ea2e16f09257f703faa41fc24a5574

Observation ee157638-d7a2-4999-a40a-7d55726c6ca0 · outbound

This paper cites Genie: A Generator of Natural Language Semantic Parsers for Virtual Assistant Commands.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Genie: A Generator of Natural Language Semantic Parsers for Virtual Assistant Commands

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:43.088698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.215072Z digest=sha256:1877d90e4c2e600b6383bddb014c5783ad5417e22b463be8be205ac3ab767edb

Observation 6512ffbd-b7dd-481e-88e9-49d88c10eb49 · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.371492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.371492Z digest=sha256:bcb8e6e76a1c9cc5e62376374b2217ef10f7fa52cdfcda72626b71268d0b3df4

Observation eaa1891e-46d6-412f-9683-9fdaeac463ce · outbound

This paper cites TinyAgent: Function Calling at the Edge.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TinyAgent: Function Calling at the Edge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.481560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.481560Z digest=sha256:55f1e819d89dc6f43d02657033db4c5589fe43e42f717af747692b943148ed75

Observation b7820be4-1746-493c-bf47-d3385e69a83e · outbound

This paper cites ToolTalk: Evaluating Tool-Usage in a Conversational Setting.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolTalk: Evaluating Tool-Usage in a Conversational Setting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.534202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.534202Z digest=sha256:c6581ae7a64b44a2d69e1fa75db774b15394d24a89f86e999c535068df12b04c

Observation f8a22051-a2c3-4fc7-9b9b-e797ebfe2b16 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.574014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.574014Z digest=sha256:8d65c16672ccf9489c5652989462238df708cf462a4795cff276448207869247

Observation f7cbef17-dec1-4b55-b70c-7db2756b1bc5 · outbound

This paper cites Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.627939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.627939Z digest=sha256:9e649ce9395154d7599a42b8073377f58058c3a3bf8f4627ab84a9f62f0a975a

Observation 140e6b49-c3e1-4f90-99d9-53875517b35e · outbound

This paper cites CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues CoSearchAgent: A Lightweight Collaborative Search Agent with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.725718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.725718Z digest=sha256:0eeb73bd9499957914fab265987bce67ea06200909f991e11aba5956b5adc7b9

Observation 64c732c0-1389-4afa-8e90-b22d0b70dc3b · outbound

This paper cites Intelligent Virtual Assistants with LLM-based Process Automation.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Intelligent Virtual Assistants with LLM-based Process Automation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.797381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.797381Z digest=sha256:3aeb801504105209ac24943ded99eecd07b7a6d0cb6a0a8f846749cd2e91840a

Observation 20756517-6596-4bec-9837-d139ff5bdd0b · outbound

This paper cites MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.861022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.861022Z digest=sha256:5eeb6222e897e85fdd2b5b972078dfdb884766a92abfc20e112e902341e9d375

Observation 17b10d36-0521-4cdb-9bc0-99c1e690b5a6 · outbound

This paper cites An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.730281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:38.931012Z digest=sha256:ac35145b74b1bc6b6a523f3a7330042cf731a6bc2e8d2a11f3a3c4e8bd39a990

Observation 1890a7c7-d408-4896-9b3f-b6f43833978d · outbound

This paper cites Mistral 7B.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.975216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.975216Z digest=sha256:499f5862ba95235683a889df789fd7c04fd9ce4432534b23ddbe29031047f256

Observation 9920ca20-b5c2-400f-a140-a0c3904cd7f0 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:45.073905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.071046Z digest=sha256:87493e16ca864a8ee693db8877a1df8876f77bc293272216906360106a3d04a9

Observation abb7d8da-005f-4b43-89b6-c9ed2ea968b6 · outbound

This paper cites Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.143917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.143917Z digest=sha256:9200225e1635fa4da6bdb7ecc161747f8701478564b9c3a3abb3f428c1303035

Observation cce720a0-115d-4a8e-88d7-e34a0c75fab4 · outbound

This paper cites SEAL: Suite for Evaluating API-use of LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues SEAL: Suite for Evaluating API-use of LLMs

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.398827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.212803Z digest=sha256:2bce8ff2b26de642c11e204d2d0ef25cca88fba5c9f6f5ea9dac40d21e0625cc

Observation 6ac36575-d17c-4731-8683-e28c29e9355e · outbound

This paper cites ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:42.211387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.273106Z digest=sha256:d9c34637ac16e4278c6552626c45d4d3337a3d43e62497efedb25219b023f5b4

Observation a4d59512-941d-4138-8afd-e64b0a39d3d4 · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.336321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.336321Z digest=sha256:803c44e14334e36800ee82a64f1cfe8791fd5eb2da13b18bdff8352b330f75ef

Observation 2f8e3684-0fda-440d-884d-98e4ea610dc6 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolACE: Winning the Points of LLM Function Calling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.409173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.409173Z digest=sha256:6ae5bf1e10b62849d05bf34a2fba3189d14aa936c0252c6d86f4fe5e68a05496

Observation 657ffe74-1b1a-4030-b055-d8fcbe58a610 · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.463112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.463112Z digest=sha256:b9c3ccb1f0c5f78cda7ab4dd3e6318e0b61f9507934beedd5ac7d96a8a004ccc

Observation 5a800961-81c6-4fc3-9866-d9d3e3ff30c3 · outbound

This paper cites Manning and Hinrich Schütze.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Manning and Hinrich Schütze

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:44.883790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.552815Z digest=sha256:c01a90cf5057ae41632dc388d7f4cd98fcc7a79e87899dde946471589c3dd6dc

Observation f0c6393b-284a-4b21-9170-17b734c0e26e · outbound

This paper cites GPT-4 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPT-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.596859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.596859Z digest=sha256:27579a3a0acd08598c8110a0bb9bb2eb95fdaa02cb57446cdf37eba9b5434893

Observation 2eaa2706-e189-4aef-9f14-8d5dd678067d · outbound

This paper cites Bernstein.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Bernstein

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:44.719098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.688152Z digest=sha256:0a9b263532d83e3d67da804b828f30ff53157839ce3ace73c7cdd45a6dfd991a

Observation ecc55d3b-2e3a-4235-9a5c-e8ae82c90c84 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Gorilla: Large Language Model Connected with Massive APIs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.734802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.734802Z digest=sha256:41623e1ca1546704d873c19d7727c92e41d7918502ab5ff119c3afaa3ee191dd

Observation 966f52ad-265e-4f2f-9823-e73d8ad708d8 · outbound

This paper cites Tool Learning with Foundation Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.781162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.781162Z digest=sha256:fbde207f9ce92d6ce3eb7900e1b7662c00afa64a10d3c508f3581da65adb5e29

Observation d8eefb7e-6129-43c4-9966-2846c830af1d · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.832816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.832816Z digest=sha256:e58b9d113475ef0d250b98869539313cd3d60d029bab33da1b45d6d6078e6f5f

Observation ebd8a6f6-34b6-4430-8e74-939421a62030 · outbound

This paper cites Tool Learning with Large Language Models: A Survey.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Tool Learning with Large Language Models: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.878785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.878785Z digest=sha256:4a194d6690d16096a4184863d40fbf61067c506c5b7b54d7dd212c1e2b372f18

Observation 32c161d1-5d57-4b99-b867-84df2740b0c3 · outbound

This paper cites Qwen2.5 Technical Report.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:39.913760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:39.913760Z digest=sha256:1d997a2f23330e0a4dbc9edd0468f74464f8e78987ae63f73cf6fb1bdd7e8a16

Observation 194f2059-bec4-4236-aa48-eaba2aac480e · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.547233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:39.971592Z digest=sha256:396bfdd3c303ce94ebd9d2070e7a7ac01be70e16a54e3da0f5accbf05fb94aee

Observation 5ad492a3-1457-44b2-90e9-cdee7a3a7618 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.010427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.010427Z digest=sha256:06fb95068efa30592763d168d366c6b1d549b21ecfe0578021cccfc135d32aa5

Observation 2fa50049-4b1e-41f4-801b-d5d2a67165f4 · outbound

This paper cites Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.051068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.051068Z digest=sha256:bda66102ca43b745b7dc0b80b459109ba0c849ccf91c11ccb6cfd5587a78e365

Observation 7584eeff-89c0-452b-a71c-7ac4129b577f · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.405014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.104042Z digest=sha256:1358b7def481fec1c75b9ae55ba6447232277f27eb46fcf0ba3a9ed952e51086

Observation ddb17d27-535d-41e0-9462-2883dd5cd10f · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.144671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.144671Z digest=sha256:93ece801a36298de1d424e238fefa00fe399be73a3f974ecaf67c294a3befafc

Observation 598b52fb-ac15-4bd5-a9b2-6d0320e92621 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.220824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.206383Z digest=sha256:7e1be038e2923c158d6946dfbe1b568cb51a0ecdf5b657698f2d65e4b8a8b31a

Observation ab839f1b-689b-47f7-a359-3fff8177c7b7 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TaskBench: Benchmarking Large Language Models for Task Automation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.257193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.257193Z digest=sha256:2c9edf92c6245ba19ae4909699ace344cc7bf91e4a26afeb85f2843b775804ba

Observation da63ca7c-fa85-4fa8-903e-c1c44a098e9e · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.291583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.291583Z digest=sha256:ddf43c81da9ebe2ca10503abc6d922db0990eacf937d77eac2b3339b4c69f66f

Observation 4c1a3027-16e2-4f09-8e63-345342b65479 · outbound

This paper cites Language Models are Few-Shot Learners.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Language Models are Few-Shot Learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.342829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.342829Z digest=sha256:b6969f45d218b9952c7eb4850df25e16f3a078bf59759c3f83f6e70e004bf78f

Observation bfc23798-d9ac-4d89-8eed-a0dc20e9b9e5 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:44.037471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.389822Z digest=sha256:59d5a6519ac8b9481f026a3df322e926055867bd274138d421453e75f38040ec

Observation 0c97b5fb-b823-42e4-a498-5d6d718bb907 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.421187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.421187Z digest=sha256:627873e72c20002467a59368600ae0f87e3bde14d6ffacb4499b2694fedb95fb

Observation 56f80d99-8a0d-437b-bc30-9babfb099cdb · outbound

This paper cites GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.459154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.459154Z digest=sha256:1c517b8c0a5ccf3457c620dd2068e488322888b7a114465b0818ce656e3cd167

Observation 65dc56bb-2c0b-442d-b3db-66e711f76d3d · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.889198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.521310Z digest=sha256:504923a5170576b78e7640824ba79f3e77c4d377d1b996d52568055faa6352d7

Observation f8e6265f-9550-45e8-9f25-2c6b399afcd5 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.705152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.565257Z digest=sha256:58ad93a4866dbe1f904310bb325bc81b003fb347d65f8b75351ed4458542cd53

Observation 5909157e-e86c-44ea-9604-ca43eb156c05 · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.606739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.606739Z digest=sha256:0d5f7770c2f5edbae1183b78648fdd60cf88f710b6140d3f46f8016220403874

Observation 223e56c2-0253-48b2-8fd1-41e74a31a625 · outbound

This paper cites MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.651835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.651835Z digest=sha256:c8f3e9b24030f147967bf665f88cda3d2bff9cf5b72be3c18077f247fd635670

Observation 97428b19-a228-4f2d-af5d-9c1cf2c353fc · outbound

This paper cites an unresolved cited work.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:02:43.580860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.699899Z digest=sha256:2e51c5b112061ee06a02943456d364b9211a8a7ce58623af262cdb0da3083214

Observation 5321c14a-8d40-4db5-a641-3352980a350c · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues On the Tool Manipulation Capability of Open-source Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.749933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.749933Z digest=sha256:70ec1ab93de791a2b5213fc1445c4d8611d25178f1efd2e67ace97594d46c1d6

Observation aa42965a-feaf-4205-9eb2-6acc7788c794 · outbound

This paper cites ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.814433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.814433Z digest=sha256:66f8231a5c7a7761e168d48d18ea1caf7ddd5e3eb3e5ad4d4591d29319d55b84

Observation 14228d02-f6a2-4b96-bd90-60acba2b90c2 · outbound

This paper cites RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:02:41.368776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.848146Z digest=sha256:6b06a01bf3b59c408e6d1086f07909755cd923d7838a25290c33d7a37d990ce5

Observation 23c8b628-30fe-47e5-a5ad-8bcf7a1084c4 · outbound

This paper cites Schweitzer, and Alison Wood Brooks.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues Schweitzer, and Alison Wood Brooks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:02:43.458492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:02:40.893204Z digest=sha256:7d0ee0ce1a0ae5082b115a348b5a08036f4216522a0b584e85f5fb90ee3ec822

Observation 381dc53e-59e5-4fd9-b493-50ebe7b64b8d · outbound

This paper cites A Survey on Multi-Turn Interaction Capabilities of Large Language Models.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues A Survey on Multi-Turn Interaction Capabilities of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.944596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.944596Z digest=sha256:3707ed203770821761b5c3f7a590b363f92e18f9ad550f0d59c9d434a1aa9967

Observation 3e5d6cd4-2439-4fcc-9d6c-a6c20b2f61bf · outbound

This paper cites ToolQA: A Dataset for LLM Question Answering with External Tools.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolQA: A Dataset for LLM Question Answering with External Tools

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.986835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.986835Z digest=sha256:5ae65fa3b7fc4a1d397fb5ce7bc1544966c251718f70f88b823a4e0f6205e79c

Observation e7d363ed-31e3-499e-9178-35297422add4 · outbound

This paper cites online" 'onlinestring :=.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues online" 'onlinestring :=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:41.035264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:41.035264Z digest=sha256:33d3b8f46a263ec663c0f5e5f24759d76a20e4014ef27ce24119b98763ce03af

Observation 6c7cba1e-c6ed-4806-8dfb-91e3dfc4747b · outbound

This paper cites write newline.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:41.100115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:41.100115Z digest=sha256:1f50a5a69b466f0d96bdb1efd04102875af72242aac7ef51646b68f759e0a13d

Pith citing papers

No inbound Pith citation observations are available.