Pith. sign in

Paper Citation Record · LEDGER

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges

As of 23 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2505.13328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13328 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:22.416314Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:56:47.089599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T17:56:51.738856Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfd89ea3-3418-41db-906e-68600786822b · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:21.964283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:21.964283Z digest=sha256:46bc43214ae7aed475ad91cc8ae601a3af8caf430084fca37a7ae41aa5f5895f

Observation 976a5420-ddfe-4684-88cc-c55c45e6e571 · outbound

This paper cites Qwen Technical Report.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:21.970313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:21.970313Z digest=sha256:7747bdcce83b60e6df69e361d885505774cd09e3696f19877219d6fc22654e80

Observation f001b702-622b-4863-bf3a-035bf3b52f45 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:23.467081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:21.975760Z digest=sha256:12af72d6ce25138aed25e3e6a75b0f116847c5ef478ef536451b4a9ec96d1f22

Observation 0434f184-a0fe-44db-a97c-f23f13e81ed4 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:21.980696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:21.980696Z digest=sha256:0aaa99cbd0e52d73739b1c64f56c8ff3ad7279f23e09cdfd2a548e3c91549401

Observation 97cf81e5-b89d-4e27-ac98-cae7a2a73a76 · outbound

This paper cites Wizard of Wikipedia: Knowledge-Powered Conversational agents.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Wizard of Wikipedia: Knowledge-Powered Conversational agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:21.985873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:21.985873Z digest=sha256:e3f3468324708d19c92eafa643167db6768a3141c7c07450878e9683923465b8

Observation 9c122317-852f-41a7-af9d-acd353c862b5 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:21.991269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:21.991269Z digest=sha256:896a82eb17a883fad1af846a0eb9d368b328b5a8282b62ce71227d1f2845e3e7

Observation e8a983d1-90a0-4947-9e68-45745d6dbc24 · outbound

This paper cites LLM as OS, Agents as Apps: Envisioning AIOS, Agents and the AIOS-Agent Ecosystem.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges LLM as OS, Agents as Apps: Envisioning AIOS, Agents and the AIOS-Agent Ecosystem

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:21.996407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:21.996407Z digest=sha256:028c3c729e75f7188656f604d436a007694bf9ab9743086aa9326a54625825a1

Observation 233a01b5-6ff7-40ec-b1ba-b3cf6dad6982 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.001578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.001578Z digest=sha256:92ea1e67522a90fbad4962411906c4ae52900e52660d588bef112066e68342f5

Observation 8152d58d-6f9f-46a9-addc-ef5a90babe9e · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:23.419766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.006135Z digest=sha256:9c621fb8e3103934a86d7fa9871265460e553355bcf2a50462a817f8c8f23f11

Observation 93261e84-8f4f-4cb4-bee5-3f8229532169 · outbound

This paper cites Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.010719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.010719Z digest=sha256:e4627a15ddd8abe85a46460388c2a613cf92905d740e523b6ee9c77430f13a63

Observation 45bb2c7e-e636-4759-b1d1-b7414dccb115 · outbound

This paper cites Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.016250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.016250Z digest=sha256:54be3c31ccb2155fb92662d13889756cdd88a9e9f88b5472a96a4adad1873918

Observation ebd0970c-d357-4967-a142-175dc3457a0c · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.020973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.020973Z digest=sha256:5f3a30958125e5b1b6c69a861d1a014efef780d28f4b56296265e05fff911d8a

Observation 5d27af6e-1617-4e5d-b691-ed7dd0ab89b6 · outbound

This paper cites Mistral 7B.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.025948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.025948Z digest=sha256:538580fa35364908934cae30bd31a6ffc83c114d8ddd697f13914661e3c6ecc7

Observation a89fcbfc-176f-4ec6-81de-d3c14286f919 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.030598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.030598Z digest=sha256:60a3add78e12f2cf8030bb9eeb0ffe017d29c1fd17c8af5b4a123044919a7dfe

Observation c69dc4e2-cbb7-42e9-8617-da2224961602 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:23.402571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.035473Z digest=sha256:66a0832271596ed3294a08ea6c232877a0cd9bff2b2cda131f81e9726924d998

Observation 1c4b58d0-74b6-49d1-aa6f-29a04d5a8ecd · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.040140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.040140Z digest=sha256:170780eb52bfe2799f6b62d3a76daffdc2f22ed9e45cf9f10c629bed16760148

Observation 3363668e-9a48-4983-87b3-8f058e75f83e · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.048321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.048321Z digest=sha256:6351246ded0062738a1744f1db1940167f88d68ac5550e45710840b45042a1ee

Observation 8c291d8c-cbf8-440a-bd2c-07e2fa8c7a0f · outbound

This paper cites Code as Policies: Language Model Programs for Embodied Control.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Code as Policies: Language Model Programs for Embodied Control

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.094572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.094572Z digest=sha256:ef239e019d9c9e2b2428434ea8d7c9f5faeb4330cdfd6d54e4caf62228ab3606

Observation 470239f6-ba28-40e9-bb14-c5994ac86121 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges AgentBench: Evaluating LLMs as Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.143383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.143383Z digest=sha256:5bcf454e8dcb60e3488f77ee8e47185cce49e293af471f49a7df3ba7f89e53d8

Observation 1873e3ea-ba02-4f5c-a9c5-b8504ac0da62 · outbound

This paper cites ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.209839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.209839Z digest=sha256:a4335348ce5792dadb4ad800fa823de9a3e83eaaddae2d9d386d45ef4aaeab18

Observation c8329736-4908-4354-9036-50ef0c05b8ce · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.234517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.234517Z digest=sha256:e16d8ea3b7337ed3675bde22b44a9be8258d9cc7f51186f255e88d52f8b24e44

Observation e12d8fcb-21d4-4a51-9490-6da9ccbc8024 · outbound

This paper cites AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.239300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.239300Z digest=sha256:c41a41bbb5a7513b36cb6edd10178fd18d7f2b72d10ecafbbb939ddfaca8795a

Observation 89341ce8-e9a6-4a0e-8d97-5bea7134ae7b · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges GAIA: a benchmark for General AI Assistants

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.243604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.243604Z digest=sha256:ff7842589306ddf42a7357b1d8b37757a732d848278675383edbe444e8859c55

Observation b2ca764e-9509-47ba-ad41-7913336dff25 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges WebGPT: Browser-assisted question-answering with human feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.248902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.248902Z digest=sha256:e7ed1125449597607a7fe682e5a1b4aa5bf5894521c6b5ba9d154ba2d948afaa

Observation 0cff8c72-4ca8-40fe-988c-902fb042ac3d · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Gorilla: Large Language Model Connected with Massive APIs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.254047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.254047Z digest=sha256:57c451511d05e343312a9f881c952c6b1b7faa2611d5b3bee739e2c404679dd4

Observation e576ddf2-50af-4364-ac8a-b7622f3a0766 · outbound

This paper cites Few-shot Natural Language Generation for Task-Oriented Dialog.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Few-shot Natural Language Generation for Task-Oriented Dialog

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.259435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.259435Z digest=sha256:faa0ce57166271dbc3d91b06b1ca004c3bea7a0300428969b3dca1d9663598e1

Observation 7217ea47-1adf-4e73-987c-fb16d574af10 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:23.387003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.264282Z digest=sha256:20a57555d6c44d2ef78135044c05f3c8c7453e3fe1ef777c983e90a970d6d3a0

Observation d0129e7a-bc08-47d9-99b9-e7527832d429 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.268926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.268926Z digest=sha256:f048c93e3e4aa0d31a6188f6ac8dec51913fd6b7470f08abfd79844f746c9c73

Observation 028a17bb-279e-49fd-90f0-4d6585a659c1 · outbound

This paper cites Tool Learning with Foundation Models.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Tool Learning with Foundation Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.274009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.274009Z digest=sha256:94d0648ed54fbd304028044ba1c6f75e12757f597ecc1b1ca5e67480b183f17e

Observation a951ee21-d661-4003-855b-0dcc2f389231 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.279250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.279250Z digest=sha256:81339143fda0d9d9933010a6801b87b3ae8f6389b58f6c17aadfec87fe6f5181

Observation 390fa33f-00f7-476e-b2d5-a0d400046213 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:23.371138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.284569Z digest=sha256:1721b37f214d474ea4bc2436684de7d2e1ef0604a2ca7cfe75d2859043920822

Observation 71c054eb-7db8-4503-af45-66ddefdeec29 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.288745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.288745Z digest=sha256:4bab7e79cce1e888976098a34d23a646855692e941c2f33d9a15a15833b2b83f

Observation a2e76e55-c5e4-47e2-b422-8af1050f183b · outbound

This paper cites Dialog2API: Task-Oriented Dialogue with API Description and Example Programs.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Dialog2API: Task-Oriented Dialogue with API Description and Example Programs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.293409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.293409Z digest=sha256:97224c354a007525ccd66b8c0e1ecf4d2fa6861ffa62f6e3975b3fc22ff5970b

Observation eb7d1599-36f0-43bb-aea7-e5fe910c9b2a · outbound

This paper cites Cognitive Architectures for Language Agents.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Cognitive Architectures for Language Agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.298603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.298603Z digest=sha256:cc7d223859d4da5e283d2eb650c1ec911219d871f95e4f227039858191a01603

Observation f6d19186-4815-4322-b422-5254acc57821 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 35

Resolution
verified exact
doi, observed 2026-08-15T20:21:22.586740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.303475Z digest=sha256:57645a216db8f6aab48c09a9c43cce07c2e190fd071d963fbf256bab487d0ecd

Observation b5e9d7ac-5cb1-482e-a8ce-676a128a783e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.307841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.307841Z digest=sha256:c1381957467e7aadf83a2ac2acfd0b7a78ba2744e863c598dba5f440b456c19d

Observation b7a12258-2f5f-4179-85ae-7ecc4f5a7048 · outbound

This paper cites CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.312614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.312614Z digest=sha256:a17350c1a2d27904981aaae3153e9025deba5a7c258af8c8412637faf8b778d8

Observation deddd860-eb66-406f-8e85-845a06c863db · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 38

Resolution
verified exact
doi, observed 2026-08-15T20:21:22.569037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.317286Z digest=sha256:38658a58a5e6b5027162422289c8a18d4e6383ee78c3b76ce1ead86b2044e5b7

Observation ab012a2b-fcda-4aec-b796-48417b1d9ad4 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 39

Resolution
verified exact
doi, observed 2026-08-15T20:21:22.546753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.322737Z digest=sha256:52ea340710c315cffc40b009eeeac315490bd8666cf48377cdf74e9334a4cb7c

Observation 99db0743-4b5f-4621-b857-4e70b2177cfc · outbound

This paper cites Pan, and Kam-Fai Wong.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Pan, and Kam-Fai Wong

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.327422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.327422Z digest=sha256:7c7516b0ebd4c32e83fd8dbdce3f90a2c83d6f92bfb740c88741121db31d136d

Observation 4ceeb9ad-05f4-47ca-8a5b-c565493e841e · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:21:23.353866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:21:22.332150Z digest=sha256:e3c4e441de3ec5e3e55c166cf57768827bdf7ae4c24c58eb52cabcd3099d8f4c

Observation 0ecedaed-9c0a-450f-9912-e53c351e171b · outbound

This paper cites A Survey of the Evolution of Language Model-Based Dialogue Systems: Data, Task and Models.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges A Survey of the Evolution of Language Model-Based Dialogue Systems: Data, Task and Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.337559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.337559Z digest=sha256:6c0749e7e758f638f18cd8a3e9895c19ef7ba923dc618e8593a8c06333e05e9e

Observation 2e29f278-418c-49e0-ad6f-1b49d6e7804d · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.342635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.342635Z digest=sha256:b169bc3c6a8077f809ea3ba76031191462b5bbabc554f7552eafb9f998883555

Observation 81ecf18d-2150-41a5-be36-529c63b84e47 · outbound

This paper cites Pan, and Kam-Fai Wong.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Pan, and Kam-Fai Wong

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.348097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.348097Z digest=sha256:829a3a16500a4779a3b9ef36aecb4ae48b45b042fdb66d0c06fcbcb502171c3b

Observation 759242c7-8a97-4167-885d-c296addf5fa3 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.353417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.353417Z digest=sha256:aebc59e21af0a5d2a78802c3ab9dcb43f3f66a16e5e5e22d67e9a4014ba7cc3a

Observation 2c9f7a41-f11c-4b04-9037-dea4e97f8840 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.357700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.357700Z digest=sha256:e3b6f404b5561d16a7b78df134dd0b903587ac6dba1f22bd7cffb0806c9ca08f

Observation 2099b230-e2c8-494b-8a8f-8fc1a85b6b63 · outbound

This paper cites MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.362158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.362158Z digest=sha256:bbe87fea97016823b15d4e5d7a226fded4dc952113cfbb1c247e1b64f6013ddb

Observation f4916298-04a0-4710-b894-8bac46be5fef · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.366756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.366756Z digest=sha256:3e49b23718f463616d7c451fde9c600e81008d8095034fb5d69316ec88511355

Observation e7c7d5a7-d410-46cd-b7f3-337db7b24c12 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.371409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.371409Z digest=sha256:55055a21a30cd2ea5ff1972cdbebf54a11a95abdfc7c0078305f60b43c6301e1

Observation c8f4ca26-2ce3-461f-8aaf-fc19f807f188 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.376398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.376398Z digest=sha256:63a96853f9c60a4b6be8fdf606641c3de73a31ac2fef9d2655cdf45f0aa1fb4a

Observation 5e6b3a31-4483-40ae-9fa7-d7be07bfd267 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.381083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.381083Z digest=sha256:085ecc6d39ef346b20a2c2fcd97521cc271b873ff2a2fc7f912490dbb58e0b6e

Observation 5faf31fe-b4b0-4e01-a865-5d3f436bd377 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.385935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.385935Z digest=sha256:aac58b747fcf73d128e6cc70e2822e301c38c695164d72ae2c9995b1cba50d61

Observation cab0b35d-329d-45fa-9469-cd064975dac4 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.390766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.390766Z digest=sha256:21353f9d6d7c57696f8d4232e86d65f2c6522356405272262b41cf0865a7dfcc

Observation 6ea6efc9-9cc3-4485-bab7-e3c71d23b362 · outbound

This paper cites CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.395180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.395180Z digest=sha256:5eaa48181e025661dc579da1372d9bb178290ed4b4b846d8a02397d1deb3178e

Observation e1490c7d-4df1-40b7-b01f-4a1550928802 · outbound

This paper cites an unresolved cited work.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.400298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.400298Z digest=sha256:b8c7e4afb2974ba37606febb38080b1e804fae971f4b67278b15a3e93ae37d34

Observation adc348c2-4bb4-48c2-873d-02333adc2639 · outbound

This paper cites ToolQA: A Dataset for LLM Question Answering with External Tools.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges ToolQA: A Dataset for LLM Question Answering with External Tools

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.405524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.405524Z digest=sha256:5e24203ad58f86f02c5745b1bbcbb56d2e116e8eb74ea8b94d639c78cc830c55

Observation 89db7830-05bd-45a3-bcf7-d1af6c9e01eb · outbound

This paper cites online" 'onlinestring :=.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges online" 'onlinestring :=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.410440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.410440Z digest=sha256:dc4fc81e01e70c10b4cb92b739d2aabb0860375b3a267808343ab80baaf1bf65

Observation 0776f5d2-6d41-49b7-8761-ea38604b646b · outbound

This paper cites write newline.

Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:22.416314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:21:22.416314Z digest=sha256:453f9a3c9c33d582f7373e58fceecb4ee3b825eb2998ba46ec0dca114f470909

Pith citing papers

Observation 24699e33-12a6-4b9e-89d9-cf52684af9ce · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:38.427855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:38.427855Z digest=sha256:2ab7cc409dc077eb06a7a8014fa58ce3d4c322a67b5f641db4e1e25a25337ac2

Observation 7b16a860-d196-4d13-ae6c-60fa27126767 · inbound

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks cites this paper.

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T17:56:51.794005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-05T17:56:47.089599Z digest=sha256:040688c3767977519e055df60d18d4ce4e893c18891354353465892037816897