Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:54.413371Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.08700.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:54.413371Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8aaef3f8-eff4-4a5b-9073-a6c8f633a1ab · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Li, M.; Zhao, Y.; Yu, B.; Song, F.; Li, H.; Yu, H.; Li, Z.; Huang, F.; and Li, Y
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2035427f-27f2-473d-beda-21829b02f35a · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e5e5bc9f-df7b-4540-88cc-0d2916e0ceb1 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0dff53-319d-4e8b-8ee3-5dc3292f1720 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Patil,S.G.;Mao,H.;Yan,F.;Ji,C.C.-J.;Suresh,V.;Stoica, I.; and Gonzalez, J
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ece5b65-2ced-4172-9e15-f43d12056726 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling tinyBenchmarks: evaluating LLMs with fewer examples
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c643d66-58b5-462d-ade7-7083ae51ab5d · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling OpenAI GPT-5 System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd76f9b-0f4a-4cec-a059-986c3f26185e · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Gemini: A Family of Highly Capable Multimodal Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a74fb33-070c-470c-a50e-5ef13408e319 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Wang,X.;Wei,J.;Schuurmans,D.;Le,Q.;Chi,E.;Narang, S.; Chowdhery, A.; and Zhou, D
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232459a5-a665-42c3-b7e7-09f51d89b5fc · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling arXiv preprint arXiv:2606.19348
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2efe832a-ca61-4f12-9d6a-f801e717b829 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405c0d3f-c382-4837-a7a2-e13c69f533fa · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InNeurIPS 2022 Foundation Models for Decision Making Workshop
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5b96ea07-d496-4c5e-bf4e-26f468aa8126 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InInternational Conference on Learning Representations, volume 2025, 102351–102390
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 107ca167-ae3e-430e-aaa1-3f37a594b991 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4242ccb-9f98-4981-abca-89cd1e644b8d · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b353d32-9f69-45bc-a1f0-83d90dc68d71 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InProceedings of the 2023 conference on empirical methods in natural language processing, 3102–3116
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1c61da75-4c2a-4d7b-ae5d-2643cc21b5db · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Kim,D.;Ren,Z.;Hao,J.;Sun,Z.;Wang,L.;Ma,X.;Ye,Z.; Han, X.; Yin, J.; Ji, H.; et al
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aef55e00-8030-495c-9321-2b2fbccf96bb · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 33526–33535
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b82de072-7f20-4029-a50b-62c3d5876d69 · outbound
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.