Pith. sign in

Paper Citation Record · LEDGER

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

As of 24 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.08700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08700 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:54.413371Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8aaef3f8-eff4-4a5b-9073-a6c8f633a1ab · outbound

This paper cites Li, M.; Zhao, Y.; Yu, B.; Song, F.; Li, H.; Yu, H.; Li, Z.; Huang, F.; and Li, Y.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Li, M.; Zhao, Y.; Yu, B.; Song, F.; Li, H.; Yu, H.; Li, Z.; Huang, F.; and Li, Y

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.305784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.305784Z digest=sha256:c7fbc88d824c48cf7aeb006e294fcc4e5317dad23035d8988c1c4c4c5c62297c

Observation 2035427f-27f2-473d-beda-21829b02f35a · outbound

This paper cites EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T04:32:55.036595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:54.324508Z digest=sha256:5bd9808b425a18af41913c2ce22deb0aea229ebc09f0132d509aeb58539ee7a3

Observation e5e5bc9f-df7b-4540-88cc-0d2916e0ceb1 · outbound

This paper cites Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.330550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.330550Z digest=sha256:ba430a0a46d42551740553afb8bdc70d38c9f019eb69ba11b319bb09681d7822

Observation 1b0dff53-319d-4e8b-8ee3-5dc3292f1720 · outbound

This paper cites Patil,S.G.;Mao,H.;Yan,F.;Ji,C.C.-J.;Suresh,V.;Stoica, I.; and Gonzalez, J.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Patil,S.G.;Mao,H.;Yan,F.;Ji,C.C.-J.;Suresh,V.;Stoica, I.; and Gonzalez, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.336410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.336410Z digest=sha256:3a605a0a86f951fb33f445d868035b731e8c62ad2c3e7c6bb271013f3247a1c4

Observation 7ece5b65-2ced-4172-9e15-f43d12056726 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling tinyBenchmarks: evaluating LLMs with fewer examples

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.342840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.342840Z digest=sha256:0ef60b9a25f94b982c0e11fc3011475ecb11a10f9d72a6de16f7158f99196728

Observation 9c643d66-58b5-462d-ade7-7083ae51ab5d · outbound

This paper cites OpenAI GPT-5 System Card.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling OpenAI GPT-5 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.350401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.350401Z digest=sha256:aa8f610d172cc18af572b21af4f19ddecdf9ee6f49d6d1baaa44711e76d38175

Observation 1cd76f9b-0f4a-4cec-a059-986c3f26185e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Gemini: A Family of Highly Capable Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.355662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.355662Z digest=sha256:abc7a8235adb0247892d6de094a16df04662c60595045d1e7eba91b031c31489

Observation 3a74fb33-070c-470c-a50e-5ef13408e319 · outbound

This paper cites Wang,X.;Wei,J.;Schuurmans,D.;Le,Q.;Chi,E.;Narang, S.; Chowdhery, A.; and Zhou, D.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Wang,X.;Wei,J.;Schuurmans,D.;Le,Q.;Chi,E.;Narang, S.; Chowdhery, A.; and Zhou, D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.362365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.362365Z digest=sha256:40b3882bb19c19fa4250503c16bc20c843c4ba956c8dcfd5179315f51dbc1ba3

Observation 232459a5-a665-42c3-b7e7-09f51d89b5fc · outbound

This paper cites arXiv preprint arXiv:2606.19348.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling arXiv preprint arXiv:2606.19348

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.374135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.374135Z digest=sha256:93ef26169e21f60633eab48ad79100dca2a6d3cc99845eb66ea24b86000a73af

Observation 2efe832a-ca61-4f12-9d6a-f801e717b829 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.386715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.386715Z digest=sha256:3c7226b3df074f25043bb973920f3046e20b8fa427a97691da427dff727ff5ef

Observation 405c0d3f-c382-4837-a7a2-e13c69f533fa · outbound

This paper cites InNeurIPS 2022 Foundation Models for Decision Making Workshop.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InNeurIPS 2022 Foundation Models for Decision Making Workshop

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:55.211557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:54.401650Z digest=sha256:bd82ce153b07a9927e3eab7e01a1fbad956ce38e112e83b00c58f5e05a3d978d

Observation 5b96ea07-d496-4c5e-bf4e-26f468aa8126 · outbound

This paper cites InInternational Conference on Learning Representations, volume 2025, 102351–102390.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InInternational Conference on Learning Representations, volume 2025, 102351–102390

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:55.185449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:54.406883Z digest=sha256:b9cf8e7459fbb7c9cc025eedf1330deaad50533dca7261f67f66599de3e844d7

Observation 107ca167-ae3e-430e-aaa1-3f37a594b991 · outbound

This paper cites Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:32:54.547984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:54.413371Z digest=sha256:cf21f155ea888c12df14eb59b034ad9f8860378016f99485a88080d76e58800b

Observation c4242ccb-9f98-4981-abca-89cd1e644b8d · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.367940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.367940Z digest=sha256:94bb2af113c1deaa4cacee19d4bc3ddb26c330b28b72655c735d2bf71d1922ce

Observation 4b353d32-9f69-45bc-a1f0-83d90dc68d71 · outbound

This paper cites InProceedings of the 2023 conference on empirical methods in natural language processing, 3102–3116.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InProceedings of the 2023 conference on empirical methods in natural language processing, 3102–3116

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:55.237855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:54.314051Z digest=sha256:d20821c5a608559da3cf0e7c2cb79d6704c0edf3677d98a418bd1a4a2beeb0ee

Observation 1c61da75-4c2a-4d7b-ae5d-2643cc21b5db · outbound

This paper cites Kim,D.;Ren,Z.;Hao,J.;Sun,Z.;Wang,L.;Ma,X.;Ye,Z.; Han, X.; Yin, J.; Ji, H.; et al.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling Kim,D.;Ren,Z.;Hao,J.;Sun,Z.;Wang,L.;Ma,X.;Ye,Z.; Han, X.; Yin, J.; Ji, H.; et al

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:55.262298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:54.295232Z digest=sha256:6fb3fe476f0d395746a96015f54b3794e8480ff1d98d3d42cc79c7551d2371e4

Observation aef55e00-8030-495c-9321-2b2fbccf96bb · outbound

This paper cites InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 33526–33535.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 33526–33535

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:55.282603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-14T04:32:54.287491Z digest=sha256:fa7956a669c83822f4d248f98f7074f1fecda23ef06f58aacfb87051c0f4c4bf

Observation b82de072-7f20-4029-a50b-62c3d5876d69 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:54.281242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:54.281242Z digest=sha256:0b6c7c10c2897d8b91f377bcc9d94d304efff97b5a0c8bdc210c6425a2a888e2

Pith citing papers

No inbound Pith citation observations are available.