Pith. sign in

Paper Citation Record · LEDGER

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2607.04686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04686 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T15:12:56.011712Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:17:57.244743Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c177f1e6-aebb-4075-a9ff-b8027fe6ad5a · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:1ab75610bdab1a1450f60b3c85ebf851c0429839f32a0714d699c5b7560c191b

Observation f6b09d4a-eaa3-46ca-a215-3baaf5d0cc6e · outbound

This paper cites and Mao, Huanzhi and Yan, Fanjia and Ji, Charlie Cheng-Jie and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents and Mao, Huanzhi and Yan, Fanjia and Ji, Charlie Cheng-Jie and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:1b3f65d248a73d16281bfb98bb228608159cf4d40df75c8ab46d727f47f874c5

Observation 12d4ccb2-8da2-4fe4-a265-1a91a2e9e564 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:6f28fd140c85a8937714269594138eaf8edc268dc7e34a70bf39daa0f913d3c6

Observation ca915570-62d2-4023-af1f-a6c894a85d2a · outbound

This paper cites and Zhang, Tianjun and Wang, Xin and Gonzalez, Joseph E.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents and Zhang, Tianjun and Wang, Xin and Gonzalez, Joseph E

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:bad749d3d136f31362669f6aed883fe6b828e8f3b3b152053bc85aa397dbfd59

Observation 060e3252-148b-4833-a3b1-498b8d40fafa · outbound

This paper cites arXiv preprint arXiv:2510.25726 , year=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents arXiv preprint arXiv:2510.25726 , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:bea95610d23a1927228abd6937dd8bf844df83fd454d5468723a51064575d8b2

Observation 826f14f1-aef0-482e-a628-34322ec0898e · outbound

This paper cites 2024 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2024 , url=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:341dfcc5f702f3cccfd6121bc2e7cfdfa1ae7bcb55cf54b0ae6c6bf9b8b69644

Observation e29d4aa5-6c83-4e9f-b6e2-71d7cd41a76e · outbound

This paper cites and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:64b459cf8fd06f46b508c045efe92dc84c9e0c75dd49d73c147f4b483206fcec

Observation ef74a8bc-1374-47d0-8009-81748a74fd5c · outbound

This paper cites Finance Agent Benchmark: Benchmarking.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents Finance Agent Benchmark: Benchmarking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:b73444cd0a6437070a2aa9a80dd3c90ee47b3e98f0efd11580fe2a6beb5e4a21

Observation 49247251-c6a0-42a8-ac97-9bde6001162e · outbound

This paper cites and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:bbce823ae02b53fb03c35b8fef1090a413b24aab3702c3956ad504229229c49b

Observation edfc4b52-4ff0-4077-a3f7-26d8fd4fe7b9 · outbound

This paper cites 2024 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2024 , url=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:4578d6cac114b5c32da070eaf77e47ae7fb656525564e72a944917e876e4cc3d

Observation e88112ad-b5a4-47ae-bd65-4936f92ee39d · outbound

This paper cites and Zhang, Hao and Gonzalez, Joseph E.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents and Zhang, Hao and Gonzalez, Joseph E

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:270b2800a293b067aa6674865c68702ab2eb23243e1b5657669c42ff38a88408

Observation 97e42aa2-bd0f-4e90-b072-10761943d95f · outbound

This paper cites an unresolved cited work.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:5942bd0c30ac363b2f3ccb7a691012f5f8889e993f017d561ae7b8780fd1005c

Observation 7f0fd84c-a0f9-419f-a5a9-1255a973dfe0 · outbound

This paper cites Qwen2.5 Technical Report.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents Qwen2.5 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:1e39efd1c8626ad7c921f2e58aa4a29e4644a4e58828d5d6c458d06c3971c307

Observation 86f1a5b3-f13c-41d4-ac14-8ac71e09c9d1 · outbound

This paper cites 2024 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2024 , url=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:0a21fecfc6a03b3827b9f7b033786119da4e3f17eebf596ed7287da68fe6de80

Observation bed24dcb-306b-48c2-a850-51490f324a17 · outbound

This paper cites 2024 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2024 , url=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:5d8ebe44f4212cb0658c76f2a3e6ee1cd84dcd822b3afe7beaf07a2087a4d454

Observation 9a1b86f5-ad7d-46b9-b450-bd93d7be622d · outbound

This paper cites Educational and Psychological Measurement , volume=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents Educational and Psychological Measurement , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:71c903822c9c35227e76c2a9320cfcfe9276523fca8f3904662fe06a14913f97

Observation 867c7751-3940-4ae1-a2de-96157fa811c2 · outbound

This paper cites Psychological Bulletin , volume=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents Psychological Bulletin , volume=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:2351281f8c0ec2812e5403023d4530da86c77f3c6ae99097586a1028dfe30c15

Observation 47fdd92f-e3d1-44d5-87f1-94dc9c1fae40 · outbound

This paper cites 2025 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2025 , url=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:581dea6ef10207bc3af513c5f32ac6beed107496bc09cd5ff3e085b1ce55ef3b

Observation 3d112f79-2203-4d70-a548-42b5ffa4cbb5 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:59c3f726dbbceb4ddf459abf6b2fb7ef724b616d76b79cc1f911e84d9faeb022

Observation 87137053-b27d-4fad-b167-b37a01310401 · outbound

This paper cites and Zhang, Hao and Stoica, Ion , booktitle=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents and Zhang, Hao and Stoica, Ion , booktitle=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:1d552c21a71a1a50e0af2384f373c91e9c7890a7545070fce6ceddb124a33e0d

Observation 1c065b3d-26d8-4c18-99ee-44b65c15f1dc · outbound

This paper cites 2023 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2023 , url=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:6d99f58daba2ce70da6c82abd6179e0afcef5c5e48d9783433816a1558b4088a

Observation 96260572-bdae-4ce1-bb7b-16a8894f0708 · outbound

This paper cites 2023 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2023 , url=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:ce9f71e7e6511f09812281fdbe087edad967e809d5731e4e24900b9fa375c470

Observation 28b440d5-65b8-4207-b4e6-cc2d36543edd · outbound

This paper cites 2024 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2024 , url=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:50998f35b33b189910729e0457149cbca29b74766d95ddc5a08a1075a82e8796

Observation d026d882-0cb0-4483-a110-44cd6461a724 · outbound

This paper cites 2024 , url=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2024 , url=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:a1e777110b2557ad69901c5bed1f8de21e4e4e10b47747ce9a449bbb2428e815

Observation d7290a79-2880-4701-93f3-6bf85cb23364 · outbound

This paper cites and Zhang, Hao and Gonzalez, Joseph E.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents and Zhang, Hao and Gonzalez, Joseph E

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:88c0ba407f4ea73a94eb8a81fa2756385b0ee8a0fa90f8e0e17592c934094d8c

Observation 84796e9f-cb2f-4007-b60e-c663e031805a · outbound

This paper cites 2026 , howpublished =.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2026 , howpublished =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:0054a536466f1c82eaba5b6308aba505c743971c1322d84eff4d7180ea73ea3b

Observation 5c5d8569-33db-467b-a7ba-41b1d18e3079 · outbound

This paper cites 2025 , month=.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents 2025 , month=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:ef41033ef3590554d105f4d88688abaf729371eff296706fb80a692518891df3

Observation 5d817d2b-7384-40db-90b2-c7f697f193e6 · outbound

This paper cites The Effect of Sampling Temperature on Problem Solving in Large Language Models.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents The Effect of Sampling Temperature on Problem Solving in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:6c6b94ec8134c41b327ca4ba3541c2334a549620ebe449a856b52a7aea5990ef

Observation 9f14c13c-6d39-44c9-b08d-16a6f6d2c51d · outbound

This paper cites The Good, The Bad, and The Greedy: Evaluation of.

ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents The Good, The Bad, and The Greedy: Evaluation of

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T15:12:56.011712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T15:12:56.011712Z digest=sha256:6d52f8456312df95221d93677a06ad0bedacff7d4618ab6af9708bb39b2763e3

Pith citing papers

Observation 8cf90e24-efa5-44e8-b18a-510e7b202ea8 · inbound

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents cites this paper.

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T12:17:57.244743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:17:57.244743Z digest=sha256:28748970276d2d786d5f0daeecb54b62b6c4c06459e849758dd45f043cd0080b