Pith. sign in

Paper Citation Record · LEDGER

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation

As of 7 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2607.02577.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02577 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:42:05.329537Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:09:21.261237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:09:21.591895Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc4f18c1-eac5-4b46-b96e-87bef83d60ae · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:0cdfb23f69118d7870a4f79d8b3d7c897931cc0e674b2ecdf5c0c7e6b631905f

Observation 0c5b989e-97e3-409d-916b-7f9f9598ef9e · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:30dd8ebcb595cc3b106186dbe1c40dfc9e0ef5311756bdde3eebca214559b032

Observation 2ed10347-f040-4e56-9800-24b1fe5ca8d7 · outbound

This paper cites MAVEN: Improving Generalization in Agentic Tool Calling.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation MAVEN: Improving Generalization in Agentic Tool Calling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:3f333bb4e5cdb2768805473a5728836799b19cdf2115b010d79ea1dba5c008e5

Observation 831b7501-61af-4166-8123-e3d52a3c65b3 · outbound

This paper cites Stabletoolbench: Towards stable large-scale benchmarking on tool learning of large language models.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation Stabletoolbench: Towards stable large-scale benchmarking on tool learning of large language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:9adc3b163f058b79619886ace8a9d03f69d3cc868a20090e852a8e1daf37a6eb

Observation ba1ca775-7917-43df-b323-7d70bdee1b3b · outbound

This paper cites API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:f1b056022a12302a20e76b711b9587fbfd0e2b0496256f995103a8b16760d2fd

Observation 467de242-e166-47fe-9dd5-12bb60678a62 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:5e01e9f294471f91f38b6cc6112ec778274f71fbe8774d5a20e97ff1e146e410

Observation 8ab01533-22df-4547-b45e-b6bda51b857b · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation Gorilla: Large Language Model Connected with Massive APIs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:718d527aff3ada10a6cfa8d1c5477aa8e47884d7a1ebb500ded310b41155684b

Observation ab7a93ac-9bd6-4d03-a5d9-60773118ae37 · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:3f0899ca32ec5a9ac23b5021b6b55d9a5e0fcb18eb5d44385d24497245f2a373

Observation 3f7d5ac6-3a9a-4c0d-99d6-7867b997d14b · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:a08e74f168ccafa7baa6a0df45412fd3810e17563dfc24dfa52134802a4abc84

Observation 76ea41cb-692c-4f97-bab5-d1455cbff6ae · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:5dba23aa2cc55b2f8043bd85a24477f73a762da4effe88068fcd8fc42ce53433

Observation 95f9596d-ada6-4c64-9c0e-85603c1955c0 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:b405b53db2b65db517c9fa9e7eefcdd027ab6330a575dda9709b7f9fdde086e3

Observation 1fbfa584-db36-476a-8a56-f138926b5c25 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T09:42:05.329537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:42:05.329537Z digest=sha256:11c21e6f60d11b3131c9280a3243f789cf05706b072a132ad28f85ef46ff541b

Pith citing papers

Observation 339d6ca2-2373-4c46-a921-acc5a25c0029 · inbound

The Bitter Lesson of Tool Calling cites this paper.

The Bitter Lesson of Tool Calling Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:09:21.711324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:09:21.261237Z digest=sha256:6988a9759b8936f5f8f0a91740d051d42ec722efeb40ed518bedef94c1e192ed