Pith. sign in

Paper Citation Record · LEDGER

Automated Benchmark Auditing for AI Agents and Large Language Models

As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2605.26079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26079 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T21:40:30.408782Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:24:59.673010Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a616ed0e-3d9d-4ff7-b6d7-487a95b25a64 · outbound

This paper cites an unresolved cited work.

Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:28ef0fdde1d2938b75dbe17fbcd67ae91eb9bf61be9264d273fe9858ac442c83

Observation b8944150-d657-4104-8334-73c68e943637 · outbound

This paper cites task" or.

Automated Benchmark Auditing for AI Agents and Large Language Models task" or

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:e2a4cefbf6dde90c0bb64055ecabdd1874cf2a9434d5993c79de51f1d7fe52b3

Observation acccf5eb-7c43-4d86-a337-7692cf14b141 · outbound

This paper cites an unresolved cited work.

Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:59688db6b31221184fc9b0acc6a58850228c484ed4a9f1c5b806d867620216ad

Observation b205c996-6a48-4e6e-ae68-5bc551b6af2b · outbound

This paper cites unscored.

Automated Benchmark Auditing for AI Agents and Large Language Models unscored

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:f5f90961a54542ed64dc9f8c93d9db18fdd62c499cff52df2561cf1f8cc13dbc

Observation 3865b20d-d8ee-496a-89ef-4a8dcdd17da5 · outbound

This paper cites It defines what counts as a finding, how to distinguish agent error from genuine task issues, and the severity scale.

Automated Benchmark Auditing for AI Agents and Large Language Models It defines what counts as a finding, how to distinguish agent error from genuine task issues, and the severity scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:41b2eeb4f95e2a717f678db0c3ca11ec337b3ced64e0a04e52f1321b1ac35d52

Observation ca4a49c5-5c35-4b77-99fd-29f8c583aa75 · outbound

This paper cites Open and read every path provided.

Automated Benchmark Auditing for AI Agents and Large Language Models Open and read every path provided

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:1f91936f14e3a4d6d62e3983941a684ac707f712ead8b75801e56baa4dc1e0dc

Observation 41c40901-569a-45fd-a918-439dd2ce7fe2 · outbound

This paper cites task_id":.

Automated Benchmark Auditing for AI Agents and Large Language Models task_id":

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:d075faf366923948317a7dac6b88d52b1799ad7cac07a4a4bfa692a9bab1ebd0

Observation 1675d99e-8656-4c00-903e-3eda34febd57 · outbound

This paper cites It defines what counts as a finding, the severity scale, and the distinction between benchmark issues and expected difficulty.

Automated Benchmark Auditing for AI Agents and Large Language Models It defines what counts as a finding, the severity scale, and the distinction between benchmark issues and expected difficulty

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:77dd01f2e17cd2350736f8205efd1c3fd6788218140382219b7438e2a7311712

Observation f6491be7-ac7a-4972-a271-af0cbdacdf01 · outbound

This paper cites an unresolved cited work.

Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:831331caec59de0af3233947e7f4b4ecfa927a08e1c253543bcc33fa8a70bb63

Observation c8b35d33-d0dd-4bb7-9f7b-fc390e9e8a8e · outbound

This paper cites an unresolved cited work.

Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:55943794f263315d482fae3e90a39f80d2245c98086419c255f8075133de4035

Observation 377efa25-9e50-4507-a2f2-75a080fe562c · outbound

This paper cites task_id":.

Automated Benchmark Auditing for AI Agents and Large Language Models task_id":

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:dc78a50e5afa713c34d65765da02986f6b766cee953bdff0ea257cf2a21e1f69

Observation e399d1f3-5d31-4c6e-9e19-b0aedfef05c0 · outbound

This paper cites source-file only.

Automated Benchmark Auditing for AI Agents and Large Language Models source-file only

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T21:40:30.408782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T21:40:30.408782Z digest=sha256:38bc4a3bf9e4f72bc3c8dadf6584d66310e19bcac32d3adac22cde0657c32a9d

Pith citing papers

Observation c5a6f6b5-3bd7-4527-8feb-e3baa75a3d7e · inbound

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks cites this paper.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Automated Benchmark Auditing for AI Agents and Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.673010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.673010Z digest=sha256:4022fe3d02094bf7062e94bbf0de7a966a30cfe377ea932f2ecbf2c0e8158475