Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T21:40:30.408782Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2605.26079.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T21:40:30.408782Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:24:59.673010Z
A source-named dated measurement, never combined with another source.
Source: cited_works
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a616ed0e-3d9d-4ff7-b6d7-487a95b25a64 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8944150-d657-4104-8334-73c68e943637 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models task" or
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acccf5eb-7c43-4d86-a337-7692cf14b141 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b205c996-6a48-4e6e-ae68-5bc551b6af2b · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models unscored
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3865b20d-d8ee-496a-89ef-4a8dcdd17da5 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models It defines what counts as a finding, how to distinguish agent error from genuine task issues, and the severity scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4a49c5-5c35-4b77-99fd-29f8c583aa75 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models Open and read every path provided
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c40901-569a-45fd-a918-439dd2ce7fe2 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models task_id":
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1675d99e-8656-4c00-903e-3eda34febd57 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models It defines what counts as a finding, the severity scale, and the distinction between benchmark issues and expected difficulty
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6491be7-ac7a-4972-a271-af0cbdacdf01 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b35d33-d0dd-4bb7-9f7b-fc390e9e8a8e · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377efa25-9e50-4507-a2f2-75a080fe562c · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models task_id":
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e399d1f3-5d31-4c6e-9e19-b0aedfef05c0 · outbound
Automated Benchmark Auditing for AI Agents and Large Language Models source-file only
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a6f6b5-3bd7-4527-8feb-e3baa75a3d7e · inbound
Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Automated Benchmark Auditing for AI Agents and Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.