Pith. sign in

Paper Citation Record · LEDGER

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.27518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27518 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:25:00.076693Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43e11603-7f26-4f4f-afbb-59b226ee71f0 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.623173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.623173Z digest=sha256:e903004ce25e2b14a38a0841660721f1001af79026f78597823974dfd5c88e2f

Observation ba71c2e3-2283-457c-ab5a-d02ceb5ff26f · outbound

This paper cites Seven simple steps for log analysis in AI systems.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Seven simple steps for log analysis in AI systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.709724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.709724Z digest=sha256:10238a51e7bf46512c8b55441b6d1c0903dc72b085b8b15592ce87bf1ef137d4

Observation 6627a295-fbbe-4c4e-b302-db9fccb0b6bb · outbound

This paper cites Measuring AI agents’ progress on multi-step cyber attack scenarios.arXiv preprint arXiv:2603.11214,.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Measuring AI agents’ progress on multi-step cyber attack scenarios.arXiv preprint arXiv:2603.11214,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.795365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.795365Z digest=sha256:322ccb79626df6e7e1419af14215399dc61a696433c843996f973627ab6f0ac4

Observation 51a99686-3a9b-4f65-b9fa-34da2d4abaa7 · outbound

This paper cites Sayash Kapoor.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Sayash Kapoor

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.886679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.886679Z digest=sha256:78318b65db76db4ca8ac7bc3655cda46c49b61b2171c69ec971cabb051d3ea39

Observation e7088df3-e805-46ef-8379-0e1d7896948d · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.053330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.053330Z digest=sha256:3595ea5ca7702356f04b59e9eec2277c9f297dce5ed288cf63d45fc7a68b8d3a

Observation 505b22e1-49f5-48b5-9e56-cafd480d89e8 · outbound

This paper cites KernelBench: Can LLMs Write Efficient GPU Kernels?.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks KernelBench: Can LLMs Write Efficient GPU Kernels?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.127654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.127654Z digest=sha256:d08255d89d8050aa0294a1d59e1907eb500c5097304d704bd35f8c1bf57d6f5a

Observation 7bfc0b18-c56c-42b3-ad83-053913b0618a · outbound

This paper cites Humanity's Last Exam.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Humanity's Last Exam

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.202308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.202308Z digest=sha256:974eed78b582a96a53b927d986795fec2a4309a73518ac1090346ab0c2b38a8b

Observation 37882ef8-f420-4dbd-af7c-29969ea7c76e · outbound

This paper cites PostTrainBench: CanLLMagentsautomateLLMpost-training?arXiv preprint arXiv:2603.08640,.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks PostTrainBench: CanLLMagentsautomateLLMpost-training?arXiv preprint arXiv:2603.08640,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.266458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.266458Z digest=sha256:60fdc37c5253f4f010953d3f766d34cfa3a64d1603a194f76366914607e5b732

Observation 3c739425-967f-4349-a8a0-56f03299ffe2 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.336807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.336807Z digest=sha256:07c9cd87217c9ec661b215f8ac1f8e29a82e5a9b6d16bfa9ac5c8147d65170a5

Observation 5332bbed-fc19-4654-9342-38f9907a4c5a · outbound

This paper cites Zachary S.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Zachary S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.409558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.409558Z digest=sha256:821ca10a285651ea21e878672005b877d9ade3386f56991799312b4c928c59d8

Observation 8f278a99-1448-4012-a90e-41ba1e0aee7c · outbound

This paper cites CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.489925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.489925Z digest=sha256:786b37d0c3474442a9b35bffa0939364dc832a46cc60a166c908f5bcdb52405b

Observation 232bb22d-d1e1-423c-958b-43a7223beb6b · outbound

This paper cites BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.579784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.579784Z digest=sha256:1ce4794c49219c0fe90d1ab5ee89c6316c8c587bd629725a92946d173065c570

Observation c5a6f6b5-3bd7-4527-8feb-e3baa75a3d7e · outbound

This paper cites Automated Benchmark Auditing for AI Agents and Large Language Models.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Automated Benchmark Auditing for AI Agents and Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.673010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.673010Z digest=sha256:4022fe3d02094bf7062e94bbf0de7a966a30cfe377ea932f2ecbf2c0e8158475

Observation 98176a1b-d30a-473d-84b3-eb363626bb36 · outbound

This paper cites ISBN 979-8-89176-251-0.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks ISBN 979-8-89176-251-0

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.745650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.745650Z digest=sha256:322e86b8040e28f5ddef6be2d2392cc573e61d994d3ec26d1dbbe1ddade7dc25

Observation df392865-68cb-4dc7-9285-56fc59986a67 · outbound

This paper cites Establishing Best Practices for Building Rigorous Agentic Benchmarks.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.912539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.912539Z digest=sha256:375fb7fe37eb21f909e21883093a2f4549d70dba1df17716461981095c286a67

Observation 57a2c2a5-dab5-4a94-92ab-b33de042f2dc · outbound

This paper cites an unresolved cited work.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.995156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.995156Z digest=sha256:8d83421a92f3f738a02cd8bc0c73da52ad2e5a4eccc066bdb8dba58277591a14

Observation 6fcf916a-ec3f-4aeb-bb44-7d465402a2ee · outbound

This paper cites write to /app/out.html.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks write to /app/out.html

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:25:00.076693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:25:00.076693Z digest=sha256:92673a8b7ba18954ce508489abde9328cc0f7db38f9d8f0f7614af9b44641f32

Observation bf6fa028-3ef7-4606-80dc-77685dfdbf4d · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.835059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.835059Z digest=sha256:09d73b9a07af6ce88b83c95b359667c1967c6c6bf19bd659fe47ebfc6887e941

Observation fa861514-1852-499f-9692-3cfa92d50724 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.973480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.973480Z digest=sha256:7b81220dfd27309257024b8571cdbfc205892260386f05c6099023a3695cfa44

Observation 5f0e2ef3-8c23-4cea-a165-92b98e898a93 · outbound

This paper cites Kao, Evangelia Spiliopoulou, and Adina Williams.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Kao, Evangelia Spiliopoulou, and Adina Williams

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.541183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.541183Z digest=sha256:b92b13ee48ce30281be2dfb44c795c4f1f4051c3de057386216483b363d1dca5

Observation 7d7df7e5-3904-428d-a5f8-0c65f2366e4e · outbound

This paper cites SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.454362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.454362Z digest=sha256:9aa6ed2355472e659e25180eb08b7f90d789f5a97c7acaae649b5284f878f012

Observation 048ba888-a8b2-4d5c-b8f5-12df41a89cb9 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:58.374494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:58.374494Z digest=sha256:3af0c54a53d7e71c4eab773f339cd0231777422a9ea43dc6290b87a0a3f808e5

Pith citing papers

No inbound Pith citation observations are available.