Pith. sign in

Paper Citation Record · LEDGER

Establishing Best Practices for Building Rigorous Agentic Benchmarks

As of 20 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 37 inbound Pith citation observations for arXiv:2507.02825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02825 v5

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:29.338123Z

measured 137 of 137 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:19:14.004821Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:08:59.930877Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 498da9e5-9da9-4859-a1f0-e2235cb99aaf · outbound

This paper cites Inspect AI: Framework for Large Language Model Evaluations, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inspect AI: Framework for Large Language Model Evaluations, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:19.476468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:19.476468Z digest=sha256:c1c8b767ce01d607cd259ed7cede10d7ed08b5625df9ed120e6332381de82527

Observation 5aea9622-5a6f-4b7e-87c6-96ad04a561b7 · outbound

This paper cites Gpt code editing benchmarks, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt code editing benchmarks, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:19.643317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:19.643317Z digest=sha256:a3939d90e62abc382939c5c0f02c96b8f1a116cc0c40e991832eb69a86e64b63

Observation e9a79027-ed54-47d7-a5ac-f01f5ccc7ccb · outbound

This paper cites o1 tops aider’s new polyglot leaderboard, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks o1 tops aider’s new polyglot leaderboard, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:19.860736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:19.860736Z digest=sha256:f34e0196266ac72707b1e2d516fb0d50c9241bd1216eec2b60e52132deda7ac4

Observation 0b486f1a-5d0e-437d-962b-35608d837b1b · outbound

This paper cites The amazon nova family of models: Technical report and model card, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The amazon nova family of models: Technical report and model card, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.062843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.062843Z digest=sha256:5c8b35684357249ac0f973ec56abd59842859c2a1a8b83938b8b1a34dd778538

Observation 1d636119-8fda-4571-83bd-dc153e8716cd · outbound

This paper cites Claude 3.5 sonnet, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Claude 3.5 sonnet, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.190930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.190930Z digest=sha256:b81ca09b2cf05fb25ab3bdc2fb4834748f2ea7763b3db1c729f604df61bcaae5

Observation 1623b2cf-dbaa-48cb-80e9-ba7df2364368 · outbound

This paper cites Claude 3.7 and claude code, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Claude 3.7 and claude code, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.342592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.342592Z digest=sha256:e4e035b34637dbfaea09d769b43fa6afb2f55d0bedbc9d0b650cf9836be311aa

Observation d89439eb-2c9e-43cd-86d2-2223c0d8a01d · outbound

This paper cites Bird minidev - corrections, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Bird minidev - corrections, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.521292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.521292Z digest=sha256:311c17120a8ee72f11fc4abe922f51038fddfb7b0d966ebfda88065d4a06aa84

Observation 59767821-da00-4625-abeb-e8755f7f14cf · outbound

This paper cites Program Synthesis with Large Language Models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.658359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.658359Z digest=sha256:fd02a79873d6743b59146c8259cfb72991d8cc468ed301d556beeecc3f65083a

Observation eb83d160-535e-4026-9385-50edf7de0464 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.815544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.815544Z digest=sha256:5de8c9436282004c9b4d2e498f15808d12166e3749ed894368ff3fa48824a62e

Observation 58917dc5-a8db-4654-8d3a-2a65a00bfc8e · outbound

This paper cites Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.949000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.949000Z digest=sha256:3c44217b6554b008a8748e217b1d446d6d49841d864ec166010d271fc86b9788

Observation c3080b94-bee8-43c4-a758-4a0b10a01473 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Establishing Best Practices for Building Rigorous Agentic Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.069258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.069258Z digest=sha256:efebb8ffbaf4b9e3d9d652e1ec8aa0d876c41a8a72924d029c9320474e57bad7

Observation 260b815f-fb0f-4913-ab43-8dc7a875ea78 · outbound

This paper cites AutoAgents: A Framework for Automatic Agent Generation.

Establishing Best Practices for Building Rigorous Agentic Benchmarks AutoAgents: A Framework for Automatic Agent Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.195372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.195372Z digest=sha256:92c8677d9596fbf8772023e9abf8dc5a675d23f741578c63c2db6c287d53f8a9

Observation 5bccefc2-874b-4a7b-9091-43df7c83c51c · outbound

This paper cites an unresolved cited work.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.316891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.316891Z digest=sha256:ed93842fc146f8ea8ae323eac3e8d85bcceab12d77de7f3ed2cb66d082ac0149

Observation 78deee31-97dc-42af-af85-ad8058691468 · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.435890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.435890Z digest=sha256:37df6cc0c09835635cd0fcd0564bc4d9b5f5dfb7f0192c803d530022d9c9d63a

Observation 0a80de66-d368-488c-ac16-16624a5bda46 · outbound

This paper cites Introducing deepseek v3, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing deepseek v3, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.523208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.523208Z digest=sha256:f6342872255208770776f8058e27bdef9fc1127df20471c28ff2f22729ca57e4

Observation 7c7a4f70-66f9-45fb-8a76-6496e199b5b6 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Imagenet: A large- scale hierarchical image database

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.651866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.651866Z digest=sha256:f2c0bcf5446c9db7e4096daef7e3a6d52da47344d805e826ecac287a3f2933c9

Observation 3e3cf797-2521-4f5a-9c09-cc752903aa01 · outbound

This paper cites Don’t label twice: Quantity beats quality when comparing binary classifiers on a budget.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Don’t label twice: Quantity beats quality when comparing binary classifiers on a budget

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.733040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.733040Z digest=sha256:b123afc82eadb5a6bbb97c257f0c036769bbaa4fa23f7bdcfcd22a0d8ec2ff74

Observation 9f23249f-8e95-4619-8525-f6f7f568a9a5 · outbound

This paper cites Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.837718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.837718Z digest=sha256:fc52920bd2673ea7e2e30c787b3933345f43459c16b432a96dacdffd985ea387

Observation 03881fb7-d8f6-4f03-b3eb-8aae63ca2da1 · outbound

This paper cites The design and operation of CloudLab.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The design and operation of CloudLab

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.946111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.946111Z digest=sha256:b7a1af75ec7080da55e935c68af16fc916ccf2e70504df67a363d1921b2638fe

Observation 06be70c0-2a6c-4b6d-9947-589ce3f0d3ab · outbound

This paper cites Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.007624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.007624Z digest=sha256:d30c7e3475bc49e681e0c66ae7f4aba532dd8e5d80ee9a3bda85cdb875a55bc3

Observation b9413ed1-7af1-4651-a911-5f97e31cd44d · outbound

This paper cites Searching for computer vision north stars.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Searching for computer vision north stars

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.076812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.076812Z digest=sha256:1bf1367bba34f707585b8e5c78c43051b7737a9b28ccb34626dbcba6066a0471

Observation 2fc04ec4-6484-4b78-9075-bd038d42f644 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Establishing Best Practices for Building Rigorous Agentic Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.159924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.159924Z digest=sha256:621c33a1ef266df3080f6a8f0d5bcb9192fef1b0c32628ee2ac8e8642a1980c0

Observation 9d379cf5-0b6c-4177-a756-93c27c27eeeb · outbound

This paper cites A classification of sql injection attacks and countermeasures.

Establishing Best Practices for Building Rigorous Agentic Benchmarks A classification of sql injection attacks and countermeasures

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.247550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.247550Z digest=sha256:dbbc136abb5d2c20b281a19ecfba53e18e844a2b40980973f25f240869bfe769

Observation 4706c5b4-9b26-4c38-8cf0-7db4d811c7d3 · outbound

This paper cites More than marketing? on the information value of ai benchmarks for practitioners.

Establishing Best Practices for Building Rigorous Agentic Benchmarks More than marketing? on the information value of ai benchmarks for practitioners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.337945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.337945Z digest=sha256:a7fd7ad003ddba9915fddb7b12a444ce0d0c2099249a574b07d5272297b825ab

Observation 8e102cf6-4a50-42e7-9ccf-abe7fefd4706 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.424552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.424552Z digest=sha256:ac383f7ba61e4c04d69feae9b2c09aeb50b5559dc8acc87b5162722bf7af1bb8

Observation fdc8141c-a335-4548-9956-ecba7fc35208 · outbound

This paper cites The extent and consequences of p-hacking in science.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The extent and consequences of p-hacking in science

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.505161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.505161Z digest=sha256:3f769bba952bb1d75e1e2088dfc0f4abb3fcb147534dd64dadedacf0fa9ac837

Observation 36a63e39-97b0-4b90-995f-a6e0a43c8980 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.595920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.595920Z digest=sha256:04660a15364fae80c025945d28e53246e0a67b3addc72db4c22a1d1e17254ff9

Observation 638e800b-6160-41d1-ba91-ddee087bb69f · outbound

This paper cites The design and analysis of benchmark experiments.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The design and analysis of benchmark experiments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.654448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.654448Z digest=sha256:ceaba26e763a17603e246e8b9807b578b81a1b75a2fd83a7f3dafcd354f15e15

Observation add6f924-0680-4f83-94e7-1440ddcdaba8 · outbound

This paper cites Preventing server-side request forgery attacks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Preventing server-side request forgery attacks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.744781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.744781Z digest=sha256:dcf64f162b6be2bf76ac480a25d1f04bdfbecb6e237e8d49bb24a727ec78378c

Observation 6739a092-f9f4-4199-be3a-b4940e5bab8f · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In ICLR, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-bench: Can language models resolve real-world github issues? In ICLR, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.816217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.816217Z digest=sha256:b17595d2cfc05f45fb07f0464d9a41286e8309323c170669fb236b1f9a832250

Observation 79952336-47d2-4a0e-a240-a8843eeec48a · outbound

This paper cites Swe-bench verified leaderboard, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-bench verified leaderboard, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.883514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.883514Z digest=sha256:70ef39f6b790c06af71e156bf8f364f5d8c8ea00b479d84fb88a6ce41aa30b05

Observation 9be6db01-506c-4190-b663-1296790ee689 · outbound

This paper cites AI Agents That Matter.

Establishing Best Practices for Building Rigorous Agentic Benchmarks AI Agents That Matter

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.949724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.949724Z digest=sha256:d5aacaea55b5f57dab654eb305969ae707f8be93bf092784fc7d47bf94bec703

Observation 122680c8-e8f9-4bcd-9ecf-6ea0d54319ac · outbound

This paper cites Gemini 2.0 is now available to everyone, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gemini 2.0 is now available to everyone, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.020325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.020325Z digest=sha256:a027104d1701ccec2e6c24cd89c0a4a64d6d1cc7e9dd60b98d4e046afbb64ca3

Observation 1c028daa-0318-4b18-9949-f93bc0f69941 · outbound

This paper cites Math-verify, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Math-verify, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.130120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.130120Z digest=sha256:9009915e6905d3ed5e7a2b382c5b7e6601c8771acd847ce229f87e289a1759d9

Observation a4b01508-7d1a-4ad4-af77-594e0d319e56 · outbound

This paper cites The ai cuda engineer: Agentic cuda kernel discovery, optimization and composition.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The ai cuda engineer: Agentic cuda kernel discovery, optimization and composition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.259606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.259606Z digest=sha256:1533b004dcf94b4eaeebdd42900c7902a0f9ff49cf91d8a363d2cc262506982f

Observation 592213bb-34ee-436c-ab07-fbb055fc2201 · outbound

This paper cites Challenges of end-to-end testing with selenium webdriver and how to face them: A survey.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Challenges of end-to-end testing with selenium webdriver and how to face them: A survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.335796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.335796Z digest=sha256:543df87c05d2f56f8b00c70dc520477ea3ceb6ae7bbaaf112d6be6947b40580f

Observation 8abab65e-cefc-4dcb-8bcf-2afa33484a64 · outbound

This paper cites Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.446008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.446008Z digest=sha256:ad224b18abe1a95c8b3b3bd052dd01ef079bb74eaa75551739693419ecdb11e5

Observation 56d94836-4c1c-4215-a5d3-0db83908b2d9 · outbound

This paper cites Leveraging large language models for nlg evaluation: Advances and challenges.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Leveraging large language models for nlg evaluation: Advances and challenges

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.563205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.563205Z digest=sha256:da374c7620a53dc9cd0329b83897e301e6c9ecea9cb48fc579fe1563c78db121

Observation 7754d152-ee3c-4cbd-9714-5689e29767f3 · outbound

This paper cites Let’s verify step by step.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Let’s verify step by step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.625816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.625816Z digest=sha256:cd946489ab222bcb289efe3e4f0ddc48860a219e1f0edad6bf1bdd96b842e824

Observation 9dce8c3f-afdf-480f-b6b9-62f1acf23957 · outbound

This paper cites Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.739707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.739707Z digest=sha256:088cc382c86743cb18ac70fd1d627255fce0d2bcfedd1f587c0c158c4c8b2db4

Observation 9c272ff5-c8e7-4443-82c4-911beb380184 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.835725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.835725Z digest=sha256:213ad8d15b07954ba99aa849499473fddfe7fc33e067edc773e9a798853ca86f

Observation e890901b-c91f-4e39-acfd-1e255458360b · outbound

This paper cites Introducing llama 3.1: Our most capable models to date, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing llama 3.1: Our most capable models to date, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.966371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.966371Z digest=sha256:472a4c85c9911682d3c994c51f2b6b427aa9804d66f92042bcf9640d4d628936

Observation 9f9b52ed-450a-4b93-9159-ef8f1d1a6967 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.051388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.051388Z digest=sha256:1d8e2c6e01b646660967b4e7e88402a27ac23f2239368592ef75228de006f5f7

Observation ec503b05-2bc5-4edb-9c9c-ed9b724b7cb4 · outbound

This paper cites Evaluating language-model agents on realistic autonomous tasks, 2023.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating language-model agents on realistic autonomous tasks, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.145765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.145765Z digest=sha256:b666c9abc2a9d40ebceda6e35c47a90cb63f4b83baa7ebb7626ee7973b9aab8f

Observation 28c66277-51aa-4e67-8bf9-db61072802e6 · outbound

This paper cites Example protocol for running an ai agent evaluation, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Example protocol for running an ai agent evaluation, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.235381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.235381Z digest=sha256:3fc74783951d7c3ace2a26c7251b0ffb9d39a0535e75765da2ccbc6c7db0d146

Observation 93b5332d-9b40-4d92-ab18-7ae10a98af38 · outbound

This paper cites Measuring automated kernel engineering, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring automated kernel engineering, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.324881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.324881Z digest=sha256:6141a59700d6d30e26f31c8b1e5ce3b4646acd729334430c899297409d0f3cc6

Observation a356301a-7ba5-4f73-98eb-a48ec47c1579 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gaia: a benchmark for general ai assistants

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.418642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.418642Z digest=sha256:b6a6854371656649a4a26a386c97afae49661cf6fea596a441d01d0e1f6c125b

Observation 6d5570e0-69c3-463d-bcd0-5dcb2af1dcfe · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Establishing Best Practices for Building Rigorous Agentic Benchmarks SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.518781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.518781Z digest=sha256:7f4657b15ec91f9c03340e58b8af9114c2e0d5b59e7a644ce89232f2f217f34c

Observation 78279a50-f525-4cc2-a122-712c5c892bb7 · outbound

This paper cites Mixtral large 2, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Mixtral large 2, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.571947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.571947Z digest=sha256:cae296fe2cee5cc105b2847f15c998df299a19470d4ca78c654b775ddf3e768e

Observation d6f9dd39-5c42-497a-98f3-d3b1cb4c5ea5 · outbound

This paper cites Preparedness framework (beta), 2023.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Preparedness framework (beta), 2023

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.632617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.632617Z digest=sha256:fc88c7f35899a1afbb50055051553b1c81901c43c2fa022919b991ba5834be76

Observation 1f21e953-864c-49ee-ab92-b677d6ad2eeb · outbound

This paper cites Gpt-4o system card, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt-4o system card, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:40.110959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:24.756415Z digest=sha256:2344cdebd9a9353ca96dd60ffc2af1964d283f7cb12e0af862bec5adad2306e3

Observation 1009fd69-fc03-4e62-ba60-ceb30aa5ad3a · outbound

This paper cites Openai o1 system card, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o1 system card, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.825783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.825783Z digest=sha256:9806c1002243f5e86b79cfa83c30c17d8aa47603e5c09c819fbcd12c01f31455

Observation 54095ed0-51e5-498c-8269-f597bd1b563a · outbound

This paper cites Openai o1-mini, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o1-mini, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.935710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:24.937917Z digest=sha256:2fde3d64120f600b85881cb7da9ab2561a6815307599ce2deebcb3e38efe51c5

Observation d2e2abee-bbef-4aa6-9b39-bb8f5c230a7b · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt-4o mini: advancing cost-efficient intelligence, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.740345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:25.016609Z digest=sha256:f691f597f4ab044b463348b41f5c127e0aefb66ce95d040eba4214dbf5351cc3

Observation aa6d0ab7-267f-4046-9da4-9fa7af0b1d45 · outbound

This paper cites Computer-user agent, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Computer-user agent, 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.565147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:25.134135Z digest=sha256:e302045633c7e8c1b1ce0a37ef12a4b7850e162c944d60fb468862d296b0ebf9

Observation 2c55a375-688b-4b5b-9074-2484f531c990 · outbound

This paper cites Introducing deep research, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing deep research, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.231381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.231381Z digest=sha256:f4c93420ce3f9950a17595fc34e95a259318a14bec4c498a928576071ad41e24

Observation dbf002cd-0fce-4028-a4b0-39c10a8248ef · outbound

This paper cites Introducing gpt-4.5, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing gpt-4.5, 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.386566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:25.310506Z digest=sha256:69832a4a660d80f4b8a357360a7c037d9aa3ccd5f30feeef37e5b21af26c9827

Observation 70b51f70-1084-44d1-a578-b7718c45ae58 · outbound

This paper cites Openai o3-mini, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o3-mini, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.423058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.423058Z digest=sha256:b23129b2394b82be52d41d893abffd409f1b65648a13aafee033fc1ad330d8ef

Observation e3e8824d-5612-467b-953a-585dfe728c2a · outbound

This paper cites Kernelbench: Can llms write efficient gpu kernels? In ICLR 2025 Third Workshop on Deep Learning for Code (Best paper award), 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Kernelbench: Can llms write efficient gpu kernels? In ICLR 2025 Third Workshop on Deep Learning for Code (Best paper award), 2025

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.204719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:25.505512Z digest=sha256:2366380a60d9ed7f3313c3b85646aaa2977a3e0958c9bdab3f40ca37692554b8

Observation e6cdc13f-34aa-4119-84d2-630125c893d2 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Bleu: a method for automatic evaluation of machine translation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.597919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.597919Z digest=sha256:954d437b737ef542cc27ac36530b523d8757cc686364a99e97ab8f498ea4d968

Observation 936e1ddd-4052-4ee0-884d-8a68c49e34e3 · outbound

This paper cites A survey of flaky tests.

Establishing Best Practices for Building Rigorous Agentic Benchmarks A survey of flaky tests

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.047902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:25.709063Z digest=sha256:c5cdacf1e566251f74856a981a87f2e67ad8bd797a248ed2cfb73700de276799

Observation fba94f27-761d-4acd-a0d0-017ad16ac8c7 · outbound

This paper cites Evaluating Cross-Domain Text-to-SQL Models and Benchmarks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating Cross-Domain Text-to-SQL Models and Benchmarks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.823326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.823326Z digest=sha256:e8386f61aca990ffd9a666a333192de5984ac3741b0d386fb468404ff86cd3cf

Observation 99e58958-ce9a-4adf-b53e-55b5051e5786 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

Establishing Best Practices for Building Rigorous Agentic Benchmarks CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.910315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.910315Z digest=sha256:02ce786ad2e12295c1e3a4989b8e27b0d47b01883e1442a2674357e952470947

Observation 0f260b87-08dc-41ae-bf50-586b86aefa6f · outbound

This paper cites Ai and the everything in the whole wide world benchmark.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Ai and the everything in the whole wide world benchmark

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.849661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.004695Z digest=sha256:a959d6de9b0dd4e3f7f9bb86deb7025840e771e180f6f4993853afe5128237bc

Observation 7a9bd51b-a180-4b3e-a5cb-1f3ab7cd18ad · outbound

This paper cites Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.647334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.112012Z digest=sha256:abfffd7bf70917874d5821039f8d38252c2d5a364653a91c17a40bf9056b47f0

Observation 1c2b8a5e-347e-4952-a339-46d484dbfae8 · outbound

This paper cites Analysis and testing of web applications.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Analysis and testing of web applications

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.455331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.208330Z digest=sha256:ce66270eeb1a2a0af9a2028d31818570fbc8c8f0e7a09706a4ebe443ae6cbd81

Observation 8bfbeba8-34bc-4404-8c84-a03a88d22991 · outbound

This paper cites A survey of unit testing practices.

Establishing Best Practices for Building Rigorous Agentic Benchmarks A survey of unit testing practices

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.245734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.295615Z digest=sha256:2127c5a568960b1bf8c2ce4ad207a87498f1d99af79f2c5de021ed9874c8b209

Observation adff7b59-e754-4dd2-b0e6-4130982c848a · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.017958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.364230Z digest=sha256:5fc95b175e230a265eafee788936088f0caf5e8aa1ab27fae2aa620417f1aad9

Observation 100925d5-acb5-4e73-a608-94591c4751d9 · outbound

This paper cites Strengthening ai agent hijacking eval- uations, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Strengthening ai agent hijacking eval- uations, 2025

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.816267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.483123Z digest=sha256:94a12c9bd333f86df8b5a5266d630db575e8fcee428e5a60dc14e901b2868fa9

Observation e2b3a539-22ca-43d0-911a-53dc5759bb66 · outbound

This paper cites Inference scaling flaws: The limits of llm resampling with imperfect verifiers.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inference scaling flaws: The limits of llm resampling with imperfect verifiers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:26.559726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:26.559726Z digest=sha256:fe95008e1f35cc2352c65e3a375c61059013e54ee1af58dd14714b3dbafd84e0

Observation 42f6f29f-a16c-448e-bbd4-d2fa8b248c44 · outbound

This paper cites Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.604621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.685912Z digest=sha256:d11e23feee622e1db9293df766385c40e34ea0434f529bbdff3e1ef6929fd7bb

Observation 7b3638e3-ecb7-4851-8325-2d8cc4dc018b · outbound

This paper cites End-to-end integration testing design.

Establishing Best Practices for Building Rigorous Agentic Benchmarks End-to-end integration testing design

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.334722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.752737Z digest=sha256:f9b639d552cdd31197c163b66a0a09a5cc5ee615ad7a6e6d0825730fab338a47

Observation a43f2b5b-cef1-4b2c-8f52-c4192b29f12c · outbound

This paper cites From imagenet to image classification: Contextualizing progress on benchmarks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks From imagenet to image classification: Contextualizing progress on benchmarks

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.019808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:26.832387Z digest=sha256:f80861e71a9ef6b84fd081f8eaf05fcda9ea7e5a57790e4f2896937aa0a16cfe

Observation fb92c734-fafe-4a5f-ba81-1a2c4dfd05a2 · outbound

This paper cites Introducing v0.5 of the AI Safety Benchmark from MLCommons.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing v0.5 of the AI Safety Benchmark from MLCommons

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:26.965009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:26.965009Z digest=sha256:7d58f82fe3dd7eb3458bf9afd66620d532f97d5b56f05c0b72f18668107c8eea

Observation 0a0fea27-b91a-4cf4-abb1-cb3662ed0c6d · outbound

This paper cites Evaluate & evaluation on the hub: Better best practices for data and model measurements.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluate & evaluation on the hub: Better best practices for data and model measurements

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.730961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:27.071959Z digest=sha256:ef021fe19fb7e88a975a3bf20de7942c16aa454cc8932548e94a660347527297

Observation bfecfd29-8d68-491c-9243-8e1afe8547f4 · outbound

This paper cites Structured testing: A testing methodology using the cyclomatic complexity metric, volume 500.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Structured testing: A testing methodology using the cyclomatic complexity metric, volume 500

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.494484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:27.172490Z digest=sha256:76d71e472257322812919e8aff958dd98d509a50d0c810181ca7d6b6377dfa1f

Observation 7e007b0c-0b86-4d42-94b7-84dd4f755e1e · outbound

This paper cites Measuring short-form factuality in large language models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring short-form factuality in large language models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.263615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.263615Z digest=sha256:8bf9a6642d94d2038b840f99576d70832748c54032aef1b4d65cd019bf663091

Observation 1211ce8f-51ba-4ab2-9e95-f279f69a42df · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Establishing Best Practices for Building Rigorous Agentic Benchmarks LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.342525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.342525Z digest=sha256:75e0f6db8377edb8c2734152db8049229cd33d605a7f1e62171f77c7444aea50

Observation 36100b82-0a6a-453d-8fd8-7252d4bb2915 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Establishing Best Practices for Building Rigorous Agentic Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.421486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.421486Z digest=sha256:0836b2167a45975227309c928cf5461eaa545099c45ae5edf717c3afb87e56bc

Observation f034431a-346d-4b7d-a505-30bc5b871529 · outbound

This paper cites Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.539316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.539316Z digest=sha256:6678fd62587cff48aeaecab4be0d81338c126544b3397099087cbc5679ad798a

Observation f10c0ec6-4971-4d89-a801-e30fc1a929be · outbound

This paper cites Grok 2 beta release, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Grok 2 beta release, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.319408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:27.664743Z digest=sha256:a6d0db1bebbeba686964240b95db2b09a036140f3b87a22d64b03d7ad1c114ee

Observation d6ebfa40-b878-4f3b-9a60-21e7ff5c8832 · outbound

This paper cites Grok 3 beta — the age of reasoning agents, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Grok 3 beta — the age of reasoning agents, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.151243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:27.758444Z digest=sha256:9466ec1b2d2e9e1d884a6feecea86bb673c239012d69f28c1210d3b09e80a3f7

Observation 9792836c-6497-490d-afd8-dc95cbbf3c57 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.822200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.822200Z digest=sha256:9e6b62b83b82e2877042319530144e682176cfa79deaa9f0e092b320eb870b2e

Observation d4bb5d99-1118-437a-a708-e2affa5fc08c · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-agent: Agent-computer interfaces enable automated software engineering

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.927688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.927688Z digest=sha256:b60296714bba3e20151083528e5c99cd88928f403628ccfc680cff6397dde5e8

Observation 2c8fb028-6fbe-4b09-90f3-b6b4cf0f194d · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks React: Synergizing reasoning and acting in language models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.984396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.984396Z digest=sha256:62f47f07b0d06bd7b4b683b565ad2a269fd98bac2b26fc53142f4f9275e8d100

Observation 198fd22f-9441-476c-a031-85aa2f20fae5 · outbound

This paper cites tau-bench: A bench- mark for tool-agent-user interaction in real-world domains.

Establishing Best Practices for Building Rigorous Agentic Benchmarks tau-bench: A bench- mark for tool-agent-user interaction in real-world domains

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.992119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:28.051730Z digest=sha256:d5712101cc946864591a574c8d3146ec577c5ab7d8d151a445e1a025a526f68b

Observation d949f052-c124-4d97-82f6-19578d246cb3 · outbound

This paper cites Utboost: Rigorous evaluation of coding agents on swe-bench.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Utboost: Rigorous evaluation of coding agents on swe-bench

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.844195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:28.132834Z digest=sha256:96fe650e2e0983be1a6320ef970fbf6ce0673eb8d7108eee31f699754688c993

Observation 1a034863-6f59-44c1-a9f0-367e722d3b8b · outbound

This paper cites Evaluating large language models at evaluating instruction following.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating large language models at evaluating instruction following

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.693140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:28.231761Z digest=sha256:fc6198f63e0823a1f3ed0c58c00e11ff39efc4637a9d5482f238e8571b0f9379

Observation cb576ed1-53fe-4da0-ab9c-9c13e0aa1f7e · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.287235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.287235Z digest=sha256:5aa8b5518a39645bb991ea2f0b096ad2b5309a4e207fa9280463e1b1bb6936f6

Observation 684b05ef-1caf-4548-8578-8538f64cf57d · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.350185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.350185Z digest=sha256:47aa5ee9b2005e34919a2dd7aa9b8e1753071f4b1bdb2b68ed722a5df3039770

Observation 11652142-6a23-4b3c-b5d8-bee15e27fe51 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.438581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.438581Z digest=sha256:b9c7eb2e51576de466cd0501ee58f29680d59bd2f874e059659c1eeb191e33c5

Observation 4171e63d-1c90-405f-bfc6-697c68ccd494 · outbound

This paper cites X-webarena-leaderboard,.

Establishing Best Practices for Building Rigorous Agentic Benchmarks X-webarena-leaderboard,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.549044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:28.512244Z digest=sha256:14f6f1554b4bb52c3d3d2dba7422e76e83c77fd17c92ea534696c30dced43c68

Observation ef5223c0-1a4a-4a77-9b95-a4220eb1bba5 · outbound

This paper cites Webarena: A realistic web environment for build- ing autonomous agents.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Webarena: A realistic web environment for build- ing autonomous agents

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.173406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:28.701453Z digest=sha256:2c31505854450fee268ef6bc0433910a03353164f6beca540ced673e73e13ee1

Observation 2e34b126-d2fb-4384-ab41-5e41bc211521 · outbound

This paper cites Software unit test coverage and adequacy.Acm computing surveys (csur), 29(4):366–427, 1997.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Software unit test coverage and adequacy.Acm computing surveys (csur), 29(4):366–427, 1997

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.775241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.775241Z digest=sha256:6e0af02e8f68db20695fca77aca3952e9495aaceb41d24ee9ab093f61ee21b59

Observation 221e343c-f593-4261-8244-8f2cd2d7c753 · outbound

This paper cites Fuzzing: a survey for roadmap.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Fuzzing: a survey for roadmap

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.002208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:28.878325Z digest=sha256:90dedfab3ea05373c8724043f2d77133236d3b34548fd4d013c04b8e97783271

Observation a67a9fdf-6975-4f4b-8497-98a52fbdd8aa · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

Establishing Best Practices for Building Rigorous Agentic Benchmarks CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.958503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.958503Z digest=sha256:686c56cdc17fffa36d4b29de56f25140233fddee9320bbd144f9d0da410ea24f

Observation 2e183348-8063-425d-854b-dd8490835ce7 · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Agent-as-a-Judge: Evaluate Agents with Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:29.053233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:29.053233Z digest=sha256:9db19e9bdfd7efd75b8a621a4c449a395bb20e8b7ba3cffcb1fe84522f2ae7e9

Observation f002e32b-3947-4bee-b6a8-e09f77291753 · outbound

This paper cites Can large language models transform computational social science? Computational Linguistics, 50 (1):237–291, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Can large language models transform computational social science? Computational Linguistics, 50 (1):237–291, 2024

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:34.833066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:29.137040Z digest=sha256:3b881a1f72b9e04338f38c37e2ced92e33c5f916fc62c83fbd73b6c1d63bc271

Observation 8e8f9bf9-c634-4687-8e2a-6c2be41be038 · outbound

This paper cites an unresolved cited work.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:24:34.697054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:29.250524Z digest=sha256:e88efef8e3a9cd0fdfb6e5664a921a8c6ea3e2c28fe545467f706f628d86cc3c

Observation 60fb9d5b-7a52-4cc3-9cd7-7be29b252f51 · outbound

This paper cites an unresolved cited work.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work

Reference 101

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:24:34.520858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:24:29.338123Z digest=sha256:b3a37a7f57a52b474d32d8010155551a7fb3eb9166b9373d3107a09bf8813905

Pith citing papers

Observation 66ecafb5-6e96-4e31-a056-9461b2744708 · inbound

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces cites this paper.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:07.992049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:7986b383546b5cd4c793c48ac156f64cde25a6a854f78dbcc54b426a6741de67

Observation a6403359-918a-4741-a1f5-4ba21eaec78d · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:15dc9a56dd874a2bd818ed05fe0aeebee818add4226f90b05b736d816e0f811e

Observation cb12bd46-df97-4812-ba29-41819a3a7b8b · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.219867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.219867Z digest=sha256:300e47f2c0be15b0ffbe9e99a96623d516657f3ee741a71eadfb951160197a3e

Observation 65b39ab3-2771-400f-8549-74a87ee7986f · inbound

From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI cites this paper.

From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:51.819572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:47:09.928166Z digest=sha256:b6a08f116519a15c3698e3dfdb7ef3a171bbd8f3cd310e7fd83aa7b53fb1410f

Observation c1e2136b-abc8-4404-9894-f3aec855f5b5 · inbound

Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI cites this paper.

Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:18.917983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T06:03:19.805349Z digest=sha256:ea3c3eaaeebb5cd0a9c1d0661ae642fe7a970e9a6b5d2cee17c1ace6079b88a8

Observation d47897a7-e545-4666-9bff-02f2187ace23 · inbound

BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists cites this paper.

BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:08.439353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T20:03:59.215134Z digest=sha256:490b105579a49c2d0653445bebe1f0681746cd3fcf1ab28881875473dc83e6ac

Observation 0890bdcf-531d-43e4-ab67-24038f357d5b · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.552635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:04bd71d8458bbebfdf7c33c34ee3f5c482a295bf81542c0f33087432b03817e4

Observation 33d5ddc6-5eda-4558-8d43-3812ad576fee · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:20:54.825579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T02:20:32.528550Z digest=sha256:bb6749df8acb5b910aaef4688907af703006f848e8e12936df6c1c5822ee893f

Observation 20d96c9f-295e-4276-9f73-5c280b826463 · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:49:29.231213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T21:49:16.239350Z digest=sha256:fbf49c3b627061880d9b88e85f004e25a3b39d2f9fec6c5b4ab9b8599419d93e

Observation 852088d3-e55a-41cd-be32-2580980085b6 · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T00:18:46.882265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:18:46.882265Z digest=sha256:e89ce99a26b2a9732f43f1f97d48e52b1f7ab3a606ccd6c4f8f2534f0993d13c

Observation de977e8c-5c6c-438f-9eed-76228711bbfc · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.103689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:56a35ff2f8ecea56c245169157be2392d327236d777068cf6a58d4b1e3f399fc

Observation 927faacd-9882-41fd-a98c-8bd197aaf590 · inbound

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation cites this paper.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:24.002297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:b6601f33a4d8b92e972ff0b07e313c23c6d02a94f195ae533058129c2b9fb9f8

Observation 54fc18c7-13b3-4f01-ba8a-8a9115d6c13d · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.775266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:d27f07829ef48eb1b27fa03038b45ac9491990575272c406becc6647d2f580c2

Observation 439b1478-b2b6-49fe-b683-513b6b3c67d9 · inbound

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle cites this paper.

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:37:35.612990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T18:34:39.997353Z digest=sha256:cfbdc62a5d83e557ac4bf2ea9f176da4c0aba8adae930f0ec6896a5fdcc3e425

Observation d7d74fb1-4748-4cab-b90e-c60c8b55a756 · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:05:06.699293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:adb78193eeddd21a75a990329b7d8164e508f8237b76a3e1c2b1ca6498b08a03

Observation b70d3b7a-b3ad-4582-9902-4c5ae2cc26c9 · inbound

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations cites this paper.

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:26:09.995466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T06:25:06.024542Z digest=sha256:e16b6876d9d2a4f5b6f8f7b4e7e695c359dea4ba74559c82aa2b4f744587a4fa

Observation 222d5930-73f8-4e41-83f8-2218daa2ec88 · inbound

Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard cites this paper.

Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:04:37.290934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T05:02:31.995894Z digest=sha256:1f3761cd563ebb04ba56f87718a7e358ae5077edee131b35414c4884cd85302b

Observation fea22811-59c0-4626-9166-913c7c9b35e2 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.934723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:c3524292ea8192564e85d539035c34f96852379d47d13b71c2feaa758d2ed600

Observation a6c31d99-04db-4f81-9ad6-23df445ec369 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.120173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:6f0471e92ca1989270daf4cbcab87ab870baeec4ca73012ebcfff91b50911b15

Observation 5ce31bad-c793-4d0b-9344-a1c96088f20e · inbound

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly cites this paper.

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:23:58.970057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T21:19:59.005348Z digest=sha256:0677e013e91895ba25ba8d3384fe4854e157e9512b6d861c36cf7f518a63e037

Observation f77ef2c1-260a-4d68-925f-e5722106874f · inbound

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation cites this paper.

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 21

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T21:33:59.135091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T21:29:57.626069Z digest=sha256:e672258d004e6c2532247deb69bc44fb27c8b6817698d7072c528158800e44d6

Observation 366d0cce-affa-4028-baef-5e83e2554545 · inbound

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents cites this paper.

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:36:52.845154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T04:58:10.803420Z digest=sha256:620fba46a38aebffdbf462f4f1576c0f348d87d14c38197f837bc2cf8a38d02b

Observation 9390dd14-4f0c-450d-b790-a0e734359b67 · inbound

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework cites this paper.

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:59.932826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T23:42:20.304205Z digest=sha256:156f3438e11d1282583d037eff694639f21ef8a8b28173045b4222ed7b4e3be6

Observation 99f7b5c4-9deb-474e-9b48-e3d3354158c1 · inbound

SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills cites this paper.

SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T00:59:52.952364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:59:52.952364Z digest=sha256:89945c01154a9ba76d664209faecfd4a344d3e0b9fe394bc4316700ccb4b3c89

Observation ad9430a8-1636-4380-a69e-32bdca129aca · inbound

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning cites this paper.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:34:10.122008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.122008Z digest=sha256:4d6a247d90a616f0674df419757e9fb9bfb8e65d18fd085dbaa9acd1ab9ffde7

Observation df467229-2c23-4b23-8ba3-c1307cd628b6 · inbound

KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch? cites this paper.

KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch? Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:38.648918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:00:38.648918Z digest=sha256:f99cd10b955240605370776660f97e080678db7cca1dc23cd97689eac538a9aa

Observation 4f0e28b2-9010-4193-a5c6-d8e3470bbf1a · inbound

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery cites this paper.

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:09.728336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:09.728336Z digest=sha256:2d5575285f84108ba4b8dc38aad87c9b508cd1052c54a28576889cc90707788a

Observation e573652e-ca9f-4e82-a458-090e7836e591 · inbound

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems cites this paper.

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:38.319817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:54:38.319817Z digest=sha256:b51c57ba670e13a745c177cc1950e1bab223606c5987149d5eaaa9691d600844

Observation 6f262067-0274-4d99-92e7-06ea467ed41a · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:28.621324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:28.621324Z digest=sha256:e60beb15ef192e479d2bb899aaa5137dddaf775b99ce1a3b5da645d66c282eab

Observation 65fb692f-d7ab-4f20-8921-75e04f349ef3 · inbound

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents cites this paper.

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T23:58:27.966370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:58:27.966370Z digest=sha256:11e0a43c3fc4f8723ae633c739d99c6eea0cbedb69f1273647c539dd7ca30dad

Observation df392865-68cb-4dc7-9285-56fc59986a67 · inbound

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks cites this paper.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.912539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.912539Z digest=sha256:375fb7fe37eb21f909e21883093a2f4549d70dba1df17716461981095c286a67

Observation 897a0c3a-b537-43ee-b4ea-d7abc981c559 · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:08.204255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:08.204255Z digest=sha256:bec8754ba92a2fcae374daa49a8964609a87e72858540f8eb590b877db86b476

Observation 1d6a3c2d-ea8a-4eed-8871-45486bfd7d0f · inbound

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation cites this paper.

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:48.271679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:48.271679Z digest=sha256:e898bc45f1530d2877ea320854bab172cc2ecb51866ca902c40bdc9ffac2f54e

Observation 7738450d-3702-4acb-a2eb-5776ee593f3e · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:18:50.856828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:18:50.856828Z digest=sha256:9edb3e4669246f9139b029680f3dcb965608f856507c9a5d9f4e2bf63d5b1c40

Observation a2274c8f-1274-4080-9b81-a586d4d723cb · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:24.726546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:16:24.726546Z digest=sha256:b9dcbc6808d07144b8afd3298939f001a20681258909350e2172a97ba3b166ad

Observation 0cf16eda-132d-4502-aa08-c8cad43c3488 · inbound

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents cites this paper.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.480747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.480747Z digest=sha256:1bc5fe09b16fdbee904eca071fde446b7f252daed5bc1b04d695c467d80e9760

Observation e2b186c1-b8e7-4440-af39-d8191b3b9363 · inbound

Agent Safety Should Be a Runtime Contract cites this paper.

Agent Safety Should Be a Runtime Contract Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:14.004821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:14.004821Z digest=sha256:4153fb890e90582dc3c724687ca7cdbd9db4e0a52812c7eb2eb7eb1d35a5d29e