Pith. sign in

Paper Citation Record · LEDGER

Establishing Best Practices for Building Rigorous Agentic Benchmarks

As of 23 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 37 inbound Pith citation observations for arXiv:2507.02825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02825 v5

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:29.338123Z

measured 137 of 137 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:19:14.004821Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:08:59.930877Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved73
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 498da9e5-9da9-4859-a1f0-e2235cb99aaf · outbound

This paper cites Inspect AI: Framework for Large Language Model Evaluations, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inspect AI: Framework for Large Language Model Evaluations, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:19.476468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:19.476468Z digest=sha256:f8419032292d4a9670a83e30ce16814fc55781d19fcdf36b3a57b255ac3a5051

Observation 5aea9622-5a6f-4b7e-87c6-96ad04a561b7 · outbound

This paper cites Gpt code editing benchmarks, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt code editing benchmarks, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:19.643317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:19.643317Z digest=sha256:96a7edbe0b08393d9e87da2dff77e63551a2338d26fb72e342e402c0a99f239d

Observation e9a79027-ed54-47d7-a5ac-f01f5ccc7ccb · outbound

This paper cites o1 tops aider’s new polyglot leaderboard, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks o1 tops aider’s new polyglot leaderboard, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:19.860736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:19.860736Z digest=sha256:62b7bff648a68cc5ec8bbab364c07f049d61f12519887e8224b6b956eb13bdee

Observation 0b486f1a-5d0e-437d-962b-35608d837b1b · outbound

This paper cites The amazon nova family of models: Technical report and model card, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The amazon nova family of models: Technical report and model card, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.062843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.062843Z digest=sha256:53716c2b3b960aa5d70949ae83019a000cde24792315e3582b4d956351e4eb00

Observation 1d636119-8fda-4571-83bd-dc153e8716cd · outbound

This paper cites Claude 3.5 sonnet, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Claude 3.5 sonnet, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.190930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.190930Z digest=sha256:78d884f6dcb0fe57ad2af5b2e877b9fbb685a7771ae2053446e8f1b25b241e55

Observation 1623b2cf-dbaa-48cb-80e9-ba7df2364368 · outbound

This paper cites Claude 3.7 and claude code, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Claude 3.7 and claude code, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.342592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.342592Z digest=sha256:ffc89ae7bc91263a1361ce80439439240c7917c34917997d9af03c476dbe64f0

Observation d89439eb-2c9e-43cd-86d2-2223c0d8a01d · outbound

This paper cites Bird minidev - corrections, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Bird minidev - corrections, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.521292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.521292Z digest=sha256:45f446ed4a81faab499cef2a701b9e118b14a753ff4a3a0aacb310526efe5693

Observation 59767821-da00-4625-abeb-e8755f7f14cf · outbound

This paper cites Program Synthesis with Large Language Models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.658359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.658359Z digest=sha256:c6092a3b71847b2a436deac603657b9239051db4557ee2c1cc7e2ce6ff9f52f2

Observation eb83d160-535e-4026-9385-50edf7de0464 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.815544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.815544Z digest=sha256:9c67c2bf92802f5d6da8d806c44a04830181641cd907f029aff7eb33439ad266

Observation 58917dc5-a8db-4654-8d3a-2a65a00bfc8e · outbound

This paper cites Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.949000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.949000Z digest=sha256:d3340dc9a518a713e1daed76a5b1efd7bd9c5b131046cc1a3ddbd4fbb54186f6

Observation c3080b94-bee8-43c4-a758-4a0b10a01473 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Establishing Best Practices for Building Rigorous Agentic Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.069258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.069258Z digest=sha256:75d7cb673478ac8a2383df7344184324eb581b655a085b9c1884b1b85b12922c

Observation 260b815f-fb0f-4913-ab43-8dc7a875ea78 · outbound

This paper cites AutoAgents: A Framework for Automatic Agent Generation.

Establishing Best Practices for Building Rigorous Agentic Benchmarks AutoAgents: A Framework for Automatic Agent Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.195372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.195372Z digest=sha256:dea761d11f4d936164eb4511efcb3c7a5586fd1d91c76eb7f8810cb0afb5dbee

Observation 5bccefc2-874b-4a7b-9091-43df7c83c51c · outbound

This paper cites an unresolved cited work.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.316891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.316891Z digest=sha256:69363c34f7d044dddbf8fa43f6198cc264e4abdf7786501692947ed0d889512f

Observation 78deee31-97dc-42af-af85-ad8058691468 · outbound

This paper cites Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.435890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.435890Z digest=sha256:2378f8b3c0e73ae18fbd2a04aec9f1bb4d4a59a9f2ccd77826c64f736100ba01

Observation 0a80de66-d368-488c-ac16-16624a5bda46 · outbound

This paper cites Introducing deepseek v3, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing deepseek v3, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.523208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.523208Z digest=sha256:dd6fcd6610003929b10f525351a22cb36654a8438241bc6061c534afc94d6cae

Observation 7c7a4f70-66f9-45fb-8a76-6496e199b5b6 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Imagenet: A large- scale hierarchical image database

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.651866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.651866Z digest=sha256:9f8f1e4ad1f287703ee10eefe6d3f8a1eb51df26273686203a9987b1e4794b07

Observation 3e3cf797-2521-4f5a-9c09-cc752903aa01 · outbound

This paper cites Don’t label twice: Quantity beats quality when comparing binary classifiers on a budget.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Don’t label twice: Quantity beats quality when comparing binary classifiers on a budget

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.733040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.733040Z digest=sha256:604d4ff07b9d91a3eb9580d79451f4bcb26476b66ffa566dc54435417eedd8c3

Observation 9f23249f-8e95-4619-8525-f6f7f568a9a5 · outbound

This paper cites Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.837718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.837718Z digest=sha256:0b2913febce1eb08778dff74b9aa8749489549cd5b636a1349688458e0781ce2

Observation 03881fb7-d8f6-4f03-b3eb-8aae63ca2da1 · outbound

This paper cites The design and operation of CloudLab.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The design and operation of CloudLab

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.946111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.946111Z digest=sha256:b88b4da5d4feb8820a81f50959555eed85c84c57ed8f4bc5d83b037b90fef3f3

Observation 06be70c0-2a6c-4b6d-9947-589ce3f0d3ab · outbound

This paper cites Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.007624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.007624Z digest=sha256:708557a8f7a267f378518f564c88b2729cb4d9aea8afc35595b5cef34bf79233

Observation b9413ed1-7af1-4651-a911-5f97e31cd44d · outbound

This paper cites Searching for computer vision north stars.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Searching for computer vision north stars

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.076812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.076812Z digest=sha256:855db5a20706b6d8bd23e3fe8d2804af3d1829392765698872d544eeb3d7e43a

Observation 2fc04ec4-6484-4b78-9075-bd038d42f644 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Establishing Best Practices for Building Rigorous Agentic Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.159924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.159924Z digest=sha256:b0d05665276f2f25d82e43313d4dbd2117841dcf9696d03bc0a4bae7ab9d257f

Observation 9d379cf5-0b6c-4177-a756-93c27c27eeeb · outbound

This paper cites A classification of sql injection attacks and countermeasures.

Establishing Best Practices for Building Rigorous Agentic Benchmarks A classification of sql injection attacks and countermeasures

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.247550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.247550Z digest=sha256:77a9438224152498c5e2bc6cabc19fce595aeeb5a63d95b5948f76e20af4b3cf

Observation 4706c5b4-9b26-4c38-8cf0-7db4d811c7d3 · outbound

This paper cites More than marketing? on the information value of ai benchmarks for practitioners.

Establishing Best Practices for Building Rigorous Agentic Benchmarks More than marketing? on the information value of ai benchmarks for practitioners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.337945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.337945Z digest=sha256:be6d622d9bfe3f5514a03c9883da4514031b09c843ec9316664970059a335537

Observation 8e102cf6-4a50-42e7-9ccf-abe7fefd4706 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.424552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.424552Z digest=sha256:9dae6206c65c0dca03ea16eb2e7cd6f6a61956b676c94238556dbe63014140eb

Observation fdc8141c-a335-4548-9956-ecba7fc35208 · outbound

This paper cites The extent and consequences of p-hacking in science.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The extent and consequences of p-hacking in science

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.505161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.505161Z digest=sha256:95a5d303b05e5f819c4966e37f4113f0731817016914e00c2958903f2842601b

Observation 36a63e39-97b0-4b90-995f-a6e0a43c8980 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.595920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.595920Z digest=sha256:0107293de2e325dea679ab0e10762b198e95d00f2aa8f7c74b8295e9dce30d01

Observation 638e800b-6160-41d1-ba91-ddee087bb69f · outbound

This paper cites The design and analysis of benchmark experiments.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The design and analysis of benchmark experiments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.654448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.654448Z digest=sha256:e53c93b2921b76575ad52fecfe30e909cfc895b64f905d9c7a530d503772a479

Observation add6f924-0680-4f83-94e7-1440ddcdaba8 · outbound

This paper cites Preventing server-side request forgery attacks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Preventing server-side request forgery attacks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.744781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.744781Z digest=sha256:e12b92e2f319cd088965bd2049f48a2ee6f6d8d8a865365e28f7f480c317ad96

Observation 6739a092-f9f4-4199-be3a-b4940e5bab8f · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In ICLR, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-bench: Can language models resolve real-world github issues? In ICLR, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.816217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.816217Z digest=sha256:65fd78c218a9735ae25cbc96d61f6420a4d8453d42c109148bc94f750a4993b0

Observation 79952336-47d2-4a0e-a240-a8843eeec48a · outbound

This paper cites Swe-bench verified leaderboard, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-bench verified leaderboard, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.883514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.883514Z digest=sha256:6790ff6b08232b9303bc5a115b9fb0a5df345b40f43624029cceb279f9381ca5

Observation 9be6db01-506c-4190-b663-1296790ee689 · outbound

This paper cites AI Agents That Matter.

Establishing Best Practices for Building Rigorous Agentic Benchmarks AI Agents That Matter

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.949724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.949724Z digest=sha256:c6aa9791a4498819f99229b0e9d62d36ea5ce54d1651c5d0966528e06e3ebd96

Observation 122680c8-e8f9-4bcd-9ecf-6ea0d54319ac · outbound

This paper cites Gemini 2.0 is now available to everyone, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gemini 2.0 is now available to everyone, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.020325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.020325Z digest=sha256:ea22366b11e222a76106bac02a4cc7a02f622d6ac09c9e0cde8f3c90a378a3cf

Observation 1c028daa-0318-4b18-9949-f93bc0f69941 · outbound

This paper cites Math-verify, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Math-verify, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.130120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.130120Z digest=sha256:eb5a42b372ebf63fb448c104b18c20f7c577d0e327c7190512ae699eb8e9b3d3

Observation a4b01508-7d1a-4ad4-af77-594e0d319e56 · outbound

This paper cites The ai cuda engineer: Agentic cuda kernel discovery, optimization and composition.

Establishing Best Practices for Building Rigorous Agentic Benchmarks The ai cuda engineer: Agentic cuda kernel discovery, optimization and composition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.259606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.259606Z digest=sha256:6373d8f9a6cc97976db54fd025d7c832265774ed58d4629ed5a5d4f2ddf64ae4

Observation 592213bb-34ee-436c-ab07-fbb055fc2201 · outbound

This paper cites Challenges of end-to-end testing with selenium webdriver and how to face them: A survey.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Challenges of end-to-end testing with selenium webdriver and how to face them: A survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.335796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.335796Z digest=sha256:a7e31b94fd99ae38d6cb435c17b471883744ef4bba5ff49e3761c9bbd9295be3

Observation 8abab65e-cefc-4dcb-8bcf-2afa33484a64 · outbound

This paper cites Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.446008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.446008Z digest=sha256:d4b7a415d362ceeb2e5ecfeff7150168a4abe3481653b49c246627cc218d849f

Observation 56d94836-4c1c-4215-a5d3-0db83908b2d9 · outbound

This paper cites Leveraging large language models for nlg evaluation: Advances and challenges.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Leveraging large language models for nlg evaluation: Advances and challenges

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.563205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.563205Z digest=sha256:6499f602a85e2da8cbf0d6269c4486f0abd84ffdba0cb01edf69edeccc94f501

Observation 7754d152-ee3c-4cbd-9714-5689e29767f3 · outbound

This paper cites Let’s verify step by step.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Let’s verify step by step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.625816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.625816Z digest=sha256:4be9f48d1f3c7977b95c89fe0160b2e7bb3db65650ec0b77e7430436869905c5

Observation 9dce8c3f-afdf-480f-b6b9-62f1acf23957 · outbound

This paper cites Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.739707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.739707Z digest=sha256:f57592f572f4e307849afc9228221a6b8fe59e36d62955d71d60c963ee6be052

Observation 9c272ff5-c8e7-4443-82c4-911beb380184 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.835725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.835725Z digest=sha256:78d8d309db90e4defb88af6679eb13069eaca3dae78b02749159a9ae79e76655

Observation e890901b-c91f-4e39-acfd-1e255458360b · outbound

This paper cites Introducing llama 3.1: Our most capable models to date, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing llama 3.1: Our most capable models to date, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.966371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.966371Z digest=sha256:3c0d92d9f99e1cab9a43da6880a6a0c1392a02ad7c15eb6d2ec4143f78db5fbc

Observation 9f9b52ed-450a-4b93-9159-ef8f1d1a6967 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.051388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.051388Z digest=sha256:ff957b5fa8dc696d058d530c49b5aa6101bec3f6ac7f514f8c53d89df6d76c9f

Observation ec503b05-2bc5-4edb-9c9c-ed9b724b7cb4 · outbound

This paper cites Evaluating language-model agents on realistic autonomous tasks, 2023.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating language-model agents on realistic autonomous tasks, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.145765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.145765Z digest=sha256:53acd9a84ebf6e5c1185d320cf51f53cc1713582606ed2398e1f9902d62cbbc6

Observation 28c66277-51aa-4e67-8bf9-db61072802e6 · outbound

This paper cites Example protocol for running an ai agent evaluation, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Example protocol for running an ai agent evaluation, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.235381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.235381Z digest=sha256:1314074e66541fdb572463660b14b8c8ef02649dc0f9ad87a62d7d4359c90ad4

Observation 93b5332d-9b40-4d92-ab18-7ae10a98af38 · outbound

This paper cites Measuring automated kernel engineering, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring automated kernel engineering, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.324881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.324881Z digest=sha256:8b76a02679e0e220d23c3e24d0b5ae42b1ba1af1aa093f0f76bc43b92742d35a

Observation a356301a-7ba5-4f73-98eb-a48ec47c1579 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gaia: a benchmark for general ai assistants

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.418642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.418642Z digest=sha256:328d155b76fda965f7a41535dd19c6d6f2d668aa97d5200d8b036ee76049d657

Observation 6d5570e0-69c3-463d-bcd0-5dcb2af1dcfe · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Establishing Best Practices for Building Rigorous Agentic Benchmarks SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.518781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.518781Z digest=sha256:1c72559e37794beee7a9ec60136cf0e69556d4cc4916b3d409497c183fc9b29b

Observation 78279a50-f525-4cc2-a122-712c5c892bb7 · outbound

This paper cites Mixtral large 2, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Mixtral large 2, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.571947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.571947Z digest=sha256:92d25f24fef976aa68daf829b8d267bcfe86ec6a62451f7fddee70ff20884f99

Observation d6f9dd39-5c42-497a-98f3-d3b1cb4c5ea5 · outbound

This paper cites Preparedness framework (beta), 2023.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Preparedness framework (beta), 2023

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.632617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.632617Z digest=sha256:57fde884f8ef40ac3ce75dbe373444a7def9edbc50565d1712fd65a34448f2fe

Observation 1f21e953-864c-49ee-ab92-b677d6ad2eeb · outbound

This paper cites Gpt-4o system card, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt-4o system card, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:40.110959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:24.756415Z digest=sha256:d70e0ba5272d6635ab013f45e4c776b5b968a1e9d8ec57c034ba8abc03d5a7f8

Observation 1009fd69-fc03-4e62-ba60-ceb30aa5ad3a · outbound

This paper cites Openai o1 system card, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o1 system card, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:24.825783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:24.825783Z digest=sha256:7102627c3d30e3048a24d14827b3049737c58d9c567d27776e152d278aae0e09

Observation 54095ed0-51e5-498c-8269-f597bd1b563a · outbound

This paper cites Openai o1-mini, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o1-mini, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.935710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:24.937917Z digest=sha256:d4a5e60c538296eb45447ec6bcd8f0c19946f4e9781d0193ff7433c4eac1a78b

Observation d2e2abee-bbef-4aa6-9b39-bb8f5c230a7b · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt-4o mini: advancing cost-efficient intelligence, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.740345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:25.016609Z digest=sha256:fe7d6cb42b4c51f8ab4892f3f6dfac07e10ae969244bd84e30b282f75e750dec

Observation aa6d0ab7-267f-4046-9da4-9fa7af0b1d45 · outbound

This paper cites Computer-user agent, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Computer-user agent, 2025

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.565147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:25.134135Z digest=sha256:a9753aaeae98e3d919a181bcc5475ed0631b127687bc528e56dba52b542491ab

Observation 2c55a375-688b-4b5b-9074-2484f531c990 · outbound

This paper cites Introducing deep research, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing deep research, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.231381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.231381Z digest=sha256:d65ddf6899642b69e53ef871b6af5dbb8dc92ace89196563ab4f86204b386b4c

Observation dbf002cd-0fce-4028-a4b0-39c10a8248ef · outbound

This paper cites Introducing gpt-4.5, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing gpt-4.5, 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.386566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:25.310506Z digest=sha256:2dc8f438354f49e628702edad9129d59f40082a837fe9a912e48391661ce4f45

Observation 70b51f70-1084-44d1-a578-b7718c45ae58 · outbound

This paper cites Openai o3-mini, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o3-mini, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.423058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.423058Z digest=sha256:b39285829d68ed1daa4d04917cc22068be98e50f6bd094b6da4add7ccc1b8007

Observation e3e8824d-5612-467b-953a-585dfe728c2a · outbound

This paper cites Kernelbench: Can llms write efficient gpu kernels? In ICLR 2025 Third Workshop on Deep Learning for Code (Best paper award), 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Kernelbench: Can llms write efficient gpu kernels? In ICLR 2025 Third Workshop on Deep Learning for Code (Best paper award), 2025

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.204719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:25.505512Z digest=sha256:63158aa6eafe62211121e687cafd746ae9842c3053fef6e912438b08369ab170

Observation e6cdc13f-34aa-4119-84d2-630125c893d2 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Bleu: a method for automatic evaluation of machine translation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.597919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.597919Z digest=sha256:e38ebe01a7f40095c16202c615d50e78317deb30a261f204136f34b2f0a0beea

Observation 936e1ddd-4052-4ee0-884d-8a68c49e34e3 · outbound

This paper cites A survey of flaky tests.

Establishing Best Practices for Building Rigorous Agentic Benchmarks A survey of flaky tests

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:39.047902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:25.709063Z digest=sha256:fd44551bf49232cfe1035070777c366c3eedea9e1c23eddc140b0767ce37dd00

Observation fba94f27-761d-4acd-a0d0-017ad16ac8c7 · outbound

This paper cites Evaluating Cross-Domain Text-to-SQL Models and Benchmarks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating Cross-Domain Text-to-SQL Models and Benchmarks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.823326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.823326Z digest=sha256:a79b5f62cdd564542415278c3d90c09cc4a150771b2537433e02a3691d0bf745

Observation 99e58958-ce9a-4adf-b53e-55b5051e5786 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

Establishing Best Practices for Building Rigorous Agentic Benchmarks CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.910315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.910315Z digest=sha256:73fb96adac3b4a28edd298b9efde7457698c071e89a5f974b79fa822fca9881e

Observation 0f260b87-08dc-41ae-bf50-586b86aefa6f · outbound

This paper cites Ai and the everything in the whole wide world benchmark.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Ai and the everything in the whole wide world benchmark

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.849661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.004695Z digest=sha256:8bb419231e5b9ca7ed8ba2de06acc15f97d416c04e864d8af45df1373c66b9e4

Observation 7a9bd51b-a180-4b3e-a5cb-1f3ab7cd18ad · outbound

This paper cites Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.647334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.112012Z digest=sha256:d5b0114c867be3a50b36dbe9f560e4d268548dad733e5eba7a80b155f2900a33

Observation 1c2b8a5e-347e-4952-a339-46d484dbfae8 · outbound

This paper cites Analysis and testing of web applications.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Analysis and testing of web applications

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.455331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.208330Z digest=sha256:70eee6ecdfe6db141e9b73156de995ac39114dcd6dbd3774a48ff807144d8952

Observation 8bfbeba8-34bc-4404-8c84-a03a88d22991 · outbound

This paper cites A survey of unit testing practices.

Establishing Best Practices for Building Rigorous Agentic Benchmarks A survey of unit testing practices

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.245734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.295615Z digest=sha256:6764332adb3eff273d13a86eaf2c0f6b53e41a968ae1973989b7dab6a0a151d5

Observation adff7b59-e754-4dd2-b0e6-4130982c848a · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:38.017958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.364230Z digest=sha256:af7fba446c988b76d2a204016b569e2f120ab6eca4238d158a0b8f32eeb990b9

Observation 100925d5-acb5-4e73-a608-94591c4751d9 · outbound

This paper cites Strengthening ai agent hijacking eval- uations, 2025.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Strengthening ai agent hijacking eval- uations, 2025

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.816267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.483123Z digest=sha256:1713522a6d2894380347df1d64b39b0f4737a4ca9336005dea5ab6c239387ba5

Observation e2b3a539-22ca-43d0-911a-53dc5759bb66 · outbound

This paper cites Inference scaling flaws: The limits of llm resampling with imperfect verifiers.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inference scaling flaws: The limits of llm resampling with imperfect verifiers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:26.559726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:26.559726Z digest=sha256:7d5d5bd77388a8e86f4ef4c867f17c43f137cb69857306108d3d9767f9dbecf9

Observation 42f6f29f-a16c-448e-bbd4-d2fa8b248c44 · outbound

This paper cites Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.604621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.685912Z digest=sha256:1004604cac742b4ac6b0f96756a72de14d7ea735a4d0f875d6d41df626afe0c3

Observation 7b3638e3-ecb7-4851-8325-2d8cc4dc018b · outbound

This paper cites End-to-end integration testing design.

Establishing Best Practices for Building Rigorous Agentic Benchmarks End-to-end integration testing design

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.334722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.752737Z digest=sha256:23cfabb87a79043882546219c152515ae2b2417fe1fed528696e479bc1a72391

Observation a43f2b5b-cef1-4b2c-8f52-c4192b29f12c · outbound

This paper cites From imagenet to image classification: Contextualizing progress on benchmarks.

Establishing Best Practices for Building Rigorous Agentic Benchmarks From imagenet to image classification: Contextualizing progress on benchmarks

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:37.019808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:26.832387Z digest=sha256:172ab56b95298717ddbff48e4adb689aa5dd3122d4b8ac1b9c09943bc9965035

Observation fb92c734-fafe-4a5f-ba81-1a2c4dfd05a2 · outbound

This paper cites Introducing v0.5 of the AI Safety Benchmark from MLCommons.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing v0.5 of the AI Safety Benchmark from MLCommons

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:26.965009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:26.965009Z digest=sha256:2c1cf421f22a249dfeb9d350319ded146f4ed0445a19c212d2e8a5e792a30732

Observation 0a0fea27-b91a-4cf4-abb1-cb3662ed0c6d · outbound

This paper cites Evaluate & evaluation on the hub: Better best practices for data and model measurements.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluate & evaluation on the hub: Better best practices for data and model measurements

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.730961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:27.071959Z digest=sha256:4c9768b3b41a05e3c0facccde2d49a1d3de8a6a43df0fa5e4d3724b9b6469918

Observation bfecfd29-8d68-491c-9243-8e1afe8547f4 · outbound

This paper cites Structured testing: A testing methodology using the cyclomatic complexity metric, volume 500.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Structured testing: A testing methodology using the cyclomatic complexity metric, volume 500

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.494484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:27.172490Z digest=sha256:0570b66dee5691acc71dc34d9518adcbf3af5ea7b5b753b7d1988cf02be23a86

Observation 7e007b0c-0b86-4d42-94b7-84dd4f755e1e · outbound

This paper cites Measuring short-form factuality in large language models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring short-form factuality in large language models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.263615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.263615Z digest=sha256:9360bde57616ad70b1eb467fce5e5b008010fccb5a61fbdaf63e1086983d10c0

Observation 1211ce8f-51ba-4ab2-9e95-f279f69a42df · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Establishing Best Practices for Building Rigorous Agentic Benchmarks LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.342525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.342525Z digest=sha256:9bad35dd2b51a23d86099f6509f2c0921401127f311567fbb56e33a556f18c1e

Observation 36100b82-0a6a-453d-8fd8-7252d4bb2915 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Establishing Best Practices for Building Rigorous Agentic Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.421486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.421486Z digest=sha256:5270e44b7f17e984c8c64da0e931510f2e935114bd0f1e9b8f986e3dbf895c12

Observation f034431a-346d-4b7d-a505-30bc5b871529 · outbound

This paper cites Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.539316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.539316Z digest=sha256:be024550e5a1cb3c590c648f229c785d74603a7a85a373a32292ca9fbb7ae36b

Observation f10c0ec6-4971-4d89-a801-e30fc1a929be · outbound

This paper cites Grok 2 beta release, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Grok 2 beta release, 2024

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.319408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:27.664743Z digest=sha256:843f10cece16630f5ca2ffd8bd8fbdcec70317c0c919cb29548de26a4c500cf6

Observation d6ebfa40-b878-4f3b-9a60-21e7ff5c8832 · outbound

This paper cites Grok 3 beta — the age of reasoning agents, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Grok 3 beta — the age of reasoning agents, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:36.151243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:27.758444Z digest=sha256:dd131764636afbf3e9d918885a3df1ef2fab8035846909b319c641897b2b9d8c

Observation 9792836c-6497-490d-afd8-dc95cbbf3c57 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.822200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.822200Z digest=sha256:af850bb23f80882051388c7b9fe9a627667a1f859ca08f116751fba752e7fb47

Observation d4bb5d99-1118-437a-a708-e2affa5fc08c · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-agent: Agent-computer interfaces enable automated software engineering

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.927688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.927688Z digest=sha256:714a5e1dfe472ed5ea73c6bca895ba0049217ebb7a9422e3f69d7c84b619fa24

Observation 2c8fb028-6fbe-4b09-90f3-b6b4cf0f194d · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks React: Synergizing reasoning and acting in language models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.984396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.984396Z digest=sha256:b06df84ead1478d0eac435205fdcd01ea44b150fdb0355518cff0ec46912037c

Observation 198fd22f-9441-476c-a031-85aa2f20fae5 · outbound

This paper cites tau-bench: A bench- mark for tool-agent-user interaction in real-world domains.

Establishing Best Practices for Building Rigorous Agentic Benchmarks tau-bench: A bench- mark for tool-agent-user interaction in real-world domains

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.992119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:28.051730Z digest=sha256:4551bccac870ba6b1d07c1c8a34d7651fcfd0d4bca2adf04a7004af867138349

Observation d949f052-c124-4d97-82f6-19578d246cb3 · outbound

This paper cites Utboost: Rigorous evaluation of coding agents on swe-bench.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Utboost: Rigorous evaluation of coding agents on swe-bench

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.844195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:28.132834Z digest=sha256:e2a0485030228a40d0260b62eb51c9788eb27f82a103c08409effdde4e8aaf53

Observation 1a034863-6f59-44c1-a9f0-367e722d3b8b · outbound

This paper cites Evaluating large language models at evaluating instruction following.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating large language models at evaluating instruction following

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.693140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:28.231761Z digest=sha256:774600317837f31f591a0cb159deb62b968beb0a08b55669987ac4cc7b0e1546

Observation cb576ed1-53fe-4da0-ab9c-9c13e0aa1f7e · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.287235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.287235Z digest=sha256:cd19fdf169bf26466b46d707679dc2ae8d1a4f932bf07e98dbe8bacdadd2b487

Observation 684b05ef-1caf-4548-8578-8538f64cf57d · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.350185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.350185Z digest=sha256:44369624188c8ce7b1cde5d132cda0c5dd530602076f6bdfcf34e5442a1a8f7f

Observation 11652142-6a23-4b3c-b5d8-bee15e27fe51 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.438581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.438581Z digest=sha256:3cc6aac084ccda26179ff78b8932a1bee7c1e54210973e0512fdb2e153b558b4

Observation 4171e63d-1c90-405f-bfc6-697c68ccd494 · outbound

This paper cites X-webarena-leaderboard,.

Establishing Best Practices for Building Rigorous Agentic Benchmarks X-webarena-leaderboard,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.549044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:28.512244Z digest=sha256:8ab3094834b397903c16763f83a282507121ea8a2c67a86d6fe1d1db9d41e79c

Observation ef5223c0-1a4a-4a77-9b95-a4220eb1bba5 · outbound

This paper cites Webarena: A realistic web environment for build- ing autonomous agents.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Webarena: A realistic web environment for build- ing autonomous agents

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.173406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:28.701453Z digest=sha256:b177136f3f6cefb1b6d210d60d001c2768a6d511cf61228a5408857eaee5fcb8

Observation 2e34b126-d2fb-4384-ab41-5e41bc211521 · outbound

This paper cites Software unit test coverage and adequacy.Acm computing surveys (csur), 29(4):366–427, 1997.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Software unit test coverage and adequacy.Acm computing surveys (csur), 29(4):366–427, 1997

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.775241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.775241Z digest=sha256:31d1139a8878d672ac7ec2de4901396e17c4272e28c5a62417e2f5bd7b5fb30d

Observation 221e343c-f593-4261-8244-8f2cd2d7c753 · outbound

This paper cites Fuzzing: a survey for roadmap.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Fuzzing: a survey for roadmap

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:35.002208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:28.878325Z digest=sha256:2b018dbb2a267770643b583c72524ec671335eeda5a14576bb506692bce75876

Observation a67a9fdf-6975-4f4b-8497-98a52fbdd8aa · outbound

This paper cites CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities.

Establishing Best Practices for Building Rigorous Agentic Benchmarks CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.958503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.958503Z digest=sha256:e4bc86d2d006c2cb895557f325e7e22f359c8dabd0789bd8bbe9bb02b571e526

Observation 2e183348-8063-425d-854b-dd8490835ce7 · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Agent-as-a-Judge: Evaluate Agents with Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:29.053233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:29.053233Z digest=sha256:79b0388729d03fb3baf4f649b23397533eb0a5f02534eaa30daef21e2e4feeb8

Observation f002e32b-3947-4bee-b6a8-e09f77291753 · outbound

This paper cites Can large language models transform computational social science? Computational Linguistics, 50 (1):237–291, 2024.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Can large language models transform computational social science? Computational Linguistics, 50 (1):237–291, 2024

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:24:34.833066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:29.137040Z digest=sha256:67d7b8c81560124951e4018aa043d787c88a1d024742b5343d6266e57458b277

Observation 8e8f9bf9-c634-4687-8e2a-6c2be41be038 · outbound

This paper cites an unresolved cited work.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:24:34.697054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:29.250524Z digest=sha256:bc7de7da3fbc5bc1b8de84dd693c0b125152c775018e35ce46938722cbb25fd9

Observation 60fb9d5b-7a52-4cc3-9cd7-7be29b252f51 · outbound

This paper cites an unresolved cited work.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work

Reference 101

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:24:34.520858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T20:24:29.338123Z digest=sha256:462816b138bff9dfa320e227edbdfb1d7e3e0fc9b00c38fcc86cf477d03976da

Pith citing papers

Observation 66ecafb5-6e96-4e31-a056-9461b2744708 · inbound

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces cites this paper.

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:37:07.992049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T03:37:07.841385Z digest=sha256:14ad3aed2471c3e1a6a8d5c779390fc41795fd9ed4656424895b30b185821da4

Observation a6403359-918a-4741-a1f5-4ba21eaec78d · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:cd5621f508ee2645ca22307e718d3f0a57853c46b8a8d5a9da9de9aa65c9dd9c

Observation cb12bd46-df97-4812-ba29-41819a3a7b8b · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.219867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.219867Z digest=sha256:b7f7291e048c8e1c60bb5c201774a1316ed9ccd43918769220bec5898ebc1295

Observation 65b39ab3-2771-400f-8549-74a87ee7986f · inbound

From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI cites this paper.

From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:51.819572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:47:09.928166Z digest=sha256:78ee43e954515dfa323fec79ae62edd0ba64e09f6a6ed02104bea307112748af

Observation c1e2136b-abc8-4404-9894-f3aec855f5b5 · inbound

Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI cites this paper.

Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:18.917983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T06:03:19.805349Z digest=sha256:e94654f0e5ac2e237bd5d287419b231a6b2f201a30e543ee9ca47d750d2bb963

Observation d47897a7-e545-4666-9bff-02f2187ace23 · inbound

BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists cites this paper.

BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:08.439353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T20:03:59.215134Z digest=sha256:07e6a0db8500511d7bb35122ae0e0934c066e3fcb73773210e021919aa0c8526

Observation 0890bdcf-531d-43e4-ab67-24038f357d5b · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.552635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:f2b87f49f9c4596f03f0122d10efa742b3961784680dcd17b43aa86fba4e12a9

Observation 33d5ddc6-5eda-4558-8d43-3812ad576fee · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:20:54.825579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T02:20:32.528550Z digest=sha256:f690cef77ebbf96f89ada96ab0ae77377d2a00235525a0a33ca37d0d463f448a

Observation 20d96c9f-295e-4276-9f73-5c280b826463 · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:49:29.231213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T21:49:16.239350Z digest=sha256:d7645fe98a8658610770268b74d3b0910eec37cb26fe3ded7cd1db59cb54ae8d

Observation 852088d3-e55a-41cd-be32-2580980085b6 · inbound

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios cites this paper.

SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T00:18:46.882265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:18:46.882265Z digest=sha256:3af2305810b9a201e685a3c7754454cf82789d77248fba529d358ae159cbee7d

Observation de977e8c-5c6c-438f-9eed-76228711bbfc · inbound

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents cites this paper.

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:26.103689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:12:37.970638Z digest=sha256:1234519a3d54b1772d095f7c3c231cfdef031d4e86a8a3322f4d3cdca4d9ddb2

Observation 927faacd-9882-41fd-a98c-8bd197aaf590 · inbound

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation cites this paper.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:24.002297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:f71965a07f40f179f990522288aae4ef9d434a066a290fb3d5c359db0b0b8019

Observation 54fc18c7-13b3-4f01-ba8a-8a9115d6c13d · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:32:56.775266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ab97188ac85048699214c571d4dac12e982140ae0923e30ef6c3b590fc95c68c

Observation 439b1478-b2b6-49fe-b683-513b6b3c67d9 · inbound

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle cites this paper.

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:37:35.612990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T18:34:39.997353Z digest=sha256:7315ae228fda052ed209aa8debc5bd6b842b2d44019490501192cfa7dd2b72ef

Observation d7d74fb1-4748-4cab-b90e-c60c8b55a756 · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:05:06.699293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:a900bf0d7f7957e9ab62fcc9779b26b02dbfc16696e2059db247090b0bf3e450

Observation b70d3b7a-b3ad-4582-9902-4c5ae2cc26c9 · inbound

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations cites this paper.

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:26:09.995466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T06:25:06.024542Z digest=sha256:ffaae8b013bf06a365bbba56334c67523cddf276c632d3c23cdacce7d62d8839

Observation 222d5930-73f8-4e41-83f8-2218daa2ec88 · inbound

Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard cites this paper.

Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:04:37.290934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T05:02:31.995894Z digest=sha256:0d3017cc85bc116bb982eae33eb421f3a8c378b23616a6402737fff4485245bf

Observation fea22811-59c0-4626-9166-913c7c9b35e2 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.934723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:d489ca58898454caf7e1e0e403940d29ecbce0bf077ac83c70add1f4be0f98e5

Observation a6c31d99-04db-4f81-9ad6-23df445ec369 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.120173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:953e0474a7f9b39f53db4191b119ddd1b805758232c499627f5fc4d7ef777f19

Observation 5ce31bad-c793-4d0b-9344-a1c96088f20e · inbound

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly cites this paper.

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:23:58.970057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T21:19:59.005348Z digest=sha256:eb5e74f2689ba388e9484a69c6a9d1f350481c750f99fbdf257ece915c41cf5d

Observation f77ef2c1-260a-4d68-925f-e5722106874f · inbound

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation cites this paper.

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 21

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T21:33:59.135091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T21:29:57.626069Z digest=sha256:ac9b92d92a35fca1527f668beae5c09e59e437564afce10576a8c9402ec8e511

Observation 366d0cce-affa-4028-baef-5e83e2554545 · inbound

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents cites this paper.

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:36:52.845154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T04:58:10.803420Z digest=sha256:158a7b5a6d44237a49ed219d8dc285089aa9190285742dbccf09ddfd863837e8

Observation 9390dd14-4f0c-450d-b790-a0e734359b67 · inbound

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework cites this paper.

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:59.932826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T23:42:20.304205Z digest=sha256:87cb5afb31f57c83cbc11a9efcdeb6389079a40b18b9edd3e344ea3a30ea52d0

Observation 99f7b5c4-9deb-474e-9b48-e3d3354158c1 · inbound

SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills cites this paper.

SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T00:59:52.952364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:59:52.952364Z digest=sha256:e728b202fa4d3fa7f936128dfb1b289c5a738c66eac007ab2605dc57193ae737

Observation ad9430a8-1636-4380-a69e-32bdca129aca · inbound

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning cites this paper.

Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-01T23:34:10.122008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:34:10.122008Z digest=sha256:68103e22fc51dbf87b79cc11c0118cee11488e34776daf420514cea65cdcf780

Observation df467229-2c23-4b23-8ba3-c1307cd628b6 · inbound

KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch? cites this paper.

KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch? Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:38.648918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:00:38.648918Z digest=sha256:75d49babe524c62d8cc5b92e8bf56e560a883339c616282d62a06399dcb96494

Observation 4f0e28b2-9010-4193-a5c6-d8e3470bbf1a · inbound

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery cites this paper.

RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:09.728336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:09.728336Z digest=sha256:cea310f71e6afcd18679257bcc3ee3729dfa9cdde68c9af147a59580b63d871b

Observation e573652e-ca9f-4e82-a458-090e7836e591 · inbound

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems cites this paper.

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:38.319817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:54:38.319817Z digest=sha256:92c5053cbd5fc62d2bb1cbc4564c5e2eb6d0ed90670d186060e117bba0129317

Observation 6f262067-0274-4d99-92e7-06ea467ed41a · inbound

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks cites this paper.

Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T06:49:28.621324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:49:28.621324Z digest=sha256:e0e07d41665c34bb907910950a8a35d94a62c985a4d4107253236f403a7bf0d9

Observation 65fb692f-d7ab-4f20-8921-75e04f349ef3 · inbound

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents cites this paper.

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T23:58:27.966370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:58:27.966370Z digest=sha256:aa927a280e235b5df5abd45f756ea9d407560e50a03f3203c3b7909d6bf92691

Observation df392865-68cb-4dc7-9285-56fc59986a67 · inbound

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks cites this paper.

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:24:59.912539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:24:59.912539Z digest=sha256:96eacbeefbba2a405ad1219156cc2b7f0758b3cc173473f788e00b22199a2e55

Observation 897a0c3a-b537-43ee-b4ea-d7abc981c559 · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:08.204255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:08.204255Z digest=sha256:ebf88d067fb1cc06bb903ed279becc623f1c1f8e2d3f6a8d1b72016eaf5df8a7

Observation 1d6a3c2d-ea8a-4eed-8871-45486bfd7d0f · inbound

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation cites this paper.

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:48.271679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:48.271679Z digest=sha256:40f719551bfb119f7dedbfc78a4b1114589f94f6e12091a99cfbdca45c2b1c5a

Observation 7738450d-3702-4acb-a2eb-5776ee593f3e · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:18:50.856828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:18:50.856828Z digest=sha256:0e39114b6e3f345c2bc9c8d0409e8af7925c5568b07e7aa92af0b3b0edacca63

Observation a2274c8f-1274-4080-9b81-a586d4d723cb · inbound

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation cites this paper.

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:24.726546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:16:24.726546Z digest=sha256:321d2e5cd831c2907a208418a74e03cb6e2cff08ff6d486efeba2b297ce8a7a7

Observation 0cf16eda-132d-4502-aa08-c8cad43c3488 · inbound

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents cites this paper.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.480747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.480747Z digest=sha256:3f4292dcc1095d7141d448f243b50045a801f520e631b75d198f7a15c5745ec9

Observation e2b186c1-b8e7-4440-af39-d8191b3b9363 · inbound

Agent Safety Should Be a Runtime Contract cites this paper.

Agent Safety Should Be a Runtime Contract Establishing Best Practices for Building Rigorous Agentic Benchmarks

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T14:19:14.004821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:19:14.004821Z digest=sha256:20cebb8cf4725c0614dbc7cade435179aace272c69fca2190ddd2f64aa646924