Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:29.338123Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 37 inbound Pith citation observations for arXiv:2507.02825.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:24:29.338123Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:19:14.004821Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T22:08:59.930877Z
100 of 118 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 498da9e5-9da9-4859-a1f0-e2235cb99aaf · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Inspect AI: Framework for Large Language Model Evaluations, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aea9622-5a6f-4b7e-87c6-96ad04a561b7 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt code editing benchmarks, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a79027-ed54-47d7-a5ac-f01f5ccc7ccb · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks o1 tops aider’s new polyglot leaderboard, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b486f1a-5d0e-437d-962b-35608d837b1b · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks The amazon nova family of models: Technical report and model card, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d636119-8fda-4571-83bd-dc153e8716cd · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Claude 3.5 sonnet, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1623b2cf-dbaa-48cb-80e9-ba7df2364368 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Claude 3.7 and claude code, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89439eb-2c9e-43cd-86d2-2223c0d8a01d · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Bird minidev - corrections, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59767821-da00-4625-abeb-e8755f7f14cf · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Program Synthesis with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb83d160-535e-4026-9385-50edf7de0464 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58917dc5-a8db-4654-8d3a-2a65a00bfc8e · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3080b94-bee8-43c4-a758-4a0b10a01473 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260b815f-fb0f-4913-ab43-8dc7a875ea78 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks AutoAgents: A Framework for Automatic Agent Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bccefc2-874b-4a7b-9091-43df7c83c51c · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78deee31-97dc-42af-af85-ad8058691468 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a80de66-d368-488c-ac16-16624a5bda46 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing deepseek v3, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7a4f70-66f9-45fb-8a76-6496e199b5b6 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Imagenet: A large- scale hierarchical image database
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3cf797-2521-4f5a-9c09-cc752903aa01 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Don’t label twice: Quantity beats quality when comparing binary classifiers on a budget
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f23249f-8e95-4619-8525-f6f7f568a9a5 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03881fb7-d8f6-4f03-b3eb-8aae63ca2da1 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks The design and operation of CloudLab
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06be70c0-2a6c-4b6d-9947-589ce3f0d3ab · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9413ed1-7af1-4651-a911-5f97e31cd44d · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Searching for computer vision north stars
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc04ec4-6484-4b78-9075-bd038d42f644 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d379cf5-0b6c-4177-a756-93c27c27eeeb · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks A classification of sql injection attacks and countermeasures
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4706c5b4-9b26-4c38-8cf0-7db4d811c7d3 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks More than marketing? on the information value of ai benchmarks for practitioners
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e102cf6-4a50-42e7-9ccf-abe7fefd4706 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc8141c-a335-4548-9956-ecba7fc35208 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks The extent and consequences of p-hacking in science
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36a63e39-97b0-4b90-995f-a6e0a43c8980 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring Massive Multitask Language Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638e800b-6160-41d1-ba91-ddee087bb69f · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks The design and analysis of benchmark experiments
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add6f924-0680-4f83-94e7-1440ddcdaba8 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Preventing server-side request forgery attacks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6739a092-f9f4-4199-be3a-b4940e5bab8f · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-bench: Can language models resolve real-world github issues? In ICLR, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79952336-47d2-4a0e-a240-a8843eeec48a · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-bench verified leaderboard, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be6db01-506c-4190-b663-1296790ee689 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks AI Agents That Matter
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122680c8-e8f9-4bcd-9ecf-6ea0d54319ac · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Gemini 2.0 is now available to everyone, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c028daa-0318-4b18-9949-f93bc0f69941 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Math-verify, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b01508-7d1a-4ad4-af77-594e0d319e56 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks The ai cuda engineer: Agentic cuda kernel discovery, optimization and composition
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 592213bb-34ee-436c-ab07-fbb055fc2201 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Challenges of end-to-end testing with selenium webdriver and how to face them: A survey
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abab65e-cefc-4dcb-8bcf-2afa33484a64 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d94836-4c1c-4215-a5d3-0db83908b2d9 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Leveraging large language models for nlg evaluation: Advances and challenges
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7754d152-ee3c-4cbd-9714-5689e29767f3 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Let’s verify step by step
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dce8c3f-afdf-480f-b6b9-62f1acf23957 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c272ff5-c8e7-4443-82c4-911beb380184 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e890901b-c91f-4e39-acfd-1e255458360b · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing llama 3.1: Our most capable models to date, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9b52ed-450a-4b93-9159-ef8f1d1a6967 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec503b05-2bc5-4edb-9c9c-ed9b724b7cb4 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating language-model agents on realistic autonomous tasks, 2023
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c66277-51aa-4e67-8bf9-db61072802e6 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Example protocol for running an ai agent evaluation, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93b5332d-9b40-4d92-ab18-7ae10a98af38 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring automated kernel engineering, 2025
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a356301a-7ba5-4f73-98eb-a48ec47c1579 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Gaia: a benchmark for general ai assistants
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5570e0-69c3-463d-bcd0-5dcb2af1dcfe · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78279a50-f525-4cc2-a122-712c5c892bb7 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Mixtral large 2, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f9dd39-5c42-497a-98f3-d3b1cb4c5ea5 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Preparedness framework (beta), 2023
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f21e953-864c-49ee-ab92-b677d6ad2eeb · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt-4o system card, 2024
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1009fd69-fc03-4e62-ba60-ceb30aa5ad3a · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o1 system card, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54095ed0-51e5-498c-8269-f597bd1b563a · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o1-mini, 2024
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d2e2abee-bbef-4aa6-9b39-bb8f5c230a7b · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Gpt-4o mini: advancing cost-efficient intelligence, 2025
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation aa6d0ab7-267f-4046-9da4-9fa7af0b1d45 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Computer-user agent, 2025
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2c55a375-688b-4b5b-9074-2484f531c990 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing deep research, 2025
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf002cd-0fce-4028-a4b0-39c10a8248ef · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing gpt-4.5, 2025
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 70b51f70-1084-44d1-a578-b7718c45ae58 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Openai o3-mini, 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3e8824d-5612-467b-953a-585dfe728c2a · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Kernelbench: Can llms write efficient gpu kernels? In ICLR 2025 Third Workshop on Deep Learning for Code (Best paper award), 2025
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e6cdc13f-34aa-4119-84d2-630125c893d2 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Bleu: a method for automatic evaluation of machine translation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936e1ddd-4052-4ee0-884d-8a68c49e34e3 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks A survey of flaky tests
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fba94f27-761d-4acd-a0d0-017ad16ac8c7 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating Cross-Domain Text-to-SQL Models and Benchmarks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e58958-ce9a-4adf-b53e-55b5051e5786 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f260b87-08dc-41ae-bf50-586b86aefa6f · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Ai and the everything in the whole wide world benchmark
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7a9bd51b-a180-4b3e-a5cb-1f3ab7cd18ad · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1c2b8a5e-347e-4952-a339-46d484dbfae8 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Analysis and testing of web applications
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8bfbeba8-34bc-4404-8c84-a03a88d22991 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks A survey of unit testing practices
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation adff7b59-e754-4dd2-b0e6-4130982c848a · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 100925d5-acb5-4e73-a608-94591c4751d9 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Strengthening ai agent hijacking eval- uations, 2025
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e2b3a539-22ca-43d0-911a-53dc5759bb66 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Inference scaling flaws: The limits of llm resampling with imperfect verifiers
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f6f29f-a16c-448e-bbd4-d2fa8b248c44 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Worldcoder, a model-based llm agent: Building world models by writing code and interacting with the environment
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7b3638e3-ecb7-4851-8325-2d8cc4dc018b · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks End-to-end integration testing design
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a43f2b5b-cef1-4b2c-8f52-c4192b29f12c · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks From imagenet to image classification: Contextualizing progress on benchmarks
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fb92c734-fafe-4a5f-ba81-1a2c4dfd05a2 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Introducing v0.5 of the AI Safety Benchmark from MLCommons
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0fea27-b91a-4cf4-abb1-cb3662ed0c6d · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluate & evaluation on the hub: Better best practices for data and model measurements
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bfecfd29-8d68-491c-9243-8e1afe8547f4 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Structured testing: A testing methodology using the cyclomatic complexity metric, volume 500
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7e007b0c-0b86-4d42-94b7-84dd4f755e1e · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Measuring short-form factuality in large language models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1211ce8f-51ba-4ab2-9e95-f279f69a42df · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36100b82-0a6a-453d-8fd8-7252d4bb2915 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f034431a-346d-4b7d-a505-30bc5b871529 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Understanding the Effects of Noise in Text-to-SQL: An Examination of the BIRD-Bench Benchmark
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f10c0ec6-4971-4d89-a801-e30fc1a929be · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Grok 2 beta release, 2024
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d6ebfa40-b878-4f3b-9a60-21e7ff5c8832 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Grok 3 beta — the age of reasoning agents, 2024
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9792836c-6497-490d-afd8-dc95cbbf3c57 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4bb5d99-1118-437a-a708-e2affa5fc08c · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Swe-agent: Agent-computer interfaces enable automated software engineering
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8fb028-6fbe-4b09-90f3-b6b4cf0f194d · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks React: Synergizing reasoning and acting in language models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198fd22f-9441-476c-a031-85aa2f20fae5 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks tau-bench: A bench- mark for tool-agent-user interaction in real-world domains
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d949f052-c124-4d97-82f6-19578d246cb3 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Utboost: Rigorous evaluation of coding agents on swe-bench
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1a034863-6f59-44c1-a9f0-367e722d3b8b · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Evaluating large language models at evaluating instruction following
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cb576ed1-53fe-4da0-ab9c-9c13e0aa1f7e · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 684b05ef-1caf-4548-8578-8538f64cf57d · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11652142-6a23-4b3c-b5d8-bee15e27fe51 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4171e63d-1c90-405f-bfc6-697c68ccd494 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks X-webarena-leaderboard,
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ef5223c0-1a4a-4a77-9b95-a4220eb1bba5 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Webarena: A realistic web environment for build- ing autonomous agents
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2e34b126-d2fb-4384-ab41-5e41bc211521 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Software unit test coverage and adequacy.Acm computing surveys (csur), 29(4):366–427, 1997
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 221e343c-f593-4261-8244-8f2cd2d7c753 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Fuzzing: a survey for roadmap
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a67a9fdf-6975-4f4b-8497-98a52fbdd8aa · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e183348-8063-425d-854b-dd8490835ce7 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Agent-as-a-Judge: Evaluate Agents with Agents
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f002e32b-3947-4bee-b6a8-e09f77291753 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Can large language models transform computational social science? Computational Linguistics, 50 (1):237–291, 2024
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8e8f9bf9-c634-4687-8e2a-6c2be41be038 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 60fb9d5b-7a52-4cc3-9cd7-7be29b252f51 · outbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Unresolved cited work
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 66ecafb5-6e96-4e31-a056-9461b2744708 · inbound
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a6403359-918a-4741-a1f5-4ba21eaec78d · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb12bd46-df97-4812-ba29-41819a3a7b8b · inbound
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b39ab3-2771-400f-8549-74a87ee7986f · inbound
From Governance Norms to Enforceable Controls: A Layered Translation Method for Runtime Guardrails in Agentic AI Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c1e2136b-abc8-4404-9894-f3aec855f5b5 · inbound
Beyond Task Success: An Evidence-Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d47897a7-e545-4666-9bff-02f2187ace23 · inbound
BioVeil MATRIX: Uncovering and categorizing vulnerabilities of agentic biological AI scientists Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0890bdcf-531d-43e4-ab67-24038f357d5b · inbound
TeamBench: Evaluating Agent Coordination under Enforced Role Separation Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 33d5ddc6-5eda-4558-8d43-3812ad576fee · inbound
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 20d96c9f-295e-4276-9f73-5c280b826463 · inbound
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 852088d3-e55a-41cd-be32-2580980085b6 · inbound
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de977e8c-5c6c-438f-9eed-76228711bbfc · inbound
PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 927faacd-9882-41fd-a98c-8bd197aaf590 · inbound
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 54fc18c7-13b3-4f01-ba8a-8a9115d6c13d · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 439b1478-b2b6-49fe-b683-513b6b3c67d9 · inbound
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d7d74fb1-4748-4cab-b90e-c60c8b55a756 · inbound
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b70d3b7a-b3ad-4582-9902-4c5ae2cc26c9 · inbound
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 222d5930-73f8-4e41-83f8-2218daa2ec88 · inbound
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fea22811-59c0-4626-9166-913c7c9b35e2 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a6c31d99-04db-4f81-9ad6-23df445ec369 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5ce31bad-c793-4d0b-9344-a1c96088f20e · inbound
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f77ef2c1-260a-4d68-925f-e5722106874f · inbound
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 366d0cce-affa-4028-baef-5e83e2554545 · inbound
Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9390dd14-4f0c-450d-b790-a0e734359b67 · inbound
AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 99f7b5c4-9deb-474e-9b48-e3d3354158c1 · inbound
SLBench: Evaluating How LLM Agents Follow Logical Relations in Skills Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad9430a8-1636-4380-a69e-32bdca129aca · inbound
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df467229-2c23-4b23-8ba3-c1307cd628b6 · inbound
KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch? Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0e28b2-9010-4193-a5c6-d8e3470bbf1a · inbound
RECEIPT: Deterministic, Reward-Hacking-Resistant Verification for White-Box Agentic XSS Discovery Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e573652e-ca9f-4e82-a458-090e7836e591 · inbound
The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f262067-0274-4d99-92e7-06ea467ed41a · inbound
Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65fb692f-d7ab-4f20-8921-75e04f349ef3 · inbound
SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df392865-68cb-4dc7-9285-56fc59986a67 · inbound
Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897a0c3a-b537-43ee-b4ea-d7abc981c559 · inbound
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6a3c2d-ea8a-4eed-8871-45486bfd7d0f · inbound
LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7738450d-3702-4acb-a2eb-5776ee593f3e · inbound
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2274c8f-1274-4080-9b81-a586d4d723cb · inbound
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf16eda-132d-4502-aa08-c8cad43c3488 · inbound
The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b186c1-b8e7-4440-af39-d8191b3b9363 · inbound
Agent Safety Should Be a Runtime Contract Establishing Best Practices for Building Rigorous Agentic Benchmarks
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.